Prognosis prediction method of peripheral T cell lymphoma and related equipment
By acquiring gene expression data of peripheral T-cell lymphoma, screening differentially expressed genes and target genes, and using deep learning models for prognostic prediction, the problem of inaccurate prognosis in existing technologies has been solved, the accuracy of prognostic prediction has been improved, and the precise formulation of treatment plans has been facilitated.
Patent Information
- Application Number
- CN202410866497.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2025-12-30
AI Technical Summary
Current technologies are insufficient to provide accurate prognostic predictions for peripheral T-cell lymphomas, leading to imprecise treatment plans and impacting patient outcomes.
By acquiring gene expression data of peripheral T-cell lymphoma, differentially expressed genes and prognostic target genes, such as CDKN2A, GATA3, TBX21, and TET2, were screened. A deep learning model was used for prognostic prediction, and a survival analysis model was constructed by combining the characteristics of differentially expressed genes and target genes.
It improves the accuracy of prognostic prediction for peripheral T-cell lymphoma, helps doctors develop more effective treatment plans, and improves disease management for patients.
Smart Images

Figure CN121237392A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a prognosis prediction method for peripheral T-cell lymphoma and related equipment. BACKGROUND
[0002] With the development of sequencing technology, the application of RNA-seq (RNA Sequencing) in analyzing disease characteristics is becoming more and more widespread. RNA-seq can analyze RNA molecules in cells and provide information on changes in gene expression.
[0003] Changes in gene expression are an important indicator of prognosis prediction. Prognosis prediction can assist clinicians in developing treatment plans for patients, helping patients clearly understand their own disease conditions, and indirectly improving the quality of the treatment process. Therefore, accurate prognosis results have important value for medical research and clinical practice. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a prognosis prediction method for peripheral T-cell lymphoma and related equipment to solve or partially solve the above problems.
[0005] To achieve the above purpose, the first aspect of the present application provides a prognosis prediction method for peripheral T-cell lymphoma, comprising:
[0006] obtaining gene expression data and target genes of the peripheral T-cell lymphoma, the target genes including genes related to the prognosis of the peripheral T-cell lymphoma;
[0007] obtaining differential genes according to the gene expression data;
[0008] performing prognosis prediction based on the differential genes and the target genes to obtain a prognosis result of the peripheral T-cell lymphoma.
[0009] The second aspect of the present application provides a prognosis prediction device for peripheral T-cell lymphoma, comprising:
[0010] an acquisition module configured to obtain gene expression data and target genes of the peripheral T-cell lymphoma, the target genes being genes that have an impact on the prognosis of the peripheral T-cell lymphoma;
[0011] a processing module configured to obtain differential genes according to the gene expression data;
[0012] a prediction module configured to perform prognosis prediction based on the differential genes and the target genes to obtain a prognosis result of the peripheral T-cell lymphoma.
[0013] In a third aspect, the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor implements the method according to the first aspect when executing the program.
[0014] In a fourth aspect, the present application provides a non-transitory computer readable storage medium storing computer instructions for causing a computer to execute the method according to the first aspect.
[0015] From the above, it can be seen that the present application provides a prognosis prediction method for peripheral T-cell lymphoma and related equipment. The method comprises: obtaining gene expression data of the peripheral T-cell lymphoma and target genes, the target genes comprising genes related to the prognosis of the peripheral T-cell lymphoma, obtaining differential genes according to the gene expression data, and performing prognosis prediction based on the differential genes and the target genes to obtain a prognosis result of the peripheral T-cell lymphoma. By using differential genes and genes related to the prognosis of peripheral T-cell lymphoma for prognosis prediction, the accuracy of the prognosis result can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present application or related art, the drawings needed in the embodiments or related art description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0017] Figure 1 A schematic diagram of an exemplary system 100 according to an embodiment of the present application is shown.
[0018] Figure 2 A flowchart of an exemplary prognosis prediction method 200 according to an embodiment of the present application is shown.
[0019] Figure 3 A structural diagram of an exemplary deep learning model 300 according to an embodiment of the present application is shown.
[0020] Figure 4A An experimental result diagram of an exemplary survival analysis model one according to an embodiment of the present application is shown.
[0021] Figure 4B An experimental result diagram of an exemplary survival analysis model two according to an embodiment of the present application is shown.
[0022] Figure 5 A flowchart of an exemplary prognosis prediction method 500 for peripheral T-cell lymphoma according to an embodiment of the present application is shown.
[0023] Figure 6 A schematic diagram of an example peripheral T-cell lymphoma prognosis prediction apparatus according to an embodiment of the present application is shown.
[0024] Figure 7 A schematic diagram of an example electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0025] For the purposes of the present application, the technical solutions and advantages thereof are more clearly apparent, the following further describes the present application with reference to the specific embodiments and with reference to the accompanying drawings.
[0026] It should be noted that, unless otherwise defined, technical terms or scientific terms used in the embodiments of the present application should be understood as their common meanings to those of ordinary skill in the art to which the present application pertains. The terms "first", "second", and similar terms used in the embodiments of the present application do not denote any order, quantity, or importance, but are merely used to distinguish different components. The terms "include", "contain", and similar terms mean that the elements or objects before the terms encompass the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms "connect" and "connected" and similar terms do not mean only physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", and the like are merely used to indicate relative positional relationships, and when the absolute positions of the described objects are changed, the relative positional relationships can also be changed accordingly.
[0027] It can be understood that, before using the technical solutions of the various embodiments of the present application, the user will be informed of the types of personal information involved, the scope of use, the use scenarios, and the like by appropriate means, and the authorization of the user will be obtained.
[0028] For example, in response to receiving an active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require the acquisition and use of personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the electronic device, application program, server, or storage medium, and the like software or hardware that performs the operation of the technical solutions of the present application according to the prompt information.
[0029] As an optional but non-limiting implementation manner, in response to accepting the active request of the user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in the form of text. In addition, the pop-up window can also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0030] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation of the present application, and other methods that meet relevant laws and regulations can also be applied to the implementation of the present application.
[0031] As described in the background, prognosis prediction can assist clinicians in developing treatment plans for patients, helping patients clearly understand their disease conditions, and indirectly improving the quality of the treatment process. Therefore, accurate prognosis results have important value for medical research and clinical practice.
[0032] Therefore, the present application provides a prognosis prediction method for peripheral T cell lymphoma and related equipment. The method comprises: obtaining gene expression data and target genes of the peripheral T cell lymphoma, the target genes including genes related to the prognosis of the peripheral T cell lymphoma, obtaining differential genes according to the gene expression data, and performing prognosis prediction based on the differential genes and the target genes to obtain a prognosis result of the peripheral T cell lymphoma. By using differential genes and genes related to the prognosis of peripheral T cell lymphoma for prognosis prediction, the accuracy of the prognosis result can be improved.
[0033] Prognosis refers to the prediction of disease development based on experience. Prognosis prediction is an important branch in the medical field, which involves the estimation of disease development process, possible results, and possible responses of patients to treatment. Such prediction helps doctors and patients make more informed treatment decisions and provides patients with better disease management plans.
[0034] Prognosis prediction estimates the development process and results of a disease based on a patient's specific health condition, disease characteristics, genetic and molecular markers, and other clinical and pathological information. Prognosis prediction usually includes prediction of disease recurrence, progression, response to treatment, and patient survival.
[0035] Peripheral T cell lymphoma (PTCL) is a heterogeneous disease originating from post-thymic T lymphocytes or mature NK (Natural Killer) cells. In Europe and the United States, it accounts for about 10-15% of non-Hodgkin's lymphoma, while in China it accounts for about 21.4%. The median age of onset is about 50-60 years old, and men are more common. Patients with peripheral T cell lymphoma often show poor prognostic factors, and only 25% of patients survive more than 5 years after diagnosis. Cell immunotyping is an important auxiliary diagnostic tool for lymphoid tumors and has important value for predicting prognosis.
[0036] The prognosis of peripheral T cell lymphoma is influenced by many factors. In order to facilitate the development of more suitable rehabilitation treatment plans for patients, accurate prognosis results of peripheral T cell lymphoma have important value.
[0037] Figure 1 A schematic diagram of an exemplary system 100 according to embodiments of the present application is shown.
[0038] As shown in Figure 1 the system 100 can include a terminal device 102, a server 104 and a database server 106. In some embodiments, the server 104 can obtain data from the database server 106 in response to a request from the terminal device 102, and return the required data to the terminal device 102 after processing the data. As an optional embodiment, the server 104 can directly respond to the request from the terminal device 102 and return the required data to the terminal device 102.
[0039] The terminal device 102, the server 104 and the database server 106 can include a medium (e.g. a network) providing a communication link therebetween. The network can include various connection types, such as wired communication links, wireless communication links or fiber optic cables, etc.
[0040] Various application programs (APPs) can be installed on the terminal device 102, such as image processing application programs, video conferencing application programs, reading application programs, video application programs, social application programs, payment application programs, web browsers and instant messaging tools, etc., which can all be used to display content.
[0041] The terminal device 102 herein can be hardware or software. When the terminal device 102 is hardware, it can be various electronic devices with display screens, including but not limited to smart phones, tablet computers, e-book readers, MP3 players, laptops and desktop computers, etc. When the terminal device 102 is software, it can be installed in the above-mentioned electronic devices. It can be implemented as multiple software or software modules (e.g. to provide distributed services), or as a single software or software module. No specific limitation is made herein.
[0042] The server 104 can be a server providing various services, such as a background server providing support for various applications displayed on the terminal device 102. The database server 106 can also be a database server providing various services. It can be understood that in the case where the server 104 can implement the relevant functions of the database server 106, the database server 106 can not be provided in the system 100.
[0043] The server 104 and the database server 106 herein can also be hardware or software. When they are hardware, they can be implemented as a distributed server cluster composed of multiple servers or as a single server. When they are software, they can be implemented as multiple software or software modules (for example, used to provide distributed services) or as a single software or software module. No specific limitation is made herein.
[0044] It should be noted that the prognosis prediction method of peripheral T-cell lymphoma provided in the embodiments of the present application can be executed by the system 100. Specifically, it can be executed by the interaction between the terminal device 102, the server 104 and the database server 106. It can be understood that, in the case where the terminal device 102 is provided with the functions of the server 104 and the database server 106, it can also be executed by the terminal device 102. It should be understood that, Figure 1 The number of terminal devices, users, servers and database servers in the system 100 is merely illustrative. According to the implementation needs, there can be any number of terminal devices, users, servers and database servers.
[0045] The user 108 can use the services provided by the application installed in the terminal device 102. In an exemplary scenario, the user 108 can open any application installed in the terminal device 102, and can generate a request corresponding to the prognosis prediction function of peripheral T-cell lymphoma in the application by clicking the prognosis prediction function of peripheral T-cell lymphoma. For example, the server 104 can be the server corresponding to the application, and the terminal device 102 can send the request to the server 104. After receiving the request, the server 104 can obtain relevant data according to the request and generate a corresponding prognosis result of peripheral T-cell lymphoma according to the data, and can return the prognosis result to the terminal device 102, so that the user 108 displays the prognosis result through the terminal device 102.
[0046] Figure 2 A flowchart of an exemplary prognosis prediction method 200 according to an embodiment of the present application is shown.
[0047] In order to make prognosis prediction of peripheral T-cell lymphoma, as shown in Figure 2 In some embodiments, the differential genes and the target genes can be used as training data of a deep learning model, so as to obtain a survival analysis model for prognosis prediction. In this way, the differential genes and the target genes are used as inputs of the deep learning model, so as to avoid the problem that the input feature dimension of the model is single and the prognosis effect of the model is not ideal.
[0048] Differentially expressed genes can be identified by screening large amounts of gene expression data (bulk RNA-seq expression data) from peripheral T-cell lymphoma. In some embodiments, gene expression data can be obtained from the GeneExpression Omnibus (GEO) database. The GEO database is a public gene expression data repository designed to store and disseminate high-throughput gene expression data, including but not limited to microarray and RNA-seq data. In addition to gene expression data, clinical information, such as the patient's age of onset and sex, can also be obtained from the GEO database.
[0049] like Figure 2 As shown, after obtaining gene expression data, differentially expressed genes are identified by screening the data. In some embodiments, a linear model can be constructed for each gene, and weighted least squares can be used to statistically test the parameters of the linear model to assess whether the differences in gene expression are significant. Since this involves screening large-scale gene expression data, Bayesian methods can be used for multiple test corrections.
[0050] like Figure 2 As shown, differentially expressed genes can be screened using the Limma (Linear Models for Microarray Data) method. Limma is a statistical method for analyzing gene expression data, capable of handling large-scale gene expression data and obtaining reliable results. In some embodiments, enrichment scores of immune cells associated with peripheral T-cell lymphoma can be calculated based on gene expression data. Cell types of immune cells can then be screened based on these enrichment scores, and cluster analysis can be performed on these cell types to obtain multiple subtypes of peripheral T-cell lymphoma. Differential expression analysis can then be performed on these subtypes to obtain multiple subtypes of peripheral T-cell lymphoma. Specifically, Gene Set Enrichment Analysis (GSEA) can be used to calculate the enrichment scores of immune cells, thereby assessing whether the expression of specific gene sets in the gene expression data is significant, and these gene sets may be associated with specific immune cell types. Based on the significance of these characteristics and clinical experience, cell types with significant expression can be screened from a large amount of gene expression data. Cluster analysis based on these screened cell types can reduce the computational burden of cluster analysis. Cell screening can be performed using methods such as univariate risk analysis and correlation analysis.
[0051] The log2FoldChange method can be used to perform differential expression analysis on multiple subtypes. Log2FoldChange is a statistical measure describing changes in gene expression; it calculates the relative change in gene expression levels under two conditions, typically with a base of log2. Using the log2FoldChange method to perform differential expression analysis on multiple subtypes converts fold changes to a logarithmic scale, facilitating statistical analysis and interpretation.
[0052] like Figure 2 As shown, differential expression analysis and deduplication of gene expression data based on multiple subtypes of peripheral T-cell lymphoma can yield differentially expressed genes. In some embodiments, during differential expression analysis, differentially expressed genes among multiple subtypes can be screened based on preset significance values corresponding to gene expression data (e.g., significance values could be p < 0.05).
[0053] The differentially expressed genes with prognostic value identified through screening can include the following 272 differentially expressed genes:
[0054] TMEM255A、CHL1、DNASE1L3、LPAR1、PDCD1、NELL2、CXCR3、SOX8、ZFPM2、OLFM1、NPY1R、SLC22A3、IL18、VSIG4、HEY1、BEX2、ZNF10、PCAT7、SDC1、LINC02245、LAMP3、TFAP2A、IRX5、LAMC2、UGT2B17、PAPLN、C10orf90、ADRA2A、RBP5、TMEM150C、C2orf88、SLC38A11、KEL、UGT8、LINC01013、CYSLTR1、LINC02397、ITM2C、CD28、CP、FAM151A、TNF、MEOX2、SH3RF2、KRT86、MYO15B、TC2N、PLS1、FCRL3、MZB1、C7、CRTAM、CLEC4M、SPINT1、DNTT、NME8、LOC101927943、CRNDE、ARPP21、CD163、SGPP2、PAGE5、ICOS、DSP、GSTT1、RGS4、ABCB4、CELSR3、CHRM3-AS2、KLHL34、NTS、SCN3A、PLA2G2D、CPA5、WFDC21P、ADGRG2、TPSB2、LINC02375、LINC00643、RASSF6、TSPAN8、BIK、XKR4、FAM216B、PLD5、PRDM13、ZFP57、GZMK、CCL20、TPSAB1、LTF、FCAMR、ARMH1、CCL22、MMP1、CHIT1、IGFL2、MAGEA6、ZIC2、MMP12、ICA1、ATP13A4、LINC00839、PYGO1、CNN1、CCR4、IL11、FBLN1、NUDT9P1、OMD、DCN、RBP4、MARC1、ITGA9-AS1、MFAP5、CHRM2、EOMES、LINC01613、PLA2G2A、GAP43、SPRR1B、SERPINB3、AKR1B10、FAM217A、CST6、GPC6、PTGIS、C5orf46、SFRP2、FUT7、SFRP4、SIGLEC6、ANGPTL4、WT1、CCL13、CD19、LYPD3、SNTB1、SRPX2、PLN、STX11、FAM129C、IFI44L、KLRC4、LINC02555、SOX5、FABP4、ETNPPL、SAMD5、CCL11、LOC100506563、LINC01122、LINC01808、MIR17HG、ANKLE1, ZIC1, PODNL1, C8orf34-AS1, CORIN, ZBED2, HMCN1, ST7-AS2, FMO1, MMP3, LINC01553, SLC6A20, ARG1, MEGF11, LRRN1, MGST1, ANKRD22, LINC02148, GNLY, IL13RA2, BNC1, LILRA4, MME, DNAH8, LINC00472, DSC1, DSG2-AS1, DNAH14, PCDH8, OGN, SERPINB4, CHGB, EIF1AY, PKHD1L1, TNFRSF17, ASPN, BMS1P20, LGALS2, COMP, IGLV6-57, PROC, LGI2, CAPN6, LMOD1, BCL2A1, TMPRSS3, SHMT1, MAGEB4, TONSL, ADAM12, COTL1, SYDE2, CLIC6, IL12B, CNTNAP3P2, CLIP3, AGT, LIF, NTN4, DUSP4, RNF180, PCDHAC2, UST, SMIM6, RNF141, KIR3DL1, MAL, KIAA1549, FGF9, CD1A, CARMN, FOXD2-AS1, IL37, ATP8A1, KBTBD11, SNAP91, GPAT2, LINC01198, LOC102723493, TFPI2, ITIH3, NOG, KIR2DL5A, PCAT18, GTSF1, PCDHB16, GREB1, SERPINA3, LINC01771, IL21, FOXG1, KIAA1109, GABRG2, IGHV1-69, CARMIL1, LINC02274, FAM110C, CSRNP3, LRRN3, FCRL4, STMN2, BCHE, KIR2DS2, ASB2, ELOVL4, ZDHHC14, GJB2, GREM1, CCL17, IL22RA2, CENPV, SOX2-OT, DPP4, SMIM24, RRAD, KCCAT333, FSTL5, CCL26。
[0055] Such as Figure 2As shown, to address the issues of limited input features and unsatisfactory prognostic results in some implementations, differentially expressed genes and target genes influencing the prognosis of peripheral T-cell lymphoma can be used as inputs to a deep learning model, thereby training the model to obtain a survival analysis model. Target genes may include CDKN2A, GATA3, TBX21, and TET2. Combining features from 276 genes (differentially expressed and target genes) for prognostic prediction improves the accuracy of the results. The input includes both differentially expressed genes identified through data analysis and genes associated with peripheral T-cell lymphoma. Increasing the model's input information enhances its prognostic prediction performance.
[0056] The CDKN2A gene encodes two proteins, p16INK4a and p14ARF, both key factors in cell cycle regulation. p16INK4a inhibits cyclin-dependent kinases CDK4 and CDK6, while p14ARF stabilizes the p53 protein by negatively regulating MDM2. These functions are crucial for suppressing tumor development. Therefore, the state of the CDKN2A gene (such as mutation, loss, or expression level) can influence tumor development; for example, in some cases, loss of CDKN2A may predict faster disease progression and shorter survival.
[0057] GATA3 is a transcription factor that plays a crucial role in the differentiation and function of T cells, particularly helper T cells (Th cells) and follicular helper T cells (Tfh cells), which are key cells in the immune response and the tumor immune microenvironment. In the tumor microenvironment, GATA3 may influence tumor development and patient prognosis by affecting the infiltration and function of immune cells. Furthermore, genomic abnormalities, such as copy number gain or mutations, may occur in the chromosomal region where the GATA3 gene is located, potentially affecting GATA3 function and tumor development.
[0058] The protein encoded by the TBX21 gene is T-bet, a transcription factor of the T-box family, which is crucial for the differentiation and function of Th1 cells. Th1 cells play a central role in cell-mediated immune responses. However, the expression level of TBX21 may affect the number and activity of Th1 cells, thereby influencing tumor immune surveillance, which in turn affects tumor development and patient prognosis.
[0059] TET2 is a protein that acts as a DNA hydroxymethylase, participating in the regulation of DNA methylation and gene expression by converting 5-methylcytosine (5mC) into 5-hydroxymethylcytosine (5hmC) and other forms of hydroxymethyl derivatives. This activity is crucial for maintaining genomic stability and normal cellular function. In certain types of cancer, including some subtypes of peripheral T-cell lymphoma, the TET2 gene has a high mutation frequency. Mutations or loss of function in TET2 can lead to decreased tumor suppressor function, affecting cell differentiation and proliferation, thereby impacting tumor development and patient prognosis.
[0060] Therefore, using CDKN2A, GATA3, TBX21, and TET2 genes, which have an impact on the prognosis of peripheral T-cell lymphoma, as inputs to the model can increase the dimensionality of the model's input features, thereby making the model's prognostic prediction for peripheral T-cell lymphoma more accurate.
[0061] Figure 3 A schematic diagram of the structure of an exemplary deep learning model 300 according to an embodiment of this application is shown.
[0062] In some embodiments, the deep learning model 300 can be a survival analysis model, which is composed of a DeepHit network. The DeepHit network is a multi-task network that can be used to estimate the joint distribution of time and space. Figure 3 As shown, the network architecture of the deep learning model 300 can consist of a single shared subnetwork 302 and K specific causal subnetworks (e.g., specific causal subnetworks 304a and 304b), and uses a softmax layer as the final output layer to output the joint distribution of the K competing events learned by the model and the marginal distribution of each cause. The loss function uses survival time and relative risk.
[0063] The shared subnetwork 302 consists of fully connected layers (FC). In a convolutional neural network, a fully connected layer acts as a classifier, mapping the input data features one-to-one to the sample label space. The forward computation of a fully connected layer is a linear weighted process; the output of a fully connected layer can be seen as the product of each neuron and its weight coefficient in the previous layer, plus a bias.
[0064] In some embodiments, differentially expressed genes and target genes can be used as a first training dataset, and the first training dataset can be input into a pre-built survival analysis model (e.g., a deep learning model 300) for training based on the k-fold cross-validation method. Training the model using the k-fold cross-validation method helps improve the model's accuracy.
[0065] k-fold cross-validation methods can include five-fold cross-validation. In some embodiments of five-fold cross-validation, the first training dataset can be divided into five equal parts to obtain a second training dataset. Four parts of the second training dataset are randomly selected as the third training dataset and input into a pre-built survival analysis model for training. In five-fold cross-validation, the training dataset is divided into five equal parts, with four parts serving as the training set and the remaining part as the test set. Therefore, since the model is trained on five equally divided datasets and then evaluated on five separate test sets, this multiple evaluation reduces the randomness that might arise from a single test set and lowers the variance of the evaluation results.
[0066] In prognostic prediction using survival analysis models, in some embodiments, covariates X can be obtained based on differentially expressed genes and target genes. These covariates X are then input into a shared subnetwork 302 to generate a vector (e.g., a first vector) of latent factors with K competing events. This vector, along with the first vector, forms the output z of the shared subnetwork 302 (output z includes the covariates and the first vector). K specific causal subnetworks take the output z as input, learn the vector representing the covariates and the latent factors, and output the probability of the first hit time for a specific cause K (e.g., the target probability). The sum of these outputs is a joint probability distribution over the first hit event and time. Etiology-specific subnetworks learn in parallel the marginal distribution of the first hit time for each cause.
[0067] Survival analysis models do not rely on the assumption of any underlying stochastic processes; network modeling depicts the evolution of the relationship between covariates X and risk over time. The output y (e.g., y0) 1,1 y 1,2 ...y 1,Tmax and y 2,1 y 2,2 ...y 2,Tmax ) represents the (estimated) probability that the patient will experience event K at time s.
[0068] In some embodiments, the concordance index (C-index) can be used to assess prognostic outcomes and evaluate the predictive power of the survival analysis model. The C-index is the proportion of patient pairs in a study where the predicted outcome matches the actual outcome. It estimates the probability that the predicted outcome is consistent with the observed outcome. The C-index is calculated by randomly pairing all subjects in the studied data. In survival analysis, for a given patient, if the predicted survival time of the patient with the longer survival time is also longer than that of the patient with the lower predicted survival time, then the predicted outcome is considered consistent with the actual outcome.
[0069] To evaluate the prognostic effect of the survival analysis model in this application embodiment, the model can be trained under the same conditions using a first training dataset (including differentially expressed genes and target genes) and by using differentially expressed genes alone as input, thereby obtaining two survival analysis models. The prognostic effect of the model is evaluated by comparing the two survival analysis models.
[0070] In some embodiments, the dataset may include 136 data points, and the dataset may be randomly divided into a training set and a test set in a certain ratio (e.g., a 9:1 ratio), with the training set consisting of 123 data points and the test set consisting of 13 data points.
[0071] The embodiments of this application evaluate the performance of the two survival analysis models described above using different model evaluation metrics.
[0072] Survival analysis model one uses differentially expressed genes as input. The training dataset may include 136 data points, which are then randomly divided into a training set (123 data points) and a test set (13 data points) at a certain ratio (e.g., a 9:1 ratio). Five-fold cross-validation is used to train the model, resulting in survival analysis model one.
[0073] Survival analysis model two uses differentially expressed genes and target genes as inputs. The training dataset may include 136 data points, which are randomly divided into a training set (123 data points) and a test set (13 data points) at a certain ratio (e.g., a 9:1 ratio). Under the same conditions as survival analysis model one, the model is trained using five-fold cross-validation to obtain survival analysis model one.
[0074] The predictive power of the models was evaluated using the C-index. The optimal C-index for survival analysis model one was 0.6992, while that for survival analysis model two was 0.7038. A higher C-index indicates a greater consistency between the model's predictions and the actual results. Therefore, based on the above C-index results, it can be seen that survival analysis model two, using differentially expressed genes and target genes as inputs, provides better predictive results.
[0075] Figure 4A The experimental results of an exemplary survival analysis model according to an embodiment of this application are shown in the figure. Figure 4B The experimental results of an exemplary survival analysis model two according to an embodiment of this application are shown in the figure.
[0076] AUC (Area Under the Curve) is a metric used to evaluate model performance. Combined with... Figure 4A and Figure 4BThe experimental results are plotted with the horizontal axis representing time (days) and the vertical axis representing the evaluation index values. As can be seen, in the results for 400, 500, 600, 700, 800, 900, and 1000 days, the evaluation index values for survival analysis model one are 0.5040, 0.4563, 0.4921, 0.4921, 0.4921, 0.6042, and 0.6250, respectively; while the evaluation index values for survival analysis model two are 0.5000, 0.4762, 0.5476, 0.4881, 0.5556, 0.7042, and 0.6792, respectively. Therefore, survival analysis model two, which uses both differentially expressed genes and target genes as inputs, generally outperforms survival analysis model one, which uses only differentially expressed genes as inputs, indicating better robustness.
[0077] Combining the results of the two model evaluation metrics above, it can be seen that the survival analysis model trained using differentially expressed genes and target genes has a better prognostic effect than the survival analysis model trained using differentially expressed genes.
[0078] Figure 5 A schematic flowchart of an exemplary prognostic prediction method 500 for peripheral T-cell lymphoma according to an embodiment of this application is shown. Method 500 can be implemented by system 100 (e.g., Figure 1 The system 100 in the middle is executed and may include the following steps.
[0079] In step 502, gene expression data and target genes of the peripheral T-cell lymphoma are obtained, the target genes including genes related to the prognosis of the peripheral T-cell lymphoma.
[0080] In some embodiments, the target genes include CDKN2A, GATA3, TBX21, and TET2. Combining features from 276 genes (including differentially expressed genes and target genes) for prognostic prediction addresses the problem of unsatisfactory prognostic results when the model has only a single input feature.
[0081] In step 504, differentially expressed genes are obtained based on the gene expression data.
[0082] In some embodiments, obtaining differentially expressed genes based on the gene expression data further includes: calculating the enrichment fraction of immune cells associated with the peripheral T-cell lymphoma based on the gene expression data; screening cell types of the immune cells based on the enrichment fraction; performing cluster analysis on the cell types to obtain multiple subtypes of the peripheral T-cell lymphoma; and performing differential expression analysis based on the multiple subtypes to obtain the differentially expressed genes. Cluster analysis based on cell types screened from a large amount of gene expression data can reduce computational load.
[0083] In some embodiments, the differential expression analysis based on the multiple subtypes to obtain the differentially expressed genes further includes: determining the significance value of the difference corresponding to the gene expression data; and screening the differentially expressed genes among the multiple subtypes based on the significance value. By screening, genes with large differences among the multiple subtypes are obtained, and the model is trained based on the characteristics of these genes.
[0084] In step 506, prognostic prediction is performed based on the differentially expressed genes and the target genes to obtain the prognostic results of the peripheral T-cell lymphoma.
[0085] In some embodiments, the prognostic prediction based on the differentially expressed genes and the target genes to obtain the prognostic result of the peripheral T-cell lymphoma further includes: inputting the differentially expressed genes and the target genes into a survival analysis model for prognostic prediction to obtain the prognostic result of the peripheral T-cell lymphoma. Prognostic prediction combines both differentially expressed genes and target genes; the input includes both differentially expressed genes obtained through data analysis and genes associated with peripheral T-cell lymphoma. Increasing the information input to the model improves its prognostic prediction effectiveness.
[0086] In some embodiments, the survival analysis model is trained by: obtaining a first training dataset; and inputting the first training dataset into a pre-built survival analysis model for training using a k-fold cross-validation method. Using k-fold cross-validation to train the model helps improve its accuracy.
[0087] In some embodiments, the k-fold cross-validation method includes a five-fold cross-validation method. The step of training the pre-built survival analysis model using the k-fold cross-validation method further includes: dividing the first training dataset into five equal parts to obtain a second training dataset; randomly selecting four equal parts of the second training dataset as a third training dataset and inputting them into the pre-built survival analysis model for training. Since the model is trained on five equally divided datasets and then evaluated on five separate test sets, this multiple evaluation can reduce the randomness that may be introduced by a single test set and decrease the variance of the evaluation results.
[0088] In some embodiments, after inputting the differentially expressed gene and the target gene into a survival analysis model for prognostic prediction to obtain the prognostic result of the peripheral T-cell lymphoma, the method further includes: using a concordance index to assess the prognostic result to evaluate the predictive ability of the survival analysis model.
[0089] In some embodiments, the survival analysis model includes a single shared subnetwork and multiple specific causal subnetworks, the shared subnetwork including fully connected layers, and the specific causal subnetworks including fully connected layers. The survival analysis model can be used to estimate the joint distribution of time and space.
[0090] In some embodiments, the prognostic prediction based on the differentially expressed gene and the target gene to obtain the prognostic result of the peripheral T-cell lymphoma further includes: obtaining covariates based on the differentially expressed gene and the target gene; inputting the covariates into the shared subnetwork to obtain a first vector; inputting the covariates and the first vector into the plurality of specific causal subnetworks to obtain a target probability of hit time; and using the target probability as the prognostic result. The survival analysis model does not rely on any assumptions about underlying stochastic processes and can provide the evolution of the relationship between covariates and risk over time.
[0091] This application provides a method and related equipment for predicting the prognosis of peripheral T-cell lymphoma. The method includes: acquiring gene expression data and target genes for the peripheral T-cell lymphoma, wherein the target genes include genes related to the prognosis of the peripheral T-cell lymphoma; obtaining differentially expressed genes based on the gene expression data; and performing prognostic prediction based on the differentially expressed genes and the target genes to obtain the prognostic result of the peripheral T-cell lymphoma. By using differentially expressed genes and genes related to the prognosis of peripheral T-cell lymphoma for prognostic prediction, the accuracy of the prognostic result can be improved.
[0092] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.
[0093] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0094] Based on the same technical concept, and corresponding to any of the above embodiments, this application also provides a prognostic prediction device for peripheral T-cell lymphoma.
[0095] refer to Figure 6 The prognostic prediction device for peripheral T-cell lymphoma includes:
[0096] The acquisition module 602 is configured to acquire gene expression data and target genes of the peripheral T-cell lymphoma, the target genes including genes related to the prognosis of the peripheral T-cell lymphoma.
[0097] The target genes include CDKN2A, GATA3, TBX21, and TET2.
[0098] Processing module 604 is configured to obtain differentially expressed genes based on the gene expression data.
[0099] The processing module 604 is further configured to calculate the enrichment fraction of immune cells associated with the peripheral T-cell lymphoma based on the gene expression data; screen the cell types of the immune cells based on the enrichment fraction; perform cluster analysis on the cell types to obtain multiple subtypes of the peripheral T-cell lymphoma; and perform differential expression analysis based on the multiple subtypes to obtain the differentially expressed genes.
[0100] The processing module 604 is further configured to determine the significance value of the difference corresponding to the gene expression data; and to screen the differentially expressed genes among the multiple subtypes based on the significance value of the difference.
[0101] The prediction module 606 is configured to perform prognostic prediction based on the differentially expressed gene and the target gene to obtain the prognostic result of the peripheral T-cell lymphoma.
[0102] The prediction module 606 is also configured to input the differentially expressed gene and the target gene into a survival analysis model for prognostic prediction in order to obtain the prognostic result of the peripheral T-cell lymphoma.
[0103] The survival analysis model is trained by the following method: obtaining a first training dataset; and inputting the first training dataset into the pre-built survival analysis model for training based on the k-fold cross-validation method.
[0104] The k-fold cross-validation method includes a five-fold cross-validation method. The step of inputting the first training dataset into the pre-built survival analysis model for training based on the k-fold cross-validation method further includes: dividing the first training dataset into five equal parts to obtain a second training dataset; randomly selecting four equal parts of the second training dataset as a third training dataset and inputting them into the pre-built survival analysis model for training.
[0105] The prediction module 606 is also configured to evaluate the prognostic outcome using a consistency index.
[0106] The survival analysis model includes a single shared subnetwork and multiple specific causal subnetworks. The shared subnetwork includes a fully connected layer, and the specific causal subnetworks include fully connected layers.
[0107] The prediction module 606 is further configured to obtain covariates based on the differentially expressed genes and the target genes; input the covariates into the shared subnetwork to obtain a first vector; input the covariates and the first vector into the plurality of specific causal subnetworks to obtain the target probability of the hit time; and use the target probability as the prognosis result.
[0108] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.
[0109] The apparatus of the above embodiments is used to implement the corresponding method 500 in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0110] Based on the same technical concept, corresponding to any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method 500 as described in any of the above embodiments.
[0111] Figure 7 A schematic diagram of an exemplary electronic device according to an embodiment of this application is shown. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are internally connected to each other via the bus 1050.
[0112] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0113] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0114] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0115] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0116] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0117] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0118] The electronic devices described above are used to implement the corresponding method 500 in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0119] Based on the same technical concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to perform the method 500 as described in any of the above embodiments.
[0120] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0121] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the method 500 as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0122] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.
[0123] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0124] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0125] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.
Claims
1. A method for prognosis prediction of peripheral T-cell lymphoma, comprising: obtaining gene expression data of the peripheral T-cell lymphoma and target genes, the target genes comprising genes related to prognosis of the peripheral T-cell lymphoma; obtaining differential genes according to the gene expression data; performing prognosis prediction based on the differential genes and the target genes to obtain a prognosis result of the peripheral T-cell lymphoma.
2. The method of claim 1, wherein, The target genes comprise CDKN2A, GATA3, TBX21 and TET2.
3. The method of claim 1, wherein, The obtaining differential genes according to the gene expression data further comprises: calculating enrichment scores of immune cells related to the peripheral T-cell lymphoma according to the gene expression data; screening cell types of the immune cells according to the enrichment scores; performing cluster analysis on the cell types to obtain multiple subtypes of the peripheral T-cell lymphoma; performing differential expression analysis based on the multiple subtypes to obtain the differential genes.
4. The method of claim 3, wherein, The performing differential expression analysis based on the multiple subtypes to obtain the differential genes further comprises: determining differential significance values corresponding to the gene expression data; screening the differential genes between the multiple subtypes according to the differential significance values.
5. The method of claim 1, wherein, The performing prognosis prediction based on the differential genes and the target genes to obtain a prognosis result of the peripheral T-cell lymphoma further comprises: inputting the differential genes and the target genes into a survival analysis model to perform prognosis prediction to obtain the prognosis result of the peripheral T-cell lymphoma.
6. The method of claim 5, wherein, The survival analysis model is trained by the following method: obtaining a first training data set; inputting the first training data set into a pre-constructed survival analysis model for training based on a k-fold cross-validation method.
7. The method of claim 6, wherein, The k-fold cross-validation method comprises a five-fold cross-validation method, and the inputting the first training data set into the pre-constructed survival analysis model for training based on the k-fold cross-validation method further comprises: dividing the first training data set into five equal parts to obtain a second training data set; randomly selecting four equal parts of the second training data set as a third training data set to input into the pre-constructed survival analysis model for training.
8. The method of claim 5, wherein, After the inputting the differential genes and the target genes into the survival analysis model to perform prognosis prediction to obtain the prognosis result of the peripheral T-cell lymphoma, the method further comprises: evaluating the prognosis result by a consistency index.
9. The method of claim 5, wherein, The survival analysis model comprises a single shared subnetwork and multiple specific causal subnetworks, the shared subnetwork comprising a fully connected layer, and the specific causal subnetworks comprising fully connected layers.
10. The method of claim 9, wherein, The performing prognosis prediction based on the differential genes and the target genes to obtain a prognosis result of the peripheral T-cell lymphoma further comprises: obtaining covariates according to the differential genes and the target genes; inputting the covariates into the shared subnetwork to obtain a first vector; inputting the covariates and the first vector into the multiple specific causal subnetworks to obtain target probabilities of time of hit; taking the target probabilities as the prognosis result.
11. A device for prognosis prediction of peripheral T-cell lymphoma, comprising: an acquisition module configured to acquire gene expression data of the peripheral T-cell lymphoma and a target gene, the target gene being a gene having influence on prognosis of the peripheral T-cell lymphoma; a processing module configured to obtain differential genes according to the gene expression data; a prediction module configured to perform prognosis prediction based on the differential genes and the target gene to obtain a prognosis result of the peripheral T-cell lymphoma.
12. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, The processor implements the method in any one of claims 1 to 10 when executing the program.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to make the computer execute the method in any one of claims 1 to 10. The computer instructions are used to make the computer execute the method in any one of claims 1 to 10.