Biomarkers for predicting, diagnosing and differentiating of posttransplant lymphoproliferative disorders, use thereof, and associated computer-implemented method, system and related computer program product for predicting and differentiating of posttransplant lymphoproliferative disorders

A computer-implemented method using biomarkers and weighted Naive Bayes Classifiers effectively predicts and differentiates PTLD, addressing the limitations of existing methods by providing robust and efficient biomarker identification in small data volumes, enhancing diagnostic accuracy and patient management.

WO2026047495A1PCT designated stage Publication Date: 2026-03-05MUCHA KRZYSZTOF +2
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2026-03-05

Smart Images

  • Figure IB2025058483_05032026_PF_FP_ABST
    Figure IB2025058483_05032026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a computer implemented method for predicting the risk of posttransplant lymphoproliferative disorder as well as a computer implemented method for simultaneous predicting the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein Barr Virus (EBV)-positive from EBV-negative patients in posttransplant lymphoproliferative disorder and associated systems. The method for predicting the risk of posttransplant lymphoproliferative disorder PTLD, comprises the following steps: receiving information representative for expression level of biomarkers, acquired from a sample to be assessed, said biomarkers being at least three selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1, classifying said information representative for expression level of said at least three biomarkers, outputting the classification results, said results being indicative of whether the assessed sample belongs to one of two classes: PTLD or non-PTLD patient. The present invention provides further biomarkers for predicting, diagnosing and differentiating of posttransplant lymphoproliferative disorders and use thereof.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] P32197PC00 / WAW 22-08-2025

[0002] Biomarkers for predicting, diagnosing and differentiating of posttransplant lymphoproliferative disorders, use thereof, and associated computer-implemented method, system and related computer program product for predicting and differentiating of posttransplant lymphoproliferative disorders

[0003] FIELD OF THE INVENTION

[0004] The present invention provides methods for predicting, diagnosing and differentiating of posttransplant lymphoproliferative disorders in a patient, monitoring the patient's response to lymphoproliferative disorders’ treatment, monitoring the progression or recurrence of lymphoproliferative disorders based on the use of a combination of disorder specific biomarkers. The subject of the invention is also the use of combinations of the above-mentioned disorder specific biomarkers. The subject of the invention is also computer-implemented method, system and related computer program product for predicting and differentiating of posttransplant lymphoproliferative disorders. The other subject of the present invention is a computer-implemented method and related computer program products for the identification of disorder specific biomarkers based on small data volume, including identification of biomarkers for posttransplant lymphoproliferative disorders.

[0005] BACKGROUND OF THE INVENTION

[0006] Posttransplant lymphoproliferative disorder (PTLD) is a life-threatening complication following transplantation [1-4], It is characterized by uncontrolled B or T cell proliferation occurring in transplant recipients, with pathological features ranging from mild polymorphic lymphocytes expansion to severe monomorphic large-cell non-Hodgkin lymphoma [3], Compared to lymphomas in the general population, PTLD is characterized by increased extranodal involvement, a more aggressive clinical course, and a poorer response to conventional therapy [5, 6], The estimated incidence of PTLD after solid organ transplantation ranges from 1% to 20% but varies with the type of allograft transplanted [1-3, 7], The highest incidence has been reported in the recipients of small bowel (17%) and lung (up to 20%) transplants [8, 9], By contrast, an analysis of more than 100,000 primary kidney transplant recipients (KTR) revealed that the 5-year incidence of PTLD was 0.84%

[0010] , In a single-center study conducted over 15 years, the incidence of PTLD was 0.62% in 2598 KTR and 1.67% in 1378 liver transplant recipients (LTR) [4], The age of the recipients also has an impact on PTLD incidence. The lifetime lymphoma risk of pediatric and adult TRs is 29 and 8 times P32197PC00 / WAW 22-08-2025 higher than the general population, respectively

[0011] , This might be partially explained by Epstein- Barr virus (EBV) involvement in PTLD pathogenesis. EBV infection is a risk factor and a cause of PTLD in more than 80% of B-lymphocyte phenotypic disorders, and less commonly in T-lymphocyte proliferation [3, 12], Therefore, pediatric patients who are frequently EBV-seronegative before transplantation are especially prone to developing PTLD. However, the multiple factors involved in PTLD development including the type of transplanted organ, type of proliferating cell, patient age, EBV status, and immunosuppression intensity or time from transplantation make the clinical and histopathologic PTLD picture highly variable. This fact has important implications. The number of PTLD patients, particularly with specific histopathologic types and in selected solid organ transplant recipients, is relatively small. Therefore, the statistical power of a single-center study is limited. The epidemiology and outcomes data from the meta-analyses and multicenter studies are highly biased due to the differences in patient populations as well as the diagnostic and therapeutic approaches between centers. Consequently, data linkage between transplant databases and largescale cancer registries allowing for better data collection and analyses are needed [13, 14, 15], Traditional modeling techniques to study various variables simultaneously can lead to overfitting of the data

[0016] , To enhance the abundance of clinical data and define questions for further research or to identify areas for diagnostic and therapeutic improvement, advanced analytical techniques based on artificial intelligence have emerged [17, 18], However, the applications of these methodologies to molecular datasets derived from transplant databases are limited [16, 19],

[0007] Method of mitigating EBV virus associated end-organ damage in patients are known in the prior art. Lor example, US patent application US20220332836 Al, discloses methods of preventing a human virus-associated disorder, in particular EBV-associated disorder, as well as to a method of controlling EBV load in a subject, in particular in a subject being at risk of developing such a disorder like solid organ transplantation patients by administering to said subject a therapeutically effective dose of a CD40 antagonist. Moreover, US20210171610 Al provides methods useful for treating or preventing post-transplant lymphoproliferative disorder (PTLD) in a mammal. In these methods, the antibodies or fragments thereof are particularly effective for PTLD in EBV-seronegative persons (i.e., not previously infected with EBV) and given within a few days of solid organ or bone marrow transplant. In these methods, multiple doses of the antibodies, or functional fragments thereof, may be given over time, for example, in persons undergoing solid organ transplants. P32197PC00 / WAW 22-08-2025

[0008] However, disclosed methods do not identify genetic signatures associated with PTLD in transplant recipients (TR), and that additionally would help distinguish EBV+ and EBV- PTLD patients.

[0009] Thus, the first object of the invention is to provide methods for predicting, diagnosing and differentiating of posttransplant lymphoproliferative disorders in a patient, monitoring the patient's response to lymphoproliferative disorders’ treatment, monitoring the progression or recurrence of lymphoproliferative disorders based on the use of a combination of disorder specific biomarkers that would be credible and simple as well as would allow early detection of risk related with more severe development of disease.

[0010] One should also remember that PTLD is also one of rare disorders. Due to restricted accessibility of data, currently available computer implemented methods for identifying disorder specific biomarkers are not efficient and credible because of highly noised data which have to be processed in order to identify biomarkers.

[0011] For example, there is known from the EP2864920 Bia computer program that can identify a specific pattern in genes that can indicate if a person has a disease or not. It works by using a training data set of gene expression data and a training class set that indicates which samples have the disease and which do not. The program then generates a classifier based on this training data. It can then take a test data set and classify it using this classifier to determine if the samples have the disease or not. The program also has a step where it transforms the training data set and the test data set by subtracting a certain amount from each sample. This helps to make the classifier more accurate. It does this transformation multiple times and generates a new classifier each time. If the new classifier is better than the previous one, it is stored and used for further classification. In the end, the program outputs the best classifier as a biomarker signature, which can be used to distinguish if a person has the disease or not based on their gene expression data. Although the method can be used for small volume data, it is not resistant to overfitting if there are tens of thousands of descriptors in the set.

[0012] Thus, another object of the invention is to provide a new computer implemented method and related computer program products for identifying disorder specific biomarkers based on small volume data relating to at least one gene sequence contained in a set of patient’s samples, which would be robust, efficient and credible.

[0013] SUMMMARY OF THE INVENTION P32197PC00 / WAW 22-08-2025

[0014] According to a first aspect of the invention there is provided a computer implemented method for predicting the risk of posttransplant lymphoproliferative disorder PTLD, comprising the following steps:

[0015] - receiving information representative for expression level of biomarkers, acquired from a sample to be assessed, said biomarkers being at least three selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1

[0016] - classifying said information representative for expression level of said at least three biomarkers

[0017] - outputting the classification results, said results being indicative of whether the assessed sample belongs to one of two classes: PTLD or non-PTLD patient.

[0018] According to another aspect of the invention there is provided a computer program product for predicting the risk of posttransplant lymphoproliferative disorder, comprising instructions which when executed by a computer causes it to perform said method.

[0019] According to another aspect of the invention there is provided a computer implemented method for predicting the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein Barr Virus (EBV)-positive from EBV-negative patients in posttransplant lymphoproliferative disorder,

[0020] - receiving information representative for expression level of biomarkers, acquired from a sample to be assessed, said biomarkers being at least five selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS

[0021] - classifying said information representative for expression level of said at least five biomarkers

[0022] - outputting the results of the classification, said results being indicative of whether the assessed sample belongs to one of three classes: non-PTLD, PTLD EBV-positive, or PTLD EBV-negative patient.

[0023] Advantageously, classifying said information representative for expression level of biomarkers is performed by using an ensemble of weighted Naive Bayes Classifiers built sing all 126 possible combinations of five markers selected from the set of nine, wherein each individual classifier casts a vote for the object’s class, and the class that gets most votes is selected.

[0024] Advantageously, said biomarkers being HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS and wherein classifying said information representative for expression level P32197PC00 / WAW 22-08-2025 of biomarkers s is performed by using an ensemble of weighted Naive Bayes Classifiers built using all 84 possible combinations of six markers selected from the set of nine, wherein each individual classifier casts a vote for the object’s class, and the class that gets most votes is selected.

[0025] According to another aspect of the invention there is provided a computer program product for predicting the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein Barr Virus (EBV)-positive from EBV-negative patients in posttransplant lymphoproliferative disorder, comprising instructions which when executed by a computer causes it to perform the method according to claim 3.

[0026] According to another aspect of the invention there is provided a system for ex vivo predicting the risk of posttransplant lymphoproliferative disorder, comprising processing means (102) and memory means (106), characterized in that the processing means (102) are configured to perform the following steps:

[0027] - receive information representative for expression level of biomarkers, acquired from a sample to be assessed, said biomarkers being at least three selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1

[0028] - classify said information representative for expression level of said at least three biomarkers

[0029] - output the classification results, said results being indicative of whether the assessed sample belongs to one of two classes: PTLD or non-PTLD patient.

[0030] According to another aspect of the invention there is provided a system for ex vivo prediction of the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein Barr Virus (EBV)- positive from EBV-negative patients in posttransplant lymphoproliferative disorder, comprising processing means (102) and memory means (106), characterized in that the processing means (102) are configured to perform the following steps:

[0031] - receive information representative for expression level of biomarkers, acquired from a sample to be assessed, said biomarkers being at least five selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS

[0032] - classify said information representative for expression level of said at least five biomarkers P32197PC00 / WAW 22-08-2025

[0033] - output the results of the classification, said results being indicative of whether the assessed sample belongs to one of three classes: non-PTLD, PTLD EBV-positive, or PTLD EBV-negative patient.

[0034] Advantageously, the processing means (102) are configured to classify said information representative for expression level of biomarkers by using an ensemble of weighted Naive Bayes Classifiers built using all 126 possible combinations of five markers selected from the set of nine, wherein each individual classifier casts a vote for the object’s class, and the class that gets most votes is selected.

[0035] Advantageously, the processing means (102) are configured to classify said information representative for expression level of biomarkers by using an ensemble of weighted Naive Bayes using all 84 possible combinations of six markers selected from the set of nine, wherein each individual classifier casts a vote for the object’s class, and the class that gets most votes is selected.

[0036] According to another aspect of the invention there is provided a method of ex vivo diagnosing posttransplant lymphoproliferative disorder, characterized in that the method comprises the following steps:

[0037] (a) in vitro assessment of an expression level of at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 biomarkers selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, and COP1, in a sample obtained from a patient, and

[0038] (b) comparison of the assessed expression level of each biomarker with its mean expression level observed in a control sample or with a threshold expression level value determined in reference to control samples, and

[0039] (c) based on the comparison made in step (b) determining the presence or absence of posttransplant lymphoproliferative disorder in a patient.

[0040] Advantageously, biomarkers, whose expression level is assessed in step (a) are selected form at least 5, at least 6, at least 7, at least 8, or 9 biomarkers selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, and GMDS, and in step (c) in addition to determining the presence or absence of the posttransplant lymphoproliferative disorder in a patient, the differentiation of EBV-positive and EBV-negative PTLD patients is made. P32197PC00 / WAW 22-08-2025

[0041] According to another aspect of the invention there is provided a use of biomarkers selected from at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, and COP1; in diagnosis of posttransplant lymphoproliferative disorder in a patient.

[0042] Advantageously, diagnosis of posttransplant lymphoproliferative disorder in the patient comprises ex vivo assessment of expression level of said biomarkers in a sample obtained from said patient.

[0043] In one preferred embodiment, the biomarkers are selected from at least 5, at least 6, at least 7, at least 8, or 9 from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, and GMDS, and wherein in addition to diagnosis of the posttransplant lymphoproliferative disorder in a patient, the differentiation of EB V-positive and EB V-negative PTLD patients is made.

[0044] According to another aspect of the invention there is provided a computer-implemented method for the identification of disorder specific biomarkers based on small volume data relating to at least one gene sequence contained in a set of patient’s samples, the method comprising:

[0045] - a step of providing a set of small volume data relating to gene expression levels labeled so as the set of small volume data forms at least two groups

[0046] - a step of identification of potential biomarkers that carry information about differences between the groups, using the all-relevant feature selection algorithms;

[0047] - a step of selection of the most representative set of biomarkers; and

[0048] - a step of analysis of the interaction network between selected biomarkers so as to identify disorder specific biomarkers.

[0049] Advantageously, the step of identification of potential biomarkers that carry information about differences between the groups, using the all-relevant feature selection algorithms comprises steps of:

[0050] - identification of relevant probes with the use of at least one filtering algorithms among: MDFS, t-test, u-test, Ward-test with the use of the same predetermined significance threshold so as to generate a list of significant variables relating to gene expression levels. P32197PC00 / WAW 22-08-2025

[0051] Advantageously, the step of selection of the most representative set of biomarkers comprises:

[0052] - applying the RAFS procedure to said list of significant variables relating to gene expression levels so as to receive a shorter list of probes relating to genes, said probes being cluster representatives that are at least once selected as a representative within the run using each of at least two hierarchical clustering algorithms,

[0053] - applying the Boruta algorithm to said list of significant variables relating to gene expression levels so as to generate a shorter list of informative probes corresponding to a number of unique genes,

[0054] - generating a final list of relevant variables relating to gene expression levels by taking a union of the most relevant probes obtained from the RAFS and Boruta algorithm.

[0055] Advantageously, the step of analysis of the interaction network between selected biomarkers comprises steps of:

[0056] - generating a sub-network of protein-protein interactions (PPIs) between seeds and their neighbors, the seeds being proteins obtained from the most relevant genes, the sub-network consisting of a number of nodes and a number of edges so as the PPIs in the network are statistically significantly enriched,

[0057] - clustering genes in the subnetwork into clusters using the MCL algorithm, with a predefined value of inflation parameter.

[0058] Preferably, the step further comprises analyzing enrichment of the interaction network based on a biological process hierarchy, advantageously, of Gene Ontology.

[0059] In one preferred embodiment, the step further comprises a step of validating the selected biomarkers by generating predictive models using a machine learning algorithm;

[0060] Advantageously, the step of generating predictive models using a machine learning algorithm comprises building classifiers, using at least the top two variables from the final list of relevant variables and data relating to probes associated with said at least two top variables, said variables being gene expression levels.

[0061] According to another aspect of the invention there is provided a use of the method for the identification of genes being biomarkers representatives for posttransplant lymphoproliferative disorder and enabling differentiation of Epstein Barr Virus (EBV)-positive from EBV-negative patients in posttransplant lymphoproliferative disorder. P32197PC00 / WAW 22-08-2025

[0062] SHORT DESCRIPTION OF FIGURES

[0063] Fig.1 present the flow diagram of the method according to the invention.

[0064] Fig.2 present an exemplary computer system for implementing the method according to the invention.

[0065] Figure 3a and 3b present networks of the most important genes at two levels of clustering with MCL algorithm: inf = 2.0 (3a) and inf = 1.5 (3b). The clusters in the top panel, starting from the top right clockwise, contained the following proteins / genes: #1 (dark tan) - IFITM1, IFIT5, and SHFL, #2 (teal) - GMDS, #3 (blue-gray) contained HMGBT, #4 (jungle green) - HSIA6. #5 (red) - CD1C, KIR2DL4, and CD300A- #6 (indigo) - GPHR, #7 (olive) -ELL3 and MLLT3-, #8 (jade) - SIG. #9 (emerald) - RFWD2 / COP1. In the lower panel, clusters #1, #2, and #3 were merged into one and cluster #4 was merged with cluster #5.

[0066] Figure 4 shows Prominent Gene Ontology terms from the biological process hierarchy: top left - G0:0009615 (response to virus); top right - G0:0006955 (immune response); middle left GO: 0002684 (positive regulation of the immune system process); middle right GO: 0043161 - proteasome-mediated ubiquitin-dependent protein catabolic process; bottom left - GO: 0050793 (regulation of developmental processes); bottom right - GO: 0010468 (regulation of gene expression).

[0067] Figure 5 presents expression differences between EBV+ and EBV- for IFIT5, IFITM1, SHFL, and KIR2DL4 markers. The top and bottom box plot in each panel corresponds to EBV+ and EBV- patients, respectively.

[0068] Figure 6 shows expression differences between EBV+ and EBV- for CD300A, CD1C, HSPA6, and SP3 markers. Each panel’s top and bottom boxplot corresponds to EBV+ and EBV- patients, respectively.

[0069] Figure 7 presents expression differences between EBV+ and EBV- for MLLT and ELL3 markers. Each panel’s top and bottom boxplot corresponds to EBV+ and EBV- patients, respectively.

[0070] Figure 8 presents expression differences between EBV+ and EBV- for HMGB1 and IRF1-AS1 markers. Each panel’s top and bottom boxplot corresponds to EBV+ and EBV- patients, respectively.

[0071] Figure 9 presents expression differences between EBV+ and EBV- for COP1, GMDS, GRHPR, TMEM163, SLC6A16, and GALNT10 markers. The top and bottom box plot in each panel corresponds to EBV+ and EBV- patients, respectively.

[0072] Figure 10 presents a plot showing graphically correlations between selected genes. P32197PC00 / WAW 22-08-2025

[0073] Figures 1 la-111 present test results for the classification tool based on Naive Bayes algorithms.

[0074] Figure 12a- 12b present exemplary test results for the classification tool based on random forest algorithm.

[0075] DETAILED DESCRIPTION OF THE INVENTION

[0076] The present invention will now be described in detail. First, appropriate definitions are provided. Only exemplary implementations of the present application are shown and described in the detailed description below. As will be appreciated by those skilled in the art, the contents of this disclosure enable those skilled in the art to make changes to the disclosed detailed embodiments without departing from the spirit and scope of the inventions to which this application relates. Accordingly, the description in the drawings and detailed description is merely exemplary and not limiting.

[0077] DEFINITIONS

[0078] The present technology is described herein using several definitions, as set forth throughout the specification. Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0079] As used herein, unless otherwise stated, the singular forms “a,” “an,” and “the” include plural reference. Thus, for example, a reference to “a gene expression level” is a reference to one or more gene expression level.

[0080] The terms "marker" and "biomarker" are used interchangeably herein and refer to a marker whose presence or specific amount in a sample collected from a subject, analyzed alone or in combination with other markers, allows for the determination of the probability of a particular course or outcome. For example, when one or more markers or biomarkers reach a certain level in samples collected from a subject, that level may signal that that individual has an increased likelihood of developing a particular disease, e.g., PTLD, or made it possible to differentiate Epstein Barr Virus (EBV)-positive from EBV-negative individuals, compared to individuals whose levels of the same markers or biomarkers are different. In some embodiments according to the invention, markers are relevant genes. In further embodiments, markers may also be proteins corresponding to the most relevant genes.

[0081] The term “prediction”, as used herein, refers to any method by which one of skill in the art can estimate and / or determine the occurrence of disease, such as PTLD, before it occurs. In accordance with the present invention, the term PTLD is preferably understood to specifically refer to PTLD occurring in P32197PC00 / WAW 22-08-2025 patients diagnosed with DLBCL, the most frequent PTLD subtype. Prediction can be made by analyzing the assessment of an expression level of specific set of markers or biomarkers in a sample from the patient. In one embodiment according to the invention, the prediction of PTLD is made by determining the assessment of an expression level of at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 markers or biomarkers selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1, in a sample from a patient, and comparison of the assessed expression level of each biomarker with its mean expression level observed in a control sample. In further embodiment according to the invention, the prediction of the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein Barr Virus (EBV)-positive from EBV- negative patients in posttransplant lymphoproliferative disorder, is made by the assessment of an expression level of at least of at least 5, at least 6, at least 7, at least 8, or 9 markers or biomarkers selected from the group comprising selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, in a sample from a patient, and comparison of the assessed expression level of each biomarker with its mean expression level observed in a control sample.

[0082] The term "diagnosis" refers to any method by which one of skill in the art can estimate and / or determine the presence or absence of a particular disease or condition (such as PTLD or Epstein Barr Virus (EBV)-positive and EBV-negative PTLD) in an individual. This term does not mean that it is possible to determine the presence or absence of a specific disease with 100% accuracy. Instead, one of skill in the art will understand that diagnosis refers to the increased likelihood that an individual has a particular disease. Diagnosis can be made by analyzing the assessment of an expression level of specific set of markers or biomarkers in a sample from the patient. In one embodiment according to the invention, the diagnosis of PTLD is made by determining the assessment of an expression level of at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 markers or biomarkers selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1, in a sample from a patient, and comparison of the assessed expression level of each biomarker with its mean expression level observed in a control sample. In further embodiment according to the invention, the diagnosis of the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein P32197PC00 / WAW 22-08-2025

[0083] Barr Virus (EBV)-positive from EB V-negative patients in posttransplant lymphoproliferative disorder, is made by determining the assessment of an expression level of at least 5, at least 6, at least 7, at least 8, or 9 markers or biomarkers selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, in a sample from a patient, and comparison of the assessed expression level of each biomarker with its mean expression level observed in a control sample.

[0084] The term “biological sample” or “test sample” as used herein, refers to, but is not limited to, any biological sample derived from, or obtained from, a subject. The sample may comprise nucleic acids, such as DNAs or RNAs and / or proteins. Examples of such samples include but are not limited to fluids, tissues, cell samples, organs, biopsies, etc. For example, a sample can be obtained directly from a blood or PTLD affected lymph node. In general, suitable samples include but are not limited to blood, plasma, saliva, serum etc. DNA or RNA can be extracted from such samples according to methods known in the art (for example using a protocol from Sambrook et al., Molecular Cloning: A Laboratory Manual, Second Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989). Next, to determine the presence and quantity of RNA in a biological sample RNA sequencing (RNA-Seq) can be performed. Additionally, changes in gene expression over time and / or differences in gene expression in different groups or treatments can be determined by any techniques known in the art, for example, by hybridization-based microarrays, serial analysis of gene expression, and nextgen sequencing of complementary DNA (cDNA).

[0085] Term “control sample” as used herein, in case of disease prediction and / or diagnosis refers to a reference biological sample, which is obtained from a non PTLD patient, preferably from transplant recipients (TR), more preferably from kidney transplant recipients (KTR) who do not present with post-transplant lymphoproliferative disorder. Term “control sample” as used herein, in case of differentiation of EBV+ PTLD TR or EBV- PTLD TR is obtained from a patient having a specific disease (for example from EBV+ PTLD TR or EBV- PTLD TR, preferably from liver, kidney, heart, lung or bone marrow transplant recipients). In the context of diagnosis and differentiation using the selected markers, said term may also be understood as a sample based on which a given parameter is determined, for example the expression level of a specific gene, which serves as a reference point for the values being compared in the test sample, or establishes values for a healthy or diseased subject, as for example threshold expression level values determined in reference to control samples. In such P32197PC00 / WAW 22-08-2025 cases, the control sample may originate from a healthy individual, a transplant recipient without PTLD (PTLD- TR, non PTLD TR), or a diseased individual, including PTLD+ TR.

[0086] The term ”patient” / “subject” refers to an individual, preferably human, who suffers from or is susceptible for development of a disease or is diagnosed in the direction of a disease, such as PTLD, preferably patients diagnosed with DLBCL, the most frequent PTLD subtype.

[0087] The term “EBV+ PTLD patient” or “EBV positive PTLD patient”, as used herein, refers to a patient who suffers from EBV infection, wherein the virus is in active or latent form. The term “EBV- patients” or “EBV negative PTLD patient” refers to patients in which the EBV infection was not identified.

[0088] The term ’’small volume data” refers to a set of data relating to a number of samples less than 30 per class.

[0089] The term “probe” refers to the RNA molecule that is complementary to the probe of the DNA microarray. Such molecules are usually specific to the unique gene, however, in some cases few homologous genes and / or pseudogenes may share the specific sequence. In most cases in the dataset there is a single probe for each gene, however, in some cases multiple probes, corresponding to different sections of RNA molecule, can be associated with a single gene. In such a case only the most informative probe is retained.

[0090] The term “variable”, or “feature” describes in general term any property of the object that can be used by the machine learning model. In the context of the current application the values of the variables (features) correspond to the level of expression of particular gene.

[0091] As mentioned previously, the number of patients with posttransplant lymphoproliferative disorder (PTLD), a life-threatening complication of transplantation, is relatively small. Therefore, the statistical power of single-center studies is limited.

[0092] Unexpectedly, it was shown that machine learning is a useful prognostic tool to identify pathogenic, risk, or prognostic factors for PTLD development and unique in terms of results that may not occur with traditional statistical techniques, especially for orphan diseases. It has been shown that there exists a set of genes for which expression profiles are consistently different between two groups of patients - these who have the variant of the disease that is related to the activity of the Epstein-Barr virus (EBV+) and those who have no activity of the Epstein-Barr virus (PTLD-). The former variant P32197PC00 / WAW 22-08-2025 is more severe with significantly more intensive development of disease and significantly worse prognosis.

[0093] The working hypothesis is that both types of PTLD differ both from each other and from non PTLD TR (lymphocytes, blood) in the levels of expression of key genes that are responsible for the disease course. For each gene X, Y, Z, that has different expression level in both types of PTLD there are three possible situations: a) a gene X has abnormal expression level in EBV+ group, and normal in EBV- group. b) a gene Y has normal expression level in EBV+ and abnormal in EBV- group. c) a gene Z has abnormal expression level in EBV+ and EBV- group, albeit different between these groups.

[0094] In all cases the genes that are relevant for discerning EBV+ and EBV- groups are also relevant for discerning between PTLD and non PTLD TR In general, the number of groups (including controls) and assumed existing differences in expression levels can be called a ‘context’ of the process of differentiation.

[0095] Based on the above mentioned assumptions genetic signatures associated with PTLD were found and a classifier that made it possible to differentiate Epstein Barr Virus (EBV)-positive from EBV- negative patients was built.

[0096] As a consequence, a set of markers or biomarkers according to the invention can be used for predicting and / or diagnosing posttransplant lymphoproliferative disorder. The set of markers or biomarkers can consist of HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1.

[0097] The inventors identified set of genes differentially associated with EBV infection. These genes carry information about differences between EBV+ and EBV- PTLD patients and shed light on possible molecular mechanisms leading to different disease development patterns, which would enable a more personalized patient approach.

[0098] An innovative approach applied in the present study enabled the inventors to identify the most important PTLD-related relevant genes, and reveal the differences between EBV+ and EBV- PTLD patients. Set of genes differentially associated with the EBV status of PTLD patients were identified. P32197PC00 / WAW 22-08-2025

[0099] Namely, IFITM1, SHFL, IFIT5, and KIR2DL4, which are associated with innate immune responses. IFN- stimulated proteins such as IFITM1 and SHFL exhibit antiviral activity against Dengue, West Nile, and hepatitis C virus

[0036] , IFITM1 may inhibit viral entry into the host cell cytoplasm

[0037] , SHFL can restrict expression of viral genes [38,39], IFIT5 mediates PPIs, forms multiprotein complexes with cellular and viral proteins and RNAs through their multiple tetratricopeptide repeat motifs, and can modulate nuclear factor kappa B signaling

[0040] , KIR2DL4 plays important roles in the functional regulation of natural killer cells

[0041] , Three genes, IFIT5, IFITM1, and KIR2DL4, are also associated with malignancy. IFITM1 is an oncogene involved in various cancers including glioma

[0042] , colorectal

[0043] , prostate

[0044] , lung

[0045] , neck

[0046] , and gastric cancers

[0047] and hepatocellular carcinoma (HCC)

[0048] , IFITM1 promotes tumor cell proliferation, inhibits cell death, and stimulates invasion and metastasis

[0048] , IFIT5 promotes the progression of renal

[0049] , prostate

[0050] , and bladder cancers

[0051] , Due to the association with human papilloma virus oncogene E6, IFIT5 may also be involved in the progression of oral squamous cell carcinoma

[0052] , Another gene, KIR2DL4, is expressed in non-small cell lung cancer (NSCLC) cells and is correlated with a poor prognosis

[0053] , It is also related to renal cell carcinoma development

[0049] ,

[0100] Next genes, CD300A, CD1C, HSPA6, and SP3, are involved in leukocyte activation and the type I IFN signaling pathway and are related to malignancy. CD300A is a transmembrane glycoprotein in leukocytes that plays important roles in regulating their activation, proliferation, differentiation, and migration

[0054] , High CD300 A expression is associated with suppression of glioblastoma

[0055] , NSCLC

[0056] , and breast cancer

[0057] , However, high CD300A expression predicts poor survival in patients with acute myeloid leukemia

[0058] , CD1C is structurally related to class I major histocompatibility complex molecules but presents lipid and glycolipid antigens to T cells

[0059] , CD1C is a marker in cervical cancer

[0060] and cervical squamous cell carcinoma

[0061] , and plays an antitumor role in NSCLS

[0062] , HSPA6 is a heat shock protein family member involved in cell cycle regulation, hormone induction, and housekeeping

[0063] , Under stress conditions, it is linked to misfolded polypeptides’ degradation machinery and protects cells against unnatural damage

[0064] , HSPA6 upregulation after treatment of bladder and colorectal cancers correlates with tumor suppression

[0065] , However, HSPA6 is a risk factor for early HCC recurrence and is related to its invasiveness and prognosis

[0065] , SP3 is a transcription factor overexpressed in multiple tumors and is a negative prognostic factor for patient survival

[0066] , P32197PC00 / WAW 22-08-2025

[0101] Further, ELL3 and MLLT3 genes, which are associated with cell differentiation and development. ELL3 is an elongation factor for RNA polymerase II 3 and plays an essential role in activating developmentally regulated genes by priming them to recruit the proper transcription initiation complex during cell differentiation

[0068] , Interestingly, it both regulates and is regulated by p53 protein, which is a key tumor suppressor

[0069] , MLLT3 regulates human hematopoietic stem cell self- renewal and engraftment

[0070] ,

[0102] Additionally, the following genes have also been identified: HMGB1 and IRF1. HMGB1 is a DNA chaperone that maintains the structure and function of chromosomes. Outside the cell, HMGB1 is a damage-associated molecular signal that plays multiple roles in tumor growth and response to cancer therapy [71,72], IRF1 regulates antiviral responses and is an important cancer suppressor cooperating with the p53 protein

[0073] ,

[0103] Other relevant genes were found. Among them, COP1 is a well-conserved E3 ubiquitin ligase that is involved in the ubiquitylation of various protein substrates and therefore regulates multiple cellular functions

[0074] , COP1 overexpression contributes to the accelerated degradation of p53 protein in cancer and attenuates its tumor suppressor function

[0075] , GMDS catalyzes the synthesis of a substrate for glycosylation and is a favorable prognostic marker in endometrial and renal cancer, whereas its expression is significantly upregulated in lung adenocarcinoma at both the mRNA and protein levels

[0076] ,

[0104] The genes identified in the present study as the most relevant for diagnosing PTLD, and differentiating between EBV+ and EBV- patients allowed the disclosure of a probable mechanism that led to differences in EBV+ and EBV- PTLD. Without being bound by any theory, the presence of EBV likely results in a strong antiviral response reflected in the increased expression of genes related to the innate immune response, particularly IFITM1 and IFIT5. However, EBV-triggered regulatory mechanisms in immunocompromised patients involve changes in the mechanisms of tumor suppression, including the suppression of IRF1 activity, by increasing expression of the antisense IRF1 gene and suppressing expression of HMGB1. This may lead to modified patterns of regulation of gene expression by MLLT3 and ELL3 genes and changes in the activity of effector genes such as COP1 involved in the degradation of proteins or GALNT10 and GMDS involved in glycosylation. Although the high expression of genes involved in the innate immune response can be expected, the involvement of HMGB1 or suppression of IRF1 are important disclosures. Very high conservation of the HMGB1 sequence in evolution clearly shows the very important role of this protein. P32197PC00 / WAW 22-08-2025

[0105] Statistical studies have confirmed that EBV+ and EBV- posttransplant DLBCL differ at the molecular level [22, 23, 35], However, the relevance of several genes that are highly informative has not been shown [22, 23], In particular, of the 18 genes identified as the most important in the current study, only 3 (IFITM1, HMGB1, and ELL3) were mentioned by Craig et al.

[0035] , Furthermore, Morscio et al.

[0022] identified 10 of 18 as differentially expressed genes and only 1 (CD300A) was discussed. Interestingly, two of the genes (IFITM1 and ELL3) mentioned by Craig et al.

[0035] were also identified by Morscio et al.

[0022] as being differentially expressed but were not discussed in detail, whereas HMGB1 was not disclosed at all.

[0106] Analyses of the classifiers built using expression levels of the selected genes unexpectedly showed that a small subset could be utilized as biomarkers in clinical procedures. While only two genes (IFIT5 and IFITM1) were sufficient to construct perfect classifiers for the current dataset, due to the small number of samples, the inventors recommended also including two of the four most relevant genes (CDIC and CD300A).

[0107] Since the disclosed genes very well separated the EBV+ and EBV- groups, with p < ICT16, just two genes were sufficient to achieve perfect separation of patients in both groups. Moreover, increasing the number of used biomarkers with different biological functions increases the robustness of the biomarker set.

[0108] In conclusion, the inventors identified all genes that carry information about differences between EBV+ and EBV- PTLD patients.

[0109] According to the invention, a method of ex vivo differentiation of Epstein Barr Virus (EBV)-positive from EBV-negative patients in posttransplant lymphoproliferative disorder, comprises several steps.

[0110] The first step is an in vitro assessment of an expression level of set of markers or biomarkers.

[0111] Regarding PTLD and non PTLD, the assessment of an expression level of at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 markers or biomarkers selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1- AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1, in a sample from a patient, is performed. Regarding Epstein Barr Virus (EBV)-positive and EBV-negative PTLD patients, the assessment of an expression level of at least of at least 5, at least 6, at least 7, at least 8, or 9 markers or biomarkers selected from the group comprising selected from the group comprising HSPA6, P32197PC00 / WAW 22-08-2025

[0112] CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, in a sample from a patient, is performed. The assessment of an expression level of specific markers or biomarkers in a sample from the patient is carried out using methods known in this field. For example, for genes, said methods comprises Real Time PCR, microarray analysis, northern blot tests and SAGE analyses, and for proteins, said methods comprises Western blot, Luminex assay, and others Protein Expression Systems.

[0113] In the second step a comparison of the assessed expression level of each biomarker with its mean expression level observed in a control sample is made.

[0114] It has been additionally noticed that all findings in relation to the method of identifying PTLD specific biomarkers might be useful also for another types of disorders, especially the ones for which only small volume data are available.

[0115] As shown in Fig. 1, the method for identifying disorder specific biomarkers according to the invention comprises several steps. The first step is a step of providing a set of small volume data relating to at least one gene sequence associated to a set of biological samples labeled so as the set forms at least two groups. Each group is assigned with a specific type of label, one type of label among at least two different types of labels being indicative of a non-presence of differentiated disorder.

[0116] The first step involves providing data relating to gene sequences relating to a set of biological samples. Data relating to a gene sequence are understood as gene expression levels and are derived from nucleic acid contained in a sample relating to a tissue or a cell.

[0117] A data object in the form of a file containing gene expression levels for the whole set of samples under investigation can be called a “descriptor file” 10. It contains a set of numerical values each value being indicative of measured specific gene expression for each probe related to a specific sample from a specific subject (human). As mentioned earlier, the level of expression gene is understood herein as a numerical value indicative of measured the intensity of light emitted by the probe after the laser illuminates the appropriate field in the microarray, or the number of mRNA molecules corresponding to the marker that have been sequenced in the case of RNA-seq technology, depending on the applied measurement method. One of typical methods for measuring gene expression levels is RT- PCR.

[0118] In one embodiment, the step mentioned above involves reading appropriate data on gene expression levels from a database, for example ArrayExpress database. Said database contains Robust Multi-array Average (RMA) log2 normalized gene expression data (Affymetrix HG-U133 Plus 2.0 array). It P32197PC00 / WAW 22-08-2025 means that in one embodiment, the measurements of gene expression levels have been done by a third entity and made publicly available as databases for computer analysis in the form that can be converted to descriptor files 10.

[0119] In another embodiment, said step of providing a set of small volume data involves a series of steps performed on at least one biological sample 1. Said step of providing data on gene expression levels is also used for any test sample. In particular, it starts by RNA extraction. However comprehensive gene expression investigation requires high-quality RNA extraction, in sufficient amounts for realtime quantitative polymerase chain reaction and next-generation sequencing. Thus, the person skilled in the art would be able to choose an appropriate method and their parameters depending on the required data quality and sample quality.

[0120] In general, a descriptor file 10 contain descriptors, namely data relating to different groups the difference between each should be found. Data are arranged in a table, in which for example each row represents a sample from a patient and each column represents a probe. Each cell in this table contains a numeric value. This numerical value is related to the expression level of the gene to which the probe is assigned. Depending on what experimental technology was used, this number may reflect a slightly different reading process. Each sample is assigned a label (a class) indicative of the type of the group to which it relates. The term group should be understood here as a type of samples that have the same known origin (tissue or cell of known properties). At least one group is a control group which means it relates to control samples. Other groups relate to subjects (patients) suffering from a specific disorder. The set of groups forms a context for the differentiation process.

[0121] Typically, next step in known methods involves providing a training data set. It means that among all data some data should become a training data set and the rest becomes a test data set. However, taking into account the fact that small volume of data is processed, here the scheme of the division of data into two sets was not used.

[0122] Next step of the method according to the invention is step of identification of potential biomarkers that carry information about differences between the groups, using at least two all-relevant feature selection algorithms, among which one is RAFS algorithm and another is Boruta algorithm;

[0123] This step involves running at least two feature selection algorithms, namely Boruta and RAFS using the provided data set to find probes that can be used to discern samples relating to at least two different groups. The input to both algorithms is the set of descriptors -values of expression of genes relating P32197PC00 / WAW 22-08-2025 to probes for each sample, as well as the class of the sample. Each algorithm then produces a ranked list of descriptors with the importance score as well as a baseline indicating the level of importance that is threshold separating descriptors that are truly relevant from those that may have acquired the relevance by chance.

[0124] Regarding details of said at least two algorithms used for feature selection, Boruta algorithm is a well- established algorithm for all-relevant feature selection. It is a wrapper over the Random Forest classification algorithm. It utilizes the importance score that Random Forest can assign to each of the variables (here descriptors being gene expression values read from probes). It extends the set of descriptors in such a way, that each original descriptor has a counterpart (a shadow) with identical values that are permuted randomly between objects. Therefore, by construction a shadow variable cannot carry any information about the decision.

[0125] Boruta then runs the Random Forest algorithm and compares importance of the original variables with that achieved by the most important shadow variable. The procedure is repeated multiple times, with different permutations of the shadow variables. The variables that consistently have higher importance than the highest scoring shadow variable are deemed relevant.

[0126] The algorithm then produces a ranked list of relevant variables, which can be used for building models.

[0127] Regarding RAFS, this is an algorithm proposed by the Inventors. The name is an acronym created from the full name: Robust Agglomerative Feature Selection. The algorithm is based on three key foundations. The first is clustering of the similar features and using only representatives of clusters. The second is basing the measure of dissimilarity on the difference of information about the decision variable. The dissimilarity measure is called STIG (Symmetrical Target Information Gain). The third is application of intensive resampling to discover the repeatable patterns.

[0128] The standard minimal- optimal feature selection algorithms are based on optimization of some target function on the training set. It can be a performance of a classifier, such as in RFE (recursive feature elimination) scheme, or joint information about the decision, such as in mRmR (maximal relevance minimal redundancy) and other algorithms based on information theory. In RAFS the alternative approach is adopted - the variables are clustered hierarchically without any optimization goal. Then, representatives of the clusters are selected using mutual information between variables and decision. Due to the hierarchical clustering one can examine several levels of clustering at the same time, by running classifiers using representatives from different clustering depths. Unfortunately, the clustering P32197PC00 / WAW 22-08-2025 procedure may not be stable for the small datasets. Therefore, the procedure is repeated several times using cross-validation scheme, with set of representatives selected at each turn. One should note, that both mutual information between variables and the dissimilarity measure (STIG) are computed individually for each sample in the cross-validation scheme, hence both the representatives of the clusters and their shape may vary significantly. The robustness of the result is achieved by selecting as the final set of variables those variables that are most often used as representatives at given clustering depth

[0129] Next step of the method according to the invention is a step of selection of the most representative set of biomarkers. This step can be called also a step of generating predictive models using a machine learning algorithm. This step involves selecting top features from at least two algorithms and combining them into a single set (a core set). The RAFS algorithm allows using 4 different methods for performing hierarchical clustering - minimal linkage, maximal linkage, average linkage and Ward linkage (MDFS, t-test, u-test, Ward-test). All 4 methods can be used here, each can be repeated many times, for example at least 100 times within the scheme of five-fold cross-validation. First the clustering depth in RAFS that is sufficient to obtain good predictive model is established. The depth depends on the strength of the signal coming from samples. Here the strength is measured by the quality of the model measured in terms of area under the receiver operator curve (AUC).

[0130] In the case of discerning between EBV+ and EBV- PTLD cases, the signal is strong, the AUC for classifiers built using with the help of two markers is already nearly 1.0 and therefore the optimal classifiers can be obtained using just two representatives of the clusters, hence the depth ‘two’ was used for said exemplary case. However different depths can be established at this step.

[0131] Therefore, the number representatives that can be collected is calculated as the multiplication of the number of the method used for performing hierarchical clustering (for example 4) and the number of repetitions and the number of folds used for cross validation and the integer number indicative of the clustering depth. For example, in the case of discerning between EBV+ and EBV- PTLD cases 4 x 500 sets of 2 representatives were collected.

[0132] The theoretical maximal number of counts for individual marker is the number calculated as the multiplication of the number of the method used for performing hierarchical clustering (for example 4) and the number of repetitions and the number of folds used for cross validation. This maximal number of counts can be used as reference for establishing the threshold value for considering the marker as truly representative. The threshold value should be set at value that is non-negligible fraction P32197PC00 / WAW 22-08-2025 of maximal number - for example one or two percent. In this way by setting a threshold value for counts proper selection of the most representative set of biomarkers is possible.

[0133] For example, in the case of discerning between EBV+ and EBV- PTLD cases the theoretical maximal number of counts for individual marker would be 2000, while the highest achieved number of counts for a single probe was 753, the sum for top four probes was 1763. The threshold value of counts for the initial inclusion into the core set was then set at 2.5 percent equivalent to 40 counts - only the biomarkers (probes) that were used as representatives more than this number were considered. Said threshold value of counts allowed for a distinct gap between markers included to the set of provisionally relevant and those that were not included. The lowest scoring marker included in the set of provisionally relevant markers had 47 counts, whereas the highest scoring non-relevant marker had 29 counts. The threshold was applied to the sum of counts of all probes that were corresponding to a single biomarker.

[0134] The step of selecting the most representative set of biomarkers comprises an additional sub-step. Namely, an additional condition for increased robustness of results is the consistency among clustering methods - only those probes that are selected at least once as representatives by all clustering methods are used

[0135] In other words, to pass from the bigger set of provisionally relevant markers to the smaller set of the most important markers, the markers selected as provisional by each clustering method (genes relating to probes) are then ranked separately for each list from the best to worst to estimate their importance using the same scale. This step starts by choosing the number of items to be on the ranked lists. Here the RAFS algorithm is crucial for this sub-step namely, the markers that are above the threshold of counts obtain the consecutive real number starting in raising order from 1 while the markers below the threshold of counts obtain the same rank, called here a ‘penalty rank’. The threshold of counts is set arbitrarily to be the one after which a specific ‘step’ (‘drop’) in the number of counts is observed, for example, after which the number of counts is less of more than 40 but not more than 57 in compare with the number of counts related to a previous item. In the case of discerning between EBV+ and EBV- PTLD cases the penalty rank number was set to 25. This penalty rank number should be set significantly higher than the number of markers included in the set of provisionally important variables in the RAFS algorithm, namely significantly higher than the score that can be obtained by multiplying the highest possible rank, for the items on the list above the threshold of counts in the RAFS, by the number of clustering algorithms. The rationale for such choice is to include in the final set only P32197PC00 / WAW 22-08-2025 markers that are highly ranked by at least one feature selection algorithm, and exclude those that are at the bottom of the provisionally relevant list in ranking by one algorithm and outside of the list for the other one.

[0136] In the case of Boruta the initial, namely the provisional set of candidates (genes relating to probes) for the second clustering method, can be received by limiting the result by a predetermined number of the highest scoring probes. Said predetermined number of the highest scoring probes should be the same as number of probes selected by RAFS algorithm. In the case of discerning between EBV+ and EBV- PTLD cases said predetermined number of the highest scoring probes was set to 16 highest scoring probes. Also, in this case the markers that are below said indicated threshold obtain the same rank, the rank being the same as used for the result of the previous ranking algorithm. Also, in this case an ‘importance threshold value’ has to be set to find the items on the shortened list to have a penalty rank. It means that in the case of discerning between EBV+ and EBV- PTLD cases all probes with lower values of importance were similarly ranked at penalty number 25.

[0137] In this step, the final score for each marker is computed as a sum of ranks by each method. The biomarkers with the summary score above a predetermined final threshold total rank score is included in the final core set. In the case of discerning between EBV+ and EBV- PTLD cases the threshold total rank score was set to 40 and based on that only those variables below this limit were included as in the final core set.

[0138] Such method is applied to allow greater diversity of the marker set - it allows for inclusion of both markers that were scored highly only in one method, and markers with average, but consistent scores in both feature selection methods. In this step all variables (probes related to genes) that were deemed relevant by Boruta and RAFS are merged into a single set of relevant variables. Additionally, the final single set (final list) of relevant variables is the one in which the probes not associated with genes are excluded. Namely, biological meaning of data is checked before passing to the next step.

[0139] Most of the selected genes carry very similar information; what is more, an analysis of such a large number of genes is likely to create a deluge of trivial results. To limit the problem, it was decided to limit the analysis to the most relevant genes and those which are known to interact directly with them. To this end, a core set of the most important genes is thus created.

[0140] Additionally, if required, the chosen set of the most important genes is validated in the step of building classifiers. This step involves building the classifier using the features (probes relating to genes) P32197PC00 / WAW 22-08-2025 selected in the previous steps. For this purpose, the Random Forest (RF) algorithm [5] can be used. RF is a machine learning classifier that can un- cover complex relationships amongst predictor variables to classify patients to an outcome while minimizing bias [6], Due to its very good performance, it is widely used in multiple applications, particularly on small samples. In most cases, the application of RF leads to classifiers that are among the best per- forming, and in many cases, the RF is the best performing one. What is more, the algorithm very rarely leads to bad classifiers. Finally, it does not require extensive adjustment of parameters and is resistant to overfitting [7],

[0141] The RF algorithm can be applied to the dataset described with two, three and up to fourteen top variables selected in the abovementioned procedure. The estimation of the quality of the classifiers can be performed for example with two hundred repeats of five-fold cross-validation, but the person skilled in the art will appreciate that those parameters can be chosen differently in each specific context. As a consequence, in each repeat, a different split of patients into different number of folds is generated randomly.

[0142] Then, in a predetermined number of iterations, for example five, a predetermined number of folds, for example four folds can be used as a training set, and one can be used as a test set. Therefore, one obtains one thousand quality tests of the classifier for each set of variables. The metric that can be used for that task is the area under the receiver operator curve (AUC). This is a robust metric that works well even for the unbalanced datasets and is not sensitive to the selection of the threshold between classes and therefore very useful for the automated analyses.

[0143] Next step of the method according to the invention is a step of analysis of the interaction network between selected biomarkers, namely biomarkers contained in the final core set. This step involves, creating a graph of connection between features (genes relating to selected probes) in a total set (namely the core set) using an appropriate database relating to known proteins that can be expressed from the genes, for example the STRING database.

[0144] The STRING database collects data about proteins from various sources, including pathways databases, gene ontology, databases on protein-protein interactions, and information about cooccurrence in the scientific literature.

[0145] When supplied with a final list of genes selected as the most important biomarkers, the STRING database identifies the proteins coded by these genes and creates the list of corresponding proteins. In the case of non-recognition, the missing proteins can be identified using other bioinformatical services, P32197PC00 / WAW 22-08-2025 such as GeneCards. As a result, the STRING database can create a graph, with proteins as nodes and verified connections from the database as edges.

[0146] Then the method passes to a sub-step of extracting, from said graph, a subgraph consisting of nodes belonging to the core set and immediate neighbours of the core set (a shell set).

[0147] While the core set consists of the most relevant markers (here considered as proteins related to selected genes), the other relevant markers that are connected to the core set are likely to perform similar biological functions and be involved in the same processes. Therefore, they are included in the analysis performed in the next sub-step which is a step of performing the enrichment analysis of the core set and shell set.

[0148] Said analysis can be performed automatically with the help of the STRING web interface. The STRING performs enrichment analysis of various types of annotations for the inputted set of proteins. Two types of annotations, namely Biological Process and Molecular Function are especially useful for the analysis of disorders, including PTLD biomarkers. For the latter, they reveal what is the biological background of the differences between two classes of PTLD.

[0149] First the terms that correspond to the core set are analyzed. The list of STRING terms contains terms at every level of detail, starting from the very general terms such as Cellular Process or Protein Binding, with thousands of proteins annotated with them, down to very specific ones such as Fc receptor mediated inhibitory signaling pathway, where only 3 proteins in the entire database are annotated with it. The goal of the analysis is to find terms that fulfil the following conditions:

[0150] 1. The terms are specific enough to be meaningful in the context of the problem. For example, the process: Defense response to virus is annotated to 252 proteins in the database, 18 of them are in the analysed set of biomarkers, the p-value for such enrichment is 5.35e-15.

[0151] 2. The terms are related to the problem. The term Defense response to virus is certainly relevant for distinction between forms of PTLD with different activity of EB virus.

[0152] 3. The proteins in the core term are annotated with this term. In the case of PTLD following members of the core are annotated with the term Defense response to virus: IFITM1, IFIT5, SHFL, IRF1.

[0153] 4. The terms are correlated with the clustering structure of the STRING database. In the case of Defense response to virus 16 out of 18 proteins annotated with it are grouped in a single cluster consisting of 20 nodes. P32197PC00 / WAW 22-08-2025

[0154] The analysis gives insight into the biological context of the studied phenomena and helps to validate the results of the previous steps. In some cases, this step can lead to the decision to make the final list of genes shorter because of lack of association with expected biological function.

[0155] An exemplary system configured to perform one or more, or all the methods according to the invention is shown in Fig. 2. The computer system 100 includes processing means 102, for example a central processing unit (CPU, also "processor" and "computer processor" herein) 102, which can be a single core or multi core processor, either through sequential processing or parallel processing. The computer system 100 also includes memory means 106, for example, a memory unit or device 106 (e.g., randomaccess memory, read-only memory, flash memory), a storage unit or device 109 (e.g., hard disk), a communication interface 111 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices, either external or internal or both, such as a printer, monitor, USB drive and / or CD-ROM drive. It also comprises the touch-screen interface, a mouse, track ball, or other type of pointing device, a keyboard, or some combination thereof, and is used to input data into the computer system 100. The memory 106, storage unit 109, communication interface 111 and peripheral devices 104 are in communication with the CPU 102 through a communication bus (solid lines), such as a motherboard. The storage unit 109 can be a data storage unit (or data repository) for storing data. The computer system 100 can be operatively coupled to a computer network ("network") 101 with the aid of the communication interface 111. The network 101 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 101 in some cases is a telecommunication and / or data network. The network 101 can include one or more computer servers, which can enable a peer-to-peer network that supports distributed computing. The network 101, in some cases with the aid of the computer system 100, can implement a client-server structure, which may enable devices coupled to the computer system 100 to behave as a client or a server. Other embodiments of the computer system 100 have different architectures.

[0156] The storage device 109 is a non-transitory computer- readable storage medium such as a hard drive, compact disk read-only memory (CD-ROM), DVD, or a solid-state memory device. The memory 106 holds instructions and data used by the processor (CPU) 102.

[0157] The computer system 100 can include or be in communication with an electronic display 108 that comprises a user interface (UI) 110. Examples of UI's include, without limitation, a graphical user interface (GUI) and web-based user interface. The graphics adapter (not shown) displays images and other information on the display 108. The network adapter (communication interface) 111 couples the P32197PC00 / WAW 22-08-2025 computer system 100 to one or more computer networks. The communication interface 111 can be configured to interface via one or more network devices with one or more networks, for example, Local Area Network (LAN), Wide Area Net-work (WAN) or the Internet through a variety of connections including, but not limited to, standard telephone lines, LAN or WAN links (for example, 802.11, Tl, T3, 56 kb, X.25), broadband connections (for example, ISDN, Frame Relay, ATM), wireless connections, controller area network (CAN), or some combination of any or all of the above. The communication interface 111 can include a built-in network adapter, network interface card, PCMCIA network card, card bus network adapter, wireless network adapter, USB network adapter, modem or any other device suitable for interfacing the computer system 100 to any type of network capable of communication and performing the operations described herein.

[0158] The network 101 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 101 in some cases is a telecommunication and / or data network. The network 101 can include one or more computer servers, which can enable a peer- to-peer network that supports distributed computing. The network, in some cases with the aid of the computer system, can implement a client-server structure, which may enable devices coupled to the computer system to behave as a client or a server.

[0159] The computer system 100 is adapted to execute computer program modules for providing functionality described herein. As used herein, the term "module" refers to computer program logic used to provide the specified functionality. Thus, a module can be implemented in hardware, firmware, and / or software. In one embodiment, program modules are stored on the storage device 109, loaded into the memory 106, and executed by the processor 102.

[0160] Types of computer systems 100 used by the entities ofFig. 1 can vary depending upon the embodiment and the processing power required by the entity. For example, the presentation model unit can run in a single computer 100 or multiple computers 100 communicating with each other through a network such as in a server farm. The computer system 100 can be a user electronic device or a remote computer system. The electronic device can be a mobile electronic device. The computer system 100 can lack some of the components described above, such as graphics adapters, and displays 108.

[0161] Provided herein is also a non-transitory computer readable medium comprising machine-executable code that, upon execution by one or more computer processors, implements anyone of computer - implemented methods mentioned in the description. P32197PC00 / WAW 22-08-2025

[0162] The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or it can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a pre-compiled or as-compiled fashion.

[0163] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0164] EXAMPLE 1 - MARKER SELECTION

[0165] 1. Methodology

[0166] 1.1 Dataset

[0167] The study was conducted using the Robust Multi-array Average log2 normalized gene expression data (Affymetrix HG-U133 Plus 2.0 array) from PTLD patient samples from the Array Express database (Accession No. E-GEOD-38885). These data were previously described by Morscio et al.

[0022] and Ferreiro et al.

[0023] to characterize the clinicopathological and molecular genetic characteristics of diffuse large B-cell lymphoma (DLBCL) and to identify the differences between EBV+ and EBV- DLBCL. The training data set included 54613 probes from 33 patients (23 EBV+ samples and 10 EBV- samples) diagnosed with DLBCL, the most frequent PTLD subtype. P32197PC00 / WAW 22-08-2025

[0168] 1.2 Analysis procedure

[0169] Due to the relatively small sample size, it became necessary to design the study in such a way as to increase its robustness. Thus, more variables than strictly necessary were included, two different feature selection methods were used, and a robust machine learning algorithm was applied. Moreover, detailed analyses of not only genes that were included in the final models but also of other genes that were indicated as relevant by the feature selection algorithms and their interactions were performed.

[0170] The analytical procedure consisted of four steps:

[0171] 1) identification of biomarkers that carry information about differences between the groups, using the all-relevant feature selection algorithms;

[0172] 2) selection of the most representative set of biomarkers with the help of robust and unbiased algorithms;

[0173] 3) building predictive models using the robust machine learning algorithm; and

[0174] 4) analysis of the interaction network between biomarkers.

[0175] 1.2.1 All-relevant feature selection

[0176] Four all-relevant feature selection algorithms were applied: / -test; / / -test; Multidimensional Feature Selection (MDFS) [24, 25] based on the information theory; and Boruta [26, 27], which is a wrapper method based on the RF machine learning algorithm

[0021] , The inventors decided not to apply the common criterium, where the average expression difference between populations should be at least k- fold for different k values used in the literature. Thus, markers with a relatively small k-fold change but small variance can not only be identified as relevant but are even ranked higher than those that are selected purely based on the k-fold change.

[0177] 1.2.2 Selection of the most relevant variables

[0178] In the case of filter algorithms, the Robust Aggregative Feature Selection (RAFS) algorithm was used, which is an extended version of a novel approach recently reported

[0028] , It aims to select a small set of relevant variables that are sufficiently numerous to build the best possible classifier. To this end, it performs hierarchical clustering of the relevant variables and performs classification using representatives of the clusters on the different levels of hierarchy. The procedure is carried out in multiple repeats of the cross-validation procedure. In the default application, the variables that appear most often in the set of representatives are then used to build the final model. However, in the current P32197PC00 / WAW 22-08-2025 study, the algorithm was used in a different manner - as a generator for the list of all variables that were selected as representatives. The relative importance of a variable was based on the number of times it was selected as a representative. The algorithm was executed 50 times for four hierarchical clustering algorithms. Each run included five-fold cross-validation, and two variables were returned each time. Therefore, each variable could be selected as the representative 4 x 50 x 5 x 2 = 2000 times. To reduce noise, only the variables identified as relevant at least once within the run using each of the four hierarchical clustering algorithms were included.

[0179] In the case of variables identified as relevant, the Boruta method was used. The Boruta relevance score is the average score from multiple rounds of RF importance, which is the estimate of the variable utility for building predictive RF models. Hence, selecting the top N variables gives a reasonable chance that the most informative variables will be contained in the selected set. Most of the genes carry very similar information, and the analyses of their large number are likely to create a deluge of trivial results. Therefore, the inventors decided to limit the analyses to the most relevant genes and those that interact directly with them. A core set of the most important genes was created. The list was obtained as a union of the 14 top genes identified by Boruta and the top 14 identified by RAFS from the results of the MDFS algorithm. Number 14 was selected arbitrarily after initial analyses of the data.

[0180] 1.2.3 Generation of the classifiers

[0181] The final selection was validated by building the classifier using the features selected in the previous steps. For this purpose, the RF algorithm

[0021] was used. RF is a machine learning classifier that can uncover complex relationships among predictor variables to classify patients to an outcome while minimizing bias

[0029] , Due to its very good performance, it is widely used for multiple applications, particularly for small samples. In most cases, the application of RF leads to classifiers that are among the best performing, does not require the extensive adjustment of parameters, and is resistant to overfitting

[0030] , The RF algorithm was applied to the dataset described with 2, 3, and up to 14 top variables selected in the abovementioned procedure. The estimation of the quality of the classifiers was performed with 200 repeats of five-fold cross-validation. In each repeat, a different split of patients into five folds was generated randomly. Then in five iterations, four folds were used as the training set, and one was used as the test set. Therefore, the inventors obtained 1000 quality tests of the classifier for each set of variables, using the area under the receiver operator curve (AUC) metric. P32197PC00 / WAW 22-08-2025

[0182] 1.2.4 Interaction network generation and analyses

[0183] The interactions between genes were analyzed using the STRING database [31, 32], In the first step, a network of interactions was created using all genes identified as relevant by Boruta and RAT'S. Then a subnetwork of the most important genes and their interacting partners were extracted from the larger network. Clustering analysis of this network was performed with the help of the Markov clustering (MCL) algorithm

[0033] ,

[0184] 2. Results

[0185] 2.1 Analysis of the most relevant probes and associated genes

[0186] A very large number of molecular probes were identified as relevant by three filtering procedures. In all three cases, filters identified more than 1000 relevant variables: 1478 for the MDFS, 1595 for the Z-test, and 2083 for the w-test. For all of these filters, the significance threshold was set at p = 0.05, with family-wise error rate correction for multiple testing

[0034] , As expected, these three lists of relevant variables were highly similar, with 968 variables present in all of them. The difference between the results of all filters lay among the less significant discoveries. Therefore, only MDFS was used in subsequent steps.

[0187] Applying the RAFS procedure to the results of the MDFS gave a shorter list of cluster representatives that were at least once selected as a representative. It consisted of 66 probes associated with 61 unique genes, and 7 probes were not associated with any gene. In six cases, there are two probes corresponding to a single gene. In two cases, probes were nonspecific, and a single probe was associated with two genes.

[0188] The Boruta algorithm identified 316 informative probes corresponding to 246 unique genes. In 185 cases, the relationship was unique - a single probe corresponded with a single gene. Thirty-one probes did not correspond with any gene, and twenty-nine probes were nonspecific and associated with more than one gene. Conversely, 34 genes were associated with two probes, 8 were associated with three probes, and 1 was associated with four probes.

[0189] The final set of markers was constructed by taking a union of the 15 most relevant probes obtained from either algorithm and excluding probes not associated with any gene. The final set included 18 probes, 1 of which was associated with one gene and one pseudogene, heat shock protein family A member 6 (HSPA6) and HSPA7, respectively. P32197PC00 / WAW 22-08-2025

[0190] A set of the most relevant variables obtained with the procedure are described in section titled “Selection of the most relevant variables” above. All of these probes, the associated genes, their order on the RAFS and Boruta lists, the difference in expression between EBV- and EBV+ patients, and information on whether these probes / genes have been reported in previous studies are presented in Table 1.

[0191] Table 1 P32197PC00 / WAW 22-08-2025

[0192] 20 biomarkers obtained from feature selection as a union of 15 biomarkers returned by Boruta and RAFS feature filters. Among them, 10 were among the 15 most relevant according to both Boruta and RAFS, 8 were discovered by both algorithms but were among the top 15 most relevant according to one algorithm and outside the top 15 according to another (three for Boruta and five for RAFS), and 2 were discovered exclusively by Boruta. AUC: the AUC of the best probe for given biomarker; FC: fold changes of the mean value of normalized gene expression EBV+ samples with respect to EBV- samples; LIT: marks genes previously reported by Morscio et al.

[0022] in the main paper (M) or in the Supplementary Material (SM), or by Craig et al.

[0035] (C); OS: ranking of genes according to their oncogenic potential from the OncoScore; Rank: the unified rank of relevance (rank in RAFS and Boruta in parentheses).

[0193] 2.2 Classifiers

[0194] The probes associated with genes selected in the previous step were used to build the classifiers, using the top two, top three genes, etc. from Table 1. In the case of the top two genes, there were two equivalent selections possible. While both RAFS and Boruta showed that gene interferon (IFN)- induced protein with tetratri copeptide repeat 5 (IFIT5, probe 203596 s) was the most relevant, there was a discrepancy for second and third place. In the case of RAFS, the gene IFITM1 (probe 201601 x) was the second most relevant, and the gene cluster of differentiation 3QQ. (C 1)30 A, probe 209933 s) was the third one. In the case of Boruta, it was the other way around. Consequently, with both algorithms, there was agreement on the best three variables as well as the fourth variable - the CD1C gene (205987 at probe). The results were nearly perfect when only the two most relevant variables were used (see Table 2). In particular, for a pair of variables selected by the RAFS algorithm, the P32197PC00 / WAW 22-08-2025 average AUC over 200 repeats of the cross-validation was a perfect 1.000. It was nearly perfect (0.996) for a pair selected by Boruta. It was a perfect 1.000 when at least three variables were used.

[0195] Table 2: Area under the receiver operating characteristic curve (AUC) averaged over 200 runs of 5-fold cross-validation, obtained for classifiers built using the most relevant variables.

[0196] 2.3 Network of interactions between most relevant genes

[0197] The subnetwork of protein-protein interactions (PPIs) from the STRING database consisting of genes from the list of the most important 18 genes and their interaction partners were extracted from the data and analyzed in depth. IFN regulatory factor 1 -antisense RNA 1 (IR J-A J) is an RNA gene, and hence is not associated with proteins. The subnetwork of interactions between the 17 proteins obtained from the most relevant genes, further referred to as seeds, and their neighbors in the network consisted of 68 nodes and 305 edges.

[0198] The expected number of edges for a random set of 68 proteins is 33. Hence, the PPIs in the network were statistically significantly enriched with p < 10-16. The genes in the network could be clustered into nine clusters using the MCL algorithm, with the recommended value of inflation parameter set at 2.0 (see Fig. 1). Three of the seeds, namely polypeptide N-acetylgalactosaminyltransferase 10 (GALNT10), transmembrane protein 163, and solute carrier family 6 member 16, had no interactions with other relevant proteins and were not included in any cluster. When the inflation parameter was set to 1.5, there were six clusters - clusters #2, #3, and #4 were merged into one cluster containing GDP-mannose 4,6-dehydratase (GMDS), HSPA6, and high mobility group box 1 (HMBG1 , and P32197PC00 / WAW 22-08-2025 clusters #1 and #9 were merged into one cluster containing IFITM1, IFIT5, shiftless antiviral inhibitor of ribosomal frameshifting (SHFF), and COP1 (Fig. 1).

[0199] 2.3.1 Functional enrichments

[0200] The interaction network enrichment analysis was focused on the biological process hierarchy of Gene Ontology (GO). There are 238 terms in this hierarchy with p<0.05. The complete list of these terms is reported in the Supplementary Materials. Enrichment revealed the following clusters (Fig. 2). Cluster #1 consisted of proteins interacting directly with IFITM1, IFIT5, and SHFL and associated with viral responses (GO .’0009615). The more general term, immune response (G0:0006955) was connected to the two biggest clusters: #1 and #5. The seven seed proteins were associated with this term: three in cluster #1, CD300A and CD1C in cluster #5, HMGB1 in cluster #3, and HSPA6 in cluster #4. Interestingly, two hub proteins, signal transducer and activator of transcription 1 and IRF1, connected those clusters. While the IRF1 protein itself does not belong to the selected list of seed proteins, the antisense gene IRF1-AS1 does. Therefore, the IRF1 protein was included in further analyses. Cluster #5 was associated with the regulation of immune processes (G0:0002682), in particular, with positive regulation (G0:0002684). All three seed proteins in this cluster, CD300A, CD1C, and killer cell immunoglobulin-like receptor 2DL4 (KIR2DL4), were associated with this GO term. On the other hand, cluster #9 was associated with the proteasome- mediated ubiquitin-dependent protein catabolic process (G0:0043161). In particular, this was the only specific GO term for protein RFWD2, corresponding to the COP1 gene. Clusters #7 and #8 were associated with the general term regulation of gene expression (G0:0010468). In particular, seed proteins elongation factor RNA polymerase Il-like 3 (ELL3) and MLLT3 in cluster #7, KIR2DL4 in cluster #5, and HMGB1 in cluster #3 were associated with more specific regulation of developmental processes (GO: 0050793).

[0201] The associations of seed proteins with selected GO terms are presented in Table 3. Analyses of associations of the most relevant genes with these terms revealed four groups of genes that were identical to gene clusters #1 to #4 and a set of other genes. Their expression differed depending on EBV status (Figs. 3-7).

[0202] Table 3: Selected Gene Ontology terms and their associations with seed proteins in the columns.

[0203] Gene Ontology terms P32197PC00 / WAW 22-08-2025

[0204] For each Gene Ontology (GO) term, the number of proteins in the network associated with that term and statistical significance levels is shown using the following scheme: < ICT15, ****p

[0205] < 10-12, ***p < IO-9, **p < IO-6, *p < IO-3, a: G0:0006952 - defense response, 34*****; b: G0:0009615 - response to virus, 21*****; c: G0:0051607 - defense response to virus, 18*****; d:

[0206] G0:0006955 - immune response, 34*****; e: G0:0002376 - immune system process, 40****; f: G0:0006950 - response to stress, 45****; g: G0:0034340 - response to type I interferon, 12****; h: G0:0045087 - innate immune response, 23****; i: G0:0034097 - response to cytokine, 27****; j: G0:0019221 - cytokine-mediated signaling pathway, 22****; k: G0:0060337 - type I interferon signaling pathway, 11***; f G0:0007166 - cell surface receptor signaling pathway, 29***; m:

[0207] G0:0045321 - leukocyte activation, 17**; n: G0:0050793 - regulation of developmental processes, P32197PC00 / WAW 22-08-2025

[0208] 28**; o: G0:0010468 - regulation of gene expression, 35*; p: G0:0045595 - regulation of cell differentiation, 20*; q: GO: 0043161 - proteasome-mediated ubiquitin-dependent protein catabolic process, 7 < 0.01. Associations identified by STRING are denoted by 1, associations inferred from analysis of the literature are denoted by *.

[0209] Table 4: Markers identified by three feature selection methods P32197PC00 / WAW 22-08-2025 P32197PC00 / WAW 22-08-2025

[0210] 3. In silico confirmation of biomarkers’ specificity

[0211] In order to confirm utility of selected biomarkers further studies in silico have been performed. Said studies were conducted using the Robust Multi-array Average log2 normalized gene expression data (Affymetrix HG-U133 Plus 2.0 array) from PTLD patient samples from the ArrayExpress database (Accession No. E-GEOD-38885). These data were previously described in (Morscio et al., 2013; Ferreiro et al., 2016) to characterize the clinicopathological and molecular genetic characteristics of diffuse large B-cell lymphoma (DLBCL) and to identify the differences between EBV+ and EBV- DLBCL. The training data set included 54613 probes from 33 patients (23 EBV+ samples and 10 EBV- samples) diagnosed with DLBCL, the most frequent PTLD subtype. In addition, the second step, microarray data from 12 healthy non-PTLD TRs (Accession No. E-GEOD-22229) studied in (Newell et al., 2010) and 40 healthy non-smoking women (Accession No. E-GEOD-18723) reported in (Pan et al., 2010) were used in the gene expression comparative studies. To obtain GE levels, all probes were mapped to genes. If multiple probes were mapped to one gene, the GE value was average. The housekeeping genes examined in (Kim et al., 2022) were used to normalize the data in different experimental datasets.

[0212] As mentioned earlier, the set of 18 genes have been identified, namely HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1, that can be used to discern non PTLD patients from PTLD patient, more specifically to discern non PTLD TR from PTLD TR, and potentially, additionally between EBV+ and EBV- subtypes of PTLD - namely the top 18 genes obtained from feature selection (union of top 15 biomarkers with Boruta and top 15 with RAFS). To confirm it, further comparisons P32197PC00 / WAW 22-08-2025 of expression levels of PTLD patients, (both EBV- and EBV+) with expression levels of control samples, namely non PTLD TR, and healthy women were performed.

[0213] Detailed analysis of the data set showed that some markers might not be specific enough to simultaneously differentiate PTLD EBV+ from PTLD EBV- patients and therefore would not be considered as the most promising to build a complex classification tool.

[0214] Nine genes (HSPA6, CD300A, ILITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS) could have been chosen as the most promising for differentiation of Epstein Barr Virus (EBV)-positive from EBV-negative patients in posttransplant lymphoproliferative disorder, where, consistently all differences in gene expression levels between EBV+ and EBV- and non PTLD TR are highly significant, differences between non PTLD TR and PTLD patients are in the same direction, and expressions of EBV- patients are further away from the expressions of the non PTLD TR patients than expressions of the EBV+ patients, see Table 5.

[0215] Among remaining 9 genes with significant expression differences between PTLD variants three genes (GALNT10, IRF1-AS1, and ILIT5) might not be suitable for discerning non PTLD TR group from EBV+ patients. Two genes (CD1C and SP3) might not suitable for discerning non PTLD TR group from EBV- patients. Lor four remaining genes (COP1, MLLT3, KIR2DL4, and SLC6A16) the differences in expression between classes, while statistically significant are less pronounced in all cases at least one p- value from statistical test was larger than 0.001. Given very small sizes of samples used in the study, and fact that selection of these genes involved feature selection they were considered initially to be not suitable for the above described purpose. What is more, in the case of MLLT3, KIR2DL4, and SLC6A16 the expression levels of non PTLD TR patients are between those of EBV+ and EBV- patients leading to additional ambiguity or the results.

[0216] Table 5 P32197PC00 / WAW 22-08-2025

[0217] Table 5. Comparison of normalized Log2 gene expression (GE) levels by the u-test in the study groups (1) patients with Epstein-Barr virus (EBV) positive post-transplant lymphoproliferative (PTLD) disease, (2) patients with EBV negative PTLD disease, and (3) kidney transplant recipients (KTR). The over-expressed genes in PTLD versus TRs are marked re marked with a scale of diagonal lines, while under-expressed genes are marked blue in grey-scale. FDR-adjusted p- value levels are indicated by the dots: **** p-val. < le-6, *** p-val. < le-4, ** p-val. < le-3, and * p-val. < 0.01.

[0218] The whole set of markers has been theoretically evaluated also in terms of correlations so as to check whether based on that further genes should be potentially excluded from the best set for building the the best complex classification tool.

[0219] TABLE 6 P32197PC00 / WAW 22-08-2025

[0220] Table 6. Weighted averages of squared correlations between expressions of nine genes. Correlations were computed independently for PTLD-EBV+, PTLD-EBV- and non PTLD TR groups, and averaged with weights proportional to the number of samples from each group.

[0221] These nine genes can be divided into following four groups based on their correlations:

[0222] 1. ELL3, GRHPR, GMDS;

[0223] 2. IFITM1, SHFL;

[0224] 3. HMGB1, TMEM163;

[0225] 4. HSPA6, CD300A.

[0226] The expressions levels of nine genes initially suggested for the test development are not entirely independent and there are significant correlations between some genes in this set, see Table 6. Nevertheless, the overall correlations are not very high, what can better be seen from the box plot presented in Fig.10. Therefore, finally all these nine genes seemed to be a good set of markers for building a complex classification tool (simultaneous discerning of non PTLD from PTLD EBV+ and from PTLD EBV- patients).

[0227] Although not helpful in this case for selecting smaller set of genes, the correlation levels can be taken into account when designing specific classification tool, for example an ensemble of weighted Naive Bayes Classifiers, as will be discussed later on. The person skilled in the art will appreciate that data relating to correlations can be used in any suitable classification algorithm to improve its performance.

[0228] In order to finally check whether theoretical presumptions are correct the study on practical implementation of the classification tools have been performed.

[0229] 3. Industrial application of the invention

[0230] This study presents the foundational work toward the development of the above-mentioned computer- implemented diagnostic method and system, referred to as the method and the system for ex vivo predicting the risk of posttransplant lymphoproliferative disorder (PTLD) and in particular, more complex method and system for simultaneous prediction of the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein Barr Virus (EBV)-positive from EBV- negative patients in posttransplant lymphoproliferative disorder. The system is designed to support industrial-scale implementation of molecular diagnostics by leveraging gene expression markers to P32197PC00 / WAW 22-08-2025 classify patient samples. The research focused on evaluating the feasibility and robustness of some exemplary classification models, particularly those built using Naive Bayes algorithms and random forest, to enhance diagnostic accuracy in the context of limited and heterogeneous datasets. The following section outlines the experimental framework, classifier architecture, and performance assessments conducted to inform the system’s design and operational parameters. Extensive testing was conducted to evaluate the predictive performance of various marker combinations.

[0231] After conducting tests with different classification models on all the data, the conclusion was that any combination of at least three genes among the top 18 is good enough to correctly discern non PTLD from PTLD patients (simple classification tool). Some exemplary results of the tests are presented in Fig. Ila to Fig. Ill and in Fig. 12a to Fig. 12b.

[0232] In general, several problematic patients were identified for whom many of the proposed individual classifiers produced incorrect decisions regarding the differentiation between PTLD EBV+ and PTLD EBV- patients. One can conclude that, based on the very limited data, it is difficult to assert that one specific set of 3, 4, 5, or even 6 markers will produce very good results on unknown data.

[0233] This is why to address this limitation, in one embodiment an ensemble classification approach is proposed. This method comprises multiple individual Naive Bayes classifiers, each constructed using a subset of K markers selected from a pool of nine available markers.

[0234] Each individual classifier generates a prediction — either a class label or class probabilities — and the final decision is derived through majority voting or probability aggregation (summing the probabilities).

[0235] The results with the classification tool based on an ensemble of Naive Bayes algorithms have been summarized in the Table 7.

[0236] Table 7 P32197PC00 / WAW 22-08-2025

[0237] The analysis of the PTLD classification tests reveals two distinct performance tiers depending on the classification objective: binary discrimination between PTLD and nonPTLD, and full three-class separation among nonPTLD, EBV-positive PTLD, and EBV-negative PTLD. For the binary classification task with the Naive Bayes algorithms — distinguishing all PTLD cases (regardless of EBV status) from nonPTLD — already a set of three genes was sufficient although the most effective configurations were those using five or six markers from the preferred nine-gene panel. These combinations achieved perfect separation, with zero false positives and zero false negatives. P32197PC00 / WAW 22-08-2025

[0238] This indicates that even relatively compact subsets of well-chosen markers can robustly identify PTLD presence, making them highly suitable for clinical screening applications where early and reliable detection is critical.

[0239] When the classification task was extended to the more complex three-class problem, the performance landscape shifted. Although the five- and six-marker subsets from the preferred panel maintained perfect binary classification, they occasionally misclassified between EBV+ and EBV- subtypes. To improve subtype resolution, larger subsets were evaluated. Seven- and eight-marker combinations from the preferred nine-gene panel, as well as eight- to ten-marker combinations from the expanded 13 -gene panel, consistently achieved near-perfect classification, with only a single misclassification occurring between EBV+ and EBV-. Notably, these larger subsets maintained zero false positives and false negatives for the nonPTLD class, underscoring their reliability in distinguishing healthy transplant recipients from those with PTLD, while also offering improved granularity in subtype differentiation.

[0240] In other words, all possible combinations of two to eight classifiers were examined. Ensembles composed of up to four classifiers still misclassified at least one patient (e.g., predicting EBV- instead of EBV+).

[0241] An ensemble of 126 classifiers, each based on five markers, demonstrated improved accuracy, although some classification decisions remained ambiguous (a bit fuzzy). The classifier built on 84 classifiers based on 6 markers was already very good.

[0242] Further analysis explored the feasibility of constructing classifiers using the nine genes excluded from the original set of 18. In other words, it was also checked whether it was possible to create classifiers on "rejected" genes - that is, 9 genes from the original set of 18 genes not selected. While classification was possible, at least one instance was consistently misclassified.

[0243] Additionally, the potential benefit of expanding the original nine-gene set with four genes that weakly or inconsistently differentiate PTLD from non PTLD TR was investigated. This resulted in a larger number of marker combinations (e.g., 1,716 combinations for selecting six or seven markers from 13, and 1,287 for selecting five or eight). However, the results did not indicate any performance improvement and the conclusion was negative.

[0244] It should be mentioned that following housekeeping genes have been used as a reference for the proposed computational models and methods: P32197PC00 / WAW 22-08-2025

[0245] TABLE 8

[0246] Table 8. Reference gene set and weights for computing the reference expression.

[0247] The use of housekeeping genes as internal references in the proposed method and system offers a significant technical advantage: it ensures that the diagnostic process is independent of the specific technology used to measure gene expression levels. This design feature means that although the validation tests were conducted using microarray-derived data, the same classification accuracy and reliability can be expected when applying the method to data obtained through alternative sequencing platforms. These platforms include, but are not limited to, quantitative PCR (qPCR), RNA sequencing (RNA-seq), NanoString technology, and digital droplet PCR (ddPCR). Such flexibility enhances the industrial applicability of the invention across diverse laboratory environments and diagnostic workflows.

[0248] Furthermore, the invention does not require the full set of housekeeping genes listed in Table 8 to achieve effective normalization. Even a reduced subset — such as ten carefully selected genes from the provided list — has been shown to provide sufficient standardization of expression values for accurate classifier performance. This allows for streamlined implementation without compromising diagnostic precision, making the system both scalable and adaptable to various clinical and research settings.

[0249] Other studies involved a classification based on the random forest algorithm.

[0250] The extended Random Forest classification tests, which included combinations of up to 18 genes from the full biomarker panel, provided a highly consistent and reinforcing picture when compared to the earlier Naive Bayes classifier results. These additional tests confirmed that the diagnostic system for PTLD — particularly for distinguishing EBV+ and EBV- subtypes — benefits most from carefully selected subsets of genes rather than from using the entire panel.

[0251] When testing combinations of 13 to 18 markers, the results showed that perfect classification was still achievable, but the number of perfect forests plateaued or even declined slightly. For example, with P32197PC00 / WAW 22-08-2025

[0252] 13 markers, 35 perfect forests were identified, and with 14 markers, only 2-3 perfect forests emerged. Interestingly, even the full 17- and 18-gene combinations did not yield perfect classification, with just one misclassification in each case and no perfect forests. This suggests that adding more genes beyond a certain threshold introduces redundancy without improving performance.

[0253] The average out-of-bag (OOB) error remained low across all configurations, hovering around 2.2- 2.5%, and false positives and false negatives were nearly eliminated, especially in the larger subsets. However, the majority of misclassifications were related to EBV+ vs EBV- subtype differentiation, not the binary PTLD vs non-PTLD classification, which remained robust even in smaller subsets.

[0254] These findings strongly align with the conclusions drawn from the Naive Bayes ensemble tests: a subset of 5-6 well-chosen genes from the preferred panel consistently delivers perfect binary classification, while larger subsets (typically 9-13 genes) are required to achieve high-fidelity three- class separation (non-PTLD, EBV+, EBV-). The Random Forest results reinforce the notion that gene selection matters more than sheer quantity, and that redundancy can dilute classifier precision.

[0255] Retesting procedures were designed to further validate the performance of classifiers previously studied and constructed as ensembles of Naive Bayes algorithms. Each ensemble consisted of multiple individual Naive Bayes models trained on different subsets of gene markers, with final predictions aggregated through majority voting. A Monte Carlo cross-validation was applied to assess the classifiers’ robustness and generalizability. In each iteration, a test set was randomly assembled by selecting two EBV-negative cases, two non-PTLD cases, and three EBV-positive cases, totaling nine samples. The remaining data formed the training set, which was used to build the classifiers and evaluate their predictive accuracy.

[0256] Three separate analyses were conducted using gene panels of 17, 12, and 9 markers, respectively. The Monte Carlo cross-validation tests performed on ensembles of Naive Bayes classifiers using gene expression data for PTLD have consistently demonstrated high classification accuracy across a wide range of subset sizes, from 5 to 14 genes. These classifiers were evaluated for their ability to distinguish between EBV-positive, EBV-negative, and non-PTLD cases, using a robust stratified sampling strategy and repeated testing over 1000 iterations per configuration.

[0257] In the first analysis, classifiers were built from combinations of K genes selected from a pool of 17, where K ranged from 5 to 14. The original set of 18 genes was reduced by one, based on prior Random Forest tests that showed one gene never appeared in any perfect classifier. This gene was excluded P32197PC00 / WAW 22-08-2025 from further consideration. The remaining 17 genes included IFIT5, IFITM1, CD300A, CD1C, SP3, HSPA6, KIR2DL4, GRHPR, TMEM163, COP1, SHFL, MLLT3, IRF1-AS1, HMGB1, ELL3, GMDS, and GALNT10. Starting with 5-gene subsets, the classifiers achieved perfect sensitivity for EBV- and non-PTLD cases, and 95.38% sensitivity for EBV+ cases. Specificity was also high, with non-PTLD reaching 100%, EBV+ at 100%, and EBV- at 96.7%. These results remained stable through the 6- and 7-gene tests, confirming that small subsets of well-chosen markers are sufficient for reliable classification.

[0258] As the subset size increased to 8 and 9 genes, performance remained strong. EBV+ sensitivity slightly improved to 95.50%, and non-PTLD sensitivity remained above 99.85%, with only a few misclassifications. The 10- and 11 -gene tests showed similar metrics, with EBV+ sensitivity at 95.52% and non-PTLD at 99.75%. These results indicate that classifier performance plateaus around 5-6 genes, with additional markers offering marginal gains.

[0259] The 12-gene test confirmed this trend, maintaining EBV- and non-PTLD sensitivity at 100% and 99.75%, respectively, and EBV+ sensitivity at 95.50%. The 13-gene test yielded nearly identical results, with only three non-PTLD samples misclassified and EBV+ sensitivity holding steady. The 14-gene test showed a slight increase in EBV+ misclassifications (226 out of 5000), but non-PTLD sensitivity remained at 99.85%, and EBV- classification was still perfect.

[0260] In the second analysis, a refined panel of 12 genes: HSPA6, CD300A, IFITM1, SHFL, COP1, HMGB1, TMEM163, ELL3, GRHPR, GMDS, MLLT3, and KIR2DL4 was used. The third analysis focused on a minimal set of 9 genes: HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, and GMDS. These panels were selected based on their prior performance in Random Forest classifiers and their frequency of appearance in perfect classification models.

[0261] The results were consistent with earlier findings and demonstrated strong diagnostic performance. Across all tests, the sensitivity for PTLD detection was 1.0, meaning no PTLD cases were missed. However, in every test, one EBV-positive case was consistently misclassified as EBV-negative. The 9-gene panel, while effective in small combinations (2-3 genes), showed diminished performance with larger subsets, where false negatives for non-PTLD began to appear from five genes onward. Even in the weakest configuration, the sensitivity for non-PTLD detection remained at 0.92625, translating to a PTLD specificity of 0.979. P32197PC00 / WAW 22-08-2025

[0262] More robust results were achieved with the 12- and 17-gene panels. For the 12-gene set, combinations of 2 to 5 genes yielded perfect separation between healthy and diseased samples. Beyond five genes, occasional misclassifications of EBV-positive instead of non-PTLD occurred, but the sensitivity for non-PTLD detection remained at or above 0.996, corresponding to a PTLD specificity of 0.999. The 17-gene panel exhibited similar behavior. Ensembles built from 5 to 8 markers achieved flawless discrimination between EBV and non-PTLD cases. Even in the least favorable configurations beyond eight markers, the specificity for non-PTLD detection was 0.9975, again translating to a PTLD specificity of 0.999.

[0263] Importantly, these results also provide clear guidance to a person skilled in the art regarding the minimal and optimal configurations for classifier construction. The consistent success of classifiers built from only two genes in correctly identifying PTLD cases — without false negatives — demonstrates that meaningful diagnostic performance can be achieved with minimal input. At the same time, the tests also reveal where classifier performance becomes optimal. Specifically, the ability to distinguish EBV-positive from EBV-negative PTLD cases improves significantly when classifiers are built from at least five genes. This was observed across both Naive Bayes and Random Forest models, where five-gene combinations consistently appeared in perfect classifiers capable of resolving EBV subtypes with high fidelity. The convergence of results from both modeling approaches confirms that five-gene classifiers strike a balance between simplicity and diagnostic granularity, making them particularly suitable for clinical implementation.

[0264] Thus, the ensemble-based testing not only confirms the feasibility of constructing classifiers from as few as two genes for PTLD detection, but also identifies five-gene configurations as the optimal threshold for achieving reliable subtype differentiation. These insights provide a strong technical foundation for their industrial applicability.

[0265] Now exemplary embodiments of the systems and methods according to the invention will be discussed. As mentioned earlier, a broad set of eighteen markers defines the claimed diagnostic space:

[0266] HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1- AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1.

[0267] This full panel establishes the scope of informative markers for use in PTLD risk assessment.

[0268] The system for ex vivo predicting the risk of posttransplant lymphoproliferative disorder according to the invention comprises memory means 106, and processing means 102 configured to P32197PC00 / WAW 22-08-2025

[0269] - receive information representative for expression level of biomarkers acquired from a sample to be assessed, said biomarkers being at least three selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1

[0270] - classify said information representative for expression level of said at least three biomarkers

[0271] - output the classification results, said results being indicative of whether the assessed sample belongs to one of two classes: PTLD or non-PTLD patient.

[0272] However, a system for ex vivo prediction of the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein Barr Virus (EBV)-positive from EB V-negative patients in posttransplant lymphoproliferative disorder (more complex classification) utilizes a subset of nine markers, selected for high discriminatory power while minimizing redundancy:

[0273] HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS.

[0274] This nine-gene panel demonstrated consistent ability to distinguish PTLD (EBV+ and EBV-) from healthy transplant recipients (HT) with the added property that EBV- cases are further separated from controls across all included markers.

[0275] A system for ex vivo prediction of the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein Barr Virus (EBV)-positive from EB V-negative patients in posttransplant lymphoproliferative disorder comprises memory means 106, and processing means 102 configured to

[0276] - receive information representative for expression level of biomarkers, acquired from a sample to be assessed, said biomarkers being at least five selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS

[0277] - classify said information representative for expression level of said at least five biomarkers

[0278] - output the results of classification, the results being indicative of belonging the sample to be assessed to the one among three classes being non PTLD or PTLD EBV + or PTLD EBV - or patient. Additionally, in one embodiment the processing means 102 are configured to:

[0279] - receive information representative for expression levels of 31 housekeeping genes that are further used as reference. transform expressions to the log2 scale. P32197PC00 / WAW 22-08-2025 compute the reference as a weighted average of logarithms of expressions of 31 housekeeping genes, using the following formula:

[0280] R = Si ^Ei (Eq. 1) where Ei is log2 expression of i-th reference gene, and Wi is its weight. - compute the standardised expressions of the 9 marker genes using the following formula:

[0281] ES= (Eq. 2) where E is a measured log2 of expression, R is a reference log2 expression and Esis a standardised log2 expression.

[0282] If the system is using Naive Bayes Classifiers the processing means 102 are configured to compute Z scores for the log2 expressions of each marker gene for each condition using formula: p , _ ..h

[0283] Zhi = p (Eq. 3) where Zhi is a Zscore for each health status h for the i-th gene, with ph, and ohi being respectively the average and standard deviation of the log2 gene expression for the i-th marker gene. classify the test sample using an ensemble of weighted Naive Bayes Classifiers built using all possible combinations of at least five markers selected from the set of nine, wherein each individual classifier casts a vote for the object’s class, and the class that gets most votes is selected, output the result of the classification

[0284] Exemplary Zscore Values are given in the Table 9.

[0285] TABLE 9 P32197PC00 / WAW 22-08-2025

[0286] The computer-implemented method of predicting the risk of posttransplant lymphoproliferative disorder according to the invention will be now presented in more details in exemplary embodiments. As mentioned earlier, a broad set of eighteen markers defines the claimed diagnostic space:

[0287] HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1- AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1.

[0288] This full panel establishes the scope of informative markers for use in PTLD risk assessment. A computer implemented method for predicting the risk of posttransplant lymphoproliferative disorder PTLD, according to the invention, comprises the following steps:

[0289] - receiving information representative for expression level of biomarkers acquired from a sample to be assessed, said biomarkers being at least three selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1

[0290] - classifying said information representative for expression level of said at least three biomarkers

[0291] - outputting the classification results, said results being indicative of whether the assessed sample belongs to one of two classes: PTLD or non-PTLD patient..

[0292] However, a computer-implemented method for prediction of the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein Barr Virus (EBV)-positive from EBV- negative patients in posttransplant lymphoproliferative disorder (more complex classification) utilizes a subset of nine markers, selected for high discriminatory power while minimizing redundancy:

[0293] HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS.

[0294] This nine-gene panel demonstrated consistent ability to distinguish PTLD (EBV+ and EBV-) from healthy transplant recipients (HT) with the added property that EBV- cases are further separated from controls across all included markers.

[0295] A computer implemented method for prediction of the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein Barr Virus (EBV)-positive from EBV-negative patients in posttransplant lymphoproliferative disorder, according to the invention, comprises the following steps P32197PC00 / WAW 22-08-2025

[0296] - receiving information representative for expression level of biomarkers, acquired from a sample to be assessed, said biomarkers being at least five selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS

[0297] - classifying said information representative for expression level of said at least five biomarkers

[0298] - outputting the results of the classification, said results being indicative of whether the assessed sample belongs to one of three classes: non-PTLD, PTLD EBV-positive, or PTLD EBV-negative patient.

[0299] Expression levels are measured for a test sample as normalized log2-transformed values (Log2GE) using standard molecular assays such as qPCR, RNA-Seq, or microarray and the results in a digital form are transferred to the system according to the invention. In the next step the system then performs normalization relative to a set of 31 housekeeping genes characterized by stable expression in lymphocytes:

[0300] SNRPD3, GUSB, CSNK2B, PSMB2, HPRT1, PSAP, ACTB, MLF2, RHOA, GABARAP, EEF2, HNRNPA2B1, CHMP2A, HSP90AB1, AP2M1, CD59, PSMB4, PFDN5, EDF1, GPI, VCP, REEP5, BSG, UBC, PCBP1, RAB7A, GAPDH, EMC7, RAB11B, TLE5, RAB1B.

[0301] Normalized expression values are calculated relative to the weighted average expression of these housekeeping genes to ensure comparability across samples. The weights for computing the reference are given in Table 8.

[0302] If the classification step is based on an ensemble of weighted Naive Bayes Classifiers, this step involves computing for each marker i and class C, a standardized Z-score: where Zhi is a Zscore for each health status h for the i-th gene, with ph, and ohi being respectively the average and standard deviation of the log2 gene expression for the i-th marker gene.

[0303] The standardized values of / i is the class-specific mean and aCi are given in Table 9.

[0304] Because gene expression markers may exhibit pairwise correlations, a first-order correction is applied to mitigate overconfidence from the independence assumption in Naive Bayes.

[0305] Thus, each marker's weight is computed as: P32197PC00 / WAW 22-08-2025 where nk is the estimated pairwise Pearson correlation coefficient (squared) between markers i and k. Correlation is computed across diagnostic classes and averaged with weighting proportional to class sizes.

[0306] This approach allocates half of the shared information in each pairwise squared correlation term to each marker, avoiding double counting and represents a first-order correction while omitting higher- order dependencies. Moreover, such approach ensures non-negative weights for interpretability and stability.

[0307] In the next step the processing means are configured to compute, for class C, the corrected loglikelihood according to formula:

[0308] Posterior probabilities of classes are calculated by the processing means by using: where P(C) represents the prior probability of class C, which can be adapted based on patient-specific clinical risk.

[0309] Next, the final probability is computed by the processing means. The result of each individual weighted Naive Bayes Classifier described above is a vector of three probabilities. The final probability is computed as average over contributions from all classifiers.

[0310] The disclosed system and method facilitate clinical decision-making in the context of post-transplant care by enabling early detection of posttransplant lymphoproliferative disorder (PTLD) in transplant recipients. It further supports the differentiation between Epstein-Barr virus positive (EBV+) and Epstein-Barr virus negative (EBV-) subtypes, which is critical for tailoring patient management strategies. The system and the method are designed for seamless integration into routine post- P32197PC00 / WAW 22-08-2025 transplant monitoring protocols, thereby allowing for timely therapeutic intervention based on molecular diagnostic insights.

[0311] A key advantage of the system and the method lies in its use of a predefined panel of genes selected for their high discriminatory power in distinguishing PTLD-related profiles. Within this panel, a preferred subset of nine genes has been optimized for practical clinical application, balancing diagnostic accuracy with implementation feasibility. The method incorporates a first-order correction mechanism to account for inter-marker correlations, which enhances both the calibration of the predictive model and the interpretability of its outputs. Furthermore, the system generates probabilistic results that can be readily aligned with clinician-defined risk thresholds, thereby supporting informed and individualized clinical decisions.

[0312] The person skilled in the art will appreciate that in the context of the described diagnostic system for posttransplant lymphoproliferative disorder (PTLD), a classification tool refers to a computational model or algorithm designed to assign biological samples to specific diagnostic categories based on gene expression data. This tool processes standardized, log-transformed expression levels of selected biomarkers — typically ranging from two to eighteen genes — and determines whether a sample corresponds to one of several health states, such as non-PTLD, PTLD EBV-positive, or PTLD EBV- negative. While Random Forest and Naive Bayes classifiers are prominently used as exemplary algorithms other suitable algorithms include Support Vector Machines (SVM), which are effective for high-dimensional data and non-linear class boundaries; Logistic Regression, which offers interpretable probabilistic classification; k-Nearest Neighbors (k-NN), which classifies based on proximity in feature space; Gradient Boosting Machines such as XGBoost or LightGBM, which build strong predictive models from weak learners; Artificial Neural Networks (ANN), which capture complex patterns in gene expression; Deep Learning models like convolutional or recurrent neural networks, which are applicable when large datasets are available; Decision Trees, which provide transparent decision-making logic; and Linear Discriminant Analysis (LDA), which is useful for dimensionality reduction and classification under certain statistical assumptions. These classification tools may be used individually or as ensembles, and their performance can be enhanced through techniques such as Z-score normalization, correlation-based weighting, and Monte Carlo cross- validation, ensuring robust and generalizable predictions for clinical applications. P32197PC00 / WAW 22-08-2025

[0313] The rationale presented herein supports also the selection of the following gene sets to serve as the fundamental components of the diagnostic assay:

[0314] • Four genes in total, including one gene from each group e.g. set consisting of HSPA6, HMGB1, IFITM1, and ELL3. In total there are 24 distinct sets of genes constructed in such manner.

[0315] • Five genes including HSPA6 and CD300Aand one gene from groups 1-3, e.g. HSPA6, CD300, HMGB1, IFITM1, and ELL3. In total there are 12 distinct sets of genes constructed in such manner.

[0316] • Five genes including HMGB1 and TMEM163, and one gene from groups 1, 2, and 4, e.g. HSPA6, HMGB1, TMEM163, IFITM1, and ELL3. In total there are 12 distinct sets of genes constructed in such manner.

[0317] • Five genes including IFITM1 and SHFL, and one gene from groups 1, 3, and 4, e.g. HSPA6, HMGB1, IFITM1, SHFL, and ELL3. In total there are 12 distinct sets of genes constructed in such manner.

[0318] • Six genes consisting of HSPA6, CD300A, HMGB1, TMEM163, and one gene from groups 1 and 2, HSPA6, CD300A, HMGB1, TMEM163, IFITM1, and ELL3. In total there are 6 distinct sets of genes constructed in such manner.

[0319] • Six genes consisting of HSPA6, CD300A, IFITM1, SHFL, and one gene from groups 1 and 3, e.g. HSPA6, CD300A, HMGB1, TMEM163, IFITM1, and ELL3. In total there are 6 distinct sets of genes constructed in such manner.

[0320] • Six genes consisting of IFITM1, SHFL, HMGB1, TMEM163, and one gene from groups 1 and 4, e.g. HSPA6, HMGB1, TMEM163, IFITM1, SHFL, and ELL3. In total there are 6 distinct sets of genes constructed in such manner.

[0321] • Seven genes consisting of all genes from groups 2-4 and one gene from the group 1, e.g. HSPA6, CD300A, HMGB1, TMEM163, IFITM1, SHFL, ELL3. There are 3 distinct sets of genes constructed in such manner.

[0322] • Eight genes consisting of all genes from groups 2-4, and two genes from the group 1. There are 3 distinct sets of genes constructed in such manner.

[0323] • All nine selected genes.

[0324] EXAMPLE 2 - SCHEME OF PROCEDURE AND USE OF THE MARKER COMBINATION ACCORDING TO THE INVENTION P32197PC00 / WAW 22-08-2025

[0325] 1. Methodology

[0326] If a person is a kidney or liver recipient, a sample is taken from the patient and biomarkers expression levels are determined.

[0327] Changes in gene expression over time and / or differences in gene expression in different groups or treatments can be determined by any techniques known in the art, for example, by hybridization-based microarrays, serial analysis of gene expression, next-gen sequencing of complementary DNA (cDNA) or RNA sequencing (RNA-Seq).

[0328] Assessment of biomarkers expression levels in a patient's sample at each stage of prediction, diagnosis monitoring and differentiation is performed using qRT-PCR method.

[0329] In particular, to determine the presence and quantity of RNA in a biological sample can be performed. Namely, the mRNA is extracted from the sample, fragmented and copied into stable ds-cDNA. The ds-cDNA is sequenced using high-throughput, short-read sequencing methods. These sequences can then be aligned to a reference genome sequence to reconstruct which genome regions were being transcribed. Obtained data can be used to annotate where expressed genes are, their relative expression levels, and any alternative splice variants. In more detail, RNA-Seq includes the following steps: RNA isolation, conversion of RNA to cDNA libraries, sequencing of cDNA into a computer-readable format, alignment of sequenced samples to a reference, and quantification for downstream analyses such as gene expression.

[0330] Other transcriptomics analysis and gene expression profiling methods, e.g. those disclosed in publication Lowe et. al (Transcriptomics technologies, PLoS Comput Biol. 2017 May; 13(5), doi: 10.1371 / journal.pcbi.1005457), namely microarrays, Serial and Cap analysis of gene expression (SAGE / CAGE), can be used to assess and determine the patient’s gene expression.

[0331] 2. Prediction, diagnosis and monitoring of PTLD

[0332] The assessment of expression level of biomarkers according to the present invention can be used to estimate and / or determine the occurrence of disease, such as PTLD, before it occurs, the presence or absence of a particular disease or condition or to monitor the effectiveness of treatment in a patient already diagnosed with the disease.

[0333] For example, in a situation where a patient who has undergone an organ transplant does not have symptoms suggesting PTLD, i.e. other analysis do not indicate the occurrence of cancer, in a sample taken from such a patient the assessment of an expression level of biomarkers according to the present P32197PC00 / WAW 22-08-2025 invention is performed, namely assessment of an expression level of set of markers or biomarkers selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1, in a sample from a patient, and comparison of the assessed expression level of each biomarker with its mean expression level observed in a control sample,. More particular, the assessment of an expression level of at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 markers or biomarkers selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1 in a sample from a patient, is performed.

[0334] In practice, a physician or patient with a full gene panel expression result, or at least 3 genes, enters it into a ‘calculator’ that calculates the risk of PTLD. In one embodiment said ‘calculator’, namely a system for ex vivo predicting the risk of posttransplant lymphoproliferative disorder according to the invention can be available online on a dedicated portal that can be accessed from a server (not shown). In another embodiment, said ‘calculator’ can be implemented on any computer system 100 as shown in Fig. 2. Details of the system for ex vivo predicting the risk of posttransplant lymphoproliferative disorder according to the invention are presented further in the description. By entering the patient's results, the computer implemented algorithm suggest the risk of PTLD, thus helping the physician decide on the intensity of immunosuppressive therapy, the implementation of targeted therapy, antiviral therapy, or prophylaxis.

[0335] After assessment of expression level of said biomarker in a sample, it is possible to indicate the presence of PTLD or it is possible to predict the occurrence of PTLD by obtaining biomarkers expression levels that are typical for PTLD, i.e. there is a significant tendency suggesting the development of the disease.

[0336] If biomarker expression levels do not indicate PTLD, the patient is referred for routine diagnostics. The next assessment of expression level of said biomarkers can be performed repeatedly, for example every 5-12 months.

[0337] However, if biomarkers expression levels indicate or predict PTLD, the patient is further diagnosed, e.g. using PET, CT, trepanobiopsy, etc. and the appropriate treatment is given. The earlier stage of PTLD is detected, the more effective treatment can be applied. For example, doses of immunosuppressive drugs are reduced so that PTLD does not develop further. P32197PC00 / WAW 22-08-2025

[0338] 3. Differentiation of Epstein Barr Virus (EBV)-positive from EBV-negative PTLD patients

[0339] If a patient is diagnosed with PTLD, or PTLD is likely to occur, a next step is performed to differentiate patients with Epstein Barr Virus (EBV) from patients who are EBV negative. Thus, in a sample taken from said patient the further assessment of an expression level of biomarkers according to the present invention is performed, namely assessment of an expression level of set of markers or biomarkers selected from the group comprising selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, in a sample from a patient, and comparison of the assessed expression level of each biomarker with its mean expression level observed in a control sample. More particular, the assessment of an expression level of at least of at least 5, at least 6, at least 7, at least 8, or 9 markers or biomarkers selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS in a sample from a patient, is performed.

[0340] In practice, a physician or patient with a full gene panel expression result, or at least five genes, enters it into the above mentioned ‘calculator’ to get this information simultaneously with the information about the risk of PTLD. In practical implementation, the aforementioned auxiliary computational module is integrated within the system for ex vivo predicting the risk of posttransplant lymphoproliferative disorder, enabling simultaneous display of the results from both diagnostic procedures on a single interface.

[0341] If, according to the output of the computer implemented system according to the invention, biomarker expression levels do not indicate EBV+ PTLD, the patient is referred for routine diagnostics. The next assessment of expression level of said biomarkers can be performed repeatedly, for example every 5- 12 months.

[0342] If, according to the output of the computer implemented system according to the invention, biomarkers expression levels indicate EBV+ PTLD, the patient is given appropriate treatment. Importantly, monitoring the expression levels of said biomarkers allow to assess the effectiveness of treatment. On this basis, a decision may be made about further treatment.

[0343] Referring the expression levels of said biomarkers to the level just after transplantation, before and during treatment allows to assess disease progression or relapse. Biomarkers assays are performed, for example, at 10-20-day intervals and the results from the computer implemented system(s) P32197PC00 / WAW 22-08-2025 according to the invention are assessed. At the same time, during this period, imaging tests (for example CT scans) are carried out at 1-6-month intervals.

[0344] If a change in biomarkers signatures is noted, indicating the development of EBV+ PTLD, the patient is referred for diagnostic imaging physician examination, which will allow further therapeutic steps to be taken.

[0345] 4. Applications of Biomarker Testing in the Management of Post-Transplant Lymphoproliferative Disorder

[0346] For symptomatic patients with suspected Post-Transplant Lymphoproliferative Disorder (PTLD), characterized by signs such as enlarged lymph nodes, splenomegaly, focal lesions in internal organs (especially the transplanted organ), or abnormal peripheral blood counts, the standard of care includes imaging (CT, MRI, PET), virological testing (including for EBV), and a biopsy of an affected lymph node or organ. In addition to these standard procedures, a non-standard blood test searching for specific markers according to the invention and their assessment by a computer implemented system according to the invention can be performed to confirm the suspicion of PTLD and to qualify the patient for treatment.

[0347] In the case of asymptomatic patients who have known risk factors for PTLD, such as EBV, CMV, or HCV infection, advanced age, a high cumulative dose of immunosuppression, a history of treatment with polyclonal antibodies like anti-thymocyte globulin, or multiple episodes of transplant rejection, a non-standard blood test for specific markers according to the invention and their assessment by a computer implemented system according to the invention is used. The goal of this test is twofold: firstly, to confirm the need to reduce immunosuppressive therapy and intensify monitoring for PTLD development, and secondly, to identify candidates for potential preemptive treatment, for instance, with cell-based therapies.

[0348] For the general population of asymptomatic post-transplant patients who do not have established risk factors, a non-standard blood test looking for specific markers according to the invention and their assessment by a computer implemented system according to the invention can be used for screening purposes. The objective is to identify individuals who may warrant more intensive surveillance to monitor for the emergence of PTLD risk factors or the development of the disease itself. P32197PC00 / WAW 22-08-2025

[0349] Finally, in patients with confirmed active EBV replication, a non-standard blood test f for specific markers according to the invention and their assessment by a computer implemented system according to the invention is performed to guide clinical decisions. This testing helps to determine a patient's eligibility for antiviral treatment aimed at preventing the development of PTLD. Furthermore, it can aid in the diagnostic workup for PTLD and support the consideration of preemptive interventions, such as cell therapies.

[0350] References

[0351] [1] Sprangers B, Riella LV, Dierickx D. Posttransplant lymphoproliferative disorder following kidney transplantation: A review. Am J Kidney Dis. 2018; 78(2): 272-281.

[0352] [2] Mumtaz K, Faisal N, Marquez M, et al. Post-transplant lymphoproliferative disorder in liver transplant recipients: Characteristics, management and outcome from a single-centre experience with > 1000 liver transplantations. Can J Gastroenterol Hepatol. 2015; 29(8): 417-422.

[0353] [3] Mucha K, Foroncewicz B, Ziarkiewicz-Wr'oblewska B, et al. Post-transplant lymphoproliferative disorder in view of the new who classification: a more rational approach to a protean disease? Nephrol Dial Transplant. 2010; 25(7): 2089-2098.

[0354] [4] K. Mucha, R. Staros, B. Foroncewicz, B. Ziarkiewicz-Wroblewska, M. Kosieradzki, S. Nazarewski, B.Naumnik, J. Raszeja-Wyszomirska, K. Zieniewicz, and L. Paczek, “Comparison of post-transplantation lymphoproliferative disorder risk and prognostic factors between kidney and liver transplant recipients,” Cancers, vol. 14, no. 8, p. 1953, 2022.

[0355] [5] R. Elstrom, C. Andreadis, N. Aqui, V Ahya, R. Bloom, S. Brozena, K. Olthoff, S. Schuster, S. Nasta, E. Stadtmauer, et al, “Treatment of PTLD with rituximab or chemotherapy,” American Journal of Transplantation, vol. 6, no. 3, pp. 569-576, 2006.

[0356] [6] R. Trappe, H. Riess, N. Babel, M. Hummel, H. Lehmkuhl, S. Jonas, I. Anagnostopoulos, M. Papp- Vary, P. Reinke, R. Hetzer, et al., “Salvage chemotherapy for refractory and relapsed posttransplant lymphoproliferative disorders (PTLD) after treatment with single-agent rituximab,” Transplantation, vol. 83, no. 7, pp. 912-918, 2007.

[0357] [7] A. Santarsieri, J. F. Rudge, I. Amin, W. Gelson, J. Parmar, S. Pettit, L. Sharkey, B. J. Uttenthal, and G. A. Follows, “Incidence and outcomes of post- transplant lymphoproliferative disease after 5365 solid-organ transplants over a 20-year period at two uk transplant centres,” British Journal of Haematology, vol. 197, no. 3, pp. 310-319, 2022.

[0358] [8] L. J. Wozniak, T. L. Mauer, R. S. Venick, J. W. Said, R. L. Kao, P. Kempert, E. A. Marcus, V. Hwang, E. Y. Cheng, R. W. Busuttil, et al., “Clinical characteristics and outcomes of ptld following intestinal transplantation,” Clinical Transplantation, vol. 32, no. 8, p. el3313, 2018. P32197PC00 / WAW 22-08-2025

[0359] [9] J. Cheng, C. A. Moore, C. J. lasella, A. R. Glanville, M. R. Morrell, R. B. Smith, J. F. McDyer, and C. R Ensor, “Systematic review and metaanalysis of post-transplant lymphoproliferative disorder in lung transplant recipients,” Clinical transplantation, vol. 32, no. 5, p. el3235, 2018.

[0360]

[0010] M. S. Sampaio, Y. W. Cho, Y. Qazi, S. Bunnapradist, I. V. Hutchinson, and T. Shah, “Posttransplant malignancies in solid organ adult recipients: an analysis of the us national transplant database,” Transplantation, vol. 94, no. 10, pp. 990-998, 2012.

[0361]

[0011] A. Francis, D. W. Johnson, A. Teixeira-Pinto, J. C. Craig, and G. Wong, “Incidence and predictors of post-transplant lymphoproliferative disease after kidney transplantation during adulthood and childhood: a registry study,” Nephrology Dialysis Transplantation, vol. 33, no. 5, pp. 881-889, 2018.

[0362]

[0012] S. Caillard, C. Lelong, F. Pessione, and B. Moulin, “Post-transplant lymphoproliferative disorders occurring after renal transplantation in adults: report of 230 cases from the french registry,” American Journal of Transplantation, vol. 6, no. 11, pp. 2735-2742, 2006.

[0363]

[0013] B. L. Kasiske, A. Kukla, D. Thomas, J. W. Ives, J. J. Snyder, Y. Qiu, Y. Peng, V. R. Dharnidharka, and A. K. Israni, “Lymphoproliferative disorders after adult kidney transplant: epidemiology and comparison of registry report with claims-based diagnoses,” American journal of kidney diseases, vol. 58, no. 6, pp. 971-980, 2011.

[0364]

[0014] E. L. Yanik, M. S. Shiels, J. M. Smith, C. A. Clarke, C. F. Lynch, A. R. Kahn, L. Koch, K. S. Pawlish, and E. A. Engels, “Contribution of solid organ transplant recipients to the pediatric non- hodgkin lymphoma burden in the united states,” Cancer, vol. 123, no. 23, pp. 4663-4671, 2017.

[0365]

[0015] P. Fernberg, G. Edgren, J. Adami, °A. Ingvar, R. Bellocco, G. Tufveson, P. H' oglund, A. Kinch, J. Simard, E. Baecklund, et al., “Time trends in risk and risk determinants of non-hodgkin lymphoma in solid organ transplant recipients,” American journal of transplantation, vol. 11, no. 11, pp. 2472- 2482, 2011.

[0366]

[0016] S. I. Wadhwani, E. K. Hsu, M. L. Shaffer, R. Anand, V. L. Ng, and J. C. Bucuvalas, “Predicting ideal outcome after pediatric liver transplantation: An exploratory study using machine learning analyses to leverage studies of pediatric liver transplantation data,” Pediatric transplantation, vol. 23, no. 7, p. el3554, 2019.

[0367]

[0017] J. Nilsson, M. Ohlsson, L. Thulin, P. H' oglund, S. A. Nashef, and J. Brandt, “Risk factor identification and mortality prediction in cardiac surgery using artificial neural networks,” The Journal of thoracic and cardiovascular surgery, vol. 132, no. 1, pp. 12-19, 2006.

[0368]

[0018] F. S. Sousa, A. D. Hummel, R F. Maciel, F. M. Cohrs, A. E. J. Falc'ao, F. Teixeira, R. Baptista, F. Mancini, T. da Costa, D. Alves, et al., “Application of the intelligent techniques in transplantation databases: a review of articles published in 2009 and 2010,” in Transplantation proceedings, vol. 43, pp. 1340-1342, Elsevier, 2011.

[0369]

[0019] T. Srinivas, D. Taber, Z. Su, J. Zhang, G. Mour, D. Northrup, A. Tripathi, J. Marsden, W. Moran, and P. Mauldin, “Big data, predictive analytics, and quality improvement in kidney P32197PC00 / WAW 22-08-2025 transplantation: a proof of concept, ” American Journal of Transplantation, vol. 17, no. 3, pp. 671- 681, 2017.

[0370]

[0020] W. R. Rudnicki, M. Wrzesien', and W. Paja, “All relevant feature selection methods and applications,” Feature Selection for Data and Pattern Recognition, pp. 11-28, 2015.

[0371]

[0021] L. Breiman, “Random forests,” Machine Learning, vol. 45, pp. 5-32, 2001.

[0372]

[0022] J. Morscio, D. Dierickx, J. Ferreiro, A. Herreman, P. Van Loo, and et al., “Gene expression profiling reveals clear differences between EBV-positive and EBV-negative posttransplant lymphoproliferative disorders.,” Am J Transplant., vol. 13, no. 5, pp. 1305-16, 2013.

[0373]

[0023] J. F. Ferreiro, J. Morscio, D. Dierickx, P. Vandenberghe, O. Gheysens, G. Verhoef, M. Zamani, T. Tousseyn, and I. Wlodarska, “EBV-positive and EBV-negative posttransplant diffuse large b cell lymphomas have distinct genomic and transcriptomic features,” American Journal of Transplantation, vol. 16, no. 2, pp. 414-425, 2016.

[0374]

[0024] K. Mnich and W. R. Rudnicki, “All-relevant feature selection using multidimensional filters with exhaustive search,” Information Sciences, vol. 524, pp. Til -297 , 2020.

[0375]

[0025] R. Piliszek, K. Mnich, S. Migacz, P. Tabaszewski, A. Sulecki, A. Polewko-Klim, and W. R. Rudnicki, “MDFS: Multidimensional feature selection in r,” RJ, vol. 11, no. 1, p. 198, 2019.

[0376]

[0026] M. B. Kursa, A. Jankowski, and W. R. Rudnicki, “Boruta-a system for feature selection,” Fundamenta Inf ormaticae, vol. 101, no. 4, pp. 271-285, 2010.

[0377]

[0027] M. B. Kursa and W. R. Rudnicki, “Feature selection with the boruta package,” Journal of Statistical Software, vol. 36, pp. 1-13, 2010.

[0378]

[0028] R. Piliszek, A. A. Brozyna, and W. R Rudnicki, “Computational Analysis Identifies Novel Biomarkers for High-Risk Bladder Cancer Patients,” International Journal of Molecular Sciences, vol. 23, p. 7057, Jan. 2022. Number: 13 Publisher: Multidisciplinary Digital Publishing Institute.

[0379]

[0029] C. Strobl, J. Malley, and G. Tutz, “An introduction to recursive partitioning: rationale, application, and characteristics of classification and regression trees, bagging, and random forests.,” Psychological methods, vol. 14, no. 4, p. 323, 2009.

[0380]

[0030] M. Fern andez-Delgado, E. Cernadas, S. Barro, and D. Amorim, “Do we need hundreds of classifiers to solve real world classification problems?,” The journal of machine learning research, vol. 15, no. 1, pp. 3133-3181, 2014.

[0381]

[0031] D. Szklarczyk, A. Franceschini, S. Wyder, K. Forslund, D. Heller, J. Huerta-Cepas, M. Simonovic, A. Roth, A. Santos, K. P. Tsafou, et al., “String vlO: protein-protein interaction networks, integrated over the tree of life,” Nucleic acids research, vol. 43, no. DI, pp. D447-D452, 2015.

[0382]

[0032] D. Szklarczyk, R. Kirsch, M. Koutrouli, K. Nastou, F. Mehryary, R. Hachilif, A. L. Gable, T. Fang, N. T. Doncheva, S. Pyysalo, et al. , “The string database in 2023 : protein-protein association P32197PC00 / WAW 22-08-2025 networks and functional enrichment analyses for any sequenced genome of interest,” Nucleic Acids Research, vol. 51, no. DI, pp. D638-D646, 2023.

[0383]

[0033] S. Van Dongen, “Graph clustering via a discrete uncoupling process,” SIAM Journal on Matrix Analysis and Applications, vol. 30, no. 1, pp. 121-141, 2008.

[0384]

[0034] S. Holm, “A simple sequentially rejective multiple test procedure,” Scandinavian journal of statistics, vol. 6, no. 2, pp. 65-70, 1979.

[0385]

[0035] F. E. Craig, L. R. Johnson, S. A. Harvey, M. A. Nalesnik, J. H. Luo, S. D. Bhattacharya, and S. H. Swerdlow, “Gene expression profiling of Epstein-Barr virus-positive and- negative monomorphic b-cell posttransplant lymphoproliferative disorders,” Diagnostic Molecular Pathology, vol. 16, no. 3, pp. 158-168, 2007.

[0386]

[0036] S. Smith, D. Busse, S. Binter, S. Weston, C. Diaz Soria, B. Laksono, S. Clare, S. Van Nieuwkoop, B. Van den Hoogen, M. Clement, et al. ,

[0387] “Interferon-induced transmembrane protein 1 restricts replication of viruses that enter cells via the plasma membrane,” Journal of virology, vol. 93, no. 6, pp. e02003-18, 2019.

[0388]

[0037] S. K. Narayana, K. J. Helbig, E. M. McCartney, N. S. Eyre, R. A. Bull, A. Eltahla, A. R. Lloyd, and M. R. Beard, “The interferon-induced transmembrane proteins, IFITM1, IFITM2, and IFITM3 inhibit hepatitis C virus entry,” Journal of Biological Chemistry, vol. 290, no. 43, pp. 25946- 25959, 2015.

[0389]

[0038] V. Kinast, A. Plociennikowska, T. Bracht, D. Todt, R J. Brown, T. Boldanova, Y. Zhang, Y. Bruggemann, M. Friesland, M. Engelmann, et al., “C19orf66 is an interferon-induced inhibitor of HCV replication that restricts formation of the viral replication organelle,” Journal of hepatology, vol. 73, no. 3, pp. 549-558, 2020.

[0390]

[0039] W. Rodriguez, K. Srivastav, and M. Muller, “C19orf66 broadly escapes virus-induced endonuclease cleavage and restricts kaposi’s sarcoma associated herpesvirus,” Journal of virology, vol. 93, no. 12, pp. e00373-19, 2019.

[0391]

[0040] C. Zheng, Z. Zheng, Z. Zhang, J. Meng, Y Liu, X. Ke, Q. Hu, and H. Wang, “IFIT5 positively regulates nf- / cb signaling through synergizing the recruitment of i / cb kinase (IKK) to tgf- / L activated kinase 1 (TAK1),” Cellular signalling, vol. 27, no. 12, pp. 2343-2354, 2015.

[0392]

[0041] X.-F. Ding, J. Chen, H.-L. Ma, Y Liang, Y.-F. Wang, H.-T. Zhang, X. Li, and G. Chen, “KIR2DL4 promotes the proliferation of RCC cell associated with PI3K / Akt signaling activation,” Life Sciences, vol. 293, p. 120320, 2022.

[0393]

[0042] F. Yu, S. S. Ng, B. K. Chow, J. Sze, G. Lu, W. S. Poon, H.-F. Kung, and M. C. Lin, “Knockdown of interferon- induced transmembrane protein 1 (IFITM1) inhibits proliferation, migration, and invasion of glioma cells,” Journal of neuro-oncology, vol. 103, pp. 187-195, 2011.

[0394]

[0043] N. I. Sari, Y.-G. Yang, L. T. H. Phi, H. Kim, M. J. Baek, D. Jeong, and H. Y. Kwon, “Interferon- induced transmembrane protein 1 (IFITM1) is required for the progression of colorectal cancer,” Oncotarget, vol. 7, no. 52, p. 86039, 2016. P32197PC00 / WAW 22-08-2025

[0395]

[0044] X. Liu, L. Chen, Y Fan, Y Hong, X. Yang, Y Li, J. Lu, J. Lv, X. Pan, F. Qu, et al., “IFIM3 promotes bone metastasis of prostate cancer cells by mediating activation of the TGF- / i signaling pathway,” Cell death & disease, vol. 10, no. 7, p. 517, 2019.

[0396]

[0045] J. Yan, Y Jiang, J. Lu, J. Wu, and M. Zhang, “Inhibiting of proliferation, migration, and invasion in lung cancer induced by silencing interferon induced transmembrane protein 1 (ifitml),” BioMed Research International, vol. 2019, 2019.

[0397]

[0046] H. Hatano, Y. Kudo, I. Ogawa, T. Tsunematsu, A. Kikuchi, Y. Abiko, and T. Takata, “Ifn-induced transmembrane protein 1 promotes invasion at early stage of head and neck cancer progression,” Clinical cancer research, vol. 14, no. 19, pp. 6097-6105, 2008.

[0398]

[0047] P.-Y. Chu, W.-C. Huang, S.-L. Tung, C.-Y. Tsai, C. J. Chen, Y.-C. Liu, C.-W. Lee, Y.-H. Lin, H- Y. Lin, C.-Y. Chen, et al., “IFITM3 promotes malignant progression, cancer sternness and chemoresistance of gastric cancer by targeting MET / AKT / FOXO3 / c-MYC axis,” Cell & Bioscience, vol. 12, no. 1, pp. 1-16, 2022.

[0399]

[0048] J. Min, J. Hu, C. Luo, J. Zhu, J. Zhao, Z. Zhu, L. Wu, and R. Yuan,

[0400] “Ifitm3 upregulates c-myc expression to promote hepatocellular carcinoma proliferation via the ERK1 / 2 signalling pathway,” BioScience Trends, vol. 13, no. 6, pp. 523-529, 2019.

[0401]

[0049] R. Liang, X. Li, and X. Zhu, “Deciphering the roles of IFITM1 in tumors,” Molecular Diagnosis A Therapy, vol. 24, pp. 433-441, 2020.

[0402]

[0050] U.-G. Lo, J. Bao, J. Cen, H.-C. Yeh, J. Luo, W. Tan, and J.-T. Hsieh, “Interferon-induced IFIT5 promotes epithelial-to-mesenchymal transition leading to renal cancer invasion,” American journal of clinical and experimental urology, vol. 7, no. 1, p. 31, 2019.

[0403]

[0051] J. Huang, U.-G. Lo, S. Wu, B. Wang, R.-C. Pong, C.-H. Lai, H. Lin, D. He, J.-T. Hsieh, and K. Wu, “The roles and mechanism of IFIT5 in bladder cancer epithelial-mesenchymal transition and progression,” Cell death & disease, vol. 10, no. 6, p. 437, 2019.

[0404]

[0052] V. K. Pidugu, H. B. Pidugu, M.-M. Wu, C.-J. Liu, and T.-C. Lee, “Emerging functions of human ifit proteins in cancer,” Frontiers in molecular biosciences, vol. 6, p. 148, 2019.

[0405]

[0053] Y. He, P. A. Bunn, C. Zhou, and D. Chan, “KIR 2D (LI, L3, L4, S4) and KIR 3DL1 protein expression in non-small cell lung cancer,” Oncotarget, vol. 7, no. 50, p. 82104, 2016.

[0406]

[0054] Y. Cao, T. Ao, X. Wang, W. Wei, J. Fan, and X. Tian, “CD300Aand CD300F molecules regulate the function of leukocytes,” International Immunopharmacology, vol. 93, p. 107373, 2021.

[0407]

[0055] X. Du, B. Liu, Q. Ding, D. He, R. Zhang, F. Yang, H. Fan, L. Teng, and T. Xin, “CD300A inhibits tumor cell growth by downregulating akt phosphorylation in human glioblastoma multiforme,” International Journal of Clinical and Experimental Pathology, vol. 11, no. 7, p. 3471, 2018.

[0408]

[0056] Z. Tang, H. Cai, R Wang, and Y. Cui, “Overexpression of CD300A inhibits progression of NSCLC through downregulating Wnt / p-catenin pathway,” OncoTargets and therapy, pp. 8875- 8883, 2018. P32197PC00 / WAW 22-08-2025

[0409]

[0057] X. Chen, J. Zhang, X. Lei, L. Yang, W. Li, L. Zheng, S. Zhang, Y Ding, J. Shi, L. Zhang, et al., “CD1C is associated with breast cancer prognosis and immune infiltrates,” BMC cancer, vol. 23, no. 1, p. 129, 2023.

[0410]

[0058] Z.-j. Xu, Y Jin, X.-1. Zhang, P.-h. Xia, X.-m. Wen, J.-c. Ma, J. Lin, and J. Qian, “Pan-cancer analysis identifies CD300 molecules as potential immune regulators and promising therapeutic targets in acute myeloid leukemia,” Cancer Medicine, vol. 12, no. 1, pp. 789-807, 2023.

[0411]

[0059] E. Layre, A. de Jong, and D. B. Moody, “Human T cells use CD1 and MR1 to recognize lipids and small molecules,” Current opinion in chemical biology, vol. 23, pp. 31-38, 2014.

[0412]

[0060] M. Uhlen, C. Zhang, S. Lee, E. Sj' ostedt, L. Fagerberg, G. Bidkhori, R Benfeitas, M. Arif, Z. Liu, F. Edfors, et al, “A pathology atlas of the human cancer transcriptome,” Science, vol. 357, no. 6352, p. eaan2507, 2017.

[0413]

[0061] J. Liu, Z. Wu, Y. Wang, S. Nie, R. Sun, J. Yang, and W. Cheng, “A prognostic signature based on immune-related genes for cervical squamous cell carcinoma and endocervical adenocarcinoma,” International Immunopharmacology, vol. 88, p. 106884, 2020.

[0414]

[0062] Y. Lu, W. Xu, Y. Gu, X. Chang, G. Wei, Z. Rong, L. Qin, X. Chen, and F. Zhou, “Non-small cell lung cancer cells modulate the development of human cdlc+ conventional dendritic cell subsets mediated by cdl03 and cd205,” Frontiers in Immunology, vol. 10, p. 2829, 2019.

[0415]

[0063] B. Song, S. Shen, S. Fu, and J. Fu, “HSPA6 and its role in cancers and other diseases,” Molecular Biology Reports, pp. 1-13, 2022.

[0416]

[0064] F. U. Hartl, “Molecular chaperones in cellular protein folding,” Nature, vol. 381, no. 6583, pp. 571-580, 1996.

[0417]

[0065] S. Safe, “Specificity proteins (sp) and cancer,” International Journal of Molecular Sciences, vol. 24, no. 6, p. 5164, 2023.

[0418]

[0066] E. Hedrick, Y Cheng, U.-H. Jin, K. Kim, and S. Safe, “Specificity protein (sp) transcription factors spl, sp3 and sp4 are non-oncogene addiction genes in cancer cells,” Oncotarget, vol. 7, no. 16, p. 22245, 2016.

[0419]

[0067] M. A. Mansour, “Sp3 is associated with migration, invasion, and akt / pkb signalling in mda-mb- 231 breast cancer cells,” Journal of biochemical and molecular toxicology, vol. 35, no. 3, p. e22657, 2021.

[0420]

[0068] J.-Y. Lee, S.-H. Lee, K.-S. Kim, K.-H. Park, and K.-S. Park, “ELL3 functions as a critical decision maker at the crossroad between stem cell senescence and apoptosis,” Stem Cell Research & Therapy, vol. 10, pp. 1-11, 2019.

[0421]

[0069] A. Kabra and J. Bushweller, “The intrinsically disordered proteins MLLT3 (AF9) and MLLT 1 (ENL)-multimodal transcriptional switches with roles in normal hematopoiesis, mil fusion leukemia, and kidney cancer,” Journal of Molecular Biology, vol. 434, no. 1, p. 167117, 2022. P32197PC00 / WAW 22-08-2025

[0422]

[0070] R. Kang, Q. Zhang, H. J. Zeh III, M. T. Lotze, and D. Tang, “Hmgbl in cancer: good, bad, or both?,” Clinical cancer research, vol. 19, no. 15, pp. 4046-4057, 2013.

[0423]

[0071] D. Tang, R. Kang, H. J. Zeh III, and M. T. Lotze, “High-mobility group box 1 and cancer,” Biochimica etBiophysica Acta (BBA)-Gene Regulatory Mechanisms, vol. 1799, no. 1-2, pp. 131— 140, 2010.

[0424]

[0072] T. Taniguchi and A. Takaoka, “The interferon-a / ? system in antiviral responses: a multimodal machinery of gene regulation by the irf family of transcription factors,” Current opinion in immunology, vol. 14, no. 1, pp. 111-116, 2002.

[0425]

[0073] R. Karki, B. R. Sharma, B. Banoth, R. S. Malireddi, P. Samir, S. Tuladhar, H. Mummareddy, A. R. Burton, P. Vogel, and T.-D. Kanneganti, “Interferon regulatory factor 1 regulates panoptosis to prevent colorectal cancer,” JCI insight, vol. 5, no. 12, 2020.

[0426]

[0074] Y. Song, Y. Liu, S. Pan, S. Xie, Z.-w. Wang, and X. Zhu, “Role of the copl protein in cancer development and therapy,” in Seminars in Cancer Biology, vol. 67, pp. 43-52, Elsevier, 2020.

[0427]

[0075] W. H. Ka, S. K. Cho, B. N. Chun, S. Y. Byun, and J. C. Ahn, “The ubiquitin ligase copl regulates cell cycle and apoptosis by affecting p53 function in human breast cancer cell lines,” Breast Cancer, vol. 25, pp. 529- 538, 2018.

[0428]

[0076] X. Wei, K. Zhang, H. Qin, J. Zhu, Q. Qin, Y. Yu, and H. Wang, “Gmds knockdown impairs cell proliferation and survival in human lung adenocarcinoma,” BMC cancer, vol. 18, no. 1, pp. 1-14, 2018.

Claims

P32197PC00 / WAW 22-08-2025CLAIMS1. A computer implemented method for predicting the risk of posttransplant lymphoproliferative disorder PTLD, comprising the following steps:- receiving information representative for expression level of biomarkers, acquired from a sample to be assessed, said biomarkers being at least three selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1- classifying said information representative for expression level of said at least three biomarkers- outputting the classification results, said results being indicative of whether the assessed sample belongs to one of two classes: PTLD or non-PTLD patient.

2. A computer program product for predicting the risk of posttransplant lymphoproliferative disorder, comprising instructions which when executed by a computer causes it to perform the method according to claim 1.

3. A computer implemented method for predicting the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein Barr Virus (EBV)-positive from EBV-negative patients in posttransplant lymphoproliferative disorder,- receiving information representative for expression level of biomarkers, acquired from a sample to be assessed, said biomarkers being at least five selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS- classifying said information representative for expression level of said at least five biomarkers- outputting the results of the classification, said results being indicative of whether the assessed sample belongs to one of three classes: non-PTLD, PTLD EBV-positive, or PTLD EBV-negative patient.

4. The method according to claim 3, wherein classifying said information representative for expression level of biomarkers is performed by using an ensemble of weighted Naive Bayes Classifiers built sing all 126 possible combinations of five markers selected from the set of nine, wherein each individual classifier casts a vote for the object’s class, and the class that gets most votes is selected.P32197PC00 / WAW 22-08-20255. The method according to claim 3, wherein said biomarkers being HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS and wherein classifying said information representative for expression level of biomarkers is performed by using an ensemble of weighted Naive Bayes Classifiers built using all 84 possible combinations of six markers selected from the set of nine, wherein each individual classifier casts a vote for the object’s class, and the class that gets most votes is selected.

6. A computer program product for predicting the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein Barr Virus (EBV)-positive from EB V-negative patients in posttransplant lymphoproliferative disorder, comprising instructions which when executed by a computer causes it to perform the method according to claim 3.

7. A system for ex vivo predicting the risk of posttransplant lymphoproliferative disorder, comprising processing means (102) and memory means (106), characterized in that the processing means (102) are configured to perform the following steps:- receive information representative for expression level of biomarkers, acquired from a sample to be assessed, said biomarkers being at least three selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, COP1- classify said information representative for expression level of said at least three biomarkers- output the classification results, said results being indicative of whether the assessed sample belongs to one of two classes: PTLD or non-PTLD patient.

8. A system for ex vivo prediction of the risk of posttransplant lymphoproliferative disorder and differentiation of Epstein Barr Virus (EBV)-positive from EB V-negative patients in posttransplant lymphoproliferative disorder, comprising processing means (102) and memory means (106), characterized in that the processing means (102) are configured to perform the following steps:- receive information representative for expression level of biomarkers, acquired from a sample to be assessed, said biomarkers being at least five selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS- classify said information representative for expression level of said at least five biomarkersP32197PC00 / WAW 22-08-2025- output the results of the classification, said results being indicative of whether the assessed sample belongs to one of three classes: non-PTLD, PTLD EBV-positive, or PTLD EBV-negative patient.

9. The system according to claim 8, wherein the processing means (102) are configured to classify said information representative for expression level of biomarkers by using an ensemble of weighted Naive Bayes Classifiers built using all 126 possible combinations of five markers selected from the set of nine, wherein each individual classifier casts a vote for the object’s class, and the class that gets most votes is selected.

10. The system according to claim 8, wherein the processing means (102) are configured to classify said information representative for expression level of biomarkers by using an ensemble of weighted Naive Bayes using all 84 possible combinations of six markers selected from the set of nine, wherein each individual classifier casts a vote for the object’s class, and the class that gets most votes is selected.

11. A method of ex vivo diagnosing posttransplant lymphoproliferative disorder, characterized in that the method comprises the following steps:(a) in vitro assessment of an expression level of at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 biomarkers selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, and COP1, in a sample obtained from a patient, and(b) comparison of the assessed expression level of each biomarker with its mean expression level observed in a control sample or with a threshold expression level value determined in reference to control samples, and(c) based on the comparison made in step (b) determining the presence or absence of posttransplant lymphoproliferative disorder in a patient.

12. The method of claim 11, wherein biomarkers, whose expression level is assessed in step (a) are selected form at least 5, at least 6, at least 7, at least 8, or 9 biomarkers selected from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, and GMDS, and in step (c) in addition to determining the presence or absence of the posttransplant lymphoproliferative disorder in a patient, the differentiation of EBV-positive and EBV-negative PTLD patients is made.P32197PC00 / WAW 22-08-202513. A use of biomarkers selected from at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, GMDS, GALNT10, IRF1-AS1, IFIT5, MLLT3, KIR2DL4, CD1C, SP3, SLC6A16, and COP1; in diagnosis of posttransplant lymphoproliferative disorder in a patient.

14. The use of claim 13, wherein diagnosis of posttransplant lymphoproliferative disorder in the patient comprises ex vivo assessment of expression level of said biomarkers in a sample obtained from said patient.

15. The use of claim 13 or 14, wherein the biomarkers are selected from at least 5, at least 6, at least 7, at least 8, or 9 from the group comprising HSPA6, CD300A, IFITM1, SHFL, HMGB1, TMEM163, ELL3, GRHPR, and GMDS, and wherein in addition to diagnosis of the posttransplant lymphoproliferative disorder in a patient, the differentiation of EB V-positive and EB V-negative PTLD patients is made.

16. A computer-implemented method for the identification of disorder specific biomarkers based on small volume data relating to at least one gene sequence contained in a set of patient’s samples, the method comprising:- a step of providing a set of small volume data relating to gene expression levels labeled so as the set of small volume data forms at least two groups- a step of identification of potential biomarkers that carry information about differences between the groups, using the all-relevant feature selection algorithms;- a step of selection of the most representative set of biomarkers; and- a step of analysis of the interaction network between selected biomarkers so as to identify disorder specific biomarkers.

17. The method according to claim 16, wherein the step of identification of potential biomarkers that carry information about differences between the groups, using the all-relevant feature selection algorithms comprises steps of:P32197PC00 / WAW 22-08-2025- identification of relevant probes with the use of at least one filtering algorithms among: MDFS, t-test, u-test, Ward-test with the use of the same predetermined significance threshold so as to generate a list of significant variables relating to gene expression levels.

18. The method according to claim 16, wherein the step of selection of the most representative set of biomarkers comprises:- applying the RAFS procedure to said list of significant variables relating to gene expression levels so as to receive a shorter list of probes relating to genes, said probes being cluster representatives that are at least once selected as a representative within the run using each of at least two hierarchical clustering algorithms,- applying the Boruta algorithm to said list of significant variables relating to gene expression levels so as to generate a shorter list of informative probes corresponding to a number of unique genes,- generating a final list of relevant variables relating to gene expression levels by taking a union of the most relevant probes obtained from the RAFS and Boruta algorithm.

19. The method according to claim 16, wherein the step of analysis of the interaction network between selected biomarkers comprises steps of:- generating a sub-network of protein-protein interactions (PPIs) between seeds and their neighbors, the seeds being proteins obtained from the most relevant genes, the sub-network consisting of a number of nodes and a number of edges so as the PPIs in the network are statistically significantly enriched,- clustering genes in the subnetwork into clusters using the MCL algorithm, with a predefined value of inflation parameter20. The method according to claim 19, wherein it further comprises analyzing enrichment of the interaction network based on a biological process hierarchy, advantageously, of Gene Ontology.

21. The method according to any of claims 16 to 20, wherein, it further comprises a step of validating the selected biomarkers by generating predictive models using a machine learning algorithm;22. The method according to claim 21, wherein the step of generating predictive models using a machine learning algorithm comprises building classifiers, using at least the top two variables from the final list of relevant variables and data relating to probes associated with said at least two top variables, said variables being gene expression levels.P32197PC00 / WAW 22-08-202523. A use of the method according to any of claims 16 to 22 for the identification of genes being biomarkers representatives for posttransplant lymphoproliferative disorder and enabling differentiation of Epstein Barr Virus (EBV)-positive from EBV-negative patients in posttransplant lymphoproliferative disorder.

Citation Information

Patent Citations

  • Systems and methods for generating biomarker signatures with integrated bias correction and class prediction

    EP2864920A1

  • Antibodies and methods for the diagnosis, prevention, and treatment of epstein barr virus infection

    US20210171610A1

  • A method for preventing human virus associated disorders in patients

    US20220332836A1