Early prediction and identification of risk of recurrent pregnancy loss

By analyzing endometrial biopsies from the proliferative phase for C10orf71 and CLEC4A gene expression, the method addresses the limitations of current RPL prediction methods, offering improved diagnostic accuracy and treatment prospects for RPL.

WO2025224154A1PCT designated stage Publication Date: 2025-10-30REGION SJAELLAND +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/061035
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-23
Filing Date
2025-04-23
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Current methods for predicting recurrent pregnancy loss (RPL) are inadequate, particularly due to the lack of reliable biomarkers and the heterogeneity of endometrial profiles during the secretory phase, which complicates the understanding of underlying pathophysiology and hinders effective diagnostic and treatment approaches.

Method used

An improved method using gene expression analysis of endometrial biopsies obtained during the proliferative phase to identify specific biomarkers, specifically C10orf71 and CLEC4A, to predict RPL with high sensitivity and specificity through a machine learning model.

Benefits of technology

The method provides a more accurate and reliable prediction of RPL risk by identifying distinct gene expression patterns in the proliferative phase, enabling better diagnostic and prognostic tools for RPL.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025061035_30102025_PF_FP_ABST
    Figure EP2025061035_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method for determining whether a female subject is at risk of recurrent pregnancy loss (RPL). In particular, the present invention relates to analysing the gene expression profile of a panel of genes, such as 2 genes, for determining if a female subject is at risk of RPL.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Early prediction and identification of risk of recurrent pregnancy loss

[0002] Technical field of the invention

[0003] The present invention relates to a method for determining whether a female subject is at risk of recurrent pregnancy loss (RPL). In particular, the present invention relates to analysing the endometrial gene expression profile of a panel of genes, such as 2 genes, for determining if a female subject is at risk of RPL.

[0004] Background of the invention

[0005] Recurrent pregnancy loss (RPL) is one of the most prevalent early reproductive disorders defined as three or more consecutive pregnancy losses in the first trimester, or two losses in the second trimester, of pregnancy. One to three percent of couples experience RPL world-wide, and 25% of all pregnancies will end in a loss. Not achieving a desired pregnancy can have a substantial impact on the health of couples. Women, who undergo several pregnancy losses, may experience high stress levels and depression.

[0006] Recurrent pregnancy loss is a complex and heterogenous disorder. Several risk factors are associated with RPL including structural uterine abnormalities, autoimmune disorders, thrombophilia disorders, parental chromosomal abnormalities, and endometrial dysfunction. Still, for about half of the women suffering from RPL, no risk factors are found reflecting the diversity of the disease and the lack of understanding of the underlying pathophysiology.

[0007] A successful pregnancy involves maternal adaption of the immune system to the semi-allogenic embryo preconceptionally and during implantation for the establishment and development of a healthy pregnancy. Therefore, both during the preconception phase and during pregnancy, immune mechanisms and interactions are likely to play an important role in RPL.

[0008] A number of risk factors is associated with RPL. However, for more than half of the women suffering from RPL, no explanations are found. The underlying pathophysiology is complex and heterogeneous. Some studies indicate that immunological dysfunction play a role even preconceptionally and during implantation. Only a few published studies have investigated the transcriptome of endometrial biopsies from women with RPL, and these have had small sample sizes and focused on the secretory phase of the menstrual cycle.

[0009] The endometrium of the uterus changes throughout the menstrual cycle, a process strictly regulated by hormones and other biochemical signals. This remodelling is important for establishing a microenvironment that supports the implantation of the embryo and the continuation of pregnancy. Therefore, impaired endometrial function might be a risk factor for RPL. Hence, useful diagnostic approaches and the ability to address an individual prognosis for couples as well as effective targeted treatment are needed.

[0010] Oth man R, et al. Microarray profiling of secretory-phase endometrium from patients with recurrent miscarriage. Reprod Biol. 2012 Jul; 12(2): 183-99) discloses a study to identify differentially expressed genes and their related biological pathways in the secretory phase endometrium from patients with recurrent miscarriage (RM) and fertile subjects.

[0011] Lucas ES, et al. Loss of Endometrial Plasticity in Recurrent Pregnancy Loss. Stem Cells 2016;34:346-356) discloses a study demonstrating that the endometrium associated with RPL exhibits loss of plasticity. The study comprises sequencing of the transcriptomes of mid-luteal (secretory phase) endometrial biopsies from RPL patients and women with no history of RPL.

[0012] Dan Liu, et al. {CSFl-associated decrease in endometrial macrophages may contribute to Asherman's syndrome') American Journal of Reproductive Immunology, Wiley- Blackwell Publishing, Inc, US, vol. 83, no. 1, 3 November 2019 (2019-11-03)) discloses a study that focuses on Asherman's syndrome, which involves endometrial fibrosis that may lead to intrauterine adhesions, and in some cases may result in recurrent pregnancy loss (RPL). In other words, Asherman's syndrome may in some cases lead to pregnancy loss, however, this is a secondary complication. Jin Hee Ahn et al. (Expression of TWIST in the first- trim ester trophoblast and decidua! tissue of women with recurrent pregnancy losses" American Journal of Reproductive Immunology, Wiley-Blackwell Publishing, Inc, US, vol. 78, no. 2, 24 March 2017 (2017-03-24)) discloses the expression of TWIST in the first trimester placenta and decidua has any association with spontaneous abortion (SAB) and recurrent pregnancy losses (RPL). Hence, Jin Hee Ahn et al. investigate abortion material (trophoblast / placenta and decidua) from RPL, cases with one spontaneous abortion, and cases of elective abortion.

[0013] Q. Hussain et al. (Endometrial glandular expression of CD138 across the menstrual cycle in women with recurrent miscarriage", RCOG National Trainees Conference 2021, vol. 128, 1 May 2021 (2021-05-01), pages 1-144, XP093208612, DOI: 10.1111 / 1471-0528.16710) discloses collection of endometrial biopsies from women, wherein some are recurrent miscarriage (RM) patients. Some biopsies are collected during the proliferative phase and some under the secretory phase. Glandular expression of the CD138 marker is investigated.

[0014] Hence, in general, endometrial biopsies have hitherto been collected from the secretory phase of the menstrual cycle because of a focus on the window of implantation. However, the profile of the molecular and cellular processes in the secretory phase may be more heterogeneous with large individual differences compared to the corresponding profile of endometrial biopsies collected during the proliferative phase.

[0015] Hence, an improved method for determining the risk of recurrent pregnancy loss in women would be advantageous, and in particular a more efficient and / or reliable method for determining the risk for RPL based on endometrial biopsies obtained in a proliferative phase of the menstrual cycle would be advantageous.

[0016] Summary of the invention

[0017] The present invention relates to determining or predicting the risk of RPL from endometrial biopsies from the proliferative phase of the menstrual cycle. The state of the endometrium in the proliferative phase is found to be more homogenous in women with a presumed normal fertility, and differences in groups of women, who later develop pregnancy complications such as RPL may be more distinct.

[0018] In order to advance the understanding of the pathophysiology leading to RPL and to improve RPL prognosis in the clinic, it is important to identify a panel of endometrial biomarkers to be used in diagnostic procedures and to further understand the pathophysiology for improvement of the treatment of women with RPL. Although a number of different pre-conceptional endometrial biomarkers have been suggested, especially with a focus on immune biomarkers, there is still no clinically relevant macroscopic, molecular or histological biomarker that is specific for RPL, and no biomarkers exist that can predict the individual outcome for RPL patients.

[0019] The inventors have characterized the gene expression profile in the preconceptional proliferative endometrium of women with RPL versus a control group of healthy women using a large cohort and RNA sequencing analysis. Furthermore, a potential machine learning predicting model for RPL based upon the endometrial transcriptome profile was designed. Further, the inventors have explored whether the preconceptional endometrial transcriptome in the RPL cohort could predict pregnancy outcome. Finally, the inventors have investigated whether there were differences in the gene expression profiles of primary versus secondary RPL, in relation to certain risk factors, and in unexplained RPL versus explained RPL.

[0020] The inventors have conducted a study comprising two cohorts; an RPL cohort and a control group of women with healthy endometria and no previous history of RPL. For the RPL cohort, an endometrial biopsy was scheduled at the proliferative phase 3-10 days after the first day of menstruation in a natural cycle. Further, peripheral progesterone and oestrogen levels were measured at the day of the biopsy. For the control group, the biopsy was scheduled in the proliferative phase in a natural cycle. Histology dating was performed on all the biopsies to validate the date of the biopsy in relation to the menstrual cycle.

[0021] Upon analysis of the endometrial biopsies, a fraction of the biopsies was in the proliferative phase and another fraction was in the secretory phase. Hence, the study also included secretory phase endometrial biopsies. Significant differences in differentially expressed genes (DEGs) between endometrial biopsies collected in the secretory phase versus the proliferative phase of the menstrual cycle were observed. The inventors focused on the proliferative phase for comparisons between RPL and controls due to few controls in the secretory phase. Surprisingly, an enrichment of DEGs in one homogenous and specific biological category, namely immunological and infectious disease pathways, appeared for proliferative phase endometrial biopsies. Additionally, a clear separation between RPL patients and control group was achieved when analysing proliferative phase endometrial biopsies.

[0022] No studies have investigated the proliferative phase in an RPL population, and the data of the present invention show that the endometrium should be assessed not only in the secretory phase. Certain endometrial dysfunctionalities (for example immunological dysfunction) are preferably better detected in the early phase of the menstrual cycle, such as the proliferative phase, due to a low degree of hormonal changes.

[0023] Thus, an object of the present invention relates to the provision of an improved method for determining whether a female subject is at risk of recurrent pregnancy loss (RPL).

[0024] In particular, it is an object of the present invention to provide a method that solves the above-mentioned problems of the prior art with an improved method using gene expression data to accurately predict if a female subject is at risk of recurrent pregnancy loss (RPL), wherein said method has a high sensitivity and specificity.

[0025] Thus, one aspect of the invention relates to a method for determining whether a female subject is at risk of recurrent pregnancy loss (RPL), the method comprising

[0026] - providing an endometrial biopsy obtained in a proliferative phase of the menstrual cycle from said subject;

[0027] - determining in the endometrial biopsy a gene expression profile for at least 2 genes comprising C10orf71 and / or CLEC4A, wherein if the gene expression profile comprises - CLEC4A being upregulated; and / or

[0028] - C10orf71 being downregulated, compared to a gene expression profile comprising C10orf71 and / or CLEC4A in a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle from control females with healthy endometria and no history of RPL, it is indicative of said subject being at risk of RPL.

[0029] Another aspect of the present invention relates to use of an endometrial biopsy obtained in a proliferative phase of the menstrual cycle for determining the risk of recurrent pregnancy loss in a female subject.

[0030] Yet another aspect of the present invention is to provide use of a gene expression profile comprising C10orf71 and / or CLEC4A for determining the risk of recurrent pregnancy loss in a female subject.

[0031] Brief description of the figures

[0032] Figure 1 shows a schematic representation of the sampling process. The RPL cohort included women from whom endometrial biopsies had been collected. Of these, some women were in the proliferative phase, and some were in the secretory phase. For the in vitro fertilisation (IVF) control group, the large majority of control women were in the proliferative phase; only one woman was in the secretory phase. In the RPL cohort, endometrial biopsies were collected in the proliferative phase at day 3-10 or in the secretory phase after day 10 after the menstruation in a natural menstrual cycle. For the IVF control group, the biopsy was scheduled in the proliferative phase in a natural cycle with no hormonal treatment, the cycle just after and up to three months after the second IVF / intracytoplasmic sperm injection (ICSI) treatment. To validate the date of the biopsy in relation to the menstrual cycle peripheral progesterone and oestrogen levels were measured at the day of the biopsy and histology dating was performed on all the biopsies. The endometrial biopsies were stored in RNA stabilization solution at -80°C. Total RNA was extracted, libraries were prepared by Novogene NGS stranded RNA library prep set, and all samples were RNA- sequenced with the use of an Illumina Novoseq 6000 platform followed by bioinformatics analyses, Figure 2 shows principal component analysis (PCA). (A) The analysis was performed after batch effect correction. X and Y axis show PCI and PC2: the proportion of variance explained by each PC is shown on the axis labels. Each point is one sample. Circle: mid-cycle. Triangle: classification not possible, Square: Proliferative phase, Cross: Secretory phase. (B) The plot shows the same PCA and data as in (A) but shows PC3 and PC4. A clear separation of RPL patients (triangles) vs IVF control women (circles) was observed in PC3 and PC4,

[0033] Figure 3 shows (A) volcano plots of secretory vs proliferative endometrium analyses. The X-axis shows to the secretory vs proliferative endometrium Iog2- fold change (log2FC) and the Y-axis to the DEseq2 -loglO(P-value). Dots (genes) are coloured according to the P-adjusted values (below or above 0.05), (B) Figure 3B is divided into a top and bottom section. Gene Set Enrichment Analysis (GSEA) based on ranking genes by their secretory vs proliferative phase log2FC values and comparing the ranks within gene ontology term sets. Plots are split into 'activated' and 'suppressed' according to the positive and negative normalized enrichment score (NES) (NES>0 or NES<0). The results are reported as dot plots ranked by the NES (X axis). Rows show GO terms. Dot size indicates the number of expressed genes that are included in the GO term (pathway or a given set of genes). Gene set enrichment analysis showed that changes in gene expression occurred in genes related to 'hormone signalling', 'chemokine signalling', 'platelet activation', 'complement and coagulation cascades', which were all upregulated in the secretory phase, while cell cycle and DNA replication term genes were downregulated. (C) Figure 3C is divided into a top and bottom section. GSEAE analysis for KEGG pathways. The plot follows the conventions of panel B, but rows show KEGG pathways,

[0034] Figure 4 shows volcano plots and gene set enrichment analysis. (A) Principal component analysis (PCA) plot showing PCI (X-axis) vs PC2 (Y-axis)) for the proliferative phase samples after batch effect correction. The proportion of variance explained by each PC is shown in the axis labels. Circles: Control subjects, Triangles: RPL subjects.

[0035] (B) Volcano plot showing comparison between RPL vs controls. The plot follows the conventions of Fig 3A, but genes are coloured according to the P-adjusted value (below or above 0.1). NA: NA p-values due to independent filtering in the DESeq2 pipeline.

[0036] (C) PCA as in (A) but subjects (dots) are coloured by the k-means subgroups defined by the 166 DEGs. Triangle: Control subjects, circle: RPL subjects. (D) Figure 4D has been separated into a top and bottom section. Both sections are to be assembled into one figure. Gene set enrichment analysis (GSEA) based on the ranking genes by their joint contribution to PCI and PC2 (so, along the diagonal that defines the subgroups) of the PCA plot in (Fig. 4C) and comparing the ranks within gene ontology term sets. The plot follows the conventions of Fig. 3B. (E) Figure 4E has been separated into a top and bottom section. Both sections are to be assembled into one figure. GSEAE analysis for KEGG pathways,

[0037] Figure 5 shows results of the prediction model in terms of average accuracy, average sensitivity, and average specificity (percentages) for the (A) top 20, (B) top 15, (C) top 10, (D) top 9, (E) top 8, (F) top 7, (G) top 6, (H) top 5, (I) top 4, (J) top 3, (K) top 2 (C10orf71 and CLEC4A), (L) top 1 (C10orf71). The variance for each prediction analysis is also presented,

[0038] Figure 6 shows multiple ROC curves overlaid having a mean Area Under the Curve (AUC) value of 0.996,

[0039] Figure 7 shows standardized gene expression of C10orf71 and / or CLEC4A in RPL and control samples. Each point corresponds to one sample. Triangles correspond to RPL samples and circles correspond to control samples, and

[0040] Figure 8 shows the probability of RPL together with the expression level of C10orf71 and / or CLEC4A. Each dot corresponds to one sample. The lighter nuance of dots corresponds to a high probability of RPL, whereas a darker nuance of dots corresponds to a low probability of RPL. The probabilities come from the training set of the original Random Forest classifier with all 157 DEGs.

[0041] The present invention will now be described in more detail in the following. Detailed description of the invention

[0042] Definitions

[0043] Prior to discussing the present invention in further details, the following terms and conventions will first be defined :

[0044] Endometrial biopsy

[0045] When used herein, endometrial biopsy refers to a tissue sample from the lining of the uterus endometrium. The endometrial biopsy may be obtained during a hysteroscopy.

[0046] It must be noted that for the present invention, the act of taking the endometrial biopsy is not part of the invention. The endometrial biopsy was obtained from the female subject prior to the present invention. Hence, the endometrial biopsy provided for the use and method steps of the invention was obtained prior to the present invention, such as retrieved from storage (e.g. from a freezer).

[0047] Control endometrial biopsies

[0048] When used herein, control endometrial biopsy refers to an endometrial biopsy obtained from a female control subject having a healthy endometrium. Hence, a control endometrial biopsy is obtained from a female not having any previous history of RPL. The female control subject from which a control endometrial biopsy is obtained, is not limited to a specific type of female control. Thus, the female control subject is neither restricted to having a specific general health status nor previous pregnancies, provided that the endometrium from said female control subject is healthy. In principle, the control endometrial biopsy may be obtained from any female control regardless of health status and whether the female has had any previous pregnancies, provided that the female control has a healthy endometrium.

[0049] The control endometrial biopsy may even be obtained from a female control, who has had previous pregnancy loss for example due to paternal causes or genetic defects of the fetus / embryo.

[0050] Panel of control endometrial biopsies

[0051] When used herein, a panel of control endometrial biopsies is to be understood as three or more, such as four or more, such as five or more, preferably six or more control endometrial biopsies from the female control group from which a general average gene expression profile of selected genes is obtained.

[0052] Thus, in an embodiment the panel of control endometrial biopsies comprises three or more, such as four or more, such as five or more, preferably six or more control endometrial biopsies obtained in a proliferative phase of the menstrual cycle from control females with healthy endometria and no (previous) history of RPL.

[0053] For the present invention the female subject may be any female subject regardless of any previous pregnancies and pregnancy losses. Preferably, the female subject is suspected of having RPL or is at risk of having RPL.

[0054] Recurrent pregnancy loss (RPL)

[0055] When used herein, recurrent pregnancy loss (RPL) is defined as three or more consecutive pregnancy losses in the first trimester, or two losses in the second trimester, of pregnancy. The number of consecutive pregnancy losses to be classified or defined as RPL varies depending on geographic region and the clinical definitions within a specific geographic location.

[0056] The most commonly identified causes of RPL include uterine problems, hormonal disorders, and genetic abnormalities. Thus, the female patient has an underlying disorder and / or abnormality, causing recurrent miscarriage.

[0057] Recurrent pregnancy loss is subdivided into primary and secondary RPL, where primary RPL is defined as RPL without a previous pregnancy, while secondary RPL is defined as RPL after one or more previous pregnancies that progressed beyond 23 weeks of gestation.

[0058] Recurrent pregnancy loss (RPL) and recurrent miscarriage (RM) may be used interchangeably.

[0059] When used herein, a history of RPL is to be understood as a previous case of three or more consecutive pregnancy losses in the first trimester, or two losses in the second trimester, of pregnancy, which leads to clinical diagnosis of RPL As mentioned in the above section, RPL can be classified as primary or secondary RPL. Common causes of RPL include female uterine problems, hormonal disorders, and genetic abnormalities, hence the female has an underlying disorder and / or abnormality, causing recurrent miscarriage.

[0060] S ecificity

[0061] When used herein, specificity refers to a measure of a number of correct negative predictions. Thereby, the specificity is the proportion of negatives that are correctly identified by the test. The specificity can be calculated by dividing the number of true negatives by the number of true negatives plus the number of false positives.

[0062] True Negatives Specificity = — - - - — — -

[0063] True Negatives + False Positives

[0064] Sensitivity

[0065] When used herein, sensitivity refers to a measure of a number of correct positive predictions^ Thereby, the sensitivity is the proportion of positives that are correctly identified by the test. The sensitivity can be calculated by dividing the number of true positives by the number of true positives plus the number of false negatives.

[0066] True Positives Sensitivity = - - - — — -

[0067] True Positives + False Negatives

[0068] Accuracy

[0069] When used herein, accuracy is to be understood as the proportion of true and correct results (both true positive and true negative) in the selected population. The accuracy can be determined by dividing the number of correct predictions by the total number of predictions.

[0070] True Positives + True Negatives

[0071] Accuracy = - - - - -

[0072] Total Predictions

[0073] Total Predictions = True Positives + True Negatives + False Positives + False Negatives Receiver operating characteristic CROC) curves

[0074] When used herein, ROC curves illustrate and evaluate the diagnostic (and / or prognostic) performance of models. Determination of an ideal cut-off value is almost always a trade-off between sensitivity (true positives) and specificity (true negatives). The ROC curve provides a graphical illustration of these trade-offs at each cut-off for any diagnostic test that uses a continuous variable. The sensitivity (true positive rate) vs (1 - specificity) (false positive rate) is plotted for each cutoff. The best cut-off value provides the most optimal values of the sensitivity and the specificity, which is located on the ROC curve at the highest point on the vertical axis and at the furthest to the left on the horizontal axis (upper left corner). If the importance or consequences of a false negative result is the same as that of a false positive result then the best cut-off will be the one, which maximizes the sum of the specificity and sensitivity. Thus, the best cut-off value results in an area under the curve (AUC) closest to 1.0. As 1.0 is the maximum value for the AUC, it indicates a (theoretically) perfect test (i.e., 100% sensitivity and 100% specificity).

[0075] In summary, ROC analysis provides important information about diagnostic test performance: the closer the apex of the curve toward the upper left corner, the greater the discriminatory ability of the test. The discriminatory ability of the test is measured quantitatively by the AUC.

[0076] Thus, as seen in the Examples 5 and 6, the performance of the prediction model is evaluated using ROC curves and determining AUC of said ROC curves.

[0077] Area Under Curve

[0078] When used herein, AUC refers to a quantitative measure of the performance and discriminatory ability of a given model.

[0079] The maximum value for the AUC is 1.0, thereby indicating a (theoretically) perfect test having 100% sensitivity and 100% specificity, respectively.

[0080] A ROC curve with a straight, diagonal line extending from the lower left corner to the upper right corner indicates the performance of a random guess classification. There are several scales for AUC value interpretation but, in general, ROC curves with an AUC of 0.6 or below are generally not clinically useful and an AUC of 0.7 or above has a clinical value. However, several factors may play a role for the applicability of AUC of a test or model in a clinical setting. Thus, in an embodiment of the present invention, the AUC is in the range of 0.7 to 1, such as 0.7 to 1, preferably 0.8 to 0.99, more preferably 0.8 to 1, such as 0.8 to 0.99.

[0081] Fold change

[0082] When used herein, fold change is used in analysis of gene expression data from microarray and RNA-Seq experiments for measuring change in the expression level of a gene. Once gene expression data is obtained, comparison of one experimental group versus a second one is performed in order to find out which genes / transcripts change between conditions. The process is called differential expression analysis.

[0083] The goal of differential expression analysis is to perform statistical analysis to discover changes in expression levels of defined features (genes, transcripts, exons) between experimental groups.

[0084] The changes in expression levels are typically reported as fold change in logarithmic scale (logz fold change, logzFC), defined as

[0085] Avererage expression in first group of samples log2FC = log2

[0086] Avererage expression in second group of samples

[0087] A positive Iog2 fold change value (above 0) indicates an increase of gene expression, while a negative fold change (below 0) indicates a decrease of gene expression.

[0088] Gene expression profiling

[0089] When used herein, gene expression profiling refers to measures of mRNA levels, showing the pattern of genes expressed in a biological sample at the transcription level at a certain time.

[0090] Techniques to measure gene expression profiles include DNA microarrays (analyzing RNA transcript expression), which measure the relative activity of previously identified target genes / gene transcripts, or sequencing technologies that allow profiling of all active genes (RNA transcripts). For the present invention, all genes mentioned are well-known and naturally occurring genes without any artificial modifications.

[0091] Comparison of gene expression profiles

[0092] Investigating the differences between diseased and / or healthy states provides information on the pathology of diseases. Comparing transcript levels between healthy and / or diseased individuals allows the identification of differentially expressed genes, which may be related to causes, consequences or mere correlates of the disease.

[0093] As an example, differentially expressed genes (DEGs) involves the identification of genes that are differentially expressed in disease. A gene in a test sample that is differentially expressed is to be understood as being either upregulated or downregulated compared to the expression of the same gene in a control sample or the average expression of the same gene in a panel of control samples. Thus, the expression of a gene in a control sample or a panel of control samples are considered to be a baseline for determining the differential expression of said same gene in a test sample.

[0094] Gene regulation is measurable using Iog2 fold change as described above, where a positive fold change value (above 0) indicates an increase of gene expression, while a negative fold change (below 0) indicates a decrease of gene expression compared to control samples.

[0095] Differentially expressed genes (DEGs)

[0096] When used herein, differentially expressed genes (DEGs) are genes, which are differentially expressed if an observed difference or change in read counts or expression levels between two experimental conditions is statistically significant. Analysis of DEGs refers to the analysis and interpretation of differences in abundance of gene transcripts within a transcriptome.

[0097] Gene regulation

[0098] When used herein, gene regulation refers to the regulation of gene expression or caused by a wide range of possible mechanisms that are used by cells to increase or decrease the production of specific gene products (RNA transcripts and proteins). Any step of gene expression can be modulated, including transcriptional initiation, RNA processing, RNA degradation, and post-translational modification of a protein.

[0099] Regulation of gene expression includes

[0100] • upregulation of gene expression, which is to be understood as an increased expression of one or more genes resulting in an increased amount of gene product (RNA and often the corresponding protein), or

[0101] • downregulation of gene expression which is to be understood as a decreased expression of one or more genes resulting in a decreased amount of gene product (RNA and of the corresponding protein).

[0102] Gene regulation is measurable using Iog2 fold change as described above, where a positive fold change value (above 0) indicates an increase or upregulation of gene expression, while a negative fold change (below 0) indicates a decrease or downregulation of gene expression.

[0103] Healthy endometrium

[0104] When used herein, the term healthy endometrium comprises epithelial, stromal, vascular and immune cells that are distributed in the two main layers of the endometrium, the stratum basalis and stratum functional. Every 28 days, on average, the functional layer grows out from the basal layer, sloughs off and is regenerated. This process is controlled by female sex hormones produced by the ovary during the proliferative and secretory phases of the menstrual cycle. Multiple cell types make up the endometrium. The epithelial layer is composed of luminal epithelial cells, which line the endometrium, and glandular epithelial cells, which are invaginations of the epithelial cells extending from the lumen into the myometrium. Stromal cells as well as the extracellular matrix (ECM) provide structural and endocrine support. Blood vessels are remodeled into spiral arteries and perivascular cells and immune cells play important roles in endometrial remodeling throughout the menstrual cycle. Interestingly, the proportions of these cell populations change during the menstrual cycle. A healthy endometrium is essential for a healthy pregnancy. Natural menstrual cycle

[0105] Menstruation is the monthly shedding of the lining of the uterus. Menstruation is also known by the terms menses, menstrual period, menstrual cycle or period. A menstrual cycle is the time from the first day of the menstrual period until the first day of the next menstrual period. The average length of a menstrual cycle is 28 days. However, a menstrual cycle can range in length from 21 days to about 35 days and still be normal.

[0106] The natural menstrual cycle comprises the phases:

[0107] - The menses phase: This phase begins on the first day of the menstrual period. It is when the lining of the uterus sheds.

[0108] - The follicular phase: This phase begins on the day of the menstrual period and ends at ovulation (it overlaps with the menses phase and ends at ovulation). Thus, the follicular phase includes the menses phase and the proliferative phase. During this time, the level of the hormone estrogen rises, which causes the lining of the uterus (the endometrium) to grow and thicken. In addition, another hormone — follicle-stimulating hormone (FSH) — causes follicles in the ovaries to grow. One of the developing follicles will form a fully mature egg (ovum).

[0109] - Ovulation: This phase occurs roughly at about day 14 in a 28-day menstrual cycle. A sudden increase in another hormone — luteinizing hormone (LH) — causes the ovary to release its egg. This event is ovulation.

[0110] - The luteal phase / secretory phase: This phase lasts from about day 15 to day 28. The egg leaves the ovary and begins to travel through the fallopian tubes to the uterus. The level of the hormone progesterone rises to help prepare the uterine lining for pregnancy.

[0111] If the egg becomes fertilized by sperm and attaches itself to the uterine wall (implantation), the process results in a pregnancy.

[0112] If pregnancy does not occur, estrogen and progesterone levels drop and the thick lining of the uterus sheds during the menstrual period.

[0113] Proliferative phase

[0114] The proliferative phase of the natural menstrual cycle is before female ovulation. The follicular phase of the menstrual cycle comprises the menstruation and the proliferative phase. The menstruation occurs from about day 1 to about day 5, where the surface of the endometrium sheds off resulting in menses. The proliferative phase occurs after the end of the menstruation, from about day 5 of the cycle, where the endometrial cells proliferate and create a thickening of the endometrial lining. The proliferative phase ends around day 14 of the menstrual cycle, where the ovulation occurs.

[0115] TOD 20 genes, top 15 genes etc.

[0116] When used herein, "top 20 genes", "top 15 genes", etc., refers to the top genes out of 157 differentially expressed genes (DEGs), which are the most important genes derived from a Random Forest model based on the 157 DEGs, which were selected based on their effect on the performance of the prediction model. The top most informative genes are selected based on the mean decrease accuracy in the first Random Forest model. The top most informative genes for the best performance of the prediction model, which are obtained by evaluating the specificity, sensitivity, accuracy, and AUC.

[0117] A method for determining the risk of recurrent pregnancy loss (RPL)

[0118] For the present invention, the method for determining the risk of RPL is an in vitro method performed outside of the human body. Hence, the invasive act of obtaining the endometrial biopsy is not part of the present invention.

[0119] An aspect of the present invention relates to a method for determining whether a female subject is at risk of recurrent pregnancy loss (RPL), the method comprising

[0120] - providing an endometrial biopsy obtained in a proliferative phase of the menstrual cycle from said subject;

[0121] - determining in the endometrial biopsy a gene expression profile for at least 2 genes comprising C10orf71 and / or CLEC4A, wherein if the gene expression profile comprises

[0122] - CLEC4A being upregulated; and / or

[0123] - C10orf71 being downregulated, compared to a gene expression profile comprising C10orf71 and / or CLEC4A in a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle from control females with healthy endometria and no (previous) history of RPL, it is indicative of said subject being at risk of RPL. As seen in Example 6, the risk of RPL can be determined or predicted based on C10orf71 and / or CLEC4A based on both a machine learning algorithm and by manually evaluating the gene expression regulation. Example 2 provides evidence of a distinct separation between RPL and control subjects when only analysing proliferative phase endometrial biopsies. Data is provided showing that the risk of RPL can be determined based on the top 2 genes, C10orf71 and CLEC4A. Further, the risk of RPL can even be accurately determined based on one out of the top 2 genes. Figure 5L shows that the risk of RPL can be determined with great accuracy, sensitivity, and specificity based on top 1 gene, such as being C10orf71, thus the risk of RPL can be determined using C10orf71 alone. The risk of RPL may also be determined based on CLEC4A alone as can be derived from the figures, such as figure 7 or 8.

[0124] Thus, an aspect of the present invention relates to a method for determining whether a female subject is at risk of recurrent pregnancy loss (RPL), the method comprising

[0125] - providing an endometrial biopsy obtained in a proliferative phase of the menstrual cycle from said subject;

[0126] - determining in the endometrial biopsy a gene expression profile comprising C10orf71 and / or CLEC4A, wherein if the gene expression profile comprises

[0127] - CLEC4A being upregulated; and / or

[0128] - C10orf71 being downregulated, compared to a gene expression profile comprising C10orf71 and / or CLEC4A in a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle from control females with healthy endometria and no (previous) history of RPL, it is indicative of said subject being at risk of RPL.

[0129] Not only does the method of the present invention determine if a female subject is at risk of having RPL. The method is also capable of determining if a female subject is not at risk of having RPL.

[0130] Thus, in an embodiment, the method further comprises - determining, in the endometrial biopsy of the female subject, the gene expression profile for at least 2 genes comprising C10orf71 and / or CLEC4A wherein if the gene expression profile comprises

[0131] - CLEC4A being downregulated or having a same expression; and / or

[0132] - C10orf71 being upregulated or having a same expression, compared to the gene expression profile comprising C10orf71 and / or CLEC4A in the panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle from control females with healthy endometria and no history of RPL, it is indicative of said subject not being at risk of RPL.

[0133] It can be assessed whether a female subject is at risk of RPL or not by using a common score. Such common score may for example be the "probability of RPL" given by a Random Forest classifier. The score ranges from 0 to 1. A score closer to 1, indicates that the subject is at risk of RPL. A score closer to 0, indicates that the subject is not at risk of RPL.

[0134] In an embodiment, the upregulation and the downregulation are determined by Iog2 fold change.

[0135] As can be seen in the examples, the regulation of gene expression is determined using Iog2 fold change.

[0136] In an embodiment, the upregulation is determined by a Iog2 fold change above 0.

[0137] In a further embodiment, the upregulation is a Iog2 fold change in the range of 0.1 to 5, such as 0.1 to 4, preferably 0.1 to 3.

[0138] As seen in Example 4, the Iog2 fold change values of the upregulated genes are within these ranges.

[0139] In an embodiment, the downregulation is determined by a Iog2 fold change below 0.

[0140] In a further embodiment, the downregulation is a Iog2 fold change in the range of -3 to -0.1, such as -2.5 to -0.1, preferably -2.3 to -0.2.

[0141] As seen in Example 4, the Iog2 fold change values of the downregulated genes are within these ranges. In an embodiment, the RPL is at least two consecutive pregnancy losses.

[0142] In an embodiment, the RPL is primary or secondary RPL.

[0143] In an embodiment, the gene expression profile comprises genes selected from the group consisting of TMEM176A, SLC22A16, ETV1, SPAG9, MGST1, UFL1, MTMR11, CYP24A1, KCNG1, SLAMF7, CDH3, CA11, TBX21, NFKB2, PTPRC, PCDHB4, SULT2B1, RGS1, ZFHX4, IL12RB1, SLC25A1, GRAP2, ZFYVE21, JAG1, CD37, CARD8, JAK3, COMP, TMEM176B, B3GAT1, CLEC4A, IK, CYTIP, TACR1, NRP2, RARRES1, TNFSF10, PAEP, ARHGAP9, C4BPA, TSPAN8, STEAP4, RAC2, GMFG, JCHAIN, PEBP4, PPHLN1, DOCK2, SDS, PKIB, DOCKIO, IL33, MMP7, FXYD2, CILP, CXCL9, PARVG, SLC46A3, NOVAI, GREB1L, PNMT, TIMM29, ABCA12, SLC26A1, MARCHF1, SSBP2, FAM217A, FGD2, DOK2, CACNA1B, SESN3, CD226, NMRAL1, GBP5, SORCS3, NCF1, MAML1, GFI1, VCAM1, SLC15A2, APEH, AMN, PLD4, PRKCB, CD3D, SCN11A, RHOH, GFRA2, TNIP2, INPP5D, CD52, GPR183, SEMA3E, KRT86, FEZ2, KRT13, CXCR6, THEMIS, NDNF, TVP23C, GAPT, KCNA3, C10orf71, CCDC184, SHISA3, ZNF852, HLA-DQB1, PLD5, CUEDC1, NQO1, TLCD5, LCN12, SHC4, EVI2B, C16orf54, ZNF566, GNG2, ARHGAP30, LITAF, TDRD7, HLA-DRB1, AKR1C3, LONP1, HLA-DQA1, SPN, AIF1, ATP10A, GPX3, IGLV1-40, TRBC1, TRBC2, IGHM, PTPRCAP, ADAT3, SIAH3, ZNF433-AS1, C4B, CASC6, GOLGA2P7, LTB, HLA-DPA1, NFAM1, CKMT1B, APOBEC3G, LRRD1, CFB, CDK11B, FZD10-AS1, SLC2A9-AS1, GPR162, GVINP1, TIFAB, MIR142HG, ADGRL1-AS1, CCL14, TRAC, and / or TGFB2-OT1.

[0144] As seen in Example 4 and 5, the above 157 genes were analyzed for DEGs and the 157 genes were used in a Random Forest prediction model.

[0145] The gene expression profile comprises at least 2 genes, wherein the at least 2 genes comprise C10orf71 and / or CLEC4A. The genes C10orf71 or CLEC4A may be combined with another gene expressed in a proliferative phase endometrial biopsy. Hence, in an embodiment, the at least 2 genes are one or more of the following combination selected from the group consisting of C10orf71 and CLEC4A, C10orf71 and SHISA3, C10orf71 and MMP7, C10orf71 and CILP, C10orf71 and LCN12, C10orf71 and SLC2A9-AS1, C10orf71 and KRT86, C10orf71 and PEBP4, C10orf71 and AIF1, C10orf71 and TSPAN8, C10orf71 and TGFB2-OT1, C10orf71 and ETV1, C10orf71 and NCF1, C10orf71 and PLD5, C10orf71 and PKIB, C10orf71 and AMN, C10orf71 and SEMA3E, C10orf71 and SHC4, C10orf71 and RHOH, CLEC4A and SHISA3, CLEC4A and MMP7, CLEC4A and CILP, CLEC4A and LCN12, CLEC4A and SLC2A9-AS1, CLEC4A and KRT86, CLEC4A and PEBP4, CLEC4A and AIF1, CLEC4A and TSPAN8, CLEC4A and TGFB2-OT1, CLEC4A and ETV1, CLEC4A and NCF1, CLEC4A and PLD5, CLEC4A and PKIB, CLEC4A and AMN, CLEC4A and SEMA3E, CLEC4A and SHC4, CLEC4A and RHOH, SHISA3 and SLC2A9-AS1, SHISA3 and PEBP4, CILP and LCN12, CILP and PEBP4, CILP and AMN, CILP and SEMA3E, CILP and RHOH, LCN12 and SLC2A9-AS1, LCN12 and PEBP4, LCN12 and AMN, LCN12 and RHOH, SLC2A9-AS1 and RHOH, PEBP4 and AIF1, PEBP4 and TSPAN8, PEBP4 and ETV1, PEBP4 and PKIB, PEBP4 and SEMA3E, PEBP4 and SHC4, PEBP4 and RHOH, and AIF1 and RHOH.

[0146] As can be seen in Example 6, the prediction of risk of RPL is successful with the pairwise gene combinations listed in Table 4.

[0147] In an embodiment, the gene expression profile comprises genes selected from the group consisting of C10orf71, CLEC4A, SHISA3, MMP7, CILP, LCN12, SLC2A9- AS1, KRT86, PEBP4, and / or AIF1 (top 10 genes), preferably selected from the group consisting of C10orf71, CLEC4A, SHISA3, MMP7, and / or CILP (top 5 genes), more preferably selected from the group consisting of C10orf71, CLEC4A, and / or SHISA3 (top 3 genes), most preferably selected from the group consisting of C10orf71 and / or CLEC4A (top 2 genes).

[0148] As seen in Example 5, the above preferred group of genes were derived from the Random Forest model. All the above top groups of genes comprise C10orf71 and / or CLEC4A.

[0149] In a further embodiment, the gene expression profile comprises genes selected from the group consisting of C10orf71, CLEC4A, SHISA3, MMP7, CILP, LCN12, SLC2A9-AS1, KRT86, PEBP4, AIF1, TSPAN8, TGFB2-OT1, ETV1, NCF1, PLD5, PKIB, AMN, SEMA3E, SHC4, and / or RHOH (top 20 genes), preferably selected from the group consisting of C10orf71, CLEC4A, SHISA3, MMP7, CILP, LCN12, SLC2A9-AS1, KRT86, PEBP4, AIF1, TSPAN8, TGFB2-OT1, ETV1, NCF1, and / or PLD5 (top 15 genes). As seen in Example 5, the above preferred group of genes were derived from the Random Forest model. Further, it is noted that all the above top group of genes comprise C10orf71 and / or CLEC4A.

[0150] In an embodiment, the proliferative phase is between 3-10 days after the first day of menstruation in a natural menstrual cycle.

[0151] In an embodiment, the proliferative phase is validated by histology dating.

[0152] In an embodiment, the control endometrial biopsy is obtained from a healthy female or a female undergoing in vitro fertilization treatment.

[0153] A healthy female and / or a female undergoing in vitro fertilization treatment are presumed to have a healthy endometrium. A female undergoing in vitro fertilization (IVF) treatment may undergo such IVF treatment for various reasons unrelated to a diseased endometrium. Such reasons may for example be due to paternal issues or the male partner. In principle, the control endometrial biopsy may be obtained from any female control regardless of health status and whether the female has had any previous pregnancies, provided that the female control has a healthy endometrium.

[0154] In an embodiment, the control endometrial biopsy is obtained from a female having a healthy endometrium.

[0155] In an embodiment, the gene expression profile is determined at cDNA level, RNA level or protein level, preferably RNA level.

[0156] As seen in the examples, the gene expression profile is determined on RNA level. However, in principle the gene expression profile can also be determined on the cDNA and / or protein level.

[0157] In a further embodiment, the RNA level is mRNA level.

[0158] In an embodiment, the gene expression profile is determined before or after diagnosis of RPL, preferably before. In an embodiment, the gene expression profile is determined by RNA sequencing, such as RNA next-generation sequencing, quantitative reverse transcription polymerase chain reaction (qRT-PCR), or microarray technology.

[0159] In a further embodiment, the microarray technology is selected from the group consisting of:

[0160] DNA microarray, such as DNA chip, biochip, spotted arrays on glass, in-situ synthesized array, self-assembled array, cDNA microarray, serial analysis of gene expression (SAGE);

[0161] RNA microarray such as serial analysis of gene expression (SAGE); and / or protein microarray, such as antibody array, enzyme-linked immunosorbent assay (ELISA).

[0162] In an embodiment, the comparison of gene expression profile of the endometrial biopsy obtained in a proliferative phase of the menstrual cycle of the female subject and a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle is performed using one or more of differential expression analysis.

[0163] As seen in Example 4, the differential expression (DE) analysis was performed with the DESeq2 package, wherein two distinct comparative analyses were conducted: (1) secretory versus proliferative phase endometrial biopsy, and (2) RPL versus control group. DESeq2 package and limma-voom package are two common methods for differential expression analysis.

[0164] An aspect of the invention relates to a method for determining whether a female subject is at risk of recurrent pregnancy loss (RPL), the method comprising

[0165] - providing an endometrial biopsy obtained in a proliferative phase of the menstrual cycle from said subject;

[0166] - determining in the endometrial biopsy a gene expression profile for at least 2 genes comprising C10orf71 and / or CLEC4A, wherein if the gene expression profile comprises

[0167] - CLEC4A being upregulated; and / or

[0168] - C10orf71 being downregulated, compared to a gene expression profile comprising C10orf71 and / or CLEC4A in a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle from control females with healthy endometria and no history of RPL, it is indicative of said subject being at risk of RPL, wherein the gene expression profile comprises one or more genes selected from the group consisting of TMEM176A, SLC22A16, ETV1, SPAG9, MGST1, UFL1, MTMR11, CYP24A1, KCNG1, SLAMF7, CDH3, CA11, TBX21, NFKB2, PTPRC, PCDHB4, SULT2B1, RGS1, ZFHX4, IL12RB1, SLC25A1, GRAP2, ZFYVE21, JAG1, CD37, CARD8, JAK3, COMP, TMEM176B, B3GAT1, CLEC4A, IK, CYTIP, TACR1, NRP2, RARRES1, TNFSF10, PAEP, ARHGAP9, C4BPA, TSPAN8, STEAP4, RAC2, GMFG, JCHAIN, PEBP4, PPHLN1, DOCK2, SDS, PKIB, DOCKIO, IL33, MMP7, FXYD2, CILP, CXCL9, PARVG, SLC46A3, NOVAI, GREB1L, PNMT, TIMM29, ABCA12, SLC26A1, MARCHF1, SSBP2, FAM217A, FGD2, DOK2, CACNA1B, SESN3, CD226, NMRAL1, GBP5, SORCS3, NCF1, MAML1, GFI1, VCAM1, SLC15A2, APEH, AMN, PLD4, PRKCB, CD3D, SCN11A, RHOH, GFRA2, TNIP2, INPP5D, CD52, GPR183, SEMA3E, KRT86, FEZ2, KRT13, CXCR6, THEMIS, NDNF, TVP23C, GAPT, KCNA3, C10orf71, CCDC184, SHISA3, ZNF852, HLA-DQB1, PLD5, CUEDC1, NQO1, TLCD5, LCN12, SHC4, EVI2B, C16orf54, ZNF566, GNG2, ARHGAP30, LITAF, TDRD7, HLA-DRB1, AKR1C3, LONP1, HLA-DQA1, SPN, AIF1, ATP10A, GPX3, IGLV1-40, TRBC1, TRBC2, IGHM, PTPRCAP, ADAT3, SIAH3, ZNF433-AS1, C4B, CASC6, GOLGA2P7, LTB, HLA-DPA1, NFAM1, CKMT1B, APOBEC3G, LRRD1, CFB, CDK11B, FZD10-AS1, SLC2A9-AS1, GPR162, GVINP1, TIFAB, MIR142HG, ADGRL1-AS1, CCL14, TRAC, TGFB2-OT1, or any combination thereof.

[0169] An aspect of the present invention relates to a method for determining whether a female subject is at risk of recurrent pregnancy loss (RPL), the method comprising

[0170] - providing an endometrial biopsy obtained in a proliferative phase of the menstrual cycle from said subject;

[0171] - determining in the endometrial biopsy a gene expression profile for genes comprising C10orf71, wherein if the gene expression profile comprises

[0172] - C10orf71 being downregulated, compared to a gene expression profile comprising C10orf71 in a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle from control females with healthy endometria and no (previous) history of RPL, it is indicative of said subject being at risk of RPL.

[0173] Figure 5L shows that the risk of RPL can be determined with great accuracy, sensitivity, and specificity based on top 1 gene being C10orf71, thus the risk of RPL can be determined using C10orf71 alone.

[0174] Use of an endometrial biopsy

[0175] A further aspect of the invention relates to use of an endometrial biopsy obtained in a proliferative phase of the menstrual cycle for determining risk of recurrent pregnancy loss (RPL) in a female subject.

[0176] As can be seen in the examples such as Example 3, significant differences in differentially expressed genes (DEGs) between endometrial biopsies collected in the secretory phase versus the proliferative phase of the menstrual cycle were observed, and this was expected. However, surprisingly, an enrichment of DEGs in immunological and infectious disease pathways, appeared for proliferative phase endometrial biopsies.

[0177] Another aspect of the invention relates to use of an endometrial biopsy obtained in a proliferative phase of the menstrual cycle for determining the risk of recurrent pregnancy loss (RPL) in a female subject, wherein the risk of RPL is determined by comparing

[0178] - a gene expression profile comprising C10orf71 and / or CLEC4A from an endometrial biopsy obtained in a proliferative phase of the menstrual cycle of the female subject, and

[0179] - a gene expression profile comprising C10orf71 and / or CLEC4A from a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle from control females with healthy endometria and no history of RPL.

[0180] In an embodiment, the gene expression profile comprises one or more genes selected from the group consisting of TMEM176A, SLC22A16, ETV1, SPAG9, MGST1, UFL1, MTMR11, CYP24A1, KCNG1, SLAMF7, CDH3, CA11, TBX21, NFKB2, PTPRC, PCDHB4, SULT2B1, RGS1, ZFHX4, IL12RB1, SLC25A1, GRAP2, ZFYVE21, JAG1, CD37, CARD8, JAK3, COMP, TMEM176B, B3GAT1, CLEC4A, IK, CYTIP, TACR1, NRP2, RARRES1, TNFSF10, PAEP, ARHGAP9, C4BPA, TSPAN8, STEAP4, RAC2, GMFG, JCHAIN, PEBP4, PPHLN1, DOCK2, SDS, PKIB, DOCKIO, IL33, MMP7, FXYD2, CILP, CXCL9, PARVG, SLC46A3, NOVAI, GREB1L, PNMT, TIMM29, ABCA12, SLC26A1, MARCHF1, SSBP2, FAM217A, FGD2, DOK2, CACNA1B, SESN3, CD226, NMRAL1, GBP5, SORCS3, NCF1, MAML1, GFI1, VCAM1, SLC15A2, APEH, AMN, PLD4, PRKCB, CD3D, SCN11A, RHOH, GFRA2, TNIP2, INPP5D, CD52, GPR183, SEMA3E, KRT86, FEZ2, KRT13, CXCR6, THEMIS, NDNF, TVP23C, GAPT, KCNA3, C10orf71, CCDC184, SHISA3, ZNF852, HLA-DQB1, PLD5, CUEDC1, NQO1, TLCD5, LCN12, SHC4, EVI2B, C16orf54, ZNF566, GNG2, ARHGAP30, LITAF, TDRD7, HLA-DRB1, AKR1C3, LONP1, HLA-DQA1, SPN, AIF1, ATP10A, GPX3, IGLV1-40, TRBC1, TRBC2, IGHM, PTPRCAP, ADAT3, SIAH3, ZNF433-AS1, C4B, CASC6, GOLGA2P7, LTB, HLA-DPA1, NFAM1, CKMT1B, APOBEC3G, LRRD1, CFB, CDK11B, FZD10-AS1, SLC2A9-AS1, GPR162, GVINP1, TIFAB, MIR142HG, ADGRL1-AS1, CCL14, TRAC, TGFB2-OT1, or any combination thereof.

[0181] Use of a gene expression profile

[0182] Yet another aspect of the invention relates to use of a gene expression profile comprising C10orf71 and / or CLEC4A for determining risk of recurrent pregnancy loss (RPL) in a female subject.

[0183] Examples 5 and 6 provide evidence that risk of RPL can be predicted and / or determined based on gene expression data comprising C10orf71 and / or CLEC4A.

[0184] The following embodiments may apply to the two above aspects relating to a use.

[0185] In an embodiment, the risk of RPL is determined by comparing

[0186] - a gene expression profile comprising C10orf71 and / or CLEC4A from an endometrial biopsy obtained in a proliferative phase of the menstrual cycle of the female subject, and

[0187] - a gene expression profile comprising C10orf71 and / or CLEC4A from a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle from control females with healthy endometria and no history of RPL. In an embodiment, the control endometrial biopsy is obtained from a healthy female or a female undergoing in vitro fertilization treatment.

[0188] A healthy female and / or a female undergoing in vitro fertilization treatment are presumed to have a healthy endometrium. A female undergoing in vitro fertilization (IVF) treatment may undergo such IVF treatment for various reasons unrelated to a diseased endometrium. Such reasons may for example be due to paternal issues, male partner, or genetic defects of the fetus. In principle, the control endometrial biopsy may be obtained from any female control regardless of health status and whether the female has had any previous pregnancies, provided that the female control has a healthy endometrium.

[0189] In an embodiment, the control endometrial biopsy is obtained from a female having a healthy endometrium.

[0190] In an embodiment, the RPL is at least two consecutive pregnancy losses.

[0191] In an embodiment, the RPL is primary or secondary RPL.

[0192] In an embodiment, the gene expression profile is obtained from an endometrial biopsy obtained in a proliferative phase of the menstrual cycle from a female subject.

[0193] In an embodiment, the proliferative phase is between 3-10 days after the first day of menstruation in a natural menstrual cycle.

[0194] In an embodiment, the proliferative phase is validated by histology dating.

[0195] In an embodiment, the gene expression profile is determined at cDNA level, RNA level or protein level, preferably RNA level.

[0196] As seen in the examples, the gene expression profile is determined on RNA level. However, in principle the gene expression profile can also be determined by cDNA and / or protein level.

[0197] In a further embodiment, the RNA level is mRNA level. In an embodiment, the gene expression profile is determined before or after diagnosis of RPL, preferably before.

[0198] In an embodiment, the gene expression profile is determined by RNA sequencing, such as RNA next-generation sequencing, quantitative reverse transcription polymerase chain reaction (qRT-PCR), or microarray technology.

[0199] In a further embodiment, the microarray technology is selected from the group consisting of:

[0200] DNA microarray, such as DNA chip, biochip, spotted arrays on glass, in-situ synthesized array, self-assembled array, cDNA microarray, serial analysis of gene expression (SAGE);

[0201] RNA microarray such as serial analysis of gene expression (SAGE); and / or protein microarray, such as antibody array, enzyme-linked immunosorbent assay (ELISA).

[0202] An aspect of the invention relates to use of a gene expression profile comprising C10orf71 and / or CLEC4A for determining the risk of recurrent pregnancy loss (RPL) in a female subject, wherein the gene expression profile comprises one or more genes selected from the group consisting of TMEM176A, SLC22A16, ETV1, SPAG9, MGST1, UFL1, MTMR11, CYP24A1, KCNG1, SLAMF7, CDH3, CA11, TBX21, NFKB2, PTPRC, PCDHB4, SULT2B1, RGS1, ZFHX4, IL12RB1, SLC25A1, GRAP2, ZFYVE21, JAG1, CD37, CARD8, JAK3, COMP, TMEM176B, B3GAT1, CLEC4A, IK, CYTIP, TACR1, NRP2, RARRES1, TNFSF10, PAEP, ARHGAP9, C4BPA, TSPAN8, STEAP4, RAC2, GMFG, JCHAIN, PEBP4, PPHLN1, DOCK2, SDS, PKIB, DOCKIO, IL33, MMP7, FXYD2, CILP, CXCL9, PARVG, SLC46A3, NOVAI, GREB1L, PNMT, TIMM29, ABCA12, SLC26A1, MARCHF1, SSBP2, FAM217A, FGD2, DOK2, CACNA1B, SESN3, CD226, NMRAL1, GBP5, SORCS3, NCF1, MAML1, GFI1, VCAM1, SLC15A2, APEH, AMN, PLD4, PRKCB, CD3D, SCN11A, RHOH, GFRA2, TNIP2, INPP5D, CD52, GPR183, SEMA3E, KRT86, FEZ2, KRT13, CXCR6, THEMIS, NDNF, TVP23C, GAPT, KCNA3, C10orf71, CCDC184, SHISA3, ZNF852, HLA-DQB1, PLD5, CUEDC1, NQO1, TLCD5, LCN12, SHC4, EVI2B, C16orf54, ZNF566, GNG2, ARHGAP30, LITAF, TDRD7, HLA-DRB1, AKR1C3, LONP1, HLA-DQA1, SPN, AIF1, ATP10A, GPX3, IGLV1-40, TRBC1, TRBC2, IGHM, PTPRCAP, ADAT3, SIAH3, ZNF433-AS1, C4B, CASC6, GOLGA2P7, LTB, HLA-DPA1, NFAM1, CKMT1B, APOBEC3G, LRRD1, CFB, CDK11B, FZD10-AS1, SLC2A9-AS1, GPR162, GVINP1, TIFAB, MIR142HG, ADGRL1-AS1, CCL14, TRAC, TGFB2-OT1, or any combination thereof.

[0203] Machine learning / Al

[0204] An aspect of the present invention relates to processor system programmed to operate according to a machine learning (ML) algorithm for determining the risk of recurrent pregnancy loss (RPL) based on gene expression of at least 2 genes from an endometrial biopsy obtained in a proliferative phase of the menstrual cycle of a female subject compared to gene expression of the same at least 2 genes from a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle, the machine learning (ML) algorithm being trained, and / or being trainable, on data obtained by a method according to the first aspect of the invention. Preferably, said at least two genes are C10orf71 and / or CLEC4A.

[0205] In yet another aspect, the invention relates to use of a machine learning (ML) algorithm trained on data obtained by a method according to the first aspect of the invention to determine the risk of recurrent pregnancy loss (RPL) based on gene expression of at least 2 genes from an endometrial biopsy obtained in a proliferative phase of the menstrual cycle of a female subject compared to gene expression of the same at least 2 genes from a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle. Preferably, said at least two genes are C10orf71 and / or CLEC4A.

[0206] An aspect of the present invention relates to a system suitable for executing an algorithm (such as machine learning (ML) algorithm) for determining the risk of recurrent pregnancy loss (RPL) based on gene expression of at least 2 genes from an endometrial biopsy obtained in a proliferative phase of the menstrual cycle of a female subject compared to gene expression of the same at least 2 genes from a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle, the (machine learning) system being trained, and / or being trainable, on data provided according to the first aspect of the invention. Advantageously, the invention may also relate to a method for training a machine learning (ML) system for determining the risk of recurrent pregnancy loss (RPL) based on gene expression of at least 2 genes from an endometrial biopsy obtained in a proliferative phase of the menstrual cycle of a female subject compared to gene expression of the same at least 2 genes from a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle, such as according to the first aspect of the invention.

[0207] In an embodiment, the endometrial biopsy further comprises genes selected from the group consisting of TMEM176A, SLC22A16, ETV1, SPAG9, MGST1, UFL1, MTMR11, CYP24A1, KCNG1, SLAMF7, CDH3, CA11, TBX21, NFKB2, PTPRC, PCDHB4, SULT2B1, RGS1, ZFHX4, IL12RB1, SLC25A1, GRAP2, ZFYVE21, JAG1, CD37, CARD8, JAK3, COMP, TMEM176B, B3GAT1, CLEC4A, IK, CYTIP, TACR1, NRP2, RARRES1, TNFSF10, PAEP, ARHGAP9, C4BPA, TSPAN8, STEAP4, RAC2, GMFG, JCHAIN, PEBP4, PPHLN1, DOCK2, SDS, PKIB, DOCKIO, IL33, MMP7, FXYD2, CILP, CXCL9, PARVG, SLC46A3, NOVAI, GREB1L, PNMT, TIMM29, ABCA12, SLC26A1, MARCHF1, SSBP2, FAM217A, FGD2, DOK2, CACNA1B, SESN3, CD226, NMRAL1, GBP5, SORCS3, NCF1, MAML1, GFI1, VCAM1, SLC15A2, APEH, AMN, PLD4, PRKCB, CD3D, SCN11A, RHOH, GFRA2, TNIP2, INPP5D, CD52, GPR183, SEMA3E, KRT86, FEZ2, KRT13, CXCR6, THEMIS, NDNF, TVP23C, GAPT, KCNA3, C10orf71, CCDC184, SHISA3, ZNF852, HLA-DQB1, PLD5, CUEDC1, NQO1, TLCD5, LCN12, SHC4, EVI2B, C16orf54, ZNF566, GNG2, ARHGAP30, LITAF, TDRD7, HLA-DRB1, AKR1C3, LONP1, HLA-DQA1, SPN, AIF1, ATP10A, GPX3, IGLV1-40, TRBC1, TRBC2, IGHM, PTPRCAP, ADAT3, SIAH3, ZNF433-AS1, C4B, CASC6, GOLGA2P7, LTB, HLA-DPA1, NFAM1, CKMT1B, APOBEC3G, LRRD1, CFB, CDK11B, FZD10-AS1, SLC2A9-AS1, GPR162, GVINP1, TIFAB, MIR142HG, ADGRL1-AS1, CCL14, TRAC, TGFB2-OT1, or any combination thereof.

[0208] In an embodiment, the at least 2 genes are one or more of the following combinations selected from the group consisting of C10orf71 and CLEC4A, C10orf71 and SHISA3, C10orf71 and MMP7, C10orf71 and CILP, C10orf71 and LCN12, C10orf71 and SLC2A9-AS1, C10orf71 and KRT86, C10orf71 and PEBP4, C10orf71 and AIF1, C10orf71 and TSPAN8, C10orf71 and TGFB2-OT1, C10orf71 and ETV1, C10orf71 and NCF1, C10orf71 and PLD5, C10orf71 and PKIB, C10orf71 and AMN, C10orf71 and SEMA3E, C10orf71 and SHC4, C10orf71 and RHOH, CLEC4A and SHISA3, CLEC4A and MMP7, CLEC4A and CILP, CLEC4A and LCN12, CLEC4A and SLC2A9-AS1, CLEC4A and KRT86, CLEC4A and PEBP4, CLEC4A and AIF1, CLEC4A and TSPAN8, CLEC4A and TGFB2-OT1, CLEC4A and ETV1, CLEC4A and NCF1, CLEC4A and PLD5, CLEC4A and PKIB, CLEC4A and AMN, CLEC4A and SEMA3E, CLEC4A and SHC4, CLEC4A and RHOH, SHISA3 and SLC2A9-AS1, SHISA3 and PEBP4, CILP and LCN12, CILP and PEBP4, CILP and AMN, CILP and SEMA3E, CILP and RHOH, LCN12 and SLC2A9-AS1, LCN12 and PEBP4, LCN12 and AMN, LCN12 and RHOH, SLC2A9-AS1 and RHOH, PEBP4 and AIF1, PEBP4 and TSPAN8, PEBP4 and ETV1, PEBP4 and PKIB, PEBP4 and SEMA3E, PEBP4 and SHC4, PEBP4 and RHOH, and AIF1 and RHOH.

[0209] Thus, yet an embodiment of the invention relates to the method comprising the steps of:

[0210] -receiving training data comprising a first set of information (1SI), such as a first database, and a second set of information (2SI), such as a second database, -training the system for estimating score variants using said training data, and

[0211] - validating the system using correlated specific score variants to the gene expression information.

[0212] Preferably the system and / or algorithm and / or method is implemented on a computer, thus being computer-implemented.

[0213] As seen in Example 5 and 6, the machine learning method used to predict the risk of RPL is Random Forest. The Random Forest (RF) machine learning method was used to predict RPL versus control status based on transcriptomics data. In Example 5, the RF method was used to predict the risk of RPL using gene expression data from the top 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, and / or 2 most important genes and down to only one most important gene.

[0214] Below is given a list of some no-limiting types of algorithms that are particularly suited for machine learning (ML) system and / or training of a (ML) system using gene expression data: 1. Deep Learning Algorithms: Deep learning is a subset of machine learning where artificial neural networks, algorithms inspired by the human brain, learn from large amounts of data. Deep learning algorithms are capable of learning to represent the world as a nested hierarchy of concepts, with each concept defined in relation to simpler concepts, and more abstract representations computed in terms of less abstract ones. The skilled reader is referred to for example University of Illinois at Urbana-Champaign; "Al predicts enzyme function better than leading tools." ScienceDaily. ScienceDaily, 30 March 2023.

[0215] < www. sciencedaily, com / releases / 2023 / 03 / 23033017212 l.htm>.

[0216] 2. Contrastive Learning : This is a type of unsupervised learning approach that trains models to learn similar features from similar data points and different features from different data points. An Al tool named 'CLEAN' was recently reported to use this algorithm to predict enzyme function, cf. Gupta, R., Srivastava, D., Sahu, M. et al. Artificial intelligence to deep learning: machine intelligence approach for drug discovery. Mol Divers 25, 1315-1360 (2021) for more details.

[0217] 3. Artificial Neural Networks (ANNs): ANNs are computing systems vaguely inspired by the biological neural networks that constitute animal brains. An ANN is based on a collection of connected units or nodes called artificial neurons, which loosely model the neurons in a biological brain.

[0218] 4. Support Vector Machines (SVMs): SVMs are supervised learning models with associated learning algorithms that analyze data for classification and regression analysis.

[0219] 5. Generative Adversarial Networks (GANs): GANs are a class of artificial intelligence algorithms used in unsupervised machine learning, implemented by a system of two neural networks contesting with each other in a zero-sum game framework.

[0220] These algorithms can be used individually or in combination, depending on the specific requirements of the type of gene expression data. The invention according to this aspect can be implemented by means of hardware, software, firmware or any combination of these. The invention or some of the features thereof can also be implemented as software running on one or more data processors and / or digital signal processors.

[0221] The individual elements of an embodiment of the invention may be physically, functionally, and logically implemented in any suitable way such as in a single unit, in a plurality of units or as part of separate functional units. The invention may be implemented in a single unit or be both physically and functionally distributed between different units and processors.

[0222] It should be noted that embodiments and features described in the context of one of the aspects of the present invention also apply to the other aspects of the invention. Thus, for example the embodiments / claims relating to the method of the invention also apply to the embodiments relating to the use- embodiments / claims. Hence, individual features mentioned in different claims, may possibly be advantageously combined, and the mentioning of these features in different claims does not exclude that a combination of features is not possible and advantageous.

[0223] Although the present invention has been described in connection with the specified embodiments, it should not be construed as being in any way limited to the presented examples. The scope of the present invention is to be interpreted in the light of the accompanying claim set. In the context of the claims, the terms "comprising" or "comprises" do not exclude other possible elements or steps. Also, the mentioning of references such as "a" or "an" etc. should not be construed as excluding a plurality.

[0224] All patent and non-patent references cited in the present application, are hereby incorporated by reference in their entirety.

[0225] The invention will now be described in further details in the following non-limiting examples. Examples

[0226] Aim of study

[0227] The aim of the present example is to characterize the RPL and control cohorts.

[0228] Materials and methods

[0229] The study consisted of two serial cohorts (Figure 1); an RPL cohort of 108 women and a control group of 27 healthy women with a presumed healthy endometrium, who were referred to a fertility clinic for IVF / ICSI treatment. The study was approved by the local Ethics Committees of The Capital Region of Denmark (H 15011157) and Region Zealand (EMN-2021-00360, SJ-788) and approved by the local data protection committee (REG-047-2019). Written informed consent and oral informed consent were obtained from all patients and controls. The study was carried out in accordance with the ethical standard of the Helsinki declaration. All women were referred to tertiary referral academic hospitals, Department of Obstetrics and Gynaecology, Rigshospitalet, or Department of Obstetrics and Gynaecology, Hvidovre Hospital, Copenhagen University Hospital, Denmark.

[0230] All patients of the RPL cohort were recruited to the project from May 2016 until February 2019. Pregnancy loss was defined as the spontaneous demise of a pregnancy before the fetus reaches viability, which includes all pregnancy losses from the time of conception until the end of second trimester. Recurrent pregnancy loss was defined as three or more consecutive pregnancy losses or two second trimester losses or still births. Recurrent pregnancy loss is subdivided into primary and secondary RPL, where primary RPL is defined as RPL without a previous pregnancy, while secondary RPL is defined as RPL after one or more previous pregnancies that progressed beyond 23 weeks of gestation (GW) (in Denmark, GW 22) (Bender Atik et al. 2023).

[0231] All women were between 18-42 years old at the time of the collection of the biopsy (two patients were 42 years old at the time of the biopsy, and all other participants was 18-40 years old). Exclusion criteria were known uterus pathology prior to the biopsy. Upon referral, all women were asked at the consultation with the medical doctor or nurse, or through a questionnaire, about information on previous pregnancies and their outcome, previous diseases, height, weight, and smoking status. All clinical information regarding the patients were obtained from clinical records and a research database.

[0232] Unexplained RPL was defined according to the ESHRE guideline 2019: no parental chromosome abnormalities (both concerning the woman and the male partner); and no antiphospholipid syndrome, defined as either a positive lupus anticoagulant test or anticardiolipin antibodies >20 klU / L in both the initial screening samples and by re-testing with 12 weeks of interval. Risk factors for RPL were defined as: (1) maternal endocrine and metabolic disease, (2) uterine abnormalities, (3) thrombophilic disorders, such as protein S deficiency, factor V Leiden heterozygocity, (4) BMI >30 kg / m2, and (5) smoking. If a woman achieved a pregnancy after referral to the Department of Obstetrics and Gynaecology, the pregnancy was confirmed by two blood samples measuring human choriogonadotropin (GW 4) with repetition after one week. Weekly ultrasound scans were performed from GW 6 to 9 and thereafter every second week until GW 16 or an incident of pregnancy loss.

[0233] In case of a normal nuchal translucency scan (GW 11-15) and ongoing pregnancy at GW 16, the women were referred to standard obstetric care at their local hospital. The cohort was followed for three years and 9 months (45 months). The pregnancy outcome was classified in four different time points referred to as A, B, C and D. Pregnancy Outcome A was described after the first menstrual cycle after the biopsy. The second registration of Outcome B was defined as the first pregnancy after the biopsy. The third Pregnancy Outcome C was the pregnancy outcome after one year, the fourth Pregnancy Outcome D was 3 years and 9 months after the biopsy. At all time points, the outcome was categorized into three categories: (1) healthy pregnancy, (2) pregnancy loss, and (3) no achieved pregnancy. For the fourth time point Pregnancy Outcome D, a measurement 'time-to-pregnancy' (TTP) was included, defined as the number of days from the biopsy to the first day of the last menstruation ending in a healthy pregnancy at term. TTP was further divided into a healthy pregnancy, defined as a healthy pregnancy at term, and an unhealthy pregnancy in cases of a loss. The control group consisted of healthy women, who were part of a couple experiencing infertility, referred to IVF or ICSI treatment. Endometrial healthy women from a RCT study at the clinic at the Department of Obstetrics and Gynaecology, Hvidovre Hospital, were included. The female subjects with healthy endometria, who became pregnant in the following menstrual cycles, were diagnosed with either male infertility (n = 13), tuba infertility (n = 3), or unknown infertility (n = 11). All controls were examined by an office hysteroscopy with biopsy as a part of the scratching procedure from March 2013 until September 2017. All women in the control group achieved a pregnancy with IVF or ICSI treatment procedures (median = 3, range 2 - 7). All biopsies were collected in a natural menstrual cycle, except for one, where the woman was treated with Nafarelin, used as a nasal spray with 200 microgram per day, while the biopsy was collected. All other controls did not take any hormonal treatment seven days prior to the biopsy.

[0234] For the RPL cohort, the endometrial biopsy was scheduled at the proliferative phase 3-10 days after the first day of menstruation in a natural menstrual cycle. The schedule was performed by the knowledge of the menstrual cycle length for the patient and the date of the first day of menstruation two to three months prior to the collection of the biopsy. Due to individual wishes of the patients in the RPL group, and sometimes change of scheduled appointments with the clinic, peripheral progesterone and oestrogen levels were measured at the day of the biopsy. Histology dating was performed afterwards to validate the date of the biopsy in relation to the menstrual cycle.

[0235] For the control group, the biopsy was scheduled in the proliferative phase in a natural menstrual cycle with no hormonal treatment, the cycle just after and up to three months after the second IVF / ICSI treatment. Histology dating was performed on all the biopsies to validate the date of the biopsy in relation to the menstrual cycle. All biopsies were collected using an office hysteroscopy in a clean procedure by a specialist in Gynaecology at the Department of Obstetrics and Gynaecology. The participants were instructed to take oral paracetamol 1000 mg and Ibuprofen 400 mg approximately one hour before the hysteroscopy, according to the normal procedure in the clinic. An office hysteroscopy with an evaluation of the uterine cavity and cervical canal was performed using an ALPHASCOPE™ hysteroscope (GMS40A) 1.9 mm with GYNECARE VERSASCOPE™ sheath (GMS805) 3.5 mm (Ethicon, Johnson & Johnson, Livingston, Scotland) and saline as distension media. The biopsy was collected using 7F forceps (GIMMI® GmbH).

[0236] Macroscopic description of the uterus was performed immediately, and if possible, two biopsies were collected in total. There was no firm strategy for precise location of the biopsies, and it was not registered. The procedure was performed by a trained clinician and was discontinued at the participant 's request in case of discomfort or pain. For both the patients and the control women, the uterine cavity was examined for possible pathology during the hysteroscopy. One endometrial biopsy was transferred into RNA stabilization solution (RNA later, RNA stabilization Reagent; QIAGEN, Venlo, The Netherlands) and stored at -80°C. Another biopsy was embedded in paraffin for histological dating and immunohistochemistry staining (IHC). For sample processing, see Figure 1.

[0237] Histology dating and immunohistochemistry staining

[0238] The second biopsy was sent for pathology examination after fixation in 10% neutral-buffered formalin. Within 24-72 hours, the tissue was paraffin-embedded, sectioned with a thickness of 4 pm and stained with haematoxylin-eosin at the local Pathology Department. The endometrium was categorized in early, mid- and late secretory phase; or in inactive and mixed phase, respectively, by experienced pathologists. If a biopsy was not available due to technical circumstances, the date of the endometrium was determined by the progesterone and oestrogen levels in peripheral blood samples taken at the day of the biopsy. A progesterone level from 0.4-5.0 nmol / L defined the biopsy to be in the proliferative phase, and >5 nmol / L in the secretory phase. This procedure was verified with 20 biopsies and peripheral blood levels of progesterone.

[0239] Immunohistochemistry (IHC) staining was performed to determine the presence and distribution of the selected biomarkers according to previously published protocols (Kofod et al. 2017). Sections of each biopsy were systematically and serially stained for the two immune cell markers: CD56 (a uterine natural killer (NK) cell marker) and CD163 (a M2 macrophage cell marker). All slides were scanned at 20x magnification in a Nanozoomer S360 scanner (Hamamatsu Oncotopix Scan, Horsholm, Denmark), and the scanned files were uploaded and analyzed in Visiopharm's VIS software (Visiopharm, Horsholm, Denmark).

[0240] Continuous counts of the assessed immune cells were performed with digital image analysis using a VIS software Brigthfield APP edited in the software's author setting. The APP was fine-tuned in an iterative process to detect the immunostaining in analysed tissue sections, dividing the measured staining intensity into low, medium and high, and to calculate IHC H-score. Immune cell counts were expressed as percentage positive cells per total number of stromal cells or gland cells in the section, or number of positive cells per mm2. All data was generated with an in-house-validated algorithm for each parameter of the two cell markers (Kofod et al. 2017; Hallager et al. 2021).

[0241] Results

[0242] The RPL cohort included 131 women with endometrial biopsies. For RNA-seq analysis, biopsies from 108 women were included. Of these, 76 women were collected in the proliferative phase, 29 in the secretory phase, and three could not be classified or were from the mid-cycle period. The IVF control group consisted of 32 women of whom 27 women had biopsies analysed by RNA-seq (Figure 1). The large majority of these (22 out of 27) were in the proliferative phase; only one woman was in the secretory phase, and four could not be classified. Background characteristics for all the data and the women with biopsies in the proliferative phase are shown in Table 1. No differences between the RPL and controls for age, BMI and smoking were observed for the women included with biopsies collected in the proliferative phase. A significant difference in the number of pregnancy losses prior to the biopsy (P < 3.0e-13; Mann-Whitney U test) were observed for the women included with biopsies collected in the proliferative phase.

[0243] Table 1: Background characteristics for the study groups. All women included and women with biopsies from the proliferative phase, respectively.

[0244] Eleven women from the control group had been pregnant prior to the biopsy.

[0245] 2Eight women from the control group (proliferative phase) had been pregnant prior to the biopsy. Pregnancies prior to the biopsy (n = 16) included 1 uncomplicated pregnancy, 1 stillbirth, 1 preterm stillbirth, 5 pregnancy terminations, 1 ectopic pregnancy and 7 pregnancy losses.

[0246] Conclusion The RPL cohort and the control group were characterized and described. of RPL vs. controls

[0247] Aim of study

[0248] To separate the RPL cohort and the control group based on RNA-sequencing data.

[0249] Materials and methods

[0250] (See Example 1 for study groups)

[0251] RNA extraction, library Dreoaration and RNA next-generation seauencing

[0252] Total RNA was isolated from individual endometrial biopsies by preparation of the tissue lysate and afterwards an RNA purification. For the tissue lysate, RLT lysis buffer (QIAGEN, Venlo, The Netherlands), p-mercaptoethanol (ThermoFisher Scientific, Bleiswijk, The Netherlands) and a mortar instrument were used (BeadBug microtube homogenizer, Benchmark Scientific, South Plainfield, USA).

[0253] The RNeasy Mini Kit (QIAGEN) was followed using the manufacturer's protocol including optional centrifugation, followed by purification steps and a DNAse treatment using the QIAGEN DNA-free kit (QIAGEN). The RNA was checked with Qubit and real-time PCR for quantification and Bioanalyzer for size distribution detection. All samples showed a high RNA quality, with clear 18S and 28S ribosomal bands and RNA integrity numbers (RIN scores) between 5 to 10. All samples with RIN score >7 were included. All the extracted RNA samples were rRNA-depleted and the paired-end strand specific libraries were prepared by Novogene NGS Stranded RNA Library Prep Set (PT044) (Novogene, Cambridge, UK), using a 150 bp paired end design. All samples were sequenced on the Illumina Novoseq 6000 platform (Novogene, Cambridge, UK).

[0254] Bioinformatics and statistical analyses

[0255] The Computerome2 platform was used for all analyses (https: / / www.computerome.dk / ), including command line versions of mapping and quality control programs as described below. For downstream analysis, including statistics, the R statistical environment version 4.1.0 was used for analyses, on the same platform using specific packages as described below. Unpaired Student's t-test, Wilcoxon test and Mann-Whitney test were used where appropriate, using R. Kruskal-Wallis test using pseudo-ranks was performed in the pseudorank package version 1.0.2 using R (Happ et al. 2020).

[0256] RNA-sea data processing and initial exploratory data analysis The quality of raw sequencing reads was assessed by FastQC (http: / / www.bioinformatics.babraham.ac.uk / projects / fastqc and MultiQC tools (Ewels et al. 2016). All reads were trimmed to remove remaining adapter sequences and unwanted nucleotide bias at the 5'-end, using Trim Galore, version 0.6.4 (https: / / www.bioinformatics.babraham.ac.uk / projects / trim_galore / ). The reads were then mapped to the human transcriptome GRCh38, annotated according to Gencode version 39, using Salmon, version 1.5.2 (Patro et al. 2017).

[0257] For this study, the quasi-mapping approach was used. To run this mode, the transcriptome index was first computed with the default settings. Secondly, transcript quantification was performed. For the latter, the options seqBias', gcBias' and validateMappings' were set. The final mean mapping rate was of 77.79%. For exploratory data analysis, gene expression was normalized (salmon transcripts per million, TPM) to variance-stabilized counts. Principal component (PC) analysis identified a technical batch effect, specifically related to date of sampling. All samples belonging to the year 2016 underwent RNA extraction in March and May 2021. Prior to the removal of the batch effect, a separation of the two endometrial phases could be observed but no clear separation between RPL and controls was evident (data not shown). This was also the case when the proliferative phase data were analysed (data not shown). This batch effect for the year of the biopsy date was removed using the ComBat() function from the sva R package, version 3.42.0 (Leek et al. 2021). After correcting for the year of biopsy a better separation between the two endometrial phases and the RPL and controls was obtained for all the data (data not shown) and for the proliferative data (data not shown). For differential gene expression analysis, this batch effect was included as part of the statistical model.

[0258] Results

[0259] Initial RNA-sea analysis shows

[0260] In total, endometrial biopsies from 108 RPL patients and 27 control women were analysed by RNA-sequencing. First, the similarity of samples was visualized using PCA. PC3 and PC4 could clearly separate RPL and IVF control subjects (Figure 2B).

[0261] Immune genes are unregulated in RPL vs controls in the Droliferative Dhase Since a noticeable separation was observed between RPL and control subjects in PC3 and PC4 (Figure 2B), genes driving this separation were searched for. Removing the subjects in the secretory phase resulted in an even more clear separation between RPL and controls (Figure 4A).

[0262] The RNA-seg data analysis did not show any separation between type of RPL (primary and secondary RPL), nor different outcome descriptions (time points for pregnancy, loss, or 'no achieved pregnancy' and TTP (data not shown).

[0263] Conclusion

[0264] A strong signal in the data (PC3-4 in Figure 2B) is due to disease classification: RPL vs IVF controls. A noticeable separation between RPL and control subjects are shown, and an even more distinct separation between RPL and control subjects are obtained when only considering proliferative phase endometrial biopsies.

[0265] Aim of study

[0266] To separate proliferative and secretory phase biopsies based on RNA-seguencing data.

[0267] Materials and methods

[0268] (See Examples 1 and 2)

[0269] Results

[0270] In total, endometrial biopsies from 108 RPL patients and 27 control women were analysed by RNA-seguencing. PCI and PC2 showed a clear separation of proliferative and secretory phase biopsies (Figure 2A).

[0271] The separation of endometrial phases (PCI and PC2 in Figure 2A) is expected due to the variation of estradiol and progesterone variation according to menstrual cycle (Emera, Romero, and Wagner 2012; Ferenczy and Bergeron 1991). The results of gene set enrichment analyses comparing the secretory and proliferative phase are consistent with this notion.

[0272] No studies have investigated the proliferative phase in an RPL population, and the data in the present study show that the endometrium should be assessed not only in the secretory phase. Certain endometrial dysfunctionalities (for example immunological dysfunction) might be better detected in the early phase of the menstrual cycle due to a lesser degree of hormonal changes.

[0273] Conclusion

[0274] The separation of endometrial phases (PCI and PC2 in Figure 2A) is identified and expected due to the variation of estradiol and progesterone variation according to the menstrual cycle.

[0275] Aim of study Materials and methods

[0276] (See Example 2)

[0277] Differential expre ana

[0278] Differential expression (DE) analysis was performed with the DESeq2 package, version 1.34.0 (Love, Huber, and Anders 2014). Two distinct comparative analyses were conducted : (1) secretory vs proliferative, and (2) RPL vs controls. For each comparative analysis, different criteria were established to determine the differentially expressed genes (DEGs). For secretory vs proliferative states, a gene was reported as DEG if False Discovery Rate (FDR) <0.05 and the absolute Iog2- fold change (log2FC) between the groups was >0.5. For RPL versus controls, a more relaxed threshold (FDR <0.1 and abs (log2FC) >0) was used due to the lower number of DEGs compared to the previous analysis. For both comparative analyses, the final log2FC estimates were shrunken by the ashr method to obtain a more stable estimation of the log2FC (Stephens 2017).

[0279] In the DeSeq2 model formula, the batch variables defined above were included together with age and BMI, which are common confounding factors. For the secretory vs proliferative analysis, the condition (RPL vs controls) and the endometrial phase (secretory vs proliferative) were also included, while only the condition 'RPL vs controls' was added in the RPL vs control model. In both models, age and BMI were treated as continuous variables, remaining variables were treated as factors. Prior to the differential expression analysis, genes with low counts were filtered out, retaining only the genes that showed 10 or more counts in at least three samples.

[0280] The 157 genes included in this study comprise TMEM176A, SLC22A16, ETV1, SPAG9, MGST1, UFL1, MTMR11, CYP24A1, KCNG1, SLAMF7, CDH3, CA11, TBX21, NFKB2, PTPRC, PCDHB4, SULT2B1, RGS1, ZFHX4, IL12RB1, SLC25A1, GRAP2, ZFYVE21, JAG1, CD37, CARD8, JAK3, COMP, TMEM176B, B3GAT1, CLEC4A, IK, CYTIP, TACR1, NRP2, RARRES1, TNFSF10, PAEP, ARHGAP9, C4BPA, TSPAN8, STEAP4, RAC2, GMFG, JCHAIN, PEBP4, PPHLN1, DOCK2, SDS, PKIB, DOCKIO, IL33, MMP7, FXYD2, CILP, CXCL9, PARVG, SLC46A3, NOVAI, GREB1L, PNMT, TIMM29, ABCA12, SLC26A1, MARCHF1, SSBP2, FAM217A, FGD2, DOK2, CACNA1B, SESN3, CD226, NMRAL1, GBP5, SORCS3, NCF1, MAML1, GFI1, VCAM1, SLC15A2, APEH, AMN, PLD4, PRKCB, CD3D, SCN11A, RHOH, GFRA2, TNIP2, INPP5D, CD52, GPR183, SEMA3E, KRT86, FEZ2, KRT13, CXCR6, THEMIS, NDNF, TVP23C, GAPT, KCNA3, C10orf71, CCDC184, SHISA3, ZNF852, HLA-DQB1, PLD5, CUEDC1, NQO1, TLCD5, LCN12, SHC4, EVI2B, C16orf54, ZNF566, GNG2, ARHGAP30, LITAF, TDRD7, HLA-DRB1, AKR1C3, LONP1, HLA-DQA1, SPN, AIF1, ATP10A, GPX3, IGLV1-40, TRBC1, TRBC2, IGHM, PTPRCAP, ADAT3, SIAH3, ZNF433-AS1, C4B, CASC6, GOLGA2P7, LTB, HLA-DPA1, NFAM1, CKMT1B, APOBEC3G, LRRD1, CFB, CDK11B, FZD10-AS1, SLC2A9-AS1, GPR162, GVINP1, TIFAB, MIR142HG, ADGRL1-AS1, CCL14, TRAC, and TGFB2-OT1.

[0281] Gene ia of Genes and Genomes ana

[0282] Gene set enrichment analysis (GSEA) for GO terms and KEGG pathways was conducted by using the clusterProfiler package, version 4.2.2 (Gene Ontology 2021; Kanehisa et al. 2016; Wu et al. 2021). For each GSEA analysis, a LFC- ranked list of all genes resulting from the appropriate DE analysis was used, as defined above. In the subgroup analysis, combination of loadings from the appropriate PCA plot was used instead (the weights that each gene had in the construction of respective PC). For clarity, genes with the highest loadings (positive or negative) on a given PC contributed the most to the specific PC (Holand 2019). The strength of the enrichment was expressed as Normalized Enrichment Scores (NES).

[0283] K-means clustering and heat maos

[0284] The variance stabilized counts after the batch effect removal were firstly scaled as follows: z = (xy - i) / 07, where XIJ refers to the expression value for gene i in sample j, pi refers to the mean counts for gene / , and a, refers to the standard deviation for gene / . The kmeansQ function with k = 4 was used to perform clustering. Results from the k-means analysis were used to display sample clustering in the heat map computed by ComplexHeatmap package, version 2.10.0, with default settings for algorithm and method (Gu, Eils, and Schlesner 2016).

[0285] Results

[0286] When comparing biopsies from women in the proliferative and secretory phase, many DEGs were identified (1736 up-regulated and 1347 down-regulated, FDR <0.05 and abs (log2FC) >0.5), visualized using a volcano plot (Figure 3A). Gene set enrichment analysis showed that genes upregulated in the secretory phase were enriched for gene ontology (GO) terms related to immune and inflammatory response, regulation of hormone levels and KEGG pathways associated to hormone signalling, chemokine signalling, platelets, complement and the coagulation cascade gene system among others (Figure 3B and Figure 3C;

[0287] Data description: Gene Set Enrichment Analysis (GSEA) based on ranking genes by their secretory vs proliferative phase log2FC values and comparing the ranks within gene ontology term sets. Plots are split into 'activated' and 'suppressed' according to the positive and negative normalized enrichment score (NES) (NES>0 or NES<0). The results are reported as dot plots ranked by the NES (X axis). Gene set enrichment analysis showed that changes in gene expression occurred in genes related to 'hormone signalling) 'chemokine signalling) 'platelet activation) 'complement and coagulation cascades) which were all upregulated in the secretory phase, while cell cycle and DNA replication term genes were downregulated). Conversely, downregulated genes were enriched in several GO terms associated with cell division and related processes, including DNA repair and organelle fission, and KEGG pathways associated to cell cycle, DNA replication, Wnt and hedgehog signalling pathways among others.

[0288] Comparing RPL cohort and control group

[0289] Analysis of differential gene expression between these groups indicated that a much smaller set of genes were differentially expressed, even with a more relaxed DE criterion (123 upregulated, 43 downregulated, FDR<0.1, abs (log2FC) >0) (Figure 4B).

[0290] Gene set enrichment analysis based on ranking genes by RPL vs controls log2FC showed that genes related to immunology and immune responses, terms and pathways were up-regulated in the endometrial biopsies from RPL women (Data description: Gene set enrichment analysis (GSEA) based on the ranking genes by RPL vs Control log2FC and comparing the ranks within gene ontology term sets. Further, GSEA analysis was conducted for KEGG pathways. Gene set enrichment analysis showed that changes in RPL-upregulated genes are enriched in immune-related terms and pathways in the upregulated genes and pathways).

[0291] The RNA-seg data analysis did not show any separation between type of RPL (primary and secondary RPL), nor different outcome descriptions (time points for pregnancy, loss, or 'no achieved pregnancy' and TTP)

[0292] Thus, within the RPL cohort, expression data could not predict pregnancy outcome from preconceptional samples based on the subgroups defined below in the time frame of 45 months that was investigated. Similarly, there were no obvious clusters that could be attributed to specific RPL risk factors, and no association between expression data for explained vs unexplained RPL (data not shown).

[0293] In the entire control cohort, 11 out of 27 subjects have had previous pregnancies, including uncomplicated pregnancy, pregnancy ending in stillbirth, pregnancy termination, ectopic pregnancy and pregnancy loss. For the control women with biopsies from the proliferative phase, 8 out of 22 women have had previous pregnancies (Table 1). PCA only including the control group with biopsies in the proliferative phase focusing on pregnancy vs no pregnancy before the biopsy showed no obvious clusters that could be attributed to previous pregnancies prior to biopsy (data not shown). Furthermore, DEGs analysis resulted in only one differentially expressed gene between the control women, who have previously been pregnant, and the control women who have not been. Functional analyses were performed and revealed no immune-related terms or pathways to be differentially expressed between the two control subgroups. Thus, these analyses suggest that the observed differences in the gene expression profile between the controls and RPL cannot be attributed to previous pregnancies.

[0294] The regulation of expression of each of the genes of all 157 genes and the top 10 most important genes were assessed between the RPL cohort and control group. The fold change value of regulation of expression of each of 157 genes are not shown, however the genes which are upregulated have a Iog2 fold change value in the range of about 0.1 to about 3, whereas the genes which are downregulated have a Iog2 fold change value in the range of about -2.2 to about -0.2.

[0295] Indeed, a heatmap visualization of the 166 DEGs strongly indicated that at least four subgroups exist with varying expression of the same genes, which was also detected using k-means clustering (data not shown). Notably, two subgroups consisted exclusively of RPL patients (Subgroup 3 and 4), one only contained controls (Subgroup 1) and the last was dominated by controls but had seven RPL patients and was though a mixed subgroup (Subgroup 2). Notably, one of the RPL subgroups (Subgroup 4), had a substantially more distinct expression pattern vs the other subgroups.

[0296] None of the clinical outcomes were associated to the subgroups defined by the expression data (data not shown). Because the heat map analysis only used the 157 DEGs, a PCA approach was also used. Here, subjects in the PCA defined above only the proliferation biopsies (Figure 4A) were colored by the k-means group defined above (Figure 4C). From this analysis, it is evident that the subjects form a continuum from controls to RPL, through a diagonal line in the PCA.

[0297] To find out what genes that were driving this diagonal separation of the groups in the PCA, genes were ranked by their joint contribution to PCI and PC2 (effectively, the mean of PCI and PC2 loading values: see methods) and a GSEA analysis was performed based on these rankings using GO terms (Figure 4D) and KEGG pathways (Figure 4E). This clearly showed that genes driving samples to be further along the diagonal (towards Subgroup 4) were enriched in immune response terms and related pathways, while genes driving samples in the opposite direction (and towards Subgroup 1 consisting only of controls) were enriched in cell division terms pathways, and also oxidative phosphorylation and carbon metabolism pathways. Thus, expression of genes related to immune response vs. cell division not only separate between RPL and controls, but also drives the separation of subgroups in the cohort.

[0298] Conclusion

[0299] The study shows a clear upregulation of genes related to the immune response and a downregulation of cell cycle genes in RPL vs controls in the proliferative phase of the endometrial changes. No studies have investigated the proliferative phase of the endometrium in an RPL population, and the data in the present study show that the endometrium should be assessed not only in the secretory phase. Certain endometrial dysfunctionalities (for example immunological dysfunction) might be better detected in the early phase of the menstrual cycle due to a lower degree of hormonal changes.

[0300] Further, it is possible to discover subgroups of RPL patients and control subjects and a continuum of gene expression patterns that distinguish the groups. The existence of subgroups may provide inroads towards better understanding of what is characteristic for the endometrium of RPL patients, or stratification of such patients.

[0301] Many different immunological processes are present and upregulated in the endometrium in the RPL cohort including human leukocyte antigen (HLA) class II genes and molecules such as HLA-DRB1. Data indicates that RPL patients may have an increase of NK cells, and to a lesser degree also macrophages and T cells, which in different ways previously have been associated with RPL.

[0302] Example 5 - Prediction model Aim of study To use a machine learning method to predict the risk of RPL.

[0303] Materials and methods

[0304] (See Examples 1-4 for study groups and DEGs)

[0305] RPL vs control classification by Random Forests

[0306] The Random Forest (RF) machine learning method was used to predict RPL vs control status based on transcriptomics data. Briefly, RF is based on building a large set of decision trees, where each classification tree is built using a binary recursive partitioning (Ho 1995). The randomForest, version 4.6-14, and caTools, version 1.18.2, packages were used for the analysis (Liaw and Wiener 2002) (https: / / cran.r-project.org / web / packages / caTools / caTools.pdf). First, a RF classification approach for transformed and scaled DEGs from RPL vs controls DE analysis was performed. From the 166 DEGs, the focus was on 157, as the remaining ones were labelled as 'None' using the gconvert function from gprofiler2 library, and the ENSEMBLE IDs did not have any gene name assigned, and were therefore ignored. The subject samples were randomly divided into two partitions, the training set, corresponding to 70% of all the subjects (68 subjects; 53 RPL patients and 15 controls) and the test set, corresponding to the remaining 30% of the subjects (30 subjects; 23 RPL patients and 7 controls).

[0307] The training set was used to train and determine the parameters of the machinelearning model, and the test set was used to assess the classification performance (this set was not used in training). The RF classifier was run 500 times and with a randomness parameter in the different decision trees (ntree = 500) trained in the model (mtry = 79) and importance = TRUE, allowing the assessment of the importance of predictive features, which is gene expression in this case. Using the test set, the average accuracy, average sensitivity, average specificity and the mean confusion matrix (CM) were computed. In each run, the importance of each gene in the training set was also assessed based on their influence on the classification results. After the 500 runs, the 20 most important genes derived from the Random Forest model with the 157 DEGs were selected based on mean decrease accuracy, which reflects how much the model accuracy decreases when a certain variable is dropped (Menze et al. 2009). As a second step, the classification procedure was repeated but only using expression data from these 20 genes. Because the training and evaluation sets were not balanced, the Random Forest was applied when down-sampling and up- sampling the cohort. Both approaches make the sample size to be the same in cases as in controls. While in down-sampling the number of individuals in each group is adjusted to the group with the fewest subjects, in up-sampling the opposite is done. For the down-sampling, 15 RPL and 15 controls with a mtry = 79 was used. For the up-sampling 53 RPL and 53 controls with a mtry = 2 where used. Using these test sets, the mean confusion matrix was computed to compare with the original test set. Finally, to explore the effect of a systematic reduction of the number of these 20 genes on the performance of the prediction model, the classification procedure was repeated only using gene expression data from the top 15, 10, 9, 8, etc most important genes down to only one most important gene, respectively.

[0308] Results

[0309] As described above, the analyses indicated that RPL and control women were separable by expression data, but also that there is stratification in the data forming subgroups.

[0310] Because differential gene expression analysis is based on assessing differential gene expression on between groups of subjects rather than on an individual basis, an important question is whether the gene expression data held sufficient information to correctly classify subjects as 'RPL' or 'control'.

[0311] A Random Forest machine-learning model was trained and evaluated based on gene expression data from the 157 DEGs derived from the 166 DEGs minus those labelled as 'None', including 98 of the endometrial biopsies from the proliferative phase. Subjects were divided into a training set (53 RPL patients and 15 controls) and an evaluation set (23 RPL and 7 controls). The procedure was repeated 500 times with the training set and evaluation set static for the 500 iterations, resulting in 96.6% average accuracy, 95.7% average sensitivity and 99.5% average specificity. An advantage of the Random Forest algorithm is that the predictive importance of each gene can be assessed, and it is thus possible to prune the number of predicative features with little or no loss of predictive power. This is particularly relevant if the goal is to use the set of genes in a targeted approach like qRT-PCR in a clinical testing scenario. Hence, the 20 most informative genes for predictions were selected based on the average mean decrease accuracy in the first Random Forest model. When the procedure was repeated with these 20 most important genes, 93.4% average accuracy, 95.7% average sensitivity and 86.2% average specificity were obtained (Table 2 and Figure 5A). Furthermore, filtering the top 15 genes (Figure 5B), top 10 genes (Figure 5C), decreasing to 1-2 top gene(s) (Figure 5D-L) were also performed. These analyses revealed that only including the top two most important genes in the prediction model results in best performance with an average accuracy of 99.9%, 100% average sensitivity and 99.7% average specificity (Figure 5K). To test this, up- and down-sampling was used to create balanced test and evaluation tests, followed by training and evaluation as described above. There were no marked differences in the confusion matrices based on the original data set compared to the ones generated from up- or down-sampling data sets (Table 2), indicating that unbalanced sets are not affecting the model training unduly.

[0312] The diagnostic performance of the present prediction model was evaluated using ROC curves (see definitions for detailed description). A plot showing multiple ROC curves overlaid is shown in Figure 6 based upon the 157 DEGs. The mean Area Under the Curve (AUC) value is 0.996, which validates the diagnostic or predictive ability of the prediction model.

[0313] Table 2. Results from the different conditions for the prediction model. For the confusion matrix the original training set was based on 15 controls and 53 recurrent pregnancy loss (RPL) patients, the up-sampling training set was based on 53 controls and 53 RPL patients and the down-sampling training set was based on 15 controls and 15 RPL patients (DEGs = differentially expressed genes).

[0314] The regulation of expression of each of the genes in all 157 genes and the top 10 most important genes were assessed between the RPL cohort and control group. The fold change regulation of expression of each of the 157 genes are not shown, however the genes which are upregulated have a Iog2 fold change value in the range of about 0.1 to about 3, whereas the genes which are downregulated have a Iog2 fold change value in the range of about -2.2 to about -0.2.

[0315] Table 3 shows up- or downregulation of the specific top 10 genes compared to controls after log fold change (LFC) shrinkage. LFC shrinkage is a method, which compensates for any uncertainties caused by genes that are expressed in a very low amount, hence after LFC shrinkage the log fold change values are more precisely estimated for genes with low expression.

[0316] Table 3: Regulation of the top 10 genes after log fold change (LFC) shrinkage between RPL cohort and control group. The log2FoldChange value is the gene expression compared to the control group.

[0317] As a summary, this demonstrates that RPL subjects are readily distinguishable from control subjects by expression data, even with a limited set of genes, which strongly indicates that clinical testing based on gene expression data may be a viable option.

[0318] Conclusion

[0319] It is demonstrated that RPL subjects are readily distinguishable from control subjects by expression data using machine learning, even with a limited set of genes, which strongly indicates that clinical testing based on gene expression data is a viable option.

[0320] Example 6 - Prediction of RPL based on two genes Aim of study

[0321] To predict the risk of RPL by combining two genes within the top 20 genes.

[0322] Materials and methods

[0323] See Example 5

[0324] Results

[0325] Possible combinations of two genes within the top 20 genes comprising C10orf71,

[0326] CLEC4A, SHISA3, MMP7, CILP, LCN12, SLC2A9-AS1, KRT86, PEBP4, AIF1,

[0327] TSPAN8, TGFB2-OT1, ETV1, NCF1, PLD5, PKIB, AMN, SEMA3E, SHC4, and RHOH were made as listed in Table 4.

[0328] A Random Forest algorithm was trained as previously explained in Example 5, and the mean accuracy, sensitivity, and specificity was returned for each combination.

[0329] Table 4: Pairwise combinations with the top 20 genes. Mean accuracy, specificity, and sensitivity are provided for each pair. A total of 190 combinations are listed.

[0330] Many of the pairs show very high accuracy (i.e. 40 pairs over 95% accuracy, 10 pairs over 97% accuracy). That fits the expectation because in differential gene expression analysis, those genes are highly correlated. For pairwise gene combinations, the thresholds of sensitivity and specificity are preferably about 90% sensitivity and / or about 90% specificity to be clinically useful, more preferably about 95% sensitivity and / or about 95% specificity to be clinically useful.

[0331] Further, the combinations reveal the importance of the C10orf71 and / or CLEC4A gene transcripts.

[0332] The diagnostic performance of the present prediction model was evaluated using ROC curves (see definitions for detailed description). A plot showing multiple ROC curves overlaid based upon the 157 DEGs is shown in Figure 6. The mean Area Under the Curve (AUC) value is 0.996, which validates the diagnostic or predictive ability of the prediction model.

[0333] To further validate the prediction model for predicting RPL based on 2 genes, the standardized gene expression level of C10orf71 and / or CLEC4A gene transcripts were plotted into a x-y scatter plot where each point corresponds to a sample (see Figure 7 and 8). As can be seen in Figure 7, the RPL samples are clustered in an area where C10orf71 is downregulated and / or CLEC4A is upregulated. On the contrary, the control samples are clustered in the opposite area, showing a distinct separation between the RPL cohort and control samples. Figure 8 visualizes the probability of RPL together with the expression level of two genes (C10orf71 and / or CLEC4A). Here it is shown that the probability of RPL is high when C10orf71 is downregulated and / or CLEC4A is upregulated. Conclusion

[0334] 190 pairwise combinations of genes show very high accuracy and reveal the importance of the C10orf71 and / or CLEC4A.

[0335] Gene listing For the present invention, the genes referred to herein are known in the art. In here, the genes are identified by their Ensembl gene IDs.

Claims

Claims1. A method for determining whether a female subject is at risk of recurrent pregnancy loss (RPL), the method comprising- providing an endometrial biopsy obtained in a proliferative phase of the menstrual cycle from said subject;- determining in the endometrial biopsy a gene expression profile for at least 2 genes comprising C10orf71 and / or CLEC4A, wherein if the gene expression profile comprises- CLEC4A being upregulated; and / or- C10orf71 being downregulated, compared to a gene expression profile comprising C10orf71 and / or CLEC4A in a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle from control females with healthy endometria and no history of RPL, it is indicative of said subject being at risk of RPL.

2. The method according to claim 1, the method further comprises- determining in the endometrial biopsy of the female subject, the gene expression profile for at least 2 genes comprising C10orf71 and / or CLEC4A wherein if the gene expression profile comprises- CLEC4A being downregulated or having a same expression; and / or- C10orf71 being upregulated or having a same expression, compared to the gene expression profile comprising C10orf71 and / or CLEC4A in the panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle from control females with healthy endometria and no history of RPL, it is indicative of said subject not being at risk of RPL.

3. The method according to any of claims 1 or 2, wherein the RPL is at least two consecutive pregnancy losses.

4. The method according to any of the preceding claims, wherein the RPL is primary or secondary RPL.

5. The method according to any of the preceding claims, wherein the upregulation and the downregulation are determined by Iog2 fold change.

6. The method according to any of the preceding claims, wherein the upregulation is determined by a Iog2 fold change being above 0.

7. The method according to claim 6, the upregulation is a Iog2 fold change in the range of 0.1 to 5, such as 0.1 to 4, preferably 0.1 to 3.

8. The method according to any of claims 1-5, wherein the downregulation is determined by a Iog2 fold change being below 0.

9. The method according to claim 8, the upregulation is a Iog2 fold change in the range of -3 to -0.1, such as -2.5 to -0.1, preferably -2.3 to -0.2.

10. The method according to any of the preceding claims, wherein the gene expression profile comprises genes selected from the group consisting of TMEM176A, SLC22A16, ETV1, SPAG9, MGST1, UFL1, MTMR11, CYP24A1, KCNG1, SI.AMF7, CDH3, CA11, TBX21, NFKB2, PTPRC, PCDHB4, SULT2B1, RGS1, ZFHX4, IL12RB1, SLC25A1, GRAP2, ZFYVE21, JAG1, CD37, CARD8, JAK3, COMP, TMEM176B, B3GAT1, CLEC4A, IK, CYTIP, TACR1, NRP2, RARRES1, TNFSF10, PAEP, ARHGAP9, C4BPA, TSPAN8, STEAP4, RAC2, GMFG, JCHAIN, PEBP4, PPHLN1, DOCK2, SDS, PKIB, DOCK10, IL33, MMP7, FXYD2, CILP, CXCL9, PARVG, SLC46A3, NOVAI, GREB1L, PNMT, TIMM29, ABCA12, SLC26A1, MARCHF1, SSBP2, FAM217A, FGD2, DOK2, CACNA1B, SESN3, CD226, NMRAL1, GBP5, SORCS3, NCF1, MAML1, GFI1, VCAM1, SLC15A2, APEH, AMN, PLD4, PRKCB, CD3D, SCN11A, RHOH, GFRA2, TNIP2, INPP5D, CD52, GPR183, SEMA3E, KRT86, FEZ2, KRT13, CXCR6, THEMIS, NDNF, TVP23C, GAPT, KCNA3, C10orf71, CCDC184, SHISA3, ZNF852, HLA-DQB1, PLD5, CUEDC1, NQO1, TLCD5, LCN12, SHC4, EVI2B, C16orf54, ZNF566, GNG2, ARHGAP30, LITAF, TDRD7, HLA-DRB1, AKR1C3, LONP1, HLA-DQA1, SPN, AIF1, ATP10A, GPX3, IGLV1-40, TRBC1, TRBC2, IGHM, PTPRCAP, ADAT3, SIAH3, ZNF433-AS1, C4B, CASC6, GOLGA2P7, LTB, HLA-DPA1, NFAM 1, CKMT1B, APOBEC3G, LRRD1, CFB, CDK11B, FZD10-AS1, SLC2A9-AS1, GPR162, GVINP1, TIFAB, MIR142HG, ADGRL1-AS1, CCL14, TRAC, and TGFB2-11. The method according to any of the preceding claims, wherein the gene expression profile comprises genes selected from the group consisting of C10orf71, CLEC4A, SHISA3, MMP7, CILP, LCN 12, SLC2A9-AS1, KRT86, PEBP4, and AIF1, preferably selected from the group consisting of C10orf71, CLEC4A, SHISA3, MMP7, and CILP, more preferably selected from the group consisting of C10orf71, CLEC4A, and SHISA3, most preferably selected from the group consisting of C10orf71 and CLEC4A.

12. The method according to any of the preceding claims, wherein the gene expression profile comprises genes selected from the group consisting of C10orf71, CLEC4A, SHISA3, MMP7, CILP, LCN 12, SLC2A9-AS1, KRT86, PEBP4, AIF1, TSPAN8, TGFB2-OT1, ETV1, NCF1, PLD5, PKIB, AMN, SEMA3E, SHC4, and / or RHOH, preferably selected from the group consisting of C10orf71, CLEC4A, SHISA3, MMP7, CILP, LCN12, SLC2A9-AS1, KRT86, PEBP4, AIF1, TSPAN8, TGFB2- OT1, ETV1, NCF1, and / or PLD5.

13. The method according to any of the preceding claims, wherein the at least 2 genes are one or more of the following combination selected from the group consisting of C10orf71 and CLEC4A, C10orf71 and SHISA3, C10orf71 and MMP7, C10orf71 and CILP, C10orf71 and LCN12, C10orf71 and SLC2A9-AS1, C10orf71 and KRT86, C10orf71 and PEBP4, C10orf71 and AIF1, C10orf71 and TSPAN8, C10orf71 and TGFB2-OT1, C10orf71 and ETV1, C10orf71 and NCF1, C10orf71 and PLD5, C10orf71 and PKIB, C10orf71 and AMN, C10orf71 and SEMA3E, C10orf71 and SHC4, C10orf71 and RHOH, CLEC4A and SHISA3, CLEC4A and MMP7, CLEC4A and CILP, CLEC4A and LCN12, CLEC4A and SLC2A9-AS1, CLEC4A and KRT86, CLEC4A and PEBP4, CLEC4A and AIF1, CLEC4A and TSPAN8, CLEC4A and TGFB2- OT1, CLEC4A and ETV1, CLEC4A and NCF1, CLEC4A and PLD5, CLEC4A and PKIB, CLEC4A and AMN, CLEC4A and SEMA3E, CLEC4A and SHC4, CLEC4A and RHOH, SHISA3 and SLC2A9-AS1, SHISA3 and PEBP4, CILP and LCN12, CILP and PEBP4, CILP and AMN, CILP and SEMA3E, CILP and RHOH, LCN12 and SLC2A9-AS1, LCN12 and PEBP4, LCN12 and AMN, LCN12 and RHOH, SLC2A9-AS1 and RHOH, PEBP4 and AIF1, PEBP4 and TSPAN8, PEBP4 and ETV1, PEBP4 and PKIB, PEBP4 and SEMA3E, PEBP4 and SHC4, PEBP4 and RHOH, and AIF1 and RHOH.

14. The method according to any of the preceding claims, wherein the proliferative phase is between 3-10 days after the first day of menstruation in a natural menstrual cycle.

15. The method according to any of the preceding claims, wherein the proliferative phase is validated by histology dating.

16. The method according to any of the preceding claims, wherein the control endometrial biopsy is obtained from a healthy female or a female undergoing in vitro fertilization treatment.

17. The method according to any of the preceding claims, wherein the control endometrial biopsy is obtained from a female having a healthy endometrium.

18. The method according to any of the preceding claims, wherein the gene expression profile is determined at cDNA level, RNA level or protein level, preferably RNA level, such as mRNA level.

19. The method according to any of the preceding claims, wherein the gene expression profile is determined before or after diagnosis of RPL, preferably before.

20. The method according to any of the preceding claims, wherein the gene expression profile is determined by RNA sequencing, such as RNA next-generation sequencing, quantitative polymerase chain reaction (qPCR), or microarray technology.

21. The method according to claim 20, wherein the microarray technology is selected from the group consisting of:DNA microarray, such as DNA chip, biochip, spotted arrays on glass, in-situ synthesized array, self-assembled array, cDNA microarray, serial analysis of gene expression (SAGE);RNA microarray such as serial analysis of gene expression (SAGE); and / orprotein microarray, such as antibody array, enzyme-linked immunosorbent assay (ELISA).

22. The method according to any of the preceding claims, wherein the comparison of gene expression profile of the endometrial biopsy obtained in a proliferative phase of the menstrual cycle of the female subject and a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle is performed using one or more of differential expression analysis.

23. Use of an endometrial biopsy obtained in a proliferative phase of the menstrual cycle for determining the risk of recurrent pregnancy loss (RPL) in a female subject.

24. Use of a gene expression profile comprising C10orf71 and / or CLEC4A for determining the risk of recurrent pregnancy loss (RPL) in a female subject.

25. The use according to any of claims 23 or 24, wherein the RPL is at least two consecutive pregnancy losses.

26. The use according to any of claims 23-25, wherein the RPL is primary or secondary RPL.

27. The use according to any of claims 24-26, wherein the gene expression profile is determined from an endometrial biopsy obtained in a proliferative phase of the menstrual cycle from a female subject.

28. The use according to any of claims 23-27, wherein the risk of RPL is determined by comparing- a gene expression profile comprising C10orf71 and / or CLEC4A from an endometrial biopsy obtained in a proliferative phase of the menstrual cycle of the female subject, and- a gene expression profile comprising C10orf71 and / or CLEC4A from a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle from control females with healthy endometria and no history of RPL.

29. The use according to claim 28, wherein the control endometrial biopsy is obtained from a healthy female or a female undergoing in vitro fertilization treatment.

30. The use according to any of claims 28 or 29, wherein the control endometrial biopsy is obtained from a female having a healthy endometrium.

31. The use according to any of claims 23 or 27-30, wherein the proliferative phase is 3-10 days after the first day of menstruation in a natural cycle.

32. The use according to any of claims 23 or 27-31, wherein the proliferative phase is validated by histology dating.

33. The use according to any of claims 24-32, wherein the gene expression profile is determined at DNA level, such as cDNA level, RNA level or protein level, preferably RNA level, such as mRNA level.

34. The use according to any of claims 24-33, wherein the gene expression profile is determined before or after diagnosis of RPL, preferably before.

35. The use according to any of claims 24-34, wherein the gene expression profile is determined by RNA sequencing, such as RNA next-generation sequencing, quantitative polymerase chain reaction (qPCR), or microarray technology.

36. The use according to claim 35, wherein the microarray technology is selected from the group consisting of:DNA microarray, such as DNA chip, biochip, spotted arrays on glass, in-situ synthesized array, self-assembled array, cDNA microarray, serial analysis of gene expression (SAGE);RNA microarray such as serial analysis of gene expression (SAGE); and / or protein microarray, such as antibody array, enzyme-linked immunosorbent assay (ELISA).

37. A processor system programmed to operate according to a machine learning (ML) algorithm for determining the risk of recurrent pregnancy loss (RPL) based on gene expression of at least 2 genes from an endometrial biopsy obtained in a proliferative phase of the menstrual cycle of a female subject compared to gene expression of the same at least 2 genes from a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle, the machine learning (ML) algorithm being trained, and / or being trainable, on data obtained by a method according to any of claims 1-22.

38. Use of a machine learning (ML) algorithm trained on data obtained by a method according to any of claims 1-22 to determine the risk of recurrent pregnancy loss (RPL) based on gene expression of at least 2 genes from an endometrial biopsy obtained in a proliferative phase of the menstrual cycle of a female subject compared to gene expression of the same at least 2 genes from a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle.

39. A system for executing an algorithm, such as machine learning (ML) algorithm, for determining the risk of recurrent pregnancy loss (RPL) based on gene expression of at least 2 genes from an endometrial biopsy obtained in a proliferative phase of the menstrual cycle of a female subject compared to gene expression of the same at least 2 genes from a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle, the system being trained, and / or being trainable, on data provided according to any of claims 1-22.

40. A method for training a machine learning (ML) system for determining the risk of recurrent pregnancy loss (RPL) based on gene expression of at least 2 genes from an endometrial biopsy obtained in a proliferative phase of the menstrual cycle of a female subject compared to gene expression of the same at least 2 genes from a panel of control endometrial biopsies obtained in a proliferative phase of the menstrual cycle, the system being trained, and / or being trainable, on data provided according to any of claims 1-22.

41. The method for training a machine learning (ML) system according to claim 40 the method comprises the steps of:-receiving training data comprising a first set of information (1SI), such as a first database, and a second set of information (2SI), such as a second database,-training the system for estimating score variants using said training data, and- validating the system using correlated specific score variants to the gene expression information.