Prediction of the foetal fraction of cffdna

The AST/ALT ratio method, combined with machine learning, addresses the challenge of unreliable cffDNA testing by estimating FF, enhancing test reliability and reducing failures in NIPT.

WO2025181337A1PCT designated stage Publication Date: 2025-09-04HVIDOVRE HOSPITAL +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/055524
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-01
Filing Date
2025-02-28
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Current non-invasive prenatal testing (NIPT) methods for foetal DNA (cffDNA) face high costs and unreliable results due to variable foetal fraction (FF), particularly in female pregnancies or pregnancy loss cases, leading to inconclusive test outcomes and increased failure rates.

Method used

A method utilizing the AST/ALT ratio in a sample, correlated with a reference level, to estimate the amount of cffDNA, combined with machine learning models to improve FF determination, enabling reliable cffDNA testing across various pregnancy scenarios.

Benefits of technology

Enhances the reliability and efficiency of cffDNA testing by providing a personalized strategy for earlier or delayed testing based on FF estimation, reducing test failures and costs, and improving diagnostic accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025055524_04092025_PF_FP_ABST
    Figure EP2025055524_04092025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method estimating the amount of circulating cell free foetal DNA (cffDNA) in an individual and apply the information in non-invasive prenatal testing (NIPT) and pregnancy loss investigation, and in particular methods estimating the amount of circulating cell free foetal DNA (cffDNA) in an individual by determining the aspartate transaminase (AST) and alanine aminotransferase (ALT) ratio (AST / ALT ratio) in a sample, and comparing the AST / ALT ratio with a cffDNA reference level, and thereby estimating the amount of cffDNA.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PREDICTION OF THE FOETAL FRACTION OF cffDNA

[0002] FIELD

[0003] The inventive concept of the present disclosure relates to a method estimating the amount of circulating cell free foetal DNA (cffDNA) in an individual and apply the information in non- invasive prenatal testing (NIPT) and pregnancy loss investigation. The present disclosure especially provides methods estimating the amount of circulating cell free foetal DNA (cffDNA) in an individual by determining the aspartate transaminase (AST) and alanine aminotransferase (ALT) ratio (AST / ALT ratio) in a sample and comparing the AST / ALT ratio with a cffDNA reference level, and thereby estimating the amount of cffDNA.

[0004] BACKGROUND

[0005] Placental trophoblasts, occasionally shedding into the maternal bloodstream via the maternal- foetal interface, release cell-free foetal DNA (cffDNA) through apoptosis. This discovery has enabled non-invasive prenatal testing (NIPT) of foetal aneuploidies from maternal blood samples as an alternative to chorionic villus sampling and amniocenteses.

[0006] In 2018, 10 million cffDNA tests were performed, and the number of annual tests is growing rapidly. In the Copenhagen Pregnancy Loss cohort (COPL), we have recently shown that the cffDNA-based test can be used for foetal genetic assessment in pregnancy loss, as low as gestational week five. However, the cffDNA-based test has a high cost; therefore, many countries do not offer it routinely. In Denmark, for instance, the cffDNA-based test is only offered to women with an increased risk of trisomy 13, 18, or 18, as determined at the first trimester risk assessment (nuchal translucency size and double test). The cffDNA test is only offered after gestational week 10 to ensure the presence of sufficient cffDNA to make a reliable diagnosis.

[0007] The major determinant of whether a cffDNA test is reliable is the fraction of foetal DNA (FF), not counting technical failures in e.g., library preparation. If the FF is too low, the cffDNA test will yield an inconclusive result (no call). The foetal fraction (FF) is easily calculated in pregnancies with a male foetus, as the number of reads from the Y chromosome (Yfetai) can be converted into the foetal fraction, considering the background number of reads from males ( maie) and females (Yfemaie). In cases of female pregnancies or pregnancies lacking a sex chromosome (e.g., Monosomy X), this is not possible, and other methods have been developed. CffDNA differs from maternal cfDNA in that it is generally shorter, which has been leveraged by many algorithms. CffDNA is rapidly cleared from the maternal plasma following delivery, with a half-life of one hour, yet the mechanism of clearance is unknown. Furthermore, there is also increasing evidence of a relationship between cffDNA and placenta-mediated life-threatening pregnancy complications, such as preeclampsia and placenta accreta.

[0008] A few biomarkers have been investigated and associated with the foetal fraction, namely £- human-chorionic gonadotrophic (£-hCG) and pregnancy-associated plasma protein-A (PAPP- A), which are also used as part of the first-trimester risk assessment for foetal trisomy 21 , 18, and 13 as the double test.

[0009] One study investigated the use of a multivariable prediction model to identify cases of ongoing pregnancies where the foetal fraction was <4% between the 12thand 24thgestational weeks in ongoing pregnancies (Hu, L. et al.) A multivariate modelling method for the prediction of low foetal fraction before non-invasive prenatal testing. (Sci. Prog. 104, 00368504211052359 (2021 )), which showed maternal weight, red blood cell count, haemoglobin, free T3, free T4, PAPP-A, alpha-fetoprotein, conjugated estriol, and p-hCG were chosen as the input variables. The model had an AUC of 0.71 .

[0010] In commercially available cffDNA tests, the no-call rate for ongoing pregnancies is between 1- 5% if the test is performed after gestational week 10. The no-call rate is between 5-50% in cases of pregnancy loss. Methods for the enrichment and precise determination of FF have gained traction in both academic and industrial settings. NIPT technologies based on microarrays and targeted sequencing employ a minimum foetal fraction cut-off. If the foetal fraction of a sample is below this threshold, no-call (a test failure) is reported. Strict foetal fraction cut-offs lead to higher test failure rates, even upon redraw. Additionally, labs are often not reimbursed when a test does not return a result even though they have incurred the costs associated with running the test.

[0011] Expanding cell-free DNA testing to more women and a broader range of conditions has transformed the prenatal screening environment, raising questions about who should be offered screening, for what, by whom, and how.

[0012] Thus, a more comprehensive understanding of the factors influencing the foetal fraction in both pregnancy loss and ongoing pregnancies can enhance the efficacy of the cffDNA test by determining whether sufficient cffDNA is sufficiently present in a specific pregnancy to undergo testing. SUMMARY

[0013] Here, we present the first step towards an in-depth understanding of the paternal and maternal mechanisms affecting FF and demonstrate for example an algorithm and a method that utilizes both maternal and paternal biomarkers drawn at the same time as the material for the cffDNA-based test. We envision that the algorithm can be used in a two-step process to determine if there is enough foetal DNA to obtain a reliable result from the cffDNA-based test in both pregnancy loss and ongoing pregnancy.

[0014] In one aspect, the present disclosure relates to a method for estimating an amount of circulating cell free foetal DNA (cffDNA) in an individual, comprising a) providing a sample from said individual b) determining an aspartate transaminase (AST) and alanine aminotransferase (ALT) ratio (AST / ALT ratio) in said sample c) correlating said AST / ALT ratio with a reference level, and thereby estimating the amount of cffDNA.

[0015] It is an object of the present disclosure to provide simpler and / or improved estimation or detection of cffDNA.

[0016] Also disclosed is a computer-implemented method for determining an amount of circulating cell free foetal DNA (cffDNA) in an individual, the method comprising obtaining sample data of a sample from the individual; optionally determining an AST content of aspartate transaminase (AST) in the sample based on the sample data; optionally determining an ALT content of alanine aminotransferase (ALT) in the sample based on the sample data; determining the amount of circulating cell free foetal DNA based on the AST content and the ALT content; and outputting the amount of circulating cell free foetal DNA.

[0017] Further, a computer-implemented method for training a machine learning model to process as inputs an AST content of aspartate transaminase (AST) of a sample and an ALT content of alanine aminotransferase (ALT) of the sample; and provide as output one or more parameters indicative of amount of circulating cell free foetal DNA of the sample is provided, the method comprising obtaining, using at least one processor, an AST content of aspartate transaminase (AST) and an ALT content of alanine aminotransferase (ALT) of a sample; performing, using the at least one processor, a training comprising: generating, using the at least one processor and a machine-learning model, cffDNA data based on the AST content and the ALT content; obtaining, using the at least one processor, training data, e.g. including cffDNA content; determining, using the at least one processor and one or more loss functions, one or more loss parameters based on the AST content, ALT content, and the training data; and training, using the at least one processor, the machine learning model based on the one or more loss parameters.

[0018] DETAILED DESCRIPTION

[0019] The present disclosure present outcome prediction models for foetal fraction using biochemistry and clinical variables. The model demonstrated strong performance metrics, that aligned well with the results from two external validation sites (Herlev University Hospital and Northern Zealand Hospital), underlining the model's generalizability. Furthermore, our association studies and machine learning models implicate not only the placenta-produced hormone £-hCG and gestational age, but also maternal levels of enzymes and lipids, thyroid function, and liver.

[0020] Taken together, this demonstrates that a blood sample can be used to estimate the levels of cffDNA and can be a powerful method of screening, which women should be offered cffDNA based testing.

[0021] The external validation performed on an independent dataset further substantiates the robustness and generalizability of our model. Despite these strengths, the cohort includes pregnancy losses, and the models may need to be further validated in ongoing pregnancies, but the present data teach that the concept is applicable in ongoing pregnancies.

[0022] The foetal fraction in ongoing pregnancies is influenced by gestational age and body mass index (BMI). A previous multivariable study created a predictive model for a foetal fraction greater than 4% and discovered that its performance was better than random. The model presented herein, including a more diverse set of biomarkers, improves upon this significantly.

[0023] Clinical guidance currently relies on gestational age and BMI for ongoing pregnancies, and a foetus still in situ for cases of pregnancy loss. The cffDNA test is currently only suggested in commercial context beyond the 10thgestational week for ongoing pregnancies, but this work lays the foundation for a personalized strategy, enabling an earlier test or postponing the test in case of a high risk of low foetal fraction.

[0024] Similarly, for pregnancy loss, the test can be used to consult with couples on the chance of success of the cffDNA based test. Furthermore, cffDNA has been associated with several complications in pregnancy. The present disclosure establishes an association with maternal levels of lipids, liver function, and thyroid function. The external validation confirms the model’s utility across different settings, making it a versatile tool for clinical practice.

[0025] Circulating cell free foetal DNA (cffDNA)

[0026] Cell-free DNA are short fragments of DNA released into the bloodstream through a natural process of cell death. During pregnancy, maternal blood contains cell-free DNA (cfDNA), both from her own tissue, and from the foetus via the placenta (cffDNA).

[0027] Approximately 2-20% of total cfDNA in maternal blood is foetal, of placental origin. cfDNA derived from the placenta can be detected as early as 4+ weeks gestation and is undetectable within hours postpartum. A non-invasive prenatal test (NIPT) analyses cfDNA from a maternal blood sample to screen for common chromosomal conditions in the foetus.

[0028] The percentage of total cell-free DNA (cfDNA) in a sample derived from the foetus or placenta (cffDNA) is called the foetal fraction — which can affect the ability of NIPT to detect foetal aneuploidy. When considering various cfDNA and / or cffDNA NIPT technologies, it’s important to understand how foetal fraction is used and how it can affect test results.

[0029] The skilled addressee knows that foetal DNA fraction from the plasma of pregnant women can be determined using sequence read counts. Thus, the cell-free foetal DNA was determined from the read counts of the total cfDNA, using a previously established algorithm, SeqFF. Read counts across the genome are binned and used as input features to a penalized regression model, which then predicts the foetal fraction.

[0030] Individual

[0031] In the present context the term "individual" relates to both the mother, father, and the unborn progeny. The individual is preferably a person in risk of carrying a foetus with an abnormal cell function or a genetic deviation of interest. Though the present examples describe the measurements in a maternal sample, the present disclosure can be adapted to measurements direct on the foetus and / or the farther.

[0032] Aspartate transaminase (AST)

[0033] Aspartate transaminase (AST) or aspartate aminotransferase, also known as AspAT / ASAT / AAT or (serum) glutamic oxaloacetic transaminase (GOT, SGOT), is a pyridoxal phosphate (PLP)-dependent transaminase enzyme (EC 2.6.1.1 ).

[0034] AST catalyses the reversible transfer of an a-amino group between aspartate and glutamate and, as such, is an important enzyme in amino acid metabolism. AST is found in the liver, heart, skeletal muscle, kidneys, brain, red blood cells and gall bladder. Serum AST level, serum ALT (alanine transaminase) level, and their ratio (AST / ALT ratio) are commonly measured clinically as biomarkers for liver health. The tests are part of blood panels.

[0035] Isoforms

[0036] Two isoenzymes are present in humans:

[0037] GOT1 / cAST, the cytosolic isoenzyme derives mainly from red blood cells and heart. GOT2 / mAST, the mitochondrial isoenzyme is present predominantly in liver.

[0038] These isoenzymes have evolved from a common ancestral AST via gene duplication, and they share a sequence homology of approximately 45%. In the Examples below, both isoforms are measured via the Siemens Atellica technology, but any means for identifying the ratio can be used including means that can differentiate between the two isoforms.

[0039] AST is similar to alanine transaminase (ALT) in that both enzymes are associated with liver parenchymal cells. The difference is that ALT is found predominantly in the liver, with clinically negligible quantities found in the kidneys, heart, and skeletal muscle, while AST is found in the liver, heart (cardiac muscle), skeletal muscle, kidneys, brain, and red blood cells.

[0040] Alanine aminotransferase (ALT)

[0041] Alanine transaminase (ALT) is a transaminase enzyme (EC 2.6.1.2). It is also called alanine aminotransferase (ALT or ALAT) and was formerly called serum glutamate-pyruvate transaminase or serum glutamic-pyruvic transaminase (SGPT). ALT is found in plasma and in various body tissues but is most common in the liver. It catalyses the two parts of the alanine cycle. Serum ALT level, serum AST (aspartate transaminase) level, and their ratio (AST / ALT ratio) are routinely measured clinically as biomarkers for liver health.

[0042] When used in diagnostics, AST and ALT are almost always measured in international units / litre (I U / L or U / L) or pkat. While sources vary on specific reference range values for patients, 0-40 I U / L is the standard reference range for experimental studies. In a clinical setting, the reference ranges for pregnant women are 16-40 U / L in gestational week 13-40 for AST and 8-36 U / L in gestational week 13 to 35 for ALT. Both AST and ALT may be elevated during pregnancy.

[0043] Fluctuation of ALT levels are normal over the course of the day, and they can also increase in response to strenuous physical exercise. When elevated ALT levels are found in the blood, the possible underlying causes can be further narrowed down by measuring other enzymes.

[0044] In 2000, the American Association for Clinical Chemistry determined that the appropriate terminology for AST and ALT are aspartate aminotransferase and alanine aminotransferase. The term transaminase is outdated and no longer used.

[0045] AST / ALT ratio

[0046] The AST / ALT ratio or De Ritis ratio is the ratio between the concentrations of the two enzymes aspartate transaminase and alanine aminotransferase in the blood of a human or an animal. The AST / ALT ratio is measured by conventional analytical methods, such as immunological methods known to the art. It is traditionally used as one of several liver function tests, and typically measured with a blood test. AST and ALT as determined for this disclosure are analysed at a Siemens Atellica CH 930 machine by enzyme spectrophotometry but can be determined by any means known to the skilled addressee.

[0047] In one or more exemplary embodiments, the AST / ALT ratio is determined by measuring RNA (cell-free or intracellular), DNA, protein, and / or metabolite levels of the aspartate transaminase and alanine aminotransferase enzymes in a sample. In one or more exemplary embodiments, the AST / ALT ratio is determined by measuring cell- free RNA levels of the aspartate transaminase and alanine aminotransferase enzymes in a sample.

[0048] In one or more exemplary embodiments, the AST / ALT ratio is determined by measuring intracellular RNA levels of the aspartate transaminase and alanine aminotransferase enzymes in a sample.

[0049] In one or more exemplary embodiments, the AST / ALT ratio is determined by measuring DNA levels of the aspartate transaminase and alanine aminotransferase enzymes in a sample.

[0050] In one or more exemplary embodiments, the AST / ALT ratio is determined by measuring protein levels of the aspartate transaminase and alanine aminotransferase enzymes in a sample.

[0051] In one or more exemplary embodiments, the AST / ALT ratio is determined by measuring metabolite levels of the aspartate transaminase and alanine aminotransferase enzymes in a sample.

[0052] In order to determine the clinical severity variations of the AST / ALT ratio in the present context, means for evaluating the detectable signal of the AST / ALT ratio measured involves a reference or reference means. The reference also makes it possible to take into account assay, kit and method variations, handling variations and other variations not related directly or indirectly to the AST / ALT ratio.

[0053] In the context of the present invention, the term "reference" relates to a standard in relation to quantity, quality or type, against which other values or characteristics can be compared, such as e.g. a standard curve.

[0054] The reference data presented in the Examples below reflects the maternal blood AST / ALT ratio from confirmed intrauterine pregnancy loss before 22 weeks of gestation but could equally be maternal blood AST / ALT ratio from all pregnant women carrying viable foetuses.

[0055] As will be generally understood by those of skill in the art, methods for screening for foetal abnormalities are processes of decision making by comparison. For any decision-making process, reference values based on individuals having the disease or condition of interest and / or individuals not having the disease or condition of interest are needed. In the present disclosure, the reference values are the maternal blood level of the measured marker or markers, for example, the AST / ALT ratio, in both pregnant women carrying for example genetically abnormal foetuses, pregnant women carrying viable foetuses, and women having a pregnancy loss. A set of reference data is established by collecting the reference values for a number of samples. As will be obvious to those of skill in the art, the set of reference data will improve by including increasing numbers of reference values.

[0056] In one or more exemplary embodiments, the reference means is an internal reference means and / or an external reference means.

[0057] In the present context the term "internal reference means" relates to a reference which is not handled by the user directly for each determination, but which is incorporated into a device for the determination of the biomarker of interest, like for example the AST / ALT ratio, whereby only the ’final result’ or the ’final measurement’ is presented. The terms the "final result" or the "final measurement" relates to the result presented to the user when the reference value has been taken into account. In one or more exemplary embodiments, the internal reference means is provided in connection to a device used for the determination of the biomarker in question.

[0058] In the present context, the term "external reference means" relates to a reference which is handled directly by the user in order to determine the biomarker, before obtaining the ’final result’ or the ’final measurement’. In one or more exemplary embodiments, the external reference means are selected from the group consisting of a table, a diagram and similar reference means where the user can compare the measured signal to selected reference means. The external reference means relates to a reference used as a calibration, value reference, information object, etc. for the AST / ALT ratio and which has been excluded from the device used.

[0059] In one or more exemplary embodiments, the reference level / predetermined value is indicative of a normal physiological condition of said individual.

[0060] In one or more exemplary embodiments, the reference level / predetermined value is indicative of a condition is a foetal abnormality.

[0061] Although any of the known analytical methods for measuring the AST / ALT ratio will function in the present invention, as obvious to one skilled in the art, the analytical method used for the AST / ALT ratio must be the same method used to generate the reference data for the AST / ALT ratio. If a new analytical method is used for the AST / ALT ratio, a new set of reference data, based on data developed with the method, must be generated.

[0062] The reference level

[0063] As shown in the Examples below, a correlation between AST / ALT ratio and cffDNA has been established. This correlation can for example be achieved through two complementary approaches: statistical modelling and machine learning modelling. The statistical model employs a Bayesian hurdle model, consisting of a binomial and log-normal component. Furthermore, in a machine learning model, specifically a gradient boosted model, we found through SHAP analysis that AST / ALT ratio was the most important feature that increased the predictive value.

[0064] To estimate the amount of cffDNA means for correlating the AST / ALT ratio to the amount of cffDNA may involve a reference or reference means. In the present context, the term "reference" relates to a standard in relation to quantity, quality or type, against which other values or characteristics can be compared, such as e.g. a standard curve.

[0065] The reference can be made from a control group.

[0066] In one or more exemplary embodiments, the reference means is an internal reference means and / or an external reference means.

[0067] In the present context the term "internal reference means" relates to a reference which is not handled by the user directly for each determination, but which is incorporated into for example a device for the determination of the AST / ALT ratio, whereby only the “final result' is presented. The term the "final result" relates to the result presented to the user when the reference value has been taken into account.

[0068] In the present context, the term "external reference means" relates to a reference which is handled directly by the user in order to determine the AST / ALT ratio, before obtaining the “final result”.

[0069] In one or more exemplary embodiments, a reference AST / ALT ratio is obtained from control groups comprising for example a woman pregnant with a normal foetus at the same gestational age. To determine whether the individual for example is at increased risk of carrying a foetus with e.g. Down syndrome, a cut-off must be established. This cut-off may be established by the laboratory, the physician or on a case-by-case basis by each individual.

[0070] The cut-off level can be based on several criteria including the number of individuals who would go on for further invasive diagnostic testing. The cut-off level could be established using a number of methods, including percentiles, mean plus or minus standard deviation(s); multiples of median value; patient specific risk or other methods known to those who are skilled in the art.

[0071] In one or more examples of the present disclosure, a computer-implemented method for determining an amount of circulating cell free foetal DNA (cffDNA) in an individual is disclosed, the method comprising obtaining sample data of a sample from the individual; determining an AST content of aspartate transaminase (AST) in the sample based on the sample data; determining an ALT content of alanine aminotransferase (ALT) in the sample based on the sample data; determining the amount of circulating cell free foetal DNA based on the AST content and the ALT content; and outputting the amount of circulating cell free foetal DNA.

[0072] In one or more examples, the AST content and / or the ALT content may be included in the sample data and therefore determining AST content and / or ALT content may be omitted.

[0073] In one or more examples, determining the amount of circulating cell free foetal DNA based on the AST content and the ALT content comprises applying a machine-learning model, such as a machine learning gradient boosted model, on the AST content and the ALT content.

[0074] In one or more examples, determining the amount of circulating cell free foetal DNA based on the AST content and the ALT content comprises determining a ratio or other relationship, such as a difference, between the AST content and the ALT content.

[0075] In one or more examples, determining the ratio between the AST content and the ALT content comprises measuring one or more of mRNA, siRNA, piRNA, ncRNA, IncRNA, miRNA RNA, DNA, methylation relating to epigenomics, protein, antibodies, metabolite, ions, lipid level of aspartate transaminase enzyme, and lipid level of alanine aminotransferase enzyme. The method optionally comprises determining the ratio between the AST content and the ALT content based on the one or more of mRNA, siRNA, piRNA, ncRNA, IncRNA, miRNA RNA, DNA, methylation relating to epigenomics, protein, antibodies, metabolite, ions, lipid level of aspartate transaminase enzyme, and lipid level of alanine aminotransferase enzyme.

[0076] In one or more examples, determining the ratio between the AST content and the ALT content, also denoted AST / ALT ratio, comprises measuring DNA levels of the aspartate transaminase and alanine aminotransferase enzymes and determining the AST / ALT ratio optionally based on the measuring DNA levels of the aspartate transaminase and alanine aminotransferase enzymes .

[0077] In one or more examples, determining the amount of circulating cell free foetal DNA based on the AST content and the ALT content comprises determining the amount of circulating cell free foetal DNA based on the AST content, the ALT content, and one or more of gestational age, beta-Human Chorionic Gonadotropin (£-hCG) level, Pregnancy-associated plasma protein A (PAPP-A) level, vascular endothelial growth factor (VEGF) level, soluble fms-like tyrosine kinase-1 (sFlt-1 ) level, Alpha-Fetoprotein (AFP) level, Disintegrin and metalloproteinase domain-containing protein 12 (ADAM 12) level, ISM2 level, TFPI2 level, ERVV-1 level, LYPD3 level, EBIB3JL27 level, CSH1 level, GDF15 level, ANGPT2 level, FBN2 level, PRG2 level, INSL4 level, LAIR2 level, TSHB level, C1 QTNF6 level, LHB level, SIGLEC6 level, MMP12 level, Placental Growth Factor (PLGF1 ) level, Alanine Transaminase (ALT) level, Aspartate transaminase (AST) level, High-density lipoprotein cholesterol (HDL) level, Low-density lipoprotein cholesterol (LDL) level, Apolipoprotein B level, Uric Acid level, Transferrin level, Bilirubin level, Creatin kinase level, and Lipoprotein A level.

[0078] Estimating the amount of cffDNA

[0079] CffDNA is determined or estimated in the Examples using a machine learning model, such as a machine learning gradient boosted model trained on historical data and validated using external cohorts. The machine learning model takes as input the AST / ALT ratio, and outputs an amount, such as a fraction between zero and one, representing the determined / estimated amount of cffDNA. The AST / ALT ratio may be combined with other markers, such as but not limited to p-hCG and gestational age, to further improve the cffDNA estimation.

[0080] Sample

[0081] In the Examples below, blood samples were collected at various times like for example either prior to the removal of pregnancy tissue or within 24 hours post-evacuation. The plasma component was separated through a dual-stage centrifugation process and then preserved at low temperatures for subsequent analysis. cfDNA was extracted from the plasma and added to library construction via PCR and sequenced and bioinformatics processing was performed as described by Hartwig et al (Schlaikjaer Hartwig, T. et al. Cell-free foetal DNA for genetic evaluation in Copenhagen Pregnancy Loss Study (COPL): a prospective cohort study. Lancet Lend. Engl. 401 , 762-771 (2023).

[0082] Briefly, reads across lanes were merged into one fastq file. Reads were aligned using bowtie2. Aligned reads with a quality < 1 were removed, and the remaining reads were sorted and deduplicated using samtools. The foetal fraction was determined using SeqFF (Kim, S. K. et al. Determination of foetal DNA fraction from the plasma of pregnant women using sequence read counts. Prenat. Diagn. 35, 810-815 (2015)) by which small differences of sequencing behaviour for maternal and foetal cell-free DNA and read counts were used for estimation.

[0083] Maternal whole blood was collected in EDTA tubes or serum clot activator tubes and separated into plasma and serum. The serum samples were used for biochemical analysis at the Department of Clinical Biochemistry, at Copenhagen University Hospital Herlev, measured using standard assays, such as £-hCG (sandwich immunoassay by Siemens Atellica IM Analyzer), creatinine and cholesterol (Enzymatic reaction and absorbance by Siemens Atellica CH 930).

[0084] In the present context, the term "sample" relates to any liquid or solid sample collected from an individual to be analysed. Preferably, the sample is liquefied at the time of assaying.

[0085] In one or more exemplary embodiments, the sample is selected from the group consisting of blood, serum, plasma, urine, faeces, rectal swab, rectal microbiome, vaginal microbiome, vaginal discharge, vaginal secretion, cervical discharge, cervical swab, vaginal swab, amniotic fluid and other secreted fluids from the vagina and / or uterus.

[0086] In one or more exemplary embodiments, the sample is selected from the group consisting of blood, urine, faeces, rectal swab, rectal microbiome, vaginal microbiome, and vaginal discharge samples.

[0087] In one or more exemplary embodiments, the sample is selected from the group consisting of blood, serum, plasma, and urine. The sample taken may be dried for transport and future analysis. Thus, the method of the present disclosure includes the analysis of both liquid and dried samples.

[0088] To increase detection efficiency, the sample data and the gestational age may be compared to a set of reference data to determine whether the individual is at increased risk of for example pregnancy loss or carrying a foetus with e.g. Down syndrome.

[0089] Blood

[0090] In one or more exemplary embodiments, the sample is a blood sample.

[0091] In one or more exemplary embodiments, the blood sample is separated into plasma and serum samples.

[0092] Plasma

[0093] In one or more exemplary embodiments, the sample is a plasma sample.

[0094] Serum

[0095] In one or more exemplary embodiments, the sample is a serum sample.

[0096] Urine

[0097] In one or more exemplary embodiments, the sample is a urine sample.

[0098] Faeces

[0099] In one or more exemplary embodiments, the sample is a faeces sample.

[0100] Rectal microbiome

[0101] In one or more exemplary embodiments, the sample is a rectal microbiome sample. Rectal swab

[0102] In one or more exemplary embodiments, the sample is a rectal swab sample.

[0103] Vaginal discharge

[0104] In one or more exemplary embodiments, the sample is a vaginal discharge sample.

[0105] Vaginal microbiome

[0106] In one or more exemplary embodiments, the sample is a vaginal microbiome sample.

[0107] Vaginal secretion

[0108] In one or more exemplary embodiments, the sample is a vaginal secretion sample.

[0109] Cervical discharge

[0110] In one or more exemplary embodiments, the sample is a cervical discharge sample.

[0111] Cervical swab

[0112] In one or more exemplary embodiments, the sample is a cervical swab sample.

[0113] Vaginal swab

[0114] In one or more exemplary embodiments, the sample is a vaginal swab sample.

[0115] Amniotic fluid

[0116] In one or more exemplary embodiments, the sample is an amniotic fluid sample.

[0117] Other secreted fluids from the vagina and / or uterus

[0118] In one or more exemplary embodiments, the sample is a secreted fluid from the vagina and / or uterus. Maternal and / or a paternal sample.

[0119] In one or more exemplary embodiments, the sample is a maternal and / or a paternal sample.

[0120] Timing of the sample taking

[0121] As described herein, the methods of this disclosure provide for more discriminatory, cheaper, less invasive, and more geographically accessible means for cffDNA based prenatal screening. Since the methods described herein can be used in both biochemical pregnancies or confirmed intrauterine pregnancy, hereunder ongoing pregnancies and pregnancy loss, the timing of the sample varies.

[0122] In one or more exemplary embodiments, the sample is taken between the 5th and 22nd gestational week.

[0123] Thus, the sample can be taken at any time during the 5th, 6th, 7th, 8th, 9th10th, 11th, 12th, 13th, 14th, 15th, 16th, 17th, 18th, 19th, 20th, 21st, and / or 22ndgestational week, such as but not limited to gestational age 5 weeks + 0 days, 5 weeks + 1 days, 5 weeks + 2 days, 5 weeks + 3 days,

[0124] 5 weeks + 4 days, 5 weeks + 5 days, 5 weeks + 6 days, 6 weeks + 0 days, 6 weeks + 1 days,

[0125] 6 weeks + 2 days, 6 weeks + 3 days, 6 weeks + 4 days, 6 weeks + 5 days, 6 weeks + 6 days,

[0126] 7 weeks + 0 days, 7 weeks + 1 days, 7 weeks + 2 days, 7 weeks + 3 days, 7 weeks + 4 days,

[0127] 7 weeks + 5 days, 7 weeks + 6 days, 8 weeks + 0 days, 8 weeks + 1 days, 8 weeks + 2 days,

[0128] 8 weeks + 3 days, 8 weeks + 4 days, 8 weeks + 5 days, 8 weeks + 6 days, 9 weeks + 0 days,

[0129] 9 weeks + 1 days, 9 weeks + 2 days, 9 weeks + 3 days, 9 weeks + 4 days, 9 weeks + 5 days,

[0130] 9 weeks + 6 days, 10 weeks + 0 days, 10 weeks + 1 days, 10 weeks + 2 days, 6 weeks + 3 days, 10 weeks + 4 days, 10 weeks + 5 days, 10 weeks + 6 days, 11 weeks + 0 days, 11 weeks + 1 days, 11 weeks + 2 days, 11 weeks + 3 days, 11 weeks + 4 days, 11 weeks + 5 days, 11 weeks + 6 days, 12 weeks + 0 days, 12 weeks + 1 days, 12 weeks + 2 days, 12 weeks + 3 days, 12 weeks + 4 days, 12 weeks + 5 days, 12 weeks + 6 days, 13 weeks + 0 days, 13 weeks + 1 days, 13 weeks + 2 days, 13 weeks + 3 days, 13 weeks + 4 days, 13 weeks + 5 days, 13 weeks + 6 days, 14 weeks + 0 days, 14 weeks + 1 days, 14 weeks + 2 days, 14 weeks + 3 days, 14 weeks + 4 days, 14 weeks + 5 days, 14 weeks + 6 days, 15 weeks + 0 days, 15 weeks + 1 days, 15 weeks + 2 days, 15 weeks + 3 days, 15 weeks + 4 days, 15 weeks + 5 days, 15 weeks + 6 days, 16 weeks + 0 days, 16 weeks + 1 days, 16 weeks + 2 days, 16 weeks + 3 days, 16 weeks + 4 days, 16 weeks + 5 days, 16 weeks + 6 days, 17 weeks + 0 days, 17 weeks + 1 days, 17 weeks + 2 days, 17 weeks + 3 days, 17 weeks + 4 days, 17 weeks + 5 days, 17 weeks + 6 days, 18 weeks + 0 days, 18 weeks + 1 days, 18 weeks + 2 days, 18 weeks + 3 days, 18 weeks + 4 days, 18 weeks + 5 days, 18 weeks + 6 days, 19 weeks + 0 days, 19 weeks + 1 days, 19 weeks + 2 days, 19 weeks + 3 days, 19 weeks + 4 days, 19 weeks + 5 days, 19 weeks + 6 days, 20 weeks + 0 days, 20 weeks + 1 days, 20 weeks + 2 days, 20 weeks + 3 days, 20 weeks + 4 days, 20 weeks + 5 days, 20 weeks + 6 days, 21 weeks + 0 days, 21 weeks + 1 days, 21 weeks + 2 days, 21 weeks + 3 day, 21 weeks + 4 days, 21 weeks + 5 days, 21 weeks + 6 days, or 22 weeks + 0.

[0131] In one or more exemplary embodiments, the sample is taken in the first trimester.

[0132] In one or more exemplary embodiments, the sample is taken in the second trimester. There may be disadvantages to second trimester testing, in that for examples delays in confirming a foetal aneuploidy diagnosis result in more traumatic abortion procedures being necessitated. Also, the emotional attachment and expectations of the pregnant woman and her family for a healthy baby, grow during the pregnancy, making the abortion decision more difficult later in the gestational term. Thus, the earlier the sample can provide a diagnostic result the better.

[0133] Trying to and / or having conceived a pregnancy / liveborn child

[0134] In one or more exemplary embodiments, the sample is taken from individual(s) in a couple trying to and / or having conceived a pregnancy through natural conception or Assisted Reproductive Technology.

[0135] In one or more exemplary embodiments, the sample is taken from individual(s) in a couple trying to and / or having a liveborn child.

[0136] In one or more exemplary embodiments, the sample is taken from individual(s) in a couple having pregnancy of unknown location, mole pregnancy, biochemical pregnancies or confirmed intrauterine pregnancy, which may be singleton or multiple gestation, hereunder ongoing pregnancies and / or even pregnancy loss.

[0137] In one or more exemplary embodiments, the sample is taken with the pregnancy tissue being in situ.

[0138] In one or more exemplary embodiments is taken prior to the removal or passage of pregnancy tissue. In one or more exemplary embodiments, the sample is taken within 96 hours after removal or spontaneous passage of the pregnancy product, such as but not limited to within 90 hours, 84 hours, 78 hours, 72 hours, 66 hours, 60 hours, 54 hours, 48 hours, 42 hours, 36 hours, 30 hours, 24 hours, 18 hours, 12 hours, 11 , hours, 10 hours, 9 hours, 8 hours, 7, hours, 6 hours, 5 hours, 4 hours, 3 hours, 2 hours or 1 hour after removal or spontaneous passage of the pregnancy product.

[0139] In one or more exemplary embodiments, the sample is taken within 96 hours of a pregnancy loss, such as but not limited to within 90 hours, 84 hours, 78 hours, 72 hours, 66 hours, 60 hours, 54 hours, 48 hours, 42 hours, 36 hours, 30 hours, 24 hours, 18 hours, 12 hours, 11 , hours, 10 hours, 9 hours, 8 hours, 7, hours, 6 hours, 5 hours, 4 hours, 3 hours, 2 hours or 1 hour after of a pregnancy loss.

[0140] Further biomarkers

[0141] Measuring AST / ALT ratio directly and / or use the AST / ALT ratio estimating the amount of circulating cell free foetal DNA (cffDNA) in combination with one or more of the following biomarkers and / or other biometric markers may reduce the number of false positive and increase the discriminatory power the methods presented herein.

[0142] In one or more exemplary embodiments, methods described herein is combined with levels of further biomarkers selected from the group consisting of gestational age, beta-Human Chorionic Gonadotropin (£-hCG), Pregnancy-associated plasma protein A (PAPP-A), Disintegrin and metalloproteinase domain-containing protein 12 (ADAM12), ISM2, TFPI2, ERW-1 , LYPD3, EBIB3JL27, CSH1 , GDF15, ANGPT2, FBN2, PRG2, INSL4, LAIR2, TSHB, C1QTNF6, LHB, SIGLEC6, MMP12, Placental Growth Factor (PLGF1 ), Alanine Transaminase (ALT), Aspartattransaminase (AST), High-density lipoprotein cholesterol (HDL), Low-density lipoprotein cholesterol (LDL), Apolipoprotein B, Uric Acid, Transferrin, Bilirubin, Creatin kinase and Lipoprotein A alpha feto-protein (AFP), unconjugated oestrol (uE3), human chorionic gonadotrophin (hCG), free alpha sub-unit of hCG (free a-hCG), free beta sub-unit of hCG (free P-hCG), beta-core hCG, hyperglycosylated hCG (ITG), placental growth hormonre (PGH), inhibin, preferably dimeric inhibin-A (inhibin A), pregnancy-associated plasma protein A (PAPP-A), Complexes of PAPP-A with proMBP (proform of major basic protein), ProMBP, ProMBP complexes with angiotensinogen and / or complement factors and split products Schwangerschaftsprotein 1 (SP1 ), Cancer antigen 125(CA125), Prostate specific antigen (PSA), Leukocyte enzymes, foetal DNA, foetal RNA, foetal cells, stem cells, oestradiol, ultrasound markers, nuchal translucency, femur length, absence of nasal bone, hyperechogen ic bowel, echogenic foci in the heart, choroids plexus cysts, hydronephrosis, foetal malformations, steroids, peptides, chemokines, interleukins (e.g. IL-6, IL-4, IL-1 ), tumour necrosis factor, transforming growth factor alpha and beta, acute phase reactants, C-reactive protein, Fibronectin, maternal or foetal single nucleotide polymorphisms, e.g. promoter region polymorphisms in TNFbeta and mannan- binding lectin, complement components, HLA-G, and / or HLA molecules in the sample.

[0143] Thus, in one embodiment, the present disclosure relates to a method as described herein, wherein the AST / ALT ratio directly and / or use the AST / ALT ratio estimating the amount of circulating cell free foetal DNA (cffDNA) amount is combined with values from at least one marker selected from the group defined above.

[0144] In one or more exemplary embodiments, the AST / ALT ratio is combined with levels of further biomarkers selected from the group consisting of gestational age, beta-Human Chorionic Gonadotropin (£-hCG), Pregnancy-associated plasma protein A (PAPP-A), vascular endothelial growth factor (VEGF), soluble fms-like tyrosine kinase-1 (sFlt-1 ), Alpha- Fetoprotein (AFP), Disintegrin and metalloproteinase domain-containing protein 12 (ADAM12), ISM2, TFPI2, ERW-1 , LYPD3, EBIB3JL27, CSH1 , GDF15, ANGPT2, FBN2, PRG2, INSL4, LAIR2, TSHB, C1QTNF6, LHB, SIGLEC6, MMP12, Placental Growth Factor (PLGF1 ), Alanine Transaminase (ALT), Aspartattransaminase (ASAT), High-density lipoprotein cholesterol (HDL), Low-density lipoprotein cholesterol (LDL), Apolipoprotein B, Uric Acid, Transferrin, Bilirubin, Creatin kinase and Lipoprotein A. in the sample.

[0145] In one or more exemplary embodiments, the AST / ALT ratio is combined with levels of further biomarkers selected from the group consisting of gestational age, p-hCG, Bilirubin, Creatin kinase and Lipoprotein A.

[0146] Gestational age

[0147] In one or more exemplary embodiments, the AST / ALT ratio is combined with the gestational age.

[0148] 13-hCG

[0149] In one or more exemplary embodiments, the AST / ALT ratio is combined with the £-hCG level in the sample. Bilirubin

[0150] In one or more exemplary embodiments, the AST / ALT ratio is combined with the Bilirubin level in the sample.

[0151] Creatin kinase

[0152] In one or more exemplary embodiments, the AST / ALT ratio is combined with the Creatin kinase level in the sample.

[0153] Lipoprotein A

[0154] In one or more exemplary embodiments, the AST / ALT ratio is combined with the Lipoprotein A level in the sample.

[0155] Statistics

[0156] The baseline characteristics and features of the machine learning models were summarized as mean (standard deviation, SD), median (interquartile range, IQR), or frequency, where appropriate. The FF was binarized, using two clinically relevant thresholds as described earlier, namely 2.5% and 4%.

[0157] Univariate association between the FF and other variables were modelled using a Bayesian Hurdle lognormal regression model with a weakly informative prior, namely the student-t distribution centred around zero with seven degrees of freedom. The exact choice of prior distribution did not affect the results.

[0158] The Hurdle lognormal regression has the likelihood, L,

[0159] In which f is the probability density function for the log-normal distribution, and F is the cumulative density function for the log-normal distribution. Consequently, the hurdle model comprises two distinct components: the first models the probability of observing a zero (i.e. FF <- 2.5%), versus a non-zero outcome (i.e. FF > 2.5%), and the second models the distribution of X conditioned on it being positive and non-zero, (i.e. the observed value for FF given FF > 2.5%). Each component is characterized by its own set of parameters.

[0160] The model was implemented in brms and ran with default parameters (5,000 warm-up iterations, 5,000 draws from the posterior from 4 independent chains). The model has converged if, and only if, (1 ) all R-hat values were < 1.01 , (2) a tree depth of ten had not been exceeded post warm-up, (3) there were no divergences. Due to the extreme values p-hCG had to be log transformed for the model to converge. Biochemistry variables were additionally adjusted for the time between sample collection and biomarker measurement (i.e., time spent at -80degrees Celsius).

[0161] In one or more exemplary embodiments, the data reported have the association statistics as the 95% Bayesian Credible Interval (BCI). BCIs excluding zero were taken as evidence that a given biomarker was associated with the outcome.

[0162] For predictive modelling we created a model with two components, resembling the hurdle lognormal model. The model consisted of two parts, namely a classifier to predict if the foetal fraction > 2.5%, and a regression model to predict the fraction. Only non-zero positive values > 2.5% were used as input for the regression model. In the modelling framework, two distinct components were employed: a classifier and a regressor. The gradient boosting framework LightGBM was utilized for both components. The LightGBM model can capture non-linearities and interactions, if present.

[0163] We also compared the performance to a model consisting of a linear regression and logistic regression. The objective function of the classification and regression model are the log loss and mean squared error, respectively. The final prediction of the foetal fraction is given by,

[0164] FF = (0 > t) • y

[0165] In which 9 is the estimated probability (i.e. a number between 0 and 1 ) of the foetal fraction > 2.5%, and t is the threshold. The parameter t is set to be equivalent to the prevalence of examples in which FF > 2.5%. y is the estimate from the regression model. Consequently, if 9 is below the threshold, t, the foetal fraction is zero. If 9 is above the threshold, t, the foetal fraction will be equal to the estimated value from the regression model (y). To identify the best set of hyperparameters, we evaluated each hyperparameter setting in a five-fold cross-validation process. We defined the loss of a hyper parameter configuration as the weighted sum of the mean squared error (MSE) and log loss (LL),

[0166] L = MSE - w + LL

[0167] The weighting factor, w, was empirically determined by fitting a linear regression model and logistic regression model with no covariates. We set the weight, w, to be 100. There may be another optimum, as will be obvious to those skilled in the art.

[0168] The optimal model was selected based on the cross-validated loss. Hyperparameters were optimized using Optuna software, run for 500 trials, with default settings. The range and distributions for hyperparameters is detailed in Table 6. The exact choice of hyperparameters will be dependent on the dataset.

[0169] We report the cross-validated Area Under the Receiver Operating Characteristic (AUC-ROC) and Area Under the Precision-Recall Curve (AUC-PRC) score as a measure of internal validation for the classification part of the model. The AUC-PRC depends on the prevalence of the outcome in the test population; thus, we also report the enrichment defined as the observed AUC-PRC divided by the prevalence in the test population. For the regression part of the model, we report the mean squared error and Pearson correlation.

[0170] We employed a bootstrap-resampling procedure for feature selection, drawing samples with replacement from the development set. In each of the 500 iterations, we trained the LightGBM model as previously described and computed SHAP values for each feature to determine their importance. These SHAP values were scaled in each bootstrap sample by dividing them by the sum of all SHAP scores, ensuring a relative comparison. The final SHAP score for each feature in each bootstrap sample was the median value across individuals. The final SHAP score for each feature, across all bootstrap samples, was the median value. We then averaged the SHAP values for the regressor and classifier and ranked them. Based on this ranking, the top 10 and top 20 features were selected. We subsequently trained the final model using either the top 10 or top 20 features from the development data set and followed the same evaluation process as initially described. Throughout this process, no external validation data sets were introduced, precluding any feature leakage.

[0171] Furthermore, we also report the sensitivity and specificity for the 2.5% threshold and 4% threshold. The sensitivity and specificity were calculated using the threshold determined in the hyperparameter optimization. Confidence intervals were generated from 1 ,000 bootstrap samples.

[0172] The analysis was performed using Snakemake, Optuna, LightGBM, and Scikit-learn. cffDNA threshold

[0173] To determine whether the amount of circulating cell free foetal DNA (cffDNA) is useful for making diagnostic predictions a cut-off must be established. This cut-off may be established by the laboratory, the physician or on a case by case basis by each individual. The primary outcome of this study was the foetal fraction (FF) in cases of pregnancy loss, estimated using SeqFF, by which small differences in sequencing behaviour for maternal and foetal cell-free DNA and read counts were used for estimation.

[0174] The foetal fraction can be established using a number of methods known to the skilled addressee, including the SNP based method (leveraging genomic variations between the mother and foetus), measuring the ratio of the Y chromosome (male foetus), or modelling the distribution of cfDNA fragment lengths.

[0175] The cut-off level can be based on several criteria including the number of individuals who would go on for further invasive diagnostic testing or other criteria known to those skilled in the art.

[0176] The cut-off level could be established using a number of methods, including percentiles, mean plus or minus standard deviation(s); multiples of median value; patient specific risk or other methods known to those who are skilled in the art.

[0177] Preferably, a specific risk can be calculated using Bayes rule, the a priori risk, and the relative frequencies for unaffected and affected pregnancies which are determined by incorporating the patient’s quantitative levels on each analyte along with the gestational age, into the probability density functions developed for the reference data using multivariate discriminant analysis or multidimensional truncated normal (or other) distributions.

[0178] The multivariate discriminant analysis and other risk assessments can be performed on the commercially available computer program statistical package Statistical Analysis system (manufactured and sold by SAS Institute Inc.) or by other methods of multivariate statistical analysis or other statistical software packages or screening software known to those skilled in the art.

[0179] In one or more exemplary embodiments, the fraction of the estimated amount of the cffDNA is more than 2.5% in the sample.

[0180] In one or more exemplary embodiments, the fraction of the estimated amount of the cffDNA is more than 4% in the sample.

[0181] In one or more exemplary embodiments, the fraction of the estimated amount of the cffDNA is more than 2.5% or more than 4% in maternal samples taken from biochemical pregnancies.

[0182] In one or more exemplary embodiments, the fraction of the estimated amount of the cffDNA is more than 2.5% or more than 4% in maternal samples taken from confirmed intrauterine pregnancy.

[0183] In one or more exemplary embodiments, the fraction of the estimated amount of the cffDNA is more than 2.5% or more than 4% in maternal samples taken from ongoing pregnancies.

[0184] In one or more exemplary embodiments, the fraction of the estimated amount of the cffDNA is more than 2.5% or more than 4% in maternal samples taken pregnancy loss between the 5th and 22nd gestational weeks.

[0185] In one or more exemplary embodiments, the fraction of the estimated amount of the cffDNA is more than 4% in maternal samples taken from biochemical pregnancies or confirmed intrauterine pregnancy, hereunder ongoing pregnancies, and pregnancy loss between the 5thand 22ndgestational weeks.

[0186] Final prediction

[0187] In one or more exemplary embodiments, the final prediction of the foetal fraction is given by FF = (0 > t) • y. 9 is an empirically determined value from the historical data, t is the probability (a number between 0 and 1 ) that the foetal fraction is > 2.5%. y is the predicted foetal fraction, which will only be greater than zero if, and only if, 9 > t. Cross-validated AUC-ROC

[0188] In one or more exemplary embodiments, the cross-validated AUC-ROC for the estimation of more than 2.5% cffDNA is higher than 0.8.

[0189] Cross-validated sensitivity

[0190] In one or more exemplary embodiments, the cross-validated sensitivity for the estimation of more than 2.5% cffDNA is higher than 0.78.

[0191] Cross-validated specificity

[0192] In one or more exemplary embodiments, the cross-validated specificity for the estimation of more than 2.5% cffDNA is higher than 0.58.

[0193] Cross-validated mean-squared error

[0194] In one or more exemplary embodiments, the cross-validated mean-squared error for the estimation of more than 4% cffDNA is less than 0.00.

[0195] In one or more exemplary embodiments, the cross-validated sensitivity for the estimation of more than 4% cffDNA is less than 0.77.

[0196] In one or more exemplary embodiments, the cross-validated specificity for the estimation of more than 4% cffDNA is less than 0.39.

[0197] Tables

[0198] Table 1: Baseline characteristics for the development (Hvidovre) and external validation cohorts (Herlev, Hille rod) Table 2: Univariate statistical analysis of the biomarkers and foetal fraction in the development cohort (Copenhagen University Hospital Hvidovre). An up-arrow indicates that the increased levels is associated with higher levels of foetal fraction.

[0199] Table 3 Performance characteristics in the external validation cohorts, using clinical characteristics and biomarkers Table 4 Performance characteristics when using Top 10 features

[0200] Table 5: Performance characteristics in the external validation cohorts using the top 20 features.

[0201] Table 6: Ranges of hyper parameters used for optimization of the machine learning model in the cross-validation step

[0202] Table 4: Features included in this study

[0203] General

[0204] The terms AST and / or ALT level, content, and amount, are used interchangeably. The following figures and examples are provided below to illustrate the present invention.

[0205] They are intended to be illustrative and are not to be construed as limiting in any way.

[0206] BRIEF DESCRIPTION OF THE FIGURES

[0207] Figure 1

[0208] Forest plot of univariate associations from the Bayesian log-normal hurdle model. The lognormal coefficient is the percentage change in the foetal fraction, and the zero component is the log-odds ratio associated with a foetal fraction > 2.5%. Results are largely in concordance, with associations to biomarkers related to the placenta function, thyroid function, liver functions, and lipids. Uric acid is not shown, due to its very large effect.

[0209] Figure 2

[0210] Receiver Operating Characteristic (ROC) Curves for University Hospital Herlev and North Zealand Hospital. The figure displays the Receiver Operating Characteristic (ROC) curves for the two external validation cohorts illustrating their performance in classifying a foetal fraction > 2.5%. The ROC curve shows the True Positive Rate (Sensitivity) against the False Positive Rate (1 -Specificity) at different classification thresholds. The Area Under the Curve (AUC) is calculated using the trapezoidal rule, and the diagonal dashed line represents a random classifier (AUC = 0.5) serving as a baseline for comparison. Both cohorts indicate that the model generalizes well and has good performance, as evidenced by their AUC-ROC values greater than 0.8.

[0211] Figure 3

[0212] Precision-Recall Curves for the model predictin Foetal Fraction > 2.5%. This figure illustrates the Precision-Recall Curves (PRC) for the two validation cohorts aimed at classifying foetal fraction > 2.5%. Precision-Recall Curves graphically represent the trade-off between Precision (Positive Predictive Value) and Recall (Sensitivity) at different classification thresholds. The baseline, l.e., random guessing, depends on the prevalence. The auPRC is, in both cases, superior to random guessing with an enrichment close to three-fold in both cohorts.

[0213] Figure 4

[0214] Feature Importance for the Classification model Ranked by Absolute Mean SHAP Values. It presents the 15 most significant features in the classification model, where the outcome is a foetal fraction > 2.5%, based on their absolute mean SHAP values. The SHAP values on the log-odds scale. These values indicate the extent to which each feature influences the model's output. The features are listed from top to bottom in descending order of their absolute mean SHAP values. Figure 5

[0215] SHAP Bar Plot for the investigated individual with the highest predicted probability of a foetal fraction > 2.5%. The probability estimated from the model was 98%, and the individual did have a foetal fraction > 2.5%. The values on the left side of the figure are the actual individual values. The major contributing factors to the prediction were £-hCG values (on the log scale), gestational age, creatine kinase, and TSH.

[0216] Figure 6

[0217] SHAP Bar Plot for the investigated individual with the lowest predicted probability of a foetal fraction > 2.5%. The probability estimated from the model was 40.7%, and the observed foetal fraction was < 2.5%. The major contributing factors to the prediction was low p-hCG, early gestational age, low CRP, and high maternal BMI.

[0218] Figure 7

[0219] Shapley Additive Explanations (SHAP) Bar Plot for an individual at 8+6 days of gestation and a P-hCG of 5039 Ul / L. The individual had an estimated probability of 88%, and the foetal fraction > 2.5%. The most important factors, apart from £-hCG and gestational age, were creatine kinase, HDL Cholesterol, and Transferrin. The individual has the same £-hCG value and gestational age as in Figure 9, yet the predicted probability and outcome is vastly different.

[0220] Figure 8

[0221] Shapley Additive Explanations (SHAP) Bar Plot for an individual at 8+6 days of gestation and a P-hCG of 5039 Ul / L. The individual had an estimated probability of 64%, despite having appropriate P-hCG levels for gestational week 8+6. ALT, prior number of pregnancy losses, and a skewed ALT I AST ratio decreased the probability. The individual has the same £-hCG value and gestational age as in Figure 9, yet the predicted probability and outcome is vastly different. The individual has the same P-hCG value and gestational age as in Figure 8, yet the predicted probability and outcome is vastly different.

[0222] Figure 9

[0223] Feature Importance Ranked by Absolute Mean SHAP Values for the log-normal part of the Hurdle prediction model. This figure presents the SHAP scores for the 15 most important features in the predictive model according to their absolute mean SHAP values. The SHAP values, which are on the log-odds scale, provide a measure of the impact of each feature on the model output, considering the magnitude feature's effect. Features are ranked from top to bottom in descending order based on their absolute mean SHAP values. The top ranked features are largely concordant with the classification model, albeit maternal BMI ranks significantly lower.

[0224] EXAMPLES

[0225] Example 1 - Study design and participants

[0226] The Copenhagen Pregnancy Loss Study (COPL) is a prospective cohort study that investigates pregnancy loss. Participants and their partners, if possible, referred with a pregnancy loss, were enrolled from three Danish hospitals (Copenhagen University Hospital Hvidovre, Herlev University Hospital, or North Zealand Hospital) from November 12, 2020, to May 1 , 2022. Eligibility criteria included age > 18 years, confirmed intrauterine pregnancy loss before 22 weeks of gestation, and the ability to comprehend information in Danish or English and provide informed consent. The exclusion criteria were pregnancies of unknown location, molar pregnancies, and inability to make informed decisions. This study adhered to the Declaration of Helsinki and received ethical approval from the appropriate health research committee. Data and personal information were handled securely. The participants had the autonomy to withdraw without impacting their care. Further details of this study are available on the COPL homepage (https: / / www.graviditetstab.dk / ).

[0227] Example 2 - Sample collection and analysis

[0228] Maternal blood was collected either prior to the removal of pregnancy tissue or within 24 hours post-evacuation. The plasma component was separated through a dual-stage centrifugation process, initially for 10 minutes at 2,500g followed by another 10 minutes at 12,500g. The plasma was then preserved at temperatures of either -18°C or -80°C, pending subsequent analysis.

[0229] Upon defrosting, cell-free DNA (cfDNA) was extracted from 2 ml of plasma utilizing the MaxWell RSC cfDNA Plasma Kit, and the resultant DNA was subsequently reconstituted in 60 pl of elution buffer. Library construction was performed with the TruSeq Nano DNA LT Library Prep Kit supplied by Illumina, which includes an eight-cycle polymerase chain reaction (PCR) amplification. Quantification of the library concentration was accomplished using QuantStudio 6 (QPCR), and library integrity was assessed by electrophoretic analysis with a Fragment Analyzer. Libraries from eight distinct samples were amalgamated in equal molar proportions to achieve a cumulative concentration of 20 pM DNA, which was then sequenced on the HiSeq1500 system utilizing the TruSeq Rapid SBS kit for 50 cycles (single-read). For demultiplexing and conversion from bcl to FASTQ file formats, CASAVA software was employed, allowing for a single mismatch within their tags. Only samples yielding 10 million or more raw reads were considered suitable for further data analysis.

[0230] Bioinformatics processing was performed as described by Hartwig et al (Schlaikjaer Hartwig, T. et al. Cell-free foetal DNA for genetic evaluation in Copenhagen Pregnancy Loss Study (COPL): a prospective cohort study. Lancet Lond. Engl. 401 , 762-771 (2023).

[0231] Briefly, reads across lanes were merged into one fastq file. Reads were aligned using bowtie2. Aligned reads with a quality < 1 were removed, and the remaining reads were sorted and deduplicated using samtools. The foetal fraction was determined using SeqFF (Kim, S. K. et al. Determination of foetal DNA fraction from the plasma of pregnant women using sequence read counts. Prenat. Diagn. 35, 810-815 (2015)) by which small differences of sequencing behaviour for maternal and foetal cell-free DNA and read counts were used for estimation.

[0232] Maternal whole blood was collected in EDTA tubes or serum clot activator tubes and separated into plasma and serum. The serum samples were used for biochemical analysis at the Department of Clinical Biochemistry, at Copenhagen University Hospital Herlev, measured using standard assays, such as £-hCG (sandwich immunoassay by Siemens Atellica IM Analyzer), creatinine and cholesterol (Enzymatic reaction and absorbance by Siemens Atellica CH 930).

[0233] Example 3 - Outcome

[0234] The primary outcome of this study was the foetal fraction (FF) in cases of pregnancy loss, estimated using SeqFF, by which small differences in sequencing behaviour for maternal and foetal cell-free DNA and read counts were used for estimation.

[0235] The SeqFF algorithm estimates a continuous unbounded value. In this study, values at or below 2.5% were set to zero, as these are noise values that most likely represent the lack of foetal DNA. The 2.5% cut-off is also used to determine if a result is reliable in pregnancy loss. For ongoing pregnancies, a foetal fraction of > 4% is used to determine if the test result is reliable. The blood sample used for the cffDNA test was collected prior to tissue passing through the uterus in >90% of the samples.

[0236] Example 4 - Predictors

[0237] Our study evaluated the predictive capabilities of variables available prior to isolation and sequencing of cffDNA (i.e., the expensive step of the test), which are broadly categorized into maternal factors (such as BMI and gestational age), maternal biochemistry (e.g., TSH, CRP, etc.), foetal factors (e.g., gestational week, p-hCG, etc.), and paternal factors (BMI, age). A comprehensive list of these factors is provided in Table 6.

[0238] In total we included nine clinical measures and 28 biochemistry measurements.

[0239] Three sets of predictors were defined,

[0240] Maternal and Foetal Clinical Measures (model Clin)

[0241] Maternal Biochemistry and Foetal Biochemistry (model Biochem)

[0242] Maternal and Foetal Clinical measures + Maternal and Foetal Biochemistry (model ClinBiochem)

[0243] Sample size

[0244] The COPL study is an ongoing prospective cohort study. Pregnancy losses with a cffDNA test at the time of writing were included. CffDNA tests with a total of analyzed reads below 8 million were excluded.

[0245] Missing data

[0246] Missing values below the detection limit were imputed from a uniform distribution ranging from zero to the laboratory defined lower limit of detection. The ratio between conjugated and unconjugated bilirubin was imputed from a uniform distribution ranging between 0 and 1.

[0247] Other missing data were imputed using mean values for continuous variables, and the most frequent value for categorical variables.

[0248] Statistical analysis

[0249] The baseline characteristics and features of the machine learning models were summarized as mean (standard deviation, SD), median (interquartile range, IQR), or frequency, where appropriate. The FF was binarized, using two clinically relevant thresholds as described earlier, namely 2.5% and 4%.

[0250] Univariate association between the FF and other variables were modelled using a Bayesian

[0251] Hurdle lognormal regression model with a weakly informative prior, namely the student-t distribution centered around zero with seven degrees of freedom. The Hurdle lognormal regression has the likelihood, L,

[0252] In which f is the probability density function for the log-normal distribution, and F is the cumulative density function for the log-normal distribution. Consequently, the hurdle model comprises two distinct components: the first models the probability of observing a zero versus a non-zero outcome, and the second models the distribution of X conditioned on it being positive and non-zero. Each component is characterized by its own set of parameters. The model was implemented in brms and ran with default parameters (5,000 warm-up iterations, 5,000 draws from the posterior from 4 independent chains). The model has converged if, and only if, (1 ) all R-hat values were < 1.01 , (2) a tree depth of ten had not been exceeded post warm-up, (3) there were no divergences. Due to the extreme values p-hCG had to be log transformed for the model to converge. Biochemistry variables were additionally adjusted for the time between sample collection and biomarker measurement (i.e. , time spent at -80degrees Celsius). We report the association statistics as the 95% Bayesian Credible Interval (BCI). BCIs excluding zero were taken as evidence that a given biomarker was associated with the outcome.

[0253] For predictive modelling we created a model with two components, resembling the hurdle lognormal model. The model consisted of two parts, namely a classifier to predict if the foetal fraction > 2.5%, and a regression model to predict the fraction. Only non-zero positive values > 2.5% were used as input for the regression model. In the modelling framework, two distinct components were employed: a classifier and a regressor. The gradient boosting framework LightGBM was utilized for both components. The LightGBM model can capture non-linearities and interactions if present. We also compared the performance to a model consisting of a linear regression and logistic regression. The objective function of the classification and regression model are the log loss and mean squared error, respectively. The final prediction of the foetal fraction is given by,

[0254] FF = (0 > t) • y In which 6 is the estimated probability of the foetal fraction > 2.5%, and t is the threshold. The parameter t is set to be equivalent of the prevalence of the major class, y is the estimate from the regression model. Consequently, if 9 is below the threshold, t, the foetal fraction is zero. To identify the best set of hyperparameters, we evaluated each hyperparameter setting in a five-fold cross-validation process. We defined the loss of a hyper parameter configuration as the weighted sum of the mean squared error (MSE) and log loss (LL),

[0255] L = MSE - w + LL

[0256] The weighting factor, w, was empirically determined by fitting a linear regression model and logistic regression model with no covariates. We set the weight, w, to be 100.

[0257] The optimal model was selected based on the cross-validated loss. Hyperparameters were optimized using Optuna software, run for 500 trials, with default settings. The range and distributions for hyperparameters is detailed in Table 6. We report the cross-validated Area Under the Receiver Operating Characteristic (AUC-ROC) and Area Under the Precision- Recall Curve (AUC-PRC) score as a measure of internal validation for the classification part of the model. The AUC-PRC depends on the prevalence of the outcome in the test population; thus, we also report the enrichment defined as the observed AUC-PRC divided by the prevalence in the test population. For the regression part of the model, we report the mean squared error and Pearson correlation.

[0258] We employed a bootstrap-resampling procedure for feature selection, drawing samples with replacement from the development set. In each of the 500 iterations, we trained the LightGBM model as previously described and computed SHAP values for each feature to determine their importance. These SHAP values were scaled in each bootstrap sample by dividing them by the sum of all SHAP scores, ensuring a relative comparison. The final SHAP score for each feature in each bootstrap sample was the median value across individuals. The final SHAP score for each feature, across all bootstrap samples, was the median value. We then averaged the SHAP values for the regressor and classifier and ranked them. Based on this ranking, the top 10 and top 20 features were selected. We subsequently trained the final model using either the top 10 or top 20 features from the development data set and followed the same evaluation process as initially described. Throughout this process, no external validation data sets were introduced, precluding any feature leakage.

[0259] Model development was done using the data collected from the main recruitment site, Copenhagen University Hospital Hvidovre. External model validation was performed using data from the two separate recruitment sites (Herlev University Hospital and North Zealand Hospital), which were not involved in the development of the model. We report the same characteristics as for the development data set. Furthermore, we also report the sensitivity and specificity. The sensitivity and specificity were calculated using the threshold determined in the hyperparameter optimization. Confidence intervals were generated from 1 ,000 bootstrap samples.

[0260] The analysis was performed using Snakemake, Optuna, LightGBM, and Scikit-learn. To ensure the reproducibility and reliability of our findings, we transparently reported each step of our statistical analysis in accordance with the TRIPOD statement, demonstrating our commitment to following best practices.

[0261] Example 5 - Development of the predictive model

[0262] Participants

[0263] A total of 978 participants recruited at Copenhagen University Hospital Hvidovre were included in the development of the predictive model. These individuals represented a diverse range of age, BMI, and biochemical value ranges, ensuring comprehensive inclusion of the target population (Table 1 ).

[0264] The model was validated using an independent dataset comprising 507 participants from two separate inclusion sites (Herlev University Hospital and Northern Zealand Hospital). These participants were not involved in the model's training or internal validation process, thus providing a robust test for the model's generalizability. This validation cohort maintained similar diversity in terms of age, BMI, and biochemical value ranges. The time a sample had spent in the freezer, prior to being used to measure the biochemical assays, was higher for the Hvidovre cohort. B-hCG levels were lower in the Northern Zealand cohort, despite experiencing the pregnancy loss at a later gestational age (Table 1 ).

[0265] Model Development

[0266] In the development cohort, the proportion of samples with a foetal fraction insufficient for the cffDNA test was 10.5% at the 2.5% threshold. The median foetal fraction was 4.5%. Similarly, in the independent validation data sets, the proportion was 8.5% and median 4.9%, respectively. The application of the Bayesian Hurdle Model provided two parameter estimations for each biomarker under investigation. This dual-parameter approach yielded an odds ratio, elucidating the likelihood for the biomarker concentration to fall below a threshold of 2.5%, and a regression coefficient derived from a log-normal distribution model, quantifying the expected change in the biomarker's log-concentration. We found ten biochemical biomarkers associated with a foetal fraction < 2.5%. The biomarkers with the largest absolute magnitude were P-hCG (OR 0.30, 95% BCI 0.18; 0.45), Creatinine (OR 1.49, 95% BCI 1.22; 1 .85), and HDL Cholesterol (OR 0.68, 95% BCI 0.54; 0.85) (Table 2). Furthermore, method of conception, gestational age, BMI, and prior number of pregnancy losses were also associated. The latter portion of the Bayesian Hurdle Model analysis, focusing on the log-normal component, exhibited a high level of and included Thyroid-Stimulating Hormone (TSH), Conjugated Bilirubin, Alanine Transaminase (ALT), and the ratio of Aspartate Transaminase (AST) to Alanine Transaminase (ALT).

[0267] Model Specification

[0268] Predictive models were developed utilizing data from the development cohort and were internally validated through five-fold cross validation. The most successful model incorporated both clinical measures and biochemistry (ClinBiochem). The models that included only clinical measure or biochemistry were inferior (data now shown). The LightGBM classifier was employed to determine if the foetal fraction exceeded 2.5% and yielded a cross-validated AUC-ROC of 0.77 and an AUPRC of 0.33. This represented a 3.2-fold increase in enrichment at a prevalence of 10.5%. The LightGBM regressor demonstrated a cross-validated mean squared error of 0.0017 and a Pearson correlation of 0.33. In comparison, the logistic and linear regression models achieved an AUC-ROC of 0.72 and AUPRC of 0.23, respectively, and their mean squared error was 0.0021 and Pearson correlation was 0.29.

[0269] Model Performance

[0270] The developed model was externally validated using data collected from two geographical distinct hospitals, namely Herlev University Hospital and North Zealand Hospital. We found that the performance metrics were highly comparable to the internal validation. The AUC-ROC scores from Herlev University Hospital (0.79; 0.62-0.76 95% Cl) and North Zealand Hospital (0.66; 0.59-0.74 95% Cl) were like the results from cross-validation. The AUC-PRC was higher at Herlev University Hospital (0.31 ; 0.17-0.48 95% Cl) compared to North Zealand Hospital (0.21 ; 0.11-0.34 95% Cl) (Figure 4). Both were comparable to the cross-validation and had a 3-fold enrichment compared to the prevalence in the test cohorts (Table 3).

[0271] In both cohorts, utilizing a probability threshold that aligns with the prevalence of samples (P > 0.895) in the training cohort where foetal fraction exceeds 2.5% resulted in achieving both a sensitivity and a specificity surpassing 60% and 70%, respectively. Similarly, predicting if the foetal fraction > 4%, yielded sensitivities and specificities that were above 45% and close to 80% in both cohorts, respectively (Table 3).

[0272] Utilizing the previously described strategy for feature selection, we found that, in addition to £- hCG and gestational age, lipid levels (HDL), liver function (ASAT / ALAT ratio), kidney (creatinine), and thyroid (thyroid stimulating hormone and Thyroid Peroxidase Antibody) were important for the cell free foetal DNA levels. The machine learning models has a highly similar performance.

[0273] Model interpretation and examples of use

[0274] To understand the decision-making process of the LightGBM model, we used SHAP (SHapley Additive exPlanations) values, a state-of-the-art method for interpreting models. By analyzing all features, we found that p-hCG and gestational age calculated from the last menstruation were the most crucial, along with creatin kinase and alanine transaminase for classification of a foetal fraction > 2.5%. Creatin kinase and alanine transaminase even surpassed BMI.

[0275] We provide four examples of how to utilize the machine learning model for individual predictions of foetal fraction.

[0276] Comparing the two extremes, we selected the woman with the highest and lowest predicted probability of foetal fraction > 2.5%. In the first case, the predicted probability was 98%, attributed to P-hCG, gestational age, creatin kinase, and low levels of thyroid stimulating hormone. Conversely, in the second case, the predicted probability was 40.7%, due to low levels of P-hCG, gestational age, CRP, and high maternal BMI.

[0277] In another example, we contrasted two individuals with the same levels of P-hCG and gestational week. The first person had a probability of 88%, while the second person had a probability of 64%. Gestational age and £-hCG remained important, but other characteristics significantly changed.

[0278] In the second part of the Hurdle model, the log-normal regression, gestational age, and £-hCG were also the most important, followed by HDL Cholesterol, ALT / AST Ratio, and uric acid.

[0279] The integrative SHAP analysis provides a way to give feedback and describe to clinicians and patients which factors contributed to a low or high prediction of euploid loss. ITEMS

[0280] 1 . A method for estimating the amount of circulating cell free foetal DNA (cffDNA) in an individual, comprising a) providing a sample from said individual b) determining an aspartate transaminase (AST) and alanine aminotransferase (ALT) ratio (AST / ALT ratio) in said sample c) correlating said AST / ALT ratio with a reference level, and thereby estimating the amount of cffDNA.

[0281] 2. A method according to item 1 , wherein the sample is selected from the group consisting of blood, urine, faeces, rectal swab, rectal microbiome, vaginal microbiome, and vaginal discharge samples.

[0282] 3. A method according to item 2, wherein in the sample is a blood sample.

[0283] 4. A method according to item 3, wherein the blood sample is separated into plasma and serum samples.

[0284] 5. A method according to item 4, wherein the sample is a serum sample.

[0285] 6. A method according to item 4, wherein the sample is a plasma sample.

[0286] 7. A method according to any of the preceding items, wherein the AST / ALT ratio is determined by measuring RNA, DNA, protein, antibodies, and / or metabolite level of the aspartate transaminase and alanine aminotransferase enzymes.

[0287] 8. A method according to item 7, wherein the AST / ALT ratio is determined by measuring protein levels of the aspartate transaminase and alanine aminotransferase enzymes.

[0288] 9. A method according to any of the preceding items, wherein the AST / ALT ratio is combined with levels of further biomarkers selected from the group consisting of gestational age, betaHuman Chorionic Gonadotropin (£-hCG), Pregnancy-associated plasma protein A (PAPP-A), vascular endothelial growth factor (VEGF), soluble fms-like tyrosine kinase-1 (sFlt-1 ), Alpha- Fetoprotein (AFP), Disintegrin and metalloproteinase domain-containing protein 12 (ADAM 12), ISM2, TFPI2, ERW-1 , LYPD3, EBIB3JL27, CSH1 , GDF15, ANGPT2, FBN2, PRG2, INSL4, LAIR2, TSHB, C1QTNF6, LHB, SIGLEC6, MMP12, Placental Growth Factor (PLGF1 ), Alanine Transaminase (ALT), Aspartattransaminase (ASAT), High-density lipoprotein cholesterol (HDL), Low-density lipoprotein cholesterol (LDL), Apolipoprotein B, Uric Acid, Transferrin, Bilirubin, Creatin kinase and Lipoprotein A.

[0289] 10. A method according to item 9, wherein the further biomarkers are selected from the group consisting of gestational age, P-hCG, Bilirubin, Creatin kinase and Lipoprotein A.

[0290] 11 . A method according to item 9, wherein the further biomarker is Bilirubin.

[0291] 12. A method according to item 9, wherein the further biomarker is Creatin kinase.

[0292] 13. A method according to item 9, wherein the further biomarker is Lipoprotein A.

[0293] 14. A method according to any of the preceding items, wherein the sample is taken from individual in a couple trying to and / or having conceived a pregnancy / liveborn child.

[0294] 15. A method according to any of the preceding items, wherein the sample is a maternal and / or a paternal sample.

[0295] 16. A method according to any of the preceding items, wherein the sample is taken in gestational age 5 weeks + 0 days to 22 weeks + 0 days.

[0296] 17. A method according to any of the preceding items, wherein the fraction of the estimated amount of the cffDNA is more than 2.5% in maternal samples taken while the pregnancy tissue being in situ, the pregnancy is biochemically confirmed (£-hCG ), being a confirmed intrauterine pregnancy, and / or is an ongoing pregnancy between the 5thand 22ndgestational week.

[0297] 18. A method according to any of the preceding items, wherein the fraction of the estimated amount of the cffDNA is more than 2.5% in a sample taken within 96 hours after removal or spontaneous passage of the pregnancy product and / or pregnancy loss between the 5thand 22ndgestational weeks. 19. A method according to any of the preceding items, wherein the fraction of the estimated amount of the cffDNA is more than 4%.

[0298] 20. A method according to any of the preceding items, wherein the final prediction of the foetal fraction is given by FF = 6 > t) • y.

[0299] 21 . A method according to any of the preceding items, wherein the cross-validated AUC-ROC for the estimation of more than 2.5% cffDNA is more than 0.8.

[0300] 22. A method according to any of the preceding items, wherein the cross-validated sensitivity for the estimation of more than 2.5% cffDNA s more than 0.78.

[0301] 23. A method according to any of the preceding items, wherein the cross-validated specificity for the estimation of more than 2.5% cffDNA is more than 0.58.

[0302] 23. A method according to any of the preceding items, wherein the cross-validated mean- squared error for the estimation of more than 4% cffDNA is less than 0.00.

[0303] 24. A method according to any of the preceding items, wherein the cross-validated sensitivity for the estimation more than 4% cffDNA is less than 0.77.

[0304] 25. A method according to any of the preceding items, wherein the cross-validated specificity for the estimation more than 4% cffDNA is less than 0.39.

[0305] 26. A computer-implemented method for determining an amount of circulating cell free foetal

[0306] DNA (cffDNA) in an individual, the method comprising: obtaining sample data of a sample from the individual; determining an AST content of aspartate transaminase (AST) in the sample based on the sample data; determining an ALT content of alanine aminotransferase (ALT) in the sample based on the sample data; determining the amount of circulating cell free foetal DNA based on the AST content and the ALT content; and outputting the amount of circulating cell free foetal DNA. 27. Method according to item 26, wherein determining the amount of circulating cell free foetal DNA based on the AST content and the ALT content comprises applying a machine-learning model on the AST content and the ALT content.

[0307] 28. Method according to any one of items 26-27, wherein determining the amount of circulating cell free foetal DNA based on the AST content and the ALT content comprises determining a ratio between the AST content and the ALT content.

[0308] 29. Method according to any one of the items 26-28, wherein determining the ratio between the AST content and the ALT content comprises measuring one or more of mRNA, siRNA, piRNA, ncRNA, IncRNA, miRNA RNA, DNA, methylation relating to epigenomics, protein, antibodies, metabolite, ions, lipid level of aspartate transaminase enzyme, and lipid level of alanine aminotransferase enzyme.

[0309] 30. Method according to item 29, wherein determining the ratio between the AST content and the ALT content comprises measuring DNA levels of the aspartate transaminase and alanine aminotransferase enzymes and determining the AST / ALT ratio.

[0310] 31 . Method according to any one of the items 26-30, wherein determining the amount of circulating cell free foetal DNA based on the AST content and the ALT content comprises determining the amount of circulating cell free foetal DNA based on the AST content, the ALT content, and one or more of gestational age, beta-Human Chorionic Gonadotropin (£-hCG) level, Pregnancy-associated plasma protein A (PAPP-A) level, vascular endothelial growth factor (VEGF) level, soluble fms-like tyrosine kinase-1 (sFlt-1 ) level, Alpha-Fetoprotein (AFP) level, Disintegrin and metalloproteinase domain-containing protein 12 (ADAM12) level, ISM2 level, TFPI2 level, ERW-1 level, LYPD3 level, EBIB3JL27 level, CSH1 level, GDF15 level, ANGPT2 level, FBN2 level, PRG2 level, INSL4 level, LAIR2 level, TSHB level, C1 QTNF6 level, LHB level, SIGLEC6 level, MMP12 level, Placental Growth Factor (PLGF1 ) level, Alanine Transaminase (ALT) level, Aspartate transaminase (AST) level, High-density lipoprotein cholesterol (HDL) level, Low-density lipoprotein cholesterol (LDL) level, Apolipoprotein B level, Uric Acid level, Transferrin level, Bilirubin level, Creatin kinase level, and Lipoprotein A level.

[0311] 32. A computer-implemented method for training a machine learning model to process as inputs an AST content of aspartate transaminase (AST) of a sample and an ALT content of alanine aminotransferase (ALT) of the sample; and provide as output one or more parameters indicative of amount of circulating cell free foetal DNA of the sample, the method comprising: obtaining, using at least one processor, an AST content of aspartate transaminase (AST) and an ALT content of alanine aminotransferase (ALT) of a sample; performing, using the at least one processor, a training comprising: generating, using the at least one processor and a machine-learning model, cffDNA data based on the AST content and the ALT content; obtaining, using the at least one processor, training data; determining, using the at least one processor and one or more loss functions, one or more loss parameters based on the AST content, ALT content, and the training data; and training, using the at least one processor, the machine learning model based on the one or more loss parameters.

Claims

CLAIMS1 . A method for estimating the amount of circulating cell free foetal DNA (cffDNA) in an individual, comprising a) providing an obtained sample from said individual b) determining an aspartate transaminase (AST) and alanine aminotransferase (ALT) ratio (AST / ALT ratio) in said sample c) correlating said AST / ALT ratio with a reference level, and thereby estimating the amount of cffDNA.

2. A method according to claim 1 , wherein the sample is selected from the group consisting of blood, urine, faeces, rectal swab, rectal microbiome, vaginal microbiome, and vaginal discharge samples.

3. A method according to any of the preceding claims, wherein the AST / ALT ratio is determined by measuring RNA, DNA, protein, antibodies, and / or metabolite level of the aspartate transaminase and alanine aminotransferase enzymes.

4. A method according to any of the preceding claims, wherein the AST / ALT ratio is combined with levels of further biomarkers selected from the group consisting of gestational age, betaHuman Chorionic Gonadotropin (£-hCG), Pregnancy-associated plasma protein A (PAPP-A), vascular endothelial growth factor (VEGF), soluble fms-like tyrosine kinase-1 (sFlt-1 ), Alpha- Fetoprotein (AFP), Disintegrin and metalloproteinase domain-containing protein 12 (ADAM12), ISM2, TFPI2, ERW-1 , LYPD3, EBIB3JL27, CSH1 , GDF15, ANGPT2, FBN2, PRG2, INSL4, LAIR2, TSHB, C1QTNF6, LHB, SIGLEC6, MMP12, Placental Growth Factor (PLGF1 ), Alanine Transaminase (ALT), Aspartattransaminase (ASAT), High-density lipoprotein cholesterol (HDL), Low-density lipoprotein cholesterol (LDL), Apolipoprotein B, Uric Acid, Transferrin, Bilirubin, Creatin kinase and Lipoprotein A.

5. A method according to any of the preceding claims, wherein the sample is taken in gestational age 5 weeks + 0 days to 22 weeks + 0 days.

6. A method according to any of the preceding claims, wherein the final prediction of the foetal fraction is given by FF = 6 > t) • y.

7. A method according to any of the preceding claims, wherein the cross-validated specificity for the estimation of more than 2.5% cffDNA is more than 0.58.

8. A method according to any of the preceding claims, wherein the cross-validated specificity for the estimation more than 4% cffDNA is less than 0.39.

9. A computer-implemented method for determining an amount of circulating cell free foetal DNA (cffDNA) in an individual, the method comprising: obtaining sample data of a sample from the individual; determining an AST content of aspartate transaminase (AST) in the sample based on the sample data; determining an ALT content of alanine aminotransferase (ALT) in the sample based on the sample data; determining the amount of circulating cell free foetal DNA based on the AST content and the ALT content; and outputting the amount of circulating cell free foetal DNA.

10. Method according to claim 9, wherein determining the amount of circulating cell free foetal DNA based on the AST content and the ALT content comprises applying a machine-learning model on the AST content and the ALT content.11 . Method according to any one of claims 9-11 , wherein determining the amount of circulating cell free foetal DNA based on the AST content and the ALT content comprises determining a ratio between the AST content and the ALT content.

12. Method according to any one of the claims 9-11 , wherein determining the ratio between the AST content and the ALT content comprises measuring one or more of mRNA, siRNA, piRNA, ncRNA, IncRNA, miRNA RNA, DNA, methylation relating to epigenomics, protein, antibodies, metabolite, ions, lipid level of aspartate transaminase enzyme, and lipid level of alanine aminotransferase enzyme.

13. Method according to claim 12, wherein determining the ratio between the AST content and the ALT content comprises measuring DNA levels of the aspartate transaminase and alanine aminotransferase enzymes and determining the AST / ALT ratio.

14. Method according to any one of the claims 9-12, wherein determining the amount of circulating cell free foetal DNA based on the AST content and the ALT content comprisesdetermining the amount of circulating cell free foetal DNA based on the AST content, the ALT content, and one or more of gestational age, beta-Human Chorionic Gonadotropin (£-hCG) level, Pregnancy-associated plasma protein A (PAPP-A) level, vascular endothelial growth factor (VEGF) level, soluble fms-like tyrosine kinase-1 (sFlt-1 ) level, Alpha-Fetoprotein (AFP) level, Disintegrin and metalloproteinase domain-containing protein 12 (ADAM12) level, ISM2 level, TFPI2 level, ERW-1 level, LYPD3 level, EBIB3JL27 level, CSH1 level, GDF15 level, ANGPT2 level, FBN2 level, PRG2 level, INSL4 level, LAIR2 level, TSHB level, C1 QTNF6 level, LHB level, SIGLEC6 level, MMP12 level, Placental Growth Factor (PLGF1 ) level, Alanine Transaminase (ALT) level, Aspartate transaminase (AST) level, High-density lipoprotein cholesterol (HDL) level, Low-density lipoprotein cholesterol (LDL) level, Apolipoprotein B level, Uric Acid level, Transferrin level, Bilirubin level, Creatin kinase level, and Lipoprotein A level.

15. A computer-implemented method for training a machine learning model to process as inputs an AST content of aspartate transaminase (AST) of a sample and an ALT content of alanine aminotransferase (ALT) of the sample; and provide as output one or more parameters indicative of amount of circulating cell free foetal DNA of the sample, the method comprising: obtaining, using at least one processor, an AST content of aspartate transaminase (AST) and an ALT content of alanine aminotransferase (ALT) of a sample; performing, using the at least one processor, a training comprising: generating, using the at least one processor and a machine-learning model, cffDNA data based on the AST content and the ALT content; obtaining, using the at least one processor, training data; determining, using the at least one processor and one or more loss functions, one or more loss parameters based on the AST content, ALT content, and the training data; and training, using the at least one processor, the machine learning model based on the one or more loss parameters.