Maternal plasma transcriptome analysis by massively parallel RNA sequencing

By employing RNA-sequencing for the analysis of maternal plasma, the challenges of detecting RNA markers are overcome, allowing for the non-invasive diagnosis and monitoring of pregnancy-related disorders through the identification of fetal-specific and maternal-specific SNPs and their expression patterns.

JP2025081680APending Publication Date: 2025-05-27THE CHINESE UNIVERSITY OF HONG KONG
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025029276
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-02-28
Filing Date
2025-02-26
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Current methods for detecting RNA markers in maternal plasma are limited by the low concentration and degradation of extracellular RNA, making it difficult to identify fetal and placental RNA markers for prenatal diagnosis and monitoring.

Method used

The use of massively parallel sequencing (MPS) for RNA analysis, also known as RNA-sequencing (RNA-seq), allows for the direct and sensitive analysis of the maternal plasma transcriptome, enabling the detection of fetal-specific and maternal-specific single nucleotide polymorphisms (SNPs) and their relative contribution rates.

Benefits of technology

This approach enables the identification of pregnancy-related disorders such as preeclampsia and intrauterine growth retardation by analyzing the allelic-specific expression patterns in maternal plasma, providing a non-invasive and high-throughput method for prenatal diagnosis and monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025081680000001_ABST
    Figure 2025081680000001_ABST
Patent Text Reader

Abstract

To provide methods, systems and apparatuses for diagnosing pregnancy-associated disorders, determining allelic ratios, determining maternal or fetal contributions to circulating transcripts, and / or identifying maternal or fetal markers using a sample from a pregnant female subject.SOLUTION: Methods are provided for diagnosing pregnancy-associated disorders, determining allelic ratios, determining maternal or fetal contributions to circulating transcripts, and / or identifying maternal or fetal markers using a sample from a pregnant female subject. Also provided is use of a gene for diagnosing a pregnancy-associated disorder in a pregnant female subject.SELECTED DRAWING: Figure 18A
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 61 / 770,985, filed Feb. 28, 2013, entitled "Maternal plasma transcriptome analysis by massively parallel RNA sequencing", which is hereby incorporated by reference in its entirety for all purposes.

[0002] Background Pregnancy-specific extracellular RNA molecules in maternal plasma have already been reported. 1 This report provides a non-invasive test tool for evaluating the fetus and monitoring the progress of pregnancy by simply using maternal peripheral blood biopsies. To date, multiple promising applications as prenatal diagnoses using maternal plasma RNA have been developed. 2-8 Researchers are desperately searching for additional RNA markers to expand their scope of application to fetal disorders and pregnancy diseases in different situations.

[0003] Theoretically, the simplest method for identifying RNA markers is to directly examine extracellular RNA molecules in maternal plasma. However, this method has not been easy because high-throughput screening techniques such as microarray analysis and SAGE (serial analysis of gene expression) have limited ability to detect extracellular RNA, which is usually present at low concentrations and is partially degraded in maternal plasma. 9 Instead, in many of the reported RNA marker screenings, an indirect strategy of comparing the expression profiles of the placenta and maternal blood cells has been taken. 10Only the expression in maternal plasma samples of transcripts that are expressed at very high levels in the mother's blood compared to the placenta has been further analyzed by sensitive but low throughput systems, such as reverse transcription polymerase chain reaction (RT-PCR). To date, there have not been many plasma RNA markers identified by this indirect method. This is probably because tissue-level mining strategies do not fully account for all biological factors that affect the levels of placental RNA present in maternal plasma. In addition, transcripts expressed and released in tissues other than the placenta in response to pregnancy could not be identified by this method. Therefore, a highly sensitive and high-throughput method that enables direct transcriptome analysis of maternal plasma is highly desirable.

[0004] Similarly, direct plasma RNA profiling may be useful in other situations where mixtures of RNA molecules from two individuals are present, such as in organ transplantation. The bloodstream of transplant recipients contains nucleic acid molecules from both the donor and the recipient. Relative changes in the profiles of RNA molecules by the donor or the recipient may reveal pathologies such as graft rejection in the transplanted organ or the recipient. SUMMARY OF THE INVENTION

[0005] Summary of the Invention Provided are methods, systems, and devices for diagnosing pregnancy-related disorders, determining allelic ratios, determining maternal or fetal contribution rates to transcripts in the blood, and / or identifying maternal or fetal markers using samples from pregnant female subjects. In some embodiments, the sample is plasma containing a mixture of RNA molecules from the mother and the fetus.

[0006] Analyze the RNA molecules (e.g., determine the sequence) to obtain a plurality of reads and identify the positions of these reads in the reference sequence (e.g., identify by aligning the sequences). Identify informative loci where either the first allele of either the mother or the fetus is homozygous and the other first allele and the second allele of the mother and the fetus are heterozygous. Then, screen the informative loci and further analyze the reads located (e.g., aligned) at the selected informative loci. In some embodiments, calculate the ratio of the reads corresponding to the first allele and the second allele and compare it with a cut-off value to diagnose pregnancy-related disorders. In some embodiments, use the reads located at the selected maternal informative loci to determine the proportion of fetal-derived RNA contained in the sample. In some embodiments, calculate the ratio of the reads corresponding to the first allele and the second allele for each of the selected informative loci, compare the ratio with a cut-off value, and designate the locus as a maternal marker or a fetal marker.

[0007] This method is not limited to prenatal diagnosis and is applicable to all biological samples containing a mixture of RNA molecules derived from two individuals. For example, a plasma sample from a patient who has received an organ transplant can be used. Since the transcripts expressed in the transplanted organ reflect the genotype of the donor, expression at a detectable level occurs in the recipient's blood. By measuring these transcripts, informative loci in the genes expressed by both the donor and the recipient can be identified, and the relative expression levels of the alleles derived from either can be measured. Transplant-related disorders can be diagnosed based on abnormal expression levels of alleles by either the donor or the recipient.

[0008] Biological samples that can be used in this method include blood, plasma, serum, urine, saliva, and tissue samples. For example, fetal nucleic acids contained in the urine of pregnant women were detected. It was shown that the urine of kidney transplant recipients contains nucleic acids without cells and cells derived from the transplanted organ. Microchimerism was observed under many conditions. Microchimerism refers to the presence in the body of an individual, such as in organs or tissues, of cells or nucleic acid precursors derived from another person. Microchimerism was observed in biopsies of the thyroid, liver, spleen, skin, bone marrow, and other tissues. Microchimerism may be caused by past pregnancy or blood transfusion.

[0009] Also provided is the use of genes for diagnosing pregnancy-related disorders in pregnant women. The expression level of the gene is compared with the value of the subject determined from one or more female subjects carrying a healthy fetus. Pregnancy-related disorders addressed in the present invention include preeclampsia, intrauterine growth retardation, invasive placentation, and preterm birth. Other pregnancy-related disorders may include conditions where the fetus is at risk of death, such as hemolytic disease of the newborn, placental insufficiency, fetal hydrops, and fetal malformations. Still other pregnancy-related diseases may include conditions that cause pregnancy complications, such as HELLP syndrome, systemic lupus erythematosus, and other maternal immune diseases.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11A

Figure 11B

Figure 12A

Figure 12B

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18A

Figure 18B

Figure 19

Figure 20

Figure 21 - 1

Figure 21 - 2

Figure 21 - 3

Figure 22

Figure 23

Figure 24A

Figure 24B

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

DETAILED DESCRIPTION OF THE INVENTION

[0011] Detailed Description I. Introduction More than 10 years ago, it was reported that cell-free fetal-derived RNA exists in maternal plasma. 11 Since this discovery until now, many studies have been conducted to detect fetal and placental blood RNAs contained in maternal plasma. 1、10、12 Interestingly, it has been found that the expression level of placenta-specific transcripts in plasma is positively correlated with the expression level in placental tissue. 10 This indicates that the analysis of plasma RNA is a clinically effective non-invasive tool for monitoring the health and development of the placenta or fetus. In fact, from experiments on maternal blood RNA, preeclampsia 2、13-15 , intrauterine growth retardation 4 and preterm birth 8 such pregnancy or placenta-related disorders, as well as non-invasive tests for fetal chromosomal aneuploidy 5、7、16Clinical application examples have been found. From such developments, RNA biomarkers are regarded as promising for the molecular evaluation of prenatal disorders.

[0012] Despite the view that plasma RNA analysis is promising for prenatal testing, the number of pregnancy- or placenta-related transcripts established in maternal plasma so far is very small. Regarding these, experiments on plasma RNA biomarkers have been performed by reverse transcription (RT)-PCR. 1、4、8、10、12、13 Although RT-PCR is a highly sensitive method, usually only a small number of RNA species can be targeted in one analysis. It should be noted that these studies were premised on the fact that most pregnancy-related genes are derived from the placenta, and mainly focused on genes with relatively higher expression levels in the placenta than in maternal blood cells. From such an approach, pregnancy-related RNAs with high expression levels in the placenta have been identified, but there is a possibility of missing other important targets. Furthermore, considering the low concentration and incomplete integrity of plasma RNA 9、15 standard high-throughput methods, such as SAGE method and microarray analysis, cannot be said to be optimal for plasma transcriptome experiments.

[0013] The above-mentioned limitations regarding plasma RNA analysis may be resolved by utilizing massively parallel sequencing (MPS) for RNA analysis, that is, by RNA-sequencing (RNA-seq). 17、18 Due to its high sensitivity and wide dynamic range, RNA-sequencing has been used for gene expression analysis in many human tissues including the placenta. 19 Furthermore, 20 healthy individuals 21、22 and pregnant women 23 The superiority of MPS has been shown by its feasibility when directly profiling plasma miRNAs. Nevertheless, a comprehensive plasma transcriptome has not been performed yet, which may be due to the lower stability of RNA species longer than short miRNAs in plasma.

[0014] In this study, we showed that by using RNA sequencing to analyze fetal-specific and maternal-specific single nucleotide polymorphisms (SNPs), fetal-derived and maternal-derived transcripts in maternal plasma can be detected, and their relative contribution rates can be predicted. In addition, it is possible to monitor the allelic-specific expression patterns of the placenta using maternal plasma. We further showed that by directly testing maternal plasma before and after childbirth, pregnancy-related transcripts can be identified. II. Diagnostic method for pregnancy-related disorders based on SNP allele ratio

[0015] Analyzing the gene expression profile is useful for detecting an individual's disease status. Aspects of the present invention include methods for analyzing individual expression profiles by analyzing a mixture of RNA molecules derived from two different individuals. This method utilizes the relative amounts of alleles, one allele specific to one of the individuals and alleles common to both individuals. Based on the relative amounts of the common and individual-specific alleles, the gene expression profiles of both individuals can be determined. One application example of this method is to analyze the RNA in a maternal plasma sample containing RNA from both the fetus and the pregnant woman to analyze the fetal gene expression profile. Another application example is the analysis of the gene expression profile of donor-derived RNA in a plasma sample containing RNA from both the donor and the recipient collected from a transplant recipient.

[0016] In a mixture containing RNA from two individuals, the individual expression profiles cannot be determined by tests based on the analysis of the total amounts of various RNA transcripts contained in the mixture. This is because it is difficult to determine the relative contribution rates of RNA from each individual to the total RNA. The relative amounts of various RNA transcripts may be useful for monitoring an individual's health status or detecting a medical condition.

[0017] In this method, the genotypes of two individuals can be determined first by performing either direct genotyping or family analysis on an individual. For example, if the genotypes of both parents are AA and TT respectively, the genotype of the fetus will be AT.

[0018] The following hypothetical examples illustrate the principle of this method. There are SNPs in the coding regions of each of the target genes, and polymorphisms can be observed in the RNA transcripts of various genes. For example, assuming that the genotypes of the fetus and the pregnant woman are AB and AA respectively, the B allele is fetus-specific, and the A allele is common to both the fetus and the mother. Table 1 shows the relative expression levels of various genes in the placenta and maternal blood cells.

[0019] In this example, it is assumed that the contribution rates of the placenta and maternal blood cells are approximately equal, and 2% of each RNA transcript is contained in the maternal plasma. In other words, if the expression level of a certain gene in the placenta is 10,000, 200 of the RNA transcripts of this gene contribute to the maternal plasma. As shown in Figure 1, there is a clear positive correlation between the allele ratio B / A and the relative expression levels of each gene contained in the placenta and maternal blood cells. In contrast, since the total expression level of the gene is affected by the fluctuations in expression in maternal blood cells, the correlation with the expression in the placenta is low (Figure 2).

[0020]

Table 1

[0021] Another advantage of using allelic ratio analysis is that the expression levels of specific genes in the placenta are normalized against the expression levels in maternal blood cells. Since expression levels can vary widely between different genes and different tissues, this normalization makes comparisons between different genes more accurate and eliminates the need to compare to a reference gene, such as a housekeeping gene. In other words, in standard gene expression analysis, the expression level of a gene in a tissue is measured either relative to a housekeeping gene or as an amount relative to total RNA in the sample. Subsequently, to identify abnormal gene expression, it is common to compare the relative values of the test sample to those of a control sample.

[0022] Here we propose a new approach for determining the expression profile of a gene by normalizing the relative amount of fetal-derived RNA contributing to the maternal contribution rate for the same gene. In pathological conditions where gene expression in the placenta or maternal organs is altered, such as preeclampsia, preterm birth, or maternal diseases (e.g., systemic lupus erythematosus), it is thought that the fetal-to-maternal ratio of that gene also changes compared to a pregnancy without such pathological conditions. This approach can be implemented by comparing fetal-specific SNP alleles to common alleles in RNA transcripts contained in maternal plasma.

[0023] This approach can also be implemented by comparing maternal-specific SNP alleles to common alleles for various RNA transcripts in maternal plasma. A pathological condition can be identified if one or more of such allelic ratios of one or more RNA gene transcripts are changed compared to allelic ratios expected in the absence of a pathological condition. By using such allelic ratios between multiple loci or among multiple loci, a pathological condition can be identified. A non-pathological condition can be represented by allelic ratios during normal pregnancies that are obtained before or after the test and are no longer in a pathological state, or by existing data obtained from normal pregnancies up to that point, i.e., the reference range obtained so far.

[0024] Pregnancy-related disorders that can be diagnosed by this method include any disorders characterized by abnormal relative gene expression levels in maternal and fetal tissues. These disorders include, but are not limited to, preeclampsia, intrauterine growth restriction, invasive placentation, preterm birth, neonatal hemolytic disease, placental insufficiency, fetal hydrops, fetal malformations, HELLP syndrome, systemic lupus erythematosus, and other maternal immune disorders. This method discriminates between RNA molecules in a sample containing a mixture of RNA molecules derived from the mother and the fetus as being either mother-derived or fetus-derived. Thus, this method can identify changes in the contribution rate of an individual (i.e., the mother or the fetus) to a specific locus or a specific gene contained in the mixture, even if the contribution rate of another individual does not change or changes in the opposite direction. Such changes cannot be easily detected when measuring the overall expression level of the gene, regardless of the tissue or individual tissue. The pregnancy-related disorders discussed herein are not characterized by chromosomal abnormalities in the fetus, such as aneuploidy.

[0025] This method can be carried out even if the genotype information of the placenta has not been obtained in advance. For example, the genotype of a pregnant woman can be determined by directly subjecting blood cells to genotype analysis. Subsequently, the RNA transcripts of a plasma sample can be analyzed, for example, by, but not limited to, next-generation sequencing. RNA transcripts showing two alleles can be identified. In this combination of transcripts, when the mother is homozygous, the fetus shows a heterozygous form of a fetus-specific allele and a maternal allele. In such a situation, allelic ratio analysis of RNA can be carried out even if the genotype information of the fetus (or placenta) has not been obtained in advance. In addition, cases where the mother is heterozygous but the fetus is homozygous can also be identified by deviation from a 1:1 ratio. Such techniques are described in more detail in U.S. Patent No. 8,467,976. A. Diagnostic method

[0026] A method 300 for diagnosing pregnancy-related disorders according to some embodiments is shown in FIG. 3. In this method, a sample collected from a female subject who is pregnant with a fetus is used. The sample may be maternal plasma and may contain a mixture of RNA molecules derived from the mother and the fetus. The sample can be desirably obtained from a female subject at any stage of pregnancy. For example, pregnant women in the first, second, and third trimesters who are pregnant with only one fetus can be targeted. Pregnant women in the first and second trimesters can be classified as "early pregnancy cases", and pregnant women in the third trimester as "late pregnancy cases". In some embodiments, the sample may be collected after the fetus is born, or samples may be collected before and after birth from the same subject. A sample obtained from a non-pregnant woman can be used as a control. The sample can be obtained at any time after birth (e.g., 24 hours later).

[0027] The sample may be obtained from peripheral blood. The maternal blood sample can be processed, for example, by centrifuging each blood cell from the plasma as desired. A stabilizer may be added to the blood sample or a portion thereof, and the sample can also be stored until use. In some embodiments, samples of fetal tissue, such as chorionic villi, amniotic fluid, or placental tissue, are also collected. The fetal tissue can be used to determine the genotype of the fetus as described below. The fetal tissue can be collected before or after birth.

[0028] Once the sample is obtained, RNA or DNA can then be extracted from the maternal blood cells and plasma, as well as from any fetal tissue such as the placental sample. The RNA samples of the placenta and blood cells may be pretreated, for example, with the Ribo-Zero Gold kit (Epicentre) to remove ribosomal RNA (rRNA) before creating a library for sequencing.

[0029] In step 301 of the method, multiple reads are obtained. Reads are obtained by analyzing RNA molecules obtained from a sample. In various embodiments, reads can be obtained by sequencing, digital PCR, RT-PCR, and mass spectrometry. Although the discussion focuses on sequencing, the aspects described are also applicable to other techniques for obtaining reads. For example, using probes (including primers) for each allele, insights regarding various alleles at a specific locus can be obtained from digital PCR (such as microfluidic PCR and droplet PCR). The probes can be labeled with various existing dyes, for example, so that they can be distinguished from each other. In the same experiment or in another experiment (e.g., on another slide or chip), probes with various labels can be made to target different loci. By detecting such labels, the presence and / or amount of RNA or cDNA molecules that are amplified and correspond to RNA or cDNA molecules can be known. In this case, such reads provide sequence information regarding the RNA or cDNA molecules at the locus corresponding to the probe. The description regarding "sequence reads" is similarly applicable to reads obtained by any suitable technique including non-sequencing techniques.

[0030] Array determination can be carried out as desired by some applicable technology. Examples of techniques and methods for determining nucleic acid sequences include ultra - parallel sequencing, next - generation sequencing, whole - genome sequencing, exome sequencing, sequencing with or without target enrichment, sequencing by synthesis (e.g., sequencing by amplification, clone amplification, bridge amplification, sequencing using reversible terminators), sequencing by ligation, sequencing by hybridization (e.g., microarray sequencing), single - molecule sequencing, real - time sequencing, nanopore sequencing, pyrosequencing, semiconductor sequencing, sequencing by mass spectrometry, shotgun sequencing, and Sanger sequencing. In some embodiments, the sequencing is carried out by an ultra - parallel technique on a cDNA library prepared from RNA contained in the sample. The cDNA library can be synthesized as desired, for example, using the mRNA - Seq Sample Preparation kit (Illumina) according to the manufacturer's instructions or with minor modifications thereto. In some embodiments, 5 - fold diluted Klenow DNA polymerase is used in the end - repair step of plasma cDNA. The QIAquick PCR Purification kit and QIAquick MinElute kit (Qiagen) can be used to purify the end - repaired product and the adenylated product, respectively. In some embodiments, 10 - fold diluted paired end adapters are used for the plasma cDNA sample, or the purification step of the adapter - ligated product is carried out twice using AMPure XP beads (Agencourt). By using the HiSeq 2000 instrument (Illumina), the sequence of the cDNA library in paired - end form can be determined in 75 - bp increments. Other formats or instruments can also be used.

[0031] In step 302, the positions of the reads in the reference array are determined. In a technique using PCR, for example, the positions can be determined by correlating with the color of the detected label attached to a probe or primer for the sequence (the probe or primer is specific to the sequence). In a sequencing technique, the positions can be determined by aligning the sequence of the read with the reference sequence. The alignment of the sequences can be performed by a computer system. For the processing of raw data (e.g., removal of highly redundant reads or poor-quality reads), alignment of data, and / or normalization of data, some bioinformatics-related processing steps can be applied. The levels of RNA transcripts in maternal blood cells, placental tissue, and maternal plasma can be calculated as the number of reads per kilobase fragment per million reads mapped (FPKM). Data processing and alignment can also be performed when determining the allele ratio, but data normalization is not necessary. Alignment can be performed using any reference sequence (e.g., the hg19 reference human genome) and any algorithm.

[0032] In one example of performing steps 301 and 302, after removing overlapping reads and rRNA reads, 3 million analyzed reads were obtained on average for plasma samples of non-pregnant women; 12 million analyzed reads were obtained on average for plasma samples of each pregnant woman. In tissue RNA-sequencing, on average 173 million and 41 million analyzed reads were obtained per sample for placenta and blood cells, respectively. The statistics for aligning RNA-sequencing are summarized in FIGS. 4 and 5. For all samples, the GC content of the sequenced reads was also examined (FIGS. 6A and Table 2).

[0033] [Table 2]

[0034] In step 303, one or more informative loci are identified. In this step, for a certain locus contained in the genome, the genotypes of the female subject and the fetus can be determined or inferred, and these genotypes can be compared. If the genotype is homozygous for the first allele in the first entity (e.g., AA) and heterozygous for the first and second alleles in the second entity (e.g., AB), then that locus is considered an informative locus. The first entity may be the pregnant subject or the fetus, and the second entity is the other one of the pregnant subject and the fetus that is different from the first entity. In other words, for each informative locus, one individual (either the mother or the fetus) is homozygous and the other individual is heterozygous.

[0035] Informative loci can be further classified according to which individual is heterozygous and which individual contributes to one allele. When the fetus is homozygous and the pregnant subject is heterozygous, that locus is considered an informative maternal locus, or similarly a maternally specific locus. When the pregnant subject is homozygous and the fetus is heterozygous, that locus is considered an informative fetal locus, or similarly a fetus-specific locus. Since the same individual serves as the second entity for all these loci among all the loci identified in step 303, they are all either maternally specific loci or fetus-specific loci.

[0036] An informative locus may represent a single nucleotide polymorphism, i.e., "SNP", where the first and second alleles differ with respect to a single nucleotide. An informative locus may also represent a short insertion or deletion where one allele has one or more nucleotides inserted or deleted compared to the other allele.

[0037] In some embodiments, the genotype of one or more informative loci or each informative locus for a pregnant female is determined by sequencing genomic DNA obtained from maternal tissue. The maternal tissue may be maternal blood cells or any other type of tissue. In some embodiments, the genotype of one or more informative loci or each informative locus for a fetus is determined by sequencing genomic DNA obtained from fetal tissue, such as placenta, chorionic villi, amniotic fluid. When a less invasive method is preferred, fetal DNA for sequencing can also be obtained from a maternal blood sample. The same sample used to obtain the RNA molecules for sequencing may be the source of such fetal DNA, or a separate sample can also be used. Of course, informative fetal loci can be identified without direct genotyping of the fetus. For example, for the first allele of a certain locus, if it is determined that a pregnant female subject is homozygous, and the results of RNA sequencing indicate the presence of a second allele in the mixture of maternal and fetal RNA, it can be inferred that the fetus is heterozygous for this locus.

[0038] In some embodiments, both the mother and the fetus are subjected to genotyping by a massively parallel method to identify informative loci. Sequencing can be performed using genomic DNA samples of maternal blood cells with enriched exomes and genomic DNA of the placenta. Genomic DNA can be extracted from placental tissue and maternal blood cells using QIAamp DNA kits and QIAamp Blood kits (both from Qiagen) according to the instructions of each manufacturer. Enrichment of exomes and preparation of sequencing libraries can be carried out using the TruSeq Exome Enrichment kit (Illumina) according to the protocol provided by the manufacturer. The library can be sequenced in 75bp increments in PE format using the HiSeq 2000 instrument (Illumina).

[0039] In step 304, one or more informative loci are selected. As used herein, "selecting" means choosing a particular informative locus from among the informative loci identified in step 303 for subsequent analysis. An "selected" informative locus is a locus so chosen. The selection can be made based on one or more criteria. The fact that the informative locus is located in the genome or reference sequence is one of such criteria. In some embodiments, only loci located in exons or expressed regions of the reference sequence are further analyzed.

[0040] Another criterion for selection is to select by the number of sequence reads from among all loci obtained in step 301, aligned to the reference sequence in step 302, and aligned to the locus and containing each allele. In some embodiments, the selected informative locus must be associated with at least a predetermined first number of sequence reads containing the first allele and / or a predetermined second number of sequence reads containing the second allele. The predetermined number can be 1, 2 or more, and can also correspond to the desired read quality or sequencing depth. As a result of the selection, one or more selected informative loci are identified.

[0041] In some embodiments, only the informative loci corresponding to the SNPs are selected. In one example of the method, approximately 1 million SNPs contained in the NCBI dbSNP Build 135 database, all of which are located in exons, were analyzed, and using the genotypes of the mother and fetus, informative SNPs were identified from these approximately 1 million SNPs. In the figure of the informative SNPs, "A" represents the common allele and "B" represents the allele specific to the mother or fetus. Such informative SNPs included in the mother-specific SNP alleles are "AA" in the fetus and "AB" in the mother for a given locus. Also, these informative SNPs that include the fetus-specific SNP allele are "AA" in the mother and "AB" in the fetus for a given locus. The genotypes were subjected to a series of bioinformatics data processing pathways (pipelines) in the laboratory. For the selection, only the informative SNPs in which both the "A" allele and the "B" allele each appeared as at least one read count were included in the analysis.

[0042] In steps 305 and 306, the sequence reads aligned to each of the selected informative loci are counted and classified according to which allele they contain. For each of the selected informative loci, two numbers are determined: the first number is the number of sequence reads that are aligned to the locus and contain the corresponding first allele (i.e., the common allele); the second number is the number of sequence reads that are aligned to the locus and contain the corresponding second allele (i.e., the mother- or fetus-specific allele). The first and second numbers are calculated by computer as desired. The sum of the first and second numbers for each locus is the total number of sequence reads aligned to the locus. The ratio of the first number to the second number for a particular selected informative locus can be expressed as the "allele ratio".

[0043] In steps 307 and 308, the sum of the number of the first alleles and the number of the second alleles of the selected informative loci is calculated. In step 307, the first sum of the first numbers is calculated. The first sum represents the total number of sequence reads containing the corresponding first allele (i.e., the common allele) for the selected informative loci. In step 308, the second sum of the second numbers is calculated. This second sum corresponds to the total number of sequence reads containing the corresponding second allele (i.e., the maternal or fetal specific allele) for the selected informative loci. This sum is calculated by a computer as desired, using all or some of the selected informative loci identified in step 304. In some embodiments, this sum is weighted. For example, the contribution rate of loci from a particular chromosome to the sum is amplified or reduced by multiplying scalars to the first and second numbers for these loci.

[0044] In step 309, the ratio of the first sum to the second sum is calculated. This ratio represents the total logarithm of the number of sequence reads of the first allele and the second allele aggregated at the selected informative loci. In some embodiments, this ratio is simply calculated by dividing the first sum by the second sum. From such a calculation, the number of sequence reads for the common allele is obtained as the product of the sequence reads of the maternal or fetal specific allele. Alternatively, the ratio can also be calculated by dividing the second sum by the total of the first sum and the second sum. In this case, this ratio represents the number of sequence reads of the maternal or fetal specific allele as a proportion in all sequence reads. Other calculation methods for the ratio are also obvious to those skilled in the art. The first sum and the second sum used to calculate the ratio may also serve as a proxy for the amount of RNA (i.e., transcript) contained in the samples derived from the common and specific alleles for the selected informative loci.

[0045] In step 310, the ratio calculated in step 309 is compared with a cut-off value to determine whether the fetus, mother, or pregnant woman has a pregnancy-related disorder. In various embodiments, the cut-off value can be determined from one or more samples obtained from pregnant women who do not have a pregnancy-related disorder (i.e., control subjects), and / or from one or more samples obtained from pregnant women who have a pregnancy-related disorder. In some embodiments, the same method as described above is performed on the samples obtained from the control subjects, and the cut-off value is determined by analyzing the same or overlapping set of selected informative loci. The cut-off value can be set to a value between the value predicted for pregnancies without the disorder and the value predicted for pregnancies with the disorder. The cut-off value may be based on a statistical difference from a normal value.

[0046] In some embodiments, the ratio calculated for a pregnant woman subject (step 309) is calculated using selected informative loci included in a specific group of genes, and the ratio (i.e., the cut-off value) is calculated for the control subjects using some or all of the same selected informative loci. If the ratio calculated for the pregnant woman subject is based on a maternal-specific allele, the cut-off value calculated for the control subjects may also be based on a maternal-specific allele. Similarly, both the ratio and the cut-off value may be based on a fetal-specific allele.

[0047] The comparison of the ratio and the cut-off value can be performed as desired. For example, a disorder can be diagnosed if the ratio exceeds or is below the cut-off value by any amount or a specific excess. The comparison may involve calculating a difference or another ratio between the pregnant woman subject and the cut-off value. In some embodiments, the comparison includes evaluating whether there is a significant difference between the pregnant woman subject and the control subjects with respect to common and non-common alleles of a specific gene.

[0048] If desired, this method can be used to predict the proportion of maternal or fetal-derived RNA in a sample. This prediction is carried out by multiplying the ratio calculated in step 309 by a scalar. The scalar represents the total expression level of selected informative loci in a heterozygous individual as a relative value to the expression of the second allele.

[0049] The following example illustrates the application method of the scalar. If the second allele counted by the method is fetal-specific, the ratio calculated in step 309 may indicate the number of sequence reads of the fetal-specific allele as the proportion of all sequence reads of the selected informative loci. Among the sequence reads containing the common allele, some are from the mother and some are from the fetus. If the relative expression levels of the fetal-specific allele and the common allele in the fetus are known or predictable, the ratio can be amplified or reduced to predict the fetal-derived proportion relative to all sequence reads of the selected informative loci.

[0050] In some embodiments, the fetal allele and the common allele are expressed contrastingly at most loci, and the scalar is predicted to be approximately 2. In other embodiments, the scalar deviates from 2 to take into account asymmetric gene expression 24 . Therefore, the maternal-derived proportion relative to the sequence reads is 1 minus the fetal-derived proportion. The same calculation can be performed when the second allele of the selected informative locus is maternal-specific.

[0051] B. Examples 1. Correlation between allele ratio and plasma concentration The expression profiles of maternal plasma samples (at 37 weeks and 2 days of gestation) were analyzed and compared with the corresponding placental and maternal blood cells of the pregnant women. RNA ultra - parallel sequencing was performed for each of these samples using the Illumina HiSeq 2000 instrument. For the placenta and maternal blood cells, genotypes were analyzed by exome sequencing by ultra - parallel sequencing. Analyses were performed for approximately 1 million SNPs included in the NCBI dbSNP Build 135 database and located in exons, and informative SNPs were classified where the mother was homozygous (genotype AA) and the fetus was heterozygous (genotype AB). Thus, the A allele is the allele common to the mother and the fetus, and the B allele is fetal - specific.

[0052] The RNA - SNP allele ratios and total plasma concentrations of various RNA transcripts in maternal plasma were plotted against the relative tissue expression levels (placenta / blood cells) (Figure 7). A clear correlation was found between the RNA - SNP allele ratio in plasma and the relative tissue expression levels (Spearman R = 0.9731679, P = 1.126e - 09). However, no correlation was found between the total plasma levels and tissue expression (Spearman R = - 0.7285714, P = 0.002927).

[0053] In addition to fetal expression analysis, this method can be used to profile maternal expression. Under this condition, informative SNPs where the mother is heterozygous (genotype AB) and the fetus is homozygous (genotype AA) can be used. Next, the RNA - SNP allele ratio can be calculated by dividing the maternal - specific allele count by the count of the common allele. Figure 8 plots the RNA - SNP allele ratios and total plasma levels of various genes containing informative SNPs against the relative tissue expression levels (blood cells / placenta). A clear positive correlation was found between the tissue expression levels and the RNA - SNP allele ratios in maternal plasma (Spearman R = 0.9386, P < 2.2e - 16), but no correlation was found between the tissue expression levels and the total plasma transcript levels (Spearman R = 0.0574, P = 0.7431).

[0054] The RNA transcripts, their RNA-SNP allele ratios, and plasma levels are shown as tables in FIGS. 9 and 10. 2. Comparison of RNA-SNP Allele Ratios in Maternal Plasma for Preeclampsia and Control Pregnancy Cases

[0055] For two pregnant female subjects, on average 117,901,334 unprocessed fragments were obtained (Table 3A). The RNA-SNP allele ratios for some of the blood transcripts retaining the maternal-specific SNP alleles were compared between 5641 preeclampsia (PET) cases that developed in the third trimester of pregnancy and 7171 control cases with the same gestational age. As shown in FIG. 11A, the RNA-SNP allele ratios of these transcripts depicted different profiles in PET cases and control cases. Importantly, as shown in FIG. 11B, the profile of the RNA-SNP allele ratio was clearly different from the profile of the plasma transcript level. This suggests that the RNA-SNP allele ratio analysis of maternal plasma may be available as another more accurate measure to distinguish PET cases from normal pregnancy cases. As illustrated by the comparison between 5641 PET cases and 9356 control cases of another normal pregnancy, with the increase in the number of informative SNPs, the number of informative transcripts may further increase (FIGS. 12A and 12B). The RNA-SNP allele ratios and plasma levels of the RNA transcripts shown in FIGS. 11 and 12 are summarized in the tables of FIGS. 13 and 14, respectively.

[0056] [Table 3]

[0057] [Table 4]

[0058] III. Method for Determining Fetal or Maternal RNA Contribution Rate A. Method Figure 15 shows method 1500 for determining the proportion of fetal-derived RNA in samples from pregnant women carrying a fetus. The samples may be collected and prepared in the same manner as method 300 described above.

[0059] In step 1501, a plurality of sequence reads are obtained. In step 1502, these sequence reads are aligned with a reference sequence. These steps can be carried out in the same manner as steps 301 and 302 described above.

[0060] In step 1503, one or more informative maternal loci are identified. Each informative maternal locus is homozygous for the corresponding first allele in the fetus and heterozygous for the corresponding first allele and the corresponding second allele in the pregnant female subject. Thus, the first allele is an allele common to the mother and the fetus, and the second allele is maternal-specific. The informative maternal loci can be identified in the same manner as described for step 303 above.

[0061] In step 1504, in the same manner as step 304, the informative maternal loci are selected to identify one or more selected informative maternal loci. Then, in steps 1505 to 1511, the sequence reads are processed at the level of each selected informative maternal locus. In steps 1505 and 1506, a first number and a second number are determined for each of the selected informative maternal loci. In step 1505, the first number is determined as the number of sequence reads that are aligned with the locus and contain the corresponding first allele. The second number is determined in step 1506 as the number of sequence reads that are aligned with the locus and contain the corresponding second allele. Steps 1505 and 1506 can be carried out in the same manner as steps 305 and 306 described above.

[0062] In step 1507, for each informative maternal locus, the sum of the first number and the second number is calculated. This sum may be equal to the total number of sequence reads aligned to the locus. In step 1508, the maternal ratio is calculated by dividing the second number by the sum. The maternal ratio is the proportion of all sequence reads aligned to a particular locus that contain the second (maternal-specific) allele.

[0063] In step 1509, a scalar is determined. The scalar represents the total expression level of the selected informative maternal loci in a pregnant female subject when compared to the corresponding second allele. If the expression in a pregnant female subject of two alleles is known or predicted to be symmetric, the scalar is approximately 2. In other embodiments, the scalar may deviate from 2 and account for asymmetric gene expression. Since most loci are expressed in a balanced manner, the prediction of symmetry is confirmed.

[0064] In step 1510, the maternal contribution rate is obtained by multiplying the maternal ratio by the scalar. The maternal contribution rate represents the proportion of sequence reads (or, by extension, the proportion of RNA in the sample) derived from the mother at the selected informative maternal loci. Subtracting the maternal contribution rate from 1 gives the fetal contribution rate, which is the proportion of sequence reads derived from the fetus. In step 1511, the fetal contribution rate for the selected informative maternal loci is calculated.

[0065] In step 1512, the proportion of fetal-derived RNA in the sample is determined. This proportion is the average of the fetal contribution rates for the selected informative maternal loci. In some embodiments, this average is weighted, for example, by the sum calculated for the selected informative maternal loci. By calculating a weighted average, the proportion determined in step 1512 may reflect the relative expression levels of various loci or genes in the sample.

[0066] The proportion of fetal-derived RNA in a sample can be calculated for one, most, or all of the selected informative maternal loci for which sequencing data are available. Of course, this proportion reflects the blood transcriptome at the time the sample was taken from a female subject during a particular pregnancy. If samples are obtained from the subject at various time points during pregnancy, the combination of selected informative maternal loci identified by method 1500 may vary from sample to sample, and similarly, the proportion of fetal-derived RNA determined for these loci may also vary. The proportion of fetal-derived RNA in a sample may vary from subject to subject, even when controlling for gestational age and other factors, and may also vary between subjects with pregnancy-related disorders and healthy subjects. Thus, such disorders can be diagnosed by comparing the proportion to a cut-off value.

[0067] In some embodiments, samples from one or more healthy subjects are used to perform method 1500 to determine a cut-off value. In some embodiments, a diagnosis is made when the proportion exceeds or is below the cut-off value by any amount or a particular margin. The comparison may involve calculating a difference or calculating a ratio of the proportion of fetal-derived RNA in a pregnant female subject to the cut-off value. In some embodiments, the comparison involves assessing whether there is a significant difference between the proportion of fetal-derived RNA in a pregnant female subject and the proportion seen in healthy pregnant females with equivalent gestational ages.

[0068] By analyzing informative fetal loci, it is also possible to determine the proportion of fetal-derived RNA in a sample. For each of the informative fetal loci, this proportion can be predicted by first calculating a fetal fraction (the value obtained by dividing the number of sequence reads containing the second (fetal-specific) allele by the total number of sequence reads). Then the fetal contribution rate is the product of a scalar representing the relative expression levels of those two alleles in the fetus and the fetal fraction. Thus, it is possible to implement a variant of method 1500 by identifying informative fetal loci, selecting those loci, calculating the fetal contribution rate for each of the selected loci, and determining the average of the fetal contribution rates of the selected loci. If desired, determining the proportion of fetal-derived RNA by method 1500 may include determining the average of the fetal contribution rate calculated for the selected informative maternal loci and the fetal contribution rate calculated for the selected informative fetal loci. To reflect differences in the number of sequence reads for various loci, the average calculated here can be weighted so that maternal or fetal-specific loci are weighted more heavily, or as desired. B. Examples 1. Identification and prediction of fetal-derived and maternal-derived transcripts in maternal plasma

[0069] Based on the data regarding genotypes, informative genes defined as genes having at least one informative SNP were first identified. For two cases in the first trimester of pregnancy, a total of 6,714 and 6,753 informative genes were available respectively for analyzing the relative proportions of fetal and maternal contribution rates (Figures 16 and 17). For two cases in the third trimester of pregnancy, a total of 7,788 and 7,761 informative genes were available respectively for analyzing the relative proportions of fetal and maternal contribution rates. To measure the relative proportion of fetal contribution rate in maternal plasma, RNA transcripts in which at least one RNA-sequencing read covering a fetal-specific allele was present in the maternal plasma sample were classified. For each of the first trimester and the third trimester of pregnancy, the proportions of such fetal-derived transcripts were 3.70% and 11.28% respectively. Using a similar approach, the relative proportion of maternal contribution rate in blood was examined using maternal-specific SNP alleles, which was predicted to be 76.90% and 78.32% respectively in the first trimester and the third trimester (Figure 18A). 2. Comparison of fetal and maternal contribution rates to RNA-SNPs in PET and control pregnancy cases

[0070] The proportions of fetal and maternal contribution rates to the GNAS transcript were examined from both PET (preeclampsia) case 5641 and control case 7171 with its gestational age combined. The GNAS transcript had an informative SNP site containing a fetal-specific SNP allele at locus rs7121 in case 5641; a maternal-specific SNP allele was present at the same SNP site in case 7171. The proportions of fetal and maternal contribution rates to this SNP site in case 5641 were calculated to be 0.09 and 0.91 respectively. On the other hand, the proportions of fetal and maternal contribution rates to the same SNP site in case 7171 were 0.08 and 0.92 respectively (Figure 19). Compared with control case 7171, in this transcript, the proportion of fetal contribution rate increased by 12.5% and the proportion of maternal contribution rate decreased by 1.09%. Interestingly, when comparing case 5641 and 7171, it was detected that the FPKM of this transcript increased by 21%. IV. Method for Designating Genomic Loci as Maternal or Fetal Markers A. Method

[0071] A method 2000 for designating a genomic locus as a maternal or fetal marker is shown in FIG. 20. This method involves analyzing a sample obtained from a female subject who is pregnant with a fetus. The sample may be maternal plasma and may contain a mixture of RNA molecules derived from the mother and RNA molecules derived from the fetus. The sample can be collected and prepared in the same manner as the above-described method 300.

[0072] In step 2001, a plurality of sequence reads are obtained. In step 2002, these sequence reads are aligned with a reference sequence. These steps can be carried out in the same manner as those described above for steps 301 and 302.

[0073] In step 2003, one or more informative loci are identified. Each locus is homozygous for the corresponding first allele in a first entity and heterozygous for the corresponding first allele and the corresponding second allele in a second entity. The first entity is a pregnant female subject or a fetus, and the second entity is the other of the pregnant female subject and the fetus that is different from the first entity. The informative loci can be identified in the same manner as the above-described step 303.

[0074] In step 2004, the informative loci are screened in the same manner as step 304, and one or more screened informative loci are identified. In some embodiments, only the informative loci that are SNPs are subjected to screening. In an example of the method, an informative SNP in which at least one sequence read containing the "A" allele and one sequence read containing the "B" allele are present in maternal plasma is included in the screened informative loci.

[0075] Next, in steps 2005 - 2008, the array reads are processed at the level of the individual selected informative loci. In steps 2005 and 2006, a first number and a second number are determined for each of the selected informative loci. In step 2005, the first number is determined as the number of array reads that are aligned to the locus and contain the corresponding first allele, and in step 2006, the second number is determined as the number of array reads that are aligned to the locus and contain the corresponding second allele. Steps 2005 and 2006 can be carried out in the same manner as steps 305 and 306 discussed above.

[0076] In step 2007, for each of the selected informative loci, the ratio of the first number to the second number is calculated. This ratio represents the relative amounts of the array reads (and their extensions, the transcripts contained in the sample) that contain the first and second alleles. In some embodiments, this ratio is simply calculated by dividing the first number by the second number. Such a calculation yields the number of array reads of the common allele as a multiple of the array reads of the maternal or fetal - specific allele. This ratio, or its reciprocal, can be considered the allele ratio, and for SNP loci, it can be regarded as the RNA - SNP allele ratio.

[0077] In some embodiments, for informative SNPs containing a maternal - specific SNP allele, the RNA - SNP allele ratio for each SNP is calculated as maternal - specific allele:common allele. On the other hand, for informative SNPs containing a fetal - specific SNP allele, the RNA - SNP allele ratio can be calculated as fetal - specific allele:common allele. Unlike gene expression analysis where data normalization is essential, data normalization is not essential for RNA - SNP allele ratio analysis, and the introduced bias is also small. For transcripts containing two or more SNPs, the RNA - SNP allele ratio can be calculated for each informative SNP site, and the average RNA - SNP allele ratio per transcript can be calculated by computer.

[0078] In Project 2007, among other things, a ratio can also be calculated by dividing a second number by the sum of the first number and the second number. In this case, the ratio is the proportion of the sequence reads of the maternal or fetal-specific alleles in all sequence reads related to that locus. In terms of the "A" and "B" alleles, the ratio is B allele ratio = B allele count / (A allele count + B allele count) and can be expressed as such. Theoretically, assuming no allele-specific expression in plasma transcripts derived from only either the fetus or the mother, the B allele ratio should be 0.5. Other methods for calculating the ratio will also be apparent to those skilled in the art.

[0079] In Project 2008, when the ratio exceeds the cut-off value, the selected informative locus is designated as a marker. In some embodiments, the cut-off value is from about 0.2 to about 0.5. In some embodiments, the cut-off value is 0.4. In an example of the method, the cut-off value for the B allele ratio for RNA transcripts with a high contribution rate from the fetus or the mother was set at 0.4 or higher. This cut-off value took into account the Poisson distribution of RNA-sequencing reads and random sample extraction. In some embodiments, when the ratio exceeds the cut-off value, it can be said that the contribution rate of the "B" allele is high.

[0080] A high allele ratio at a particular informative locus may indicate (i) that in a heterozygous individual, the expression of the second, i.e., "B" allele, is higher than that of the "A" allele (the alleles are asymmetrically expressed), (ii) that most of the total RNA aligned to the locus contained in the maternal plasma is derived from heterozygous individuals, or both (i) and (ii). A high allele ratio may also indicate a pregnancy-related disorder if the gene associated with the informative locus is pathologically overexpressed. Thus, some aspects of the present method also include diagnosing a pregnancy-related disorder based on whether the selected informative locus can be designated as a marker for a second entity, i.e., a heterozygous individual. In making such a diagnosis, the cut-off value used for comparison with the allele ratio may be based on the transcript level of a plasma sample obtained from a healthy pregnant subject.

[0081] In some aspects, method 2000 can also be used to predict the proportion of RNA (derived from the mother or the fetus) of a particular selected informative locus contained in a sample. Such prediction is made by multiplying the ratio calculated in step 2007 by a scalar. The scalar represents the total expression level at that locus in a heterozygous individual as a relative value to the expression of the second allele.

[0082] For purposes of illustration, if the second (“B”) allele counted by the method is fetal-specific, then the B allele ratio calculated as described above represents the number of sequence reads of this allele as a proportion of the total sequence reads of the selected informative locus. Some of the sequence reads containing the common (“A”) allele are maternal in origin and some are fetal in origin. If the relative expression levels of the fetal-specific and common alleles in the fetus are known or predictable, the B allele ratio can be amplified or reduced to predict the proportion of the fetal contribution rate to the total sequence reads of that locus. If the “A” and “B” alleles are equally expressed in the fetus, the scalar is approximately 2. Therefore, the proportion of the maternal contribution rate to the sequence reads of that locus is 1 minus the proportion of the fetal contribution rate. Similar calculations can be performed when the “B” allele of the selected informative locus is maternal-specific.

[0083] B. Examples: High Fetal and Maternal Contribution Rates As described above, when the cut-off value of the B allele is set to 0.4, it was found that 0.91% of the blood transcripts had a high fetal contribution rate in the early pregnancy (i.e., the first and second trimesters). This percentage increased to 2.52% in the late pregnancy (i.e., the third trimester). Conversely, it was found that the maternal contribution rate was high for 42.58% and 50.98% of the transcripts in the early and late pregnancies, respectively (Figs. 16, 17, and 18A).

[0084] V. Use of Genes for Diagnosing Pregnancy-Related Disorders A. Pregnancy-Related Genes A method for identifying pregnancy-related genes is also provided. This method includes obtaining a plurality of first sequence reads and a plurality of second sequence reads. The first sequence reads are obtained from sequencing of RNA molecules from a plasma sample of a pregnant woman. The second sequence reads are obtained from sequencing of RNA molecules from a plasma sample of a non-pregnant woman. The first sequence reads and the second sequence reads are aligned with a reference sequence to designate a series of candidate genes.

[0085] According to this method, next, using the array reads, the expression level of each candidate gene included in samples obtained from pregnant women and non-pregnant women is determined. Specifically, for each of the candidate genes, a first number corresponding to the candidate gene is determined using a first array read; and a second number of transcripts corresponding to the candidate gene is determined using a second array read. The first number of transcripts and the second number of transcripts may be normalized. Thereafter, the transcript ratio of the candidate gene is calculated by dividing the first number of transcripts by the second number of transcripts. The transcript ratio is compared with a cut-off value. When the transcript ratio exceeds the cut-off value, the candidate gene is identified as a pregnancy-related gene.

[0086] In some embodiments of this method, the normalization of the first number of transcripts corresponds to scaling (amplifying or reducing) the first number of transcripts by the total number of the first array reads. Similarly, the normalization of the second number of transcripts may correspond to scaling by the total number of the second array reads of the second number of transcripts. In other embodiments, the normalization of the first number of transcripts of each candidate gene corresponds to scaling the first number of transcripts of that candidate gene by the total number of the first transcripts for all candidate genes. In some cases, the normalization of the second number of transcripts of each candidate gene corresponds to scaling the second number of transcripts of that candidate gene by the total number of the second transcripts for all candidate genes.

[0087] Basically, in this method, when a gene is expressed at a higher level in a pregnant woman compared to a non-pregnant woman and other genes are at the same level in both, the gene is identified as a pregnancy-related gene. The pregnant woman and the non-pregnant woman may be the same person, that is, samples obtained from an individual before and after childbirth can be used as the materials for the first array read and the second array read, respectively.

[0088] We have demonstrated that several of the blood RNA transcripts carrying fetal-specific alleles completely disappeared from maternal plasma after parturition (Figure 18B). These transcripts in maternal plasma were considered to be fetal-specific. On the other hand, some of the maternal-specific alleles also became undetectable after parturition. Therefore, we identified genes that showed increased expression during pregnancy, which we named pregnancy-related genes, by directly comparing their presence or absence in maternal plasma before and after parturition. Genes that were detected in the plasma of pregnant women in the third trimester of pregnancy and whose plasma levels decreased to less than one-half in all cases after delivery were defined as pregnancy-related genes. Using bioinformatics algorithms for data normalization and analysis of differential gene expression, a list of 131 pregnancy-related genes was compiled (Figure 21). Fifteen of these genes had already been reported to be specific to maternal plasma during pregnancy 1、2、4、5、9、10、12 . By one-step real-time RT-PCR, we further identified five novel pregnancy-related transcripts that were highly abundant in prepartum maternal plasma, namely STAT1, GBP1, and HSD17B1 from an additional 10 plasma samples obtained from women in the third trimester of pregnancy, and KRT18 and GADD45G from 10 plasma samples obtained from another cohort of women in the third trimester of pregnancy (Figure 22).

[0089] To evaluate the association of these 131 genes with pregnancy, hierarchical cluster analysis was performed on all plasma samples. A clear difference was observed between plasma samples obtained from pregnant women (i.e., in the first and third trimesters of pregnancy) and plasma samples obtained from currently non-pregnant women (i.e., non-pregnant controls and postpartum subjects) (Figure 23).

[0090] Interestingly, when comparing the expression patterns in plasma samples obtained from two pregnant women in the third trimester of these 131 genes with the corresponding expression patterns in the placenta and maternal blood cells, close similarities were observed between the placenta and prepartum plasma samples, and between maternal blood cells and postpartum plasma samples (Figure 24A). This observation supports the claim that most pregnancy-related genes are preferentially expressed in the placenta rather than in maternal blood cells. Furthermore, a positive correlation was observed between the expression levels of these pregnancy-related transcripts in the placenta and maternal plasma (P<0.05, Spearman correlation) (Figure 24B).

[0091] B. Differential gene expression in the placenta and maternal blood From direct experiments using maternal plasma samples before and after delivery, a list of a panel of pregnancy-related genes could be created, and at the same time, the RNA-sequencing data on the placenta and blood cells were explored for comparison. As previously reported, assuming that the expression of pregnancy-related genes is high in the placenta and low in maternal blood cells 10、15 , in the analysis using tissues, arbitrarily, a 20-fold difference was set as the minimum value of the cutoff. From this analysis using tissues, a total of 798 candidate genes were obtained, and the ratios of the fetal contribution rate to the maternal contribution rate in maternal plasma were calculated for these genes. Compared with the ratio of the whole transcriptome (Figure 18A), a group of genes with relatively high ratios, mainly derived from the fetus, were identified (Figure 25). However, since genes mainly derived from the fetus could be identified at a higher ratio, the plasma-based strategy was superior to the tissue-based strategy (Figure 25).

[0092] C. Genes related to diseases or disorders This specification also provides a method for identifying genes associated with pregnancy-related disorders. This method includes obtaining a plurality of first sequence reads and a plurality of second sequence reads. The first sequence reads are obtained as a result of determining the sequences of RNA molecules in plasma samples obtained from healthy pregnant women, and the second sequence reads are obtained as a result of determining the sequences of RNA molecules in plasma samples obtained from pregnant women carrying a fetus affected by or suffering from a pregnancy-related disorder. The first sequence reads and the second sequence reads are aligned with a reference sequence to specify a series of candidate genes.

[0093] Next, by this method, the expression levels of each candidate gene contained in the samples obtained from two pregnant women are determined using the sequence reads. Specifically, for each candidate gene, the first number of transcripts corresponding to the candidate gene is determined using the first sequence reads, and the second number of transcripts corresponding to the candidate gene is determined using the second sequence reads. The first number of transcripts and the second number of transcripts may be normalized. Thereafter, the transcript ratio for the candidate gene is calculated by dividing the first number of transcripts by the second number of transcripts. The transcript ratio is compared with a reference value. If the transcript ratio deviates from the reference value, the candidate gene is identified as a gene associated with the disorder.

[0094] In some embodiments of this method, normalizing the first number of transcripts corresponds to scaling the first number of transcripts by the total number of the first sequence reads. Similarly, normalizing the second number of transcripts may correspond to scaling the second number of transcripts by the total number of the second sequence reads. In other embodiments, normalizing the first number of transcripts for each candidate gene corresponds to scaling the first number of transcripts for that candidate gene by the total number of the first transcripts for all candidate genes. In some cases, normalizing the second number of transcripts for each candidate gene corresponds to scaling the second number of transcripts for that candidate gene by the total number of the second transcripts for all candidate genes.

[0095] According to this method for identifying genes associated with pregnancy-related disorders, in some embodiments, the reference value is 1. In some embodiments, the transcript ratio is away from the reference value if the ratio of the transcript ratio to the reference value exceeds or is below a cut-off value. In some embodiments, the transcript ratio is away from the reference value if the difference between the transcript ratio and the reference value exceeds a cut-off value.

[0096] Basically, in this method, when there is a significant difference in the expression level of the gene between pregnant women with disorders and pregnant women without disorders, and the expression of other genes is equivalent, the gene is identified as a pregnancy-related disorder-related gene.

[0097] Diagnosis and monitoring of fetal disorders and pregnancy-related disorders have hitherto been carried out using disease-related RNAs contained in maternal plasma. For example, it has been found that maternal plasma levels of corticotropin-releasing hormone (CRH) mRNA are useful for non-invasive detection and prediction of preeclampsia. 2、3 Detection of interleukin 1 receptor-like 1 (IL1RL1) mRNA in maternal plasma has also been shown to be useful for identifying women who will have spontaneous preterm birth. 8 In addition, searches for growth-related maternal plasma RNA marker panels have been carried out for the purpose of non-invasively evaluating fetal growth and intrauterine growth retardation. 4 In this study, we discuss the fact that novel disease-related blood RNA markers could be identified by directly comparing the transcriptomes of maternal plasma from healthy pregnant women and women with pregnancy complications such as preeclampsia, intrauterine growth retardation, preterm birth, and fetal aneuploidy. To demonstrate the feasibility of this approach, RNA-sequencing was performed on maternal plasma samples obtained from three pregnant women with preeclampsia and seven pregnant women without complications with a matched gestational age. We identified 98 transcripts that showed a significant increase in the plasma of pregnant women with preeclampsia (Figure 26). These newly identified preeclampsia-related transcripts may be useful for disease prediction, prognosis determination, and monitoring.

[0098] By utilizing this technology, preterm birth can be predicted and monitored. This technology can also be used to predict impending fetal death. Furthermore, if the gene in question is transcribed in fetal or placental tissue and the transcript can be detected in maternal plasma, this technology can also be used to detect diseases caused by gene mutations.

[0099] Plasma RNA sequencing can also be applied to other clinical situations. For example, the plasma RNA sequencing method developed in this study has been reported to have abnormal RNA concentrations in plasma 13、14 It may also be useful for the analysis of other pathological conditions such as cancer. For example, by comparing the plasma transcriptomes of cancer patients before and after treatment, tumor-related blood RNA markers for non-invasive diagnostic use may be identified. VI. Allele expression patterns for specific genes

[0100] RNA sequencing has been used to examine allele expression patterns. 25 We inferred that the allele expression pattern of a given gene could be detected from plasma because it is retained even when the RNA transcript is released from tissue into the blood. In this study, allele counts were analyzed for two RNA transcripts, namely PAPPA of pregnancy-specific genes and H19 of imprinted maternal-expressed genes. 26、27

[0101] For the PAPPA gene, rs386088, a SNP containing a fetal-specific SNP allele, was analyzed (Figure 27). The plasma sample after delivery did not contain PAPPA RNA sequencing reads, indicating that this gene is truly pregnancy-specific. 4 In particular, no statistically significant difference was found in the ratio of fetal-allele read counts between pre-delivery maternal plasma and placental RNA samples (P = 0.320, χ 2(Assay). This indicates that the data obtained from maternal plasma reflected the biallelic expression pattern of PAPPA in the placenta.

[0102] We recently reported that we were able to detect the DNA methylation patterns of the maternally expressed H19 gene in the placenta and maternal blood cells by bisulfite DNA sequencing of maternal plasma DNA. 28 In this study, we further investigated whether we could explore the genomic imprinting status of the H19 gene at the RNA level. First, we focused on the SNP site, rs2839698, in exon 1 of the H19 gene. This SNP is a maternal-specific allele, i.e., the allele that is AA in the fetus and AG in the mother. As shown in Figure 28, only the G-allele was detected from postpartum maternal plasma (Figure 28). Such a single-allele pattern was linked to the unmethylated G-allele of the rs4930098 SNP site in the imprinting control region (Figure 29). The G-allele was seen in prenatal maternal plasma, but the A-allele derived from the placenta was also detected (Figure 28). Similar allele patterns, i.e., biallelic in prenatal maternal plasma and monoallelic in postpartum maternal plasma, were observed for the other three SNP sites that retain the maternal-specific allele, namely rs2839701, rs2839702, and rs3741219 (Figure 28). The maternal-specific alleles found at these SNP sites belong to the same maternal haplotype, are not methylated, and are thus thought to be transcribed. In particular, the fact that H19 RNA was not expressed in maternal blood cells (Figure 28) suggests that the H19 RNA molecules contained in the plasma are derived from maternal tissues / organs rather than blood cells. Non-placental and non-fetal tissues in which H19 has been reported to be expressed include the adrenal gland, skeletal muscle, uterus, adipocytes, liver, and pancreas. 29 。

[0103] VII. Discussion In this study, we aimed to develop a technique to show the overall profile of transcriptome activity in maternal plasma by using RNA-sequencing. We have previously shown the proportion of fetal DNA contained in maternal plasma by targeting one or more fetal-specific loci. This is because the entire fetal genome is uniformly present in maternal plasma. 30 . Unlike DNA in blood, measuring the proportion of fetal-derived RNA transcripts in maternal plasma is not straightforward because the proportion is complicated by differential gene expression in fetal and maternal tissues and their release into the blood. By performing RNA-sequencing on maternal plasma and examining the polymorphic differences between the fetus and the mother, we were able to predict the proportion of plasma transcripts derived from the fetus. As expected, maternal-derived transcripts were dominant in the plasma transcriptome, but 3.70% and 11.28% of the blood transcripts contained in maternal plasma in the first and third trimesters of pregnancy, respectively, were derived from the fetus. These fetal-derived transcripts include RNA molecules derived from both the fetus and the mother as well as RNA molecules derived only from the fetus. Furthermore, we found that RNA molecules derived only from the fetus were 0.90% and 2.52% of the maternal blood transcripts in the first and third trimesters of pregnancy, respectively. The greater presence of such fetal-specific genes in the third trimester is probably correlated with the increasing size of the fetus and placenta as pregnancy progresses.

[0104] In this study, we showed that in the placenta, balanced RNA allelic expression of the pregnancy-specific PAPPA gene was observed, and in maternal plasma, single allelic expression of the imprinted H19 gene expressed in the mother could be observed. These data suggest the possibility that maternal plasma can be used as a non-invasive sample source for analyzing allelic expression patterns.

[0105] By quantitatively comparing the RNA transcripts contained in maternal plasma samples before and after childbirth, we compiled a list of 131 genes whose expression increased during pregnancy, as is also evident from the decreased expression in plasma samples after delivery. As expected, using the profiles of these genes, we were able to distinguish plasma samples from pregnant and non-pregnant women. Thus, direct comparison of maternal plasma samples before and after childbirth has shown in previous studies that placental expression is not necessarily at very high levels compared to maternal blood cells. 10、15 Pregnancy-related genes could be screened in a high-throughput manner. In short, this direct plasma assay provides another means of discovering pregnancy-related RNA transcripts in blood without the need to pre-profile the transcriptomes of placental tissue and blood cells.

[0106] So far, RNA-sequencing has been shown to be a viable method for profiling plasma transcriptomes, but several technical issues can be further improved. First, by further optimizing the protocol for sequencing, particularly by removing highly transcribed rRNA and globin genes from plasma, it may be possible to further increase the information obtained from plasma RNA-sequencing. Second, we have only focused on reference transcripts and have not yet explored individual isoforms. Future studies could include the detection of novel transcripts and differential analysis of splicing variants and their isoforms by increasing the depth of sequencing reads. Third, we did not perform allelic-specific expression (ASE) screening in the analysis of the proportion of fetal- and maternal-derived transcripts in early pregnancy samples because we had exhausted chorionic villi and amniotic fluid for genotyping. Nevertheless, in late pregnancy, we showed that ASE screening had no significant effect on the identification of genes with high fetal and maternal contribution rates in maternal plasma (Figures 16 and 17). 31-33

[0107] In summary, we have shown that by using RNA-sequencing technology, we can measure the proportion of fetal-derived transcripts contained in maternal plasma and identify pregnancy-related genes in the blood. This test has provided a more comprehensive understanding of the overall transcriptome of maternal plasma and thus facilitated the means for identifying biomarker candidates related to pregnancy-related diseases. We believe that this technology will open up new means for molecular diagnosis of diseases related to pregnancy or placenta, as well as other diseases such as cancer. 34 We think it will become a new means for molecular diagnosis of other diseases such as cancer.

[0108] VIII. Transplantation As described above, the methods described herein are also applicable to transplantation. The methods for transplantation can be carried out in a similar manner as used for fetal analysis. For example, the genotype of the transplanted tissue can be obtained. In one aspect, the transplanted tissue is heterozygous and the homozygous locus in the host organism (e.g., male or female) can be identified. In another aspect, the transplanted tissue is homozygous and the heterozygous locus in the host organism (e.g., male or female) can also be identified. The same ratio can be calculated by a computer and compared with a cut-off value to determine whether there is a disorder.

[0109] IX. Computer System Any computer systems referred to herein may use a suitable number of subsystems. An example of such a subsystem included in computer device 3000 is shown in FIG. 30. In some aspects, the computer system includes a single computer device, in which case the subsystems may be components of the computer device. In other aspects, the computer system includes a plurality of computer devices, and each computer device is a subsystem and has internal components.

[0110] The subsystems shown in FIG. 30 are interconnected via a system bus 3075. Other subsystems are also shown, such as a printing device 3074, a keyboard 3078, a storage device 3079, a monitor 3076 (connected to a display adapter 3082), etc. Peripheral devices and input / output (I / O) devices (connected to an I / O control device 3071) can be connected to the computer system by any number of means known in the art, such as a serial port 3077. For example, the computer system 3000 can be connected to a wide area network such as the Internet, a mouse input device, or a scanner using the serial port 3077 or an external interface 3081 (such as Ethernet, Wi-Fi, etc.). Through the interconnection using the system bus 3075, the central processing unit 3073 can control communication with each subsystem and execution of instructions from the system memory 3072 or the storage device 3079 (such as a fixed disk or an optical disk like a hard drive), and information can be exchanged between each subsystem. The system memory 3072 and / or the storage device 3079 may incorporate a readable medium. All data mentioned in this specification can be output from one component to another and can also be output to the user.

[0111] A computer system may include a plurality of identical components or subsystems, which are connected, for example, by an external interface 3081 or an internal interface. In some embodiments, a computer system, a subsystem, or a device may communicate over a network. In such an example, one computer can be considered a client and another computer can be considered a server, and both may be part of the same computer. The client and the server may each include a plurality of systems, subsystems, or components.

[0112] All aspects of the present invention can be implemented in the form of control theory using computer software with hardware (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or modular or integrated normally programmable processors. The processors used herein include multi-core processors on the same integrated chip, or multiple processing devices that may be on a single circuit board or connected by a network. Those skilled in the art will be able to understand and recognize other means and / or methods for implementing aspects of the present invention using hardware and combinations of hardware and software based on the techniques and teachings provided herein.

[0113] Any of the software components or functions described in this application are software code executed by a processor, and may be implemented as software code programmed by, for example, any suitable computer language such as Java, C++, or Perl, using, for example, standard or purpose-suited techniques. The software code may be stored on a computer-readable medium that can be read as a series of instructions or commands for storage and / or transmission. Suitable media include random access memory (RAM), read-only memory (ROM), magnetic media such as hard drives or floppy disks, or optical media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, etc. A computer-readable medium may be any combination of such storage or transmission devices.

[0114] Such programs may be encoded and transmitted using carrier signals adapted to various protocols over wired, optical and / or wireless networks, such as the Internet. That is, computer-readable media according to aspects of the present invention may be created using data signals encoded with such programs. Computer-readable media encoded with program code may be packaged with compatible devices or provided separately from other devices (e.g., by downloading over the Internet). Any such computer-readable media may be placed or incorporated into a single computer product (e.g., a hard drive, CD, or an entire computer system), or may also be placed or incorporated into multiple computer products on a system or network. A computer system may also include a monitor, a printing device, or any other suitable display for providing a user with any of the results mentioned herein.

[0115] Any of the methods described herein may be implemented, in whole or in part, using a computer system that includes one or more processors, and the computer system may be configured to perform the steps. Accordingly, aspects may relate to a computer system configured to perform the steps in any of the methods described herein, and the computer system may include different components for performing each step or each group of steps. Although the steps are numbered, the steps of the methods herein may be performed simultaneously or in a different order. In addition, some of these steps may be used in conjunction with some of the steps from other methods derived from other methods. Also, all or some of the steps may be optional. In addition, any of the steps in any of the methods may be performed using a module, a circuit, or any other means for performing these steps.

[0116] The specific items of a particular aspect can be combined in any suitable manner without departing from the spirit and scope of the aspects of the present invention. However, other aspects of the present invention may relate to specific aspects of each individual aspect, or specific combinations of these individual aspects.

[0117] The foregoing description of exemplary aspects of the present invention has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the forms described, and many modifications and variations are possible in light of the above teachings. The aspects were chosen and described in order to best explain the principles of the invention and its practical application, to thereby enable others skilled in the art to best utilize the invention in various aspects and to make various changes suitable for the particular use contemplated.

[0118] The recitation of "a", "one", or "the" is intended to mean "one or more" unless specifically indicated otherwise.

[0119] All patents, patent applications, publications, and descriptions referred to herein are hereby incorporated by reference in their entirety for all purposes. None of them are admitted to be prior art.

[0120] X. References 1. Ng EKO, Tsui NBY, Lau TK, Leung TN, Chiu RWK, Panesar NS, Lit LC, Chan KW, Lo YMD. mRNA of placental origin is readily detectable in maternal plasma. Proc Natl Acad Sci U S A 2003; 100: 4748-53. 2. Ng EKO, Leung TN, Tsui NBY, Lau TK, Panesar NS, Chiu RWK, Lo YMD. The concentration of circulating corticotropin-releasing hormone mRNA in maternal plasma is increased in preeclampsia. Clin Chem 2003; 49: 727-31. 3. Farina A, Sekizawa A, Sugito Y, Iwasaki M, Jimbo M, Saito H, Okai T. Fetal DNA in maternal plasma as a screening variable for preeclampsia. A preliminary nonparametric analysis of detection rate in low-risk nonsymptomatic patients. Prenat Diagn 2004; 24: 83-6. 4. Pang WW, Tsui MH, Sahota D, Leung TY, Lau TK, Lo YM, Chiu RW. A strategy for identifying circulating placental RNA markers for fetal growth assessment. Prenat Diagn 2009; 29: 495-504. 5. Lo YMD, Tsui NBY, Chiu RWK, Lau TK, Leung TN, Heung MM, Gerovassili A, Jin Y, Nicolaides KH, Cantor CR, Ding C. Plasma placental RNA allelic ratio permits noninvasive prenatal chromosomal aneuploidy detection. Nat Med 2007; 13: 218-23. 6. Tsui NBY, Akolekar R, Chiu RWK, Chow KCK, Leung TY, Lau TK, Nicolaides KH, Lo YMD. Synergy of total PLAC4 RNA concentration and measurement of the RNA single - nucleotide polymorphism allelic ratio for the noninvasive prenatal detection of trisomy 21. Clin Chem 2010; 56: 73 - 81. 7. Tsui NBY, Wong BCK, Leung TY, Lau TK, Chiu RWK, Lo YMD. Non - invasive prenatal detection of fetal trisomy 18 by RNA - SNP allelic ratio analysis using maternal plasma SERPINB2 mRNA: a feasibility study. Prenat Diagn 2009; 29: 1031 - 7. 8. Chim SS, Lee WS, Ting YH, Chan OK, Lee SW, Leung TY. Systematic identification of spontaneous preterm birth - associated RNA transcripts in maternal plasma. PLoS One 2012; 7: e34328. 9. Wong BCK, Chiu RWK, Tsui NBY, Chan KCA, Chan LW, Lau TK, Leung TN, Lo YMD. Circulating placental RNA in maternal plasma is associated with a preponderance of 5' mRNA fragments: implications for noninvasive prenatal diagnosis and monitoring. Clin Chem 2005;51: 1786-95. 10. Tsui NBY, Chim SSC, Chiu RWK, Lau TK, Ng EKO, Leung TN, Tong YK, Chan KCA, Lo YMD. Systematic micro-array based identification of placental mRNA in maternal plasma: towards non-invasive prenatal gene expression profiling. J Med Genet 2004; 41: 461-7. 11. Poon LL, Leung TN, Lau TK, Lo YMD. Presence of fetal RNA in maternal plasma. Clin Chem 2000; 46: 1832-4. 12. Go AT, Visser A, Mulders MA, Blankenstein MA, Van Vugt JM, Oudejans CB. Detection of placental transcription factor mRNA in maternal plasma. Clin Chem 2004; 50: 1413-4. 13 Smets EM, Visser A, Go AT, van Vugt JM, Oudejans CB. Novel biomarkers in preeclampsia. Clin Chim Acta 2006; 364: 22-32. 14. Purwosunu Y, Sekizawa A, Koide K, Farina A, Wibowo N, Wiknjosastro GH, et al. Cell-free mRNA concentrations of plasminogen activator inhibitor-1 and tissue-type plasminogen activator are increased in the plasma of pregnant women with preeclampsia. Clin Chem 2007; 53: 399-404. 15. Miura K, Miura S, Yamasaki K, Shimada T, Kinoshita A, Niikawa N, et al. The possibility of microarray-based analysis using cell-free placental mRNA in maternal plasma. Prenat Diagn 2010; 30: 849-61. 16. Ng EKO, El-Sheikhah A, Chiu RWK, Chan KC, Hogg M, Bindra R, et al. Evaluation of human chorionic gonadotropin beta-subunit mRNA concentrations in maternal serum in aneuploid pregnancies: a feasibility study. Clin Chem 2004; 50: 1055-7. 17. Mortazavi A, Williams BA, McCue K, Schaeffer L, Wold B. Mapping and quantifying mammalian transcriptomes by RNA-Seq. Nat Methods 2008; 5: 621-8. 18. Sultan M, Schulz MH, Richard H, Magen A, Klingenhoff A, Scherf M, et al. A global view of gene activity and alternative splicing by deep sequencing of the human transcriptome. Science 2008; 321: 956-60. 19. Kim J, Zhao K, Jiang P, Lu ZX, Wang J, Murray JC, Xing Y. Transcriptome landscape of the human placenta. BMC Genomics 2012; 13: 115. 20. Wang K, Li H, Yuan Y, Etheridge A, Zhou Y, Huang D, et al. The complex exogenous RNA spectra in human plasma: an interface with human gut biota? PLoS One 2012; 7: e51009. 21. Li H, Guo L, Wu Q, Lu J, Ge Q, Lu Z. A comprehensive survey of maternal plasma miRNAs expression profiles using high-throughput sequencing. Clin Chim Acta 2012; 413: 568-76. 22. Williams Z, Ben-Dov IZ, Elias R, Mihailovic A, Brown M, Rosenwaks Z, Tuschl T. Comprehensive profiling of circulating microRNA via small RNA sequencing of cDNA libraries reveals biomarker potential and limitations. Proc Natl Acad Sci U S A 2013; 110: 4255-60. 23. Chim SSC, Shing TK, Hung EC, Leung TY, Lau TK, Chiu RWK, Lo YMD. Detection and characterization of placental microRNAs in maternal plasma. Clin Chem 2008; 54 :482-90. 24. Pickrell JK et al. Understanding mechanisms underlying human gene expression variation with RNA sequencing. Nature 2010; 464: 768-772. 25. Smith RM, Webb A, Papp AC, Newman LC, Handelman SK, Suhy A, et al. Whole transcriptome RNA-Seq allelic expression in human brain. BMC Genomics 2013; 14: 571. 26. Frost JM, Monk D, Stojilkovic-Mikic T, Woodfine K, Chitty LS, Murrell A, et al. Evaluation of allelic expression of imprinted genes in adult human blood. PLoS One 2010; 5: e13556. 27. Daelemans C, Ritchie ME, Smits G, Abu-Amero S, Sudbery IM, Forrest MS, et al. High-throughput analysis of candidate imprinted genes and allele-specific gene expression in the human term placenta. BMC Genet 2010; 11: 25. 28. Lun FMF, Chiu RWK, Sun K, Leung TY, Jiang P, Chan KCA, et al. Noninvasive prenatal methylomic analysis by genomewide bisulfite sequencing of maternal plasma DNA. Clin Chem 2013; 59: 1583-94. 29. Wu C, Orozco C, Boyer J, Leglise M, Goodale J, Batalov S, et al. BioGPS: an extensible and customizable portal for querying and organizing gene annotation resources. Genome Biol 2009; 10: R130. 30. Lo YMD, Chan KCA, Sun H, Chen EZ, Jiang P, Lun FMF, et al. Maternal plasma DNA sequencing reveals the genome-wide genetic and mutational profile of the fetus. Sci Transl Med 2010; 2: 61ra91. 31. Cabili MN, Trapnell C, Goff L, Koziol M, Tazon-Vega B, Regev A, Rinn JL. Integrative annotation of human large intergenic noncoding RNAs reveals global properties and specific subclasses. Genes Dev 2011; 25: 1915-27. 32. St Laurent G, Shtokalo D, Tackett MR, Yang Z, Eremina T, Wahlestedt C, et al. Intronic RNAs constitute the major fraction of the non-coding RNA in mammalian cells. BMC Genomics 2012; 13: 504. 33. Anders S, Reyes A, Huber W. Detecting differential usage of exons from RNA-seq data. Genome Res 2012; 22: 2008-17. 34. Fleischhacker M, Schmidt B. Circulating nucleic acids (CNAs) and cancer--a survey. Biochim Biophys Acta 2007; 1775: 181-232.

Claims

1. 1. A method for diagnosing a pregnancy-related disorder using a sample obtained from a female subject carrying a fetus, comprising: obtaining a plurality of reads, wherein the reads are obtained from analysis of RNA molecules obtained from a sample comprising a mixture of maternal and fetal RNA molecules; identifying, by a computer system, the position of said read in a reference sequence; identifying one or more informative loci, each of which is homozygous for a corresponding first allele in a first entity and heterozygous for said corresponding first allele and a corresponding second allele in a second entity, wherein said first entity is said pregnant female subject or said fetus, and said second entity is the other of said pregnant female subject and said fetus; screening the one or more informative loci to identify one or more selected informative loci, wherein the informative loci are located in expressed regions of the reference sequence; the informative loci are located among the plurality of reads in which at least a predetermined first number of reads contain the corresponding first allele; and the informative loci are located among the plurality of reads in which at least a predetermined second number of reads contain the corresponding second allele. for each of the selected informative loci, determining a first number of reads that are located at the selected informative locus and contain the corresponding first allele; and determining a second number of reads that are located at the selected informative locus and contain the corresponding second allele. calculating a first sum of the first numbers and a second sum of the second numbers; calculating a ratio of the first sum to the second sum; and comparing said ratio to a cut-off value to determine whether said fetus, mother, or pregnancy has a pregnancy-related disorder, said cut-off value being determined from one or more samples obtained from pregnant female subjects not suffering from said pregnancy-related disorder. A method for diagnosing a pregnancy-related disorder, comprising:

2. The method of claim 1 , wherein said sample from said pregnant female is a plasma sample.

3. The method of claim 1 , wherein said second entity is said pregnant female subject.

4. The method of claim 1 , wherein the second entity is the fetus.

5. The method of claim 1 , wherein calculating a ratio of the first sum to the second sum comprises dividing the second sum by a sum of the first sum and the second sum.

6. The method of claim 1 , wherein said analysis of RNA molecules obtained from said sample comprises sequencing said RNA molecules or cDNA copies thereof.

7. 2. The method of claim 1, wherein said analyzing RNA molecules obtained from said sample comprises performing digital PCR.

8. The method of claim 1 , wherein identifying the position of the read in the reference sequence comprises aligning the read to the reference sequence.

9. 2. The method of claim 1, wherein said first predetermined number of leads is one and said second predetermined number of leads is one.

10. 10. The method of claim 1, further comprising sequencing genomic DNA obtained from maternal tissue to determine the genotype of each informative locus in said pregnant female subject.

11. 10. The method of claim 1, further comprising sequencing genomic DNA obtained from the placenta, chorionic villi, amniotic fluid, or maternal plasma to determine the genotype of each informative locus of the fetus.

12. 2. The method of claim 1, further comprising predicting the proportion of maternal or fetal origin of RNA in said sample, wherein said proportion is said ratio multiplied by a scalar.

13. 1. A method for determining a proportion of fetal RNA in a sample obtained from a female subject carrying a fetus, comprising: obtaining a plurality of reads, wherein said reads are obtained from analysis of RNA obtained from said sample containing a mixture of maternal and fetal RNA molecules; identifying, by a computer system, the position of said read in a reference sequence; identifying one or more informative maternal loci, each of which is homozygous for a corresponding first allele in said fetus and heterozygous for said corresponding first allele and a corresponding second allele in said pregnant female subject; screening the one or more informative loci to identify one or more selected informative loci, wherein the informative loci are located in expressed regions of the reference sequence; the informative loci are located at reads among the plurality of reads that contain at least a predetermined first number of the corresponding first allele; and the informative loci are located at reads among the plurality of reads that contain at least a predetermined second number of the corresponding second allele. determining, for each of the selected informative maternal loci, a first number of reads located at the selected informative maternal locus and containing the corresponding first allele; and determining a second number of reads located at the selected informative maternal locus and containing the corresponding second allele; calculating a sum of the first number and the second number; calculating a maternal ratio by dividing the second number by the sum; determining a scalar that represents the total expression of the selected informative maternal loci in the pregnant female subject relative to the expression of the corresponding second allele; multiplying the maternal ratio by the scalar to obtain a maternal contribution; and calculating a fetal contribution that is 1 minus the maternal contribution; and determining a proportion of RNA of fetal origin in said sample, said proportion being an average of said fetal contributions of said selected informative maternal loci; A method for determining the proportion of fetal RNA in a sample, comprising:

14. 14. The method of claim 13, wherein the average is weighted by the sum of the selected informative matrix loci.

15. 14. The method of claim 13, wherein the scalar for a selected informative maternal locus is estimated to be approximately 2.

16. The method of claim 13, wherein said sample from said pregnant female is a plasma sample.

17. The method of claim 13, wherein said analysis of RNA molecules obtained from said sample comprises sequencing said RNA molecules or cDNA copies thereof.

18. 14. The method of claim 13, wherein said analyzing RNA molecules obtained from said sample comprises performing digital PCR.

19. The method of claim 13 , wherein identifying the position of the read in the reference sequence comprises aligning the read to the reference sequence.

20. 14. The method of claim 13, wherein the first predetermined number of reads for the one or more selected informative matrix loci is one and the second predetermined number of reads for the one or more selected informative matrix loci is one.

21. 14. The method of claim 13, further comprising sequencing genomic DNA obtained from maternal tissue to determine the genotype of each informative maternal locus in said pregnant female subject.

22. 14. The method of claim 13, further comprising sequencing genomic DNA obtained from the placenta, the chorionic villi, the amniotic fluid, or the maternal plasma to determine the genotype of each informative maternal locus of the fetus.

23. identifying one or more informative fetal loci, each of which is homozygous for a corresponding first allele in said pregnant female subject and heterozygous for said corresponding first allele and a corresponding second allele in said fetus; screening the one or more informative fetal loci to identify one or more selected informative fetal loci, wherein the fetal informative loci are located in expressed regions of the reference sequence; the fetal informative loci are located among the plurality of reads in which at least a predetermined first number of reads contain the corresponding first allele; and the fetal informative loci are located among the plurality of reads in which at least a predetermined second number of reads contain the corresponding second allele. for each of the selected informative fetal loci, determining a first number of reads that are located at the selected fetal informative locus and contain the corresponding first allele; and determining a second number of reads that are located at the selected fetal informative locus and contain the corresponding second allele; calculating the sum of the first number and the second number; calculating a fetal ratio by dividing the second number by the sum; determining a scalar that represents the total expression of the selected informative fetal loci in the fetus relative to the expression of the corresponding second allele; and multiplying the fetal ratio by the scalar to obtain a fetal contribution ratio; determining a proportion of RNA of fetal origin in said sample, said proportion being the average of said fractional fetal contributions for said selected informative maternal loci and said selected informative fetal loci; The method of claim 13, further comprising:

24. 24. The method of claim 23, wherein the average is weighted by the sums for the selected informative maternal loci and the selected informative fetal loci.

25. 24. The method of claim 23, wherein the scalar for a selected informative fetal locus is estimated to be approximately 2.

26. 24. The method of claim 23, wherein the first number of pre-determined reads for the one or more selected informative fetal loci is one and the second number of pre-determined reads for the one or more selected informative fetal loci is one.

27. 24. The method of claim 23, further comprising sequencing genomic DNA obtained from maternal tissue to determine the genotype of each informative fetal locus in said pregnant female subject.

28. 24. The method of claim 23, further comprising sequencing genomic DNA obtained from the placenta, the chorionic cilium, the amniotic fluid, or the maternal plasma to determine the genotype of each informative fetal locus of the fetus.

29. 24. The method of claim 13 or 23, further comprising diagnosing a pregnancy-related disorder by comparing the percentage of fetal RNA in the sample with a cutoff value.

30. 1. A method for designating genomic loci as maternal or fetal markers by analyzing a sample obtained from a female subject carrying a fetus, comprising: obtaining a plurality of reads from analysis of RNA molecules obtained from said sample containing a mixture of maternal and fetal RNA molecules; identifying, by a computer system, the position of said read in a reference sequence; identifying one or more informative loci, each of which is homozygous for a corresponding first allele in a first entity and heterozygous for said corresponding first allele and a corresponding second allele in a second entity, wherein said first entity is said pregnant female subject or said fetus, and said second entity is the other of said pregnant female subject and said fetus; screening the one or more informative loci to identify one or more selected informative loci, wherein the informative loci are located in expressed regions of the reference sequence; the informative loci are located among the plurality of reads in which at least a predetermined first number of reads contain the corresponding first allele; and the informative loci are located among the plurality of reads in which at least a predetermined second number of reads contain the corresponding second allele. for each of the selected informative loci, determining a first number of reads that map to the selected informative locus and contain the corresponding first allele; and determining a second number of reads that map to the selected informative locus and contain the corresponding second allele; calculating a ratio of the first number to the second number; and designating the selected informative locus as a marker for the second entity if the ratio exceeds a cutoff value; A method for designating genomic loci as maternal or fetal markers, comprising:

31. 31. The method of claim 30, wherein said sample from said pregnant female is a plasma sample.

32. 31. The method of claim 30, wherein said second entity is said pregnant female subject.

33. 31. The method of claim 30, wherein the second entity is the fetus.

34. 31. The method of claim 30, wherein calculating the ratio of the first number to the second number comprises dividing the second number by a sum of the first number and the second number.

35. The method of claim 30, wherein the cutoff value is from about 0.2 to about 0.

5.

36. 36. The method of claim 35, wherein the cutoff value is 0.

4.

37. 31. The method of claim 30, wherein said analysis of RNA molecules obtained from said sample comprises sequencing said RNA molecules or cDNA copies thereof.

38. 31. The method of claim 30, wherein said analyzing RNA molecules obtained from said sample comprises performing digital PCR.

39. 31. The method of claim 30, wherein identifying the position of the read in a reference sequence comprises aligning the read to the reference sequence.

40. 31. The method of claim 30, wherein the first predetermined number of leads is one and the second predetermined number of leads is one.

41. 31. The method of claim 30, further comprising sequencing genomic DNA obtained from maternal tissue to determine the genotype of each informative locus in said pregnant female subject.

42. 31. The method of claim 30, further comprising sequencing genomic DNA obtained from the placenta, the chorionic cilium, the amniotic fluid, or the maternal plasma to determine the genotype of each informative locus of the fetus.

43. 31. The method of claim 30, further comprising diagnosing a pregnancy-related disorder based on whether a selected informative locus is designated as a marker for said second entity.

44. 31. The method of claim 30, further comprising determining the proportion of said RNA in said sample that is derived from said second entity, with respect to RNA in said sample that contains a selected informative locus, wherein said proportion is determined by multiplying said ratio by a scalar.

45. 45. The method of claim 44, wherein the scalar represents the total expression of the selected informative locus in the second entity relative to the expression of the corresponding second allele.

46. 45. The method of claim 44, wherein the scalar is estimated to be approximately two.

47. 45. The method of claim 44, further comprising determining the percentage of RNA derived from the first entity by subtracting the percentage of RNA derived from the second entity from 1.

48. Use of the genes listed in Figure 21 or Figure 26 to diagnose a pregnancy-related disorder in a female subject carrying a fetus by comparing the expression levels of said genes with control values ​​determined from one or more other female subjects, each of whom is carrying a healthy fetus.

49. 49. The use of claim 48, wherein the medical disorder is associated with the female subject.

50. 49. The use of claim 48, wherein the medical disorder is associated with the fetus.

51. 51. The use of claim 50, wherein the medical disorder is pre-eclampsia or preterm birth.

52. 49. The use of claim 48, wherein the medical disorder is diagnosed when the expression level exceeds the control value.

53. 49. The use of claim 48, wherein the medical disorder is diagnosed when the expression level is below the control value.

54. 49. The use of claim 48, wherein the female subject exhibits a higher expression level of the gene during pregnancy compared to after birth.

55. determining the expression level of the gene in a female subject; obtaining a plasma sample from said female subject; obtaining a plurality of reads from analysis of RNA molecules obtained from said sample; identifying a position of said read in a reference sequence using a computer system; and Counting the reads that map to the gene; 49. The use according to claim 48, wherein the measurement is by

56. A system comprising one or more processors configured to carry out the method of any one of claims 1 to 55.

57. A computer product comprising a computer readable medium storing a plurality of instructions for controlling a processor to perform the operations of the method according to any one of claims 1 to 56.

Citation Information

Patent Citations

  • Markers for prenatal diagnosis and monitoring

    JP2008524993A

  • Biomarkers for preeclampsia

    US20080233583A1

  • Gene expression related to preeclampsia

    US20110171650A1

  • Detection of genetic abnormalities and infectious disease

    US20120077185A1

  • Integrated Versatile and Systems Preparation of Specimens

    US20120164648A1