Methods and systems for estimating fetal nucleic acid concentration in non-invasive prenatal genetic testing data

CN122551879APending Publication Date: 2026-08-11SHENZHEN HUADA GENE INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

现有技术中,如论文《Maternal plasmafetalDNA fractions in pregnancies with low and high risks for fetalchromosomalaneuploidies》公开了利用Y染色体数据估算胎儿DNA浓度的方法,但仅适用于男性胎儿;论文《Size-based molecular diagnostics using plasma DNA fornoninvasive prenatal testing》公开了利用孕妇和胎儿游离核酸片段差异估计胎儿DNA浓度的方法,但需要双端测序或电泳实验,不适用于单端测序数据;论文《Determinationof fetal DNA fraction from the plasma ofpregnant women using sequence readcounts》描述的seqFF方法,将全基因组划分窗口统计测序读段数量,但其分辨率较低,不适用于胎儿核酸浓度在5%以内的样本;论文《Maternal plasma DNA sequencing revealsthegenome-wide genetic and mutational profile ofthe fetus》和《Noninvasiveprenatal methylomic analysis by genomewide bisulfite sequencingof maternal plasma DNA》分别基于等位基因和甲基化特征计算胎儿DNA浓度,但需要高深度测序或甲基化测序

Benefits of technology

[0066]本发明的方法不需要高深度测序,不需要PE测序,不需要甲基化测序,也不需要对胎儿的父母进行额外测序。本发明的方法仅用NIPT数据就能估计胎儿核酸浓度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551879A_ABST
    Figure CN122551879A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of biotechnology and discloses a method and system for estimating fetal nucleic acid concentration in non-invasive prenatal genetic testing (NIPT) data. The method includes: (1) sequencing the cell-free nucleic acid of the pregnant woman to be tested to obtain sequencing data, wherein the sequencing data includes several reads; (2) aligning the reads to a reference genome and calculating the copy ratio of the aligned reads in gene regions and / or regulatory regions of the reference genome; (3) inputting the copy ratio into a trained machine learning model to obtain the fetal nucleic acid concentration of the pregnant woman to be tested, wherein the trained machine learning model is trained using data from multiple pregnant women carrying male fetuses, and the data from the multiple pregnant women includes the fetal nucleic acid concentration calculated based on Y chromosome sequencing depth and the copy ratio of each read from the multiple pregnant women in gene regions and / or regulatory regions of the reference genome. The method of this invention can estimate fetal nucleic acid concentration using only NIPT data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, and more specifically, this invention provides a method and system for estimating the concentration of fetal nucleic acids in non-invasive prenatal genetic testing data. Background Technology

[0002] Cell-free DNA in maternal plasma can be used to analyze fetal health. Non-invasive prenatal testing (NIPT) primarily uses this cell-free DNA to infer whether the fetus has genetic diseases, such as trisomy 13. Fetal DNA concentration is a key parameter in NIPT data analysis. For male fetuses, fetal DNA concentration can be inferred from the proportion of Y chromosome data, but for female fetuses, other algorithms need to be developed. In existing technologies, for example, the paper "Maternal plasma fetal DNA fractions in pregnancies with low and high risks for fetal chromosome laneuploidies" discloses a method for estimating fetal DNA concentration using Y chromosome data, but it is only applicable to male fetuses; the paper "Size-based molecular diagnostics using plasma DNA for noninvasive prenatal testing" discloses a method for estimating fetal DNA concentration using differences in cell-free nucleic acid fragments between pregnant women and fetuses, but it requires paired-end sequencing or electrophoresis experiments and is not applicable to single-end sequencing data; the paper "Determination of fetal DNA fraction from the plasma of pregnant women using sequence read counts" describes the seqFF method, which divides the whole genome into windows to count the number of sequencing reads, but its resolution is low and it is not applicable to samples with fetal nucleic acid concentrations below 5%; the papers "Maternal plasma DNA sequencing reveals the genome-wide genetic and mutational profile of the fetus" and "Noninvasive prenatal methylomic analysis by genome-wide bisulfite sequencing of maternal plasma DNA" calculate fetal DNA concentration based on allele and methylation characteristics, respectively, but they require high-depth sequencing or methylation sequencing. Therefore, the present invention requires a method for estimating the concentration of fetal nucleic acids in cell-free DNA in maternal plasma. Summary of the Invention

[0003] The present invention aims to provide a method for estimating the concentration of fetal nucleic acids in cell-free DNA in maternal plasma, in order to solve the problems existing in the prior art.

[0004] Therefore, in a first aspect, the present invention provides a method for estimating the concentration of fetal nucleic acids in non-invasive prenatal genetic testing data, the method comprising:

[0005] (1) Sequencing the cell-free nucleic acid of the pregnant woman to be tested to obtain sequencing data, wherein the sequencing data includes several reads;

[0006] Preferably, the free nucleic acid is free DNA;

[0007] (2) Align the reads to the reference genome and calculate the copy ratio of the reads in the gene regions and / or regulatory regions of the reference genome.

[0008] (3) Input the copy ratio into the trained machine learning model to obtain the fetal nucleic acid concentration of the pregnant woman to be tested.

[0009] The trained machine learning model is trained using data from multiple pregnant women carrying male fetuses. The data from these multiple pregnant women includes fetal nucleic acid concentrations calculated based on Y chromosome sequencing depth and the copy ratio of each read from each of the multiple pregnant women in the gene regions and / or regulatory regions of the reference genome.

[0010] In one implementation, in (1), the free nucleic acid is derived from the pregnant woman's peripheral blood plasma, the pregnant woman's liver, and / or placenta.

[0011] In one embodiment, the gene region and / or regulatory region originates from an autosome, preferably an autosome other than chromosomes 13, 18, and 21.

[0012] In one embodiment, the length of the gene region is 7–2473538 bp, and the length of the regulatory region is 199–43798 bp.

[0013] In one embodiment, the number of the gene regions and / or regulatory regions is more than 10,000, preferably more than 50,000, more preferably more than 100,000, and most preferably more than 200,000.

[0014] In one implementation, the control region is a promoter sub-region.

[0015] In one implementation, the machine learning model is a machine learning regression model.

[0016] In one implementation, the machine learning regression model includes a linear regression model and a nonlinear regression model.

[0017] In one implementation, the machine learning regression model is a ridge regression model, a lasso regression model, a least squares linear regression model, a regression model based on a random forest algorithm, or a regression model based on a deep neural network.

[0018] In one implementation, the gene region and / or regulatory region are derived from ENSEMBLE.

[0019] In one implementation, the fetal nucleic acid concentration (Fraction) is calculated based on Y chromosome depth. fetal )for:

[0020]

[0021] Depth Y The average coverage depth of the Y chromosome, Depth autosomes The average coverage depth of sequencing data on autosomes is used, and preferably, the autosomes do not include chromosomes 13, 18, and 21.

[0022] In one implementation, the copy ratio of the reads compared to the reference genome in gene regions and / or regulatory regions is:

[0023]

[0024] Where copy_ratio ip The number of reads represents the copy ratio of region p of sample i. ip The length is the number of read segments in region p of sample i. p For the total length of region p, reads_number i The length is the number of read segments for sample i. ref The total length of the reference genome is n, and there are m regions in total.

[0025] In a second aspect, the present invention provides a method for constructing a machine learning model for estimating fetal nucleic acid concentrations in non-invasive prenatal genetic testing data, as described in the first aspect of the present invention, the method comprising:

[0026] (a) Obtaining sequencing data of cell-free nucleic acids from multiple pregnant women carrying male fetuses, wherein the sequencing data includes several reads;

[0027] Preferably, the free nucleic acid is free DNA;

[0028] (b) Align the reads to a reference genome, and for each of the plurality of pregnant women, calculate the fetal nucleic acid concentration based on the Y chromosome sequencing depth, and calculate the copy ratio of the reads to be aligned to the reference genome in the gene region and / or regulatory region.

[0029] (c) The fetal nucleic acid concentration calculated based on the Y chromosome sequencing depth and the copy ratio are input into the machine learning model for training to obtain the trained machine learning model.

[0030] In one implementation, in (a), the free nucleic acid fragment is derived from the pregnant woman's peripheral blood plasma, the pregnant woman's liver, and / or placenta.

[0031] In one embodiment, the gene region and / or regulatory region originates from an autosome, preferably an autosome other than chromosomes 13, 18, and 21.

[0032] In one embodiment, the length of the gene region is 7–2473538 bp, and the length of the regulatory region is 199–43798 bp.

[0033] In one embodiment, the number of the gene regions and / or regulatory regions is more than 10,000, preferably more than 50,000, more preferably more than 100,000, and most preferably more than 200,000.

[0034] In one implementation, the control region is a promoter sub-region.

[0035] In one implementation, the machine learning model is a machine learning regression model.

[0036] In one implementation, the machine learning model includes a linear regression model and a nonlinear regression model.

[0037] In one implementation, the machine learning model is a ridge regression model, a lasso regression model, a least squares linear regression model, a regression model based on a random forest algorithm, or a regression model based on a deep neural network.

[0038] In one implementation, the gene region and / or regulatory region are derived from ENSEMBLE.

[0039] In one implementation, the fetal nucleic acid concentration (Fraction) is calculated based on Y chromosome depth. fetal )for:

[0040]

[0041] Depth Y The average coverage depth of the Y chromosome, Depth autosomes The average coverage depth of sequencing data on autosomes is used, and preferably, the autosomes do not include chromosomes 13, 18, and 21.

[0042] In one implementation, the copy ratio of the reads compared to the reference genome in gene regions and / or regulatory regions is:

[0043]

[0044] Where copy_ratio ip The number of reads represents the copy ratio of region p of sample i. ip The length is the number of read segments in region p of sample i. p For the total length of region p, reads_number i The length is the number of read segments for sample i. ref The total length of the reference genome is n, and there are m regions in total.

[0045] In a third aspect, the present invention provides a machine learning model for estimating fetal nucleic acid concentration in non-invasive prenatal genetic testing data in the first aspect of the present invention, the machine learning model being constructed according to the method of the second aspect of the present invention.

[0046] In a fourth aspect, the present invention provides a system for estimating fetal nucleic acid concentrations in non-invasive prenatal genetic testing data, the system comprising:

[0047] The sequencing data acquisition module is used to sequence the cell-free nucleic acid of the pregnant woman to be tested and obtain sequencing data, wherein the sequencing data includes several reads; preferably, the cell-free nucleic acid is cell-free DNA;

[0048] The copy ratio calculation module is used to align the read to a reference genome and calculate the copy ratio of the read aligned to the reference genome in gene regions and / or regulatory regions.

[0049] The prediction module is used to input the copy ratio into a trained machine learning model to obtain the fetal nucleic acid concentration of the pregnant woman to be tested.

[0050] The trained machine learning model is trained using data from multiple pregnant women carrying male fetuses. The training includes inputting the fetal nucleic acid concentration calculated based on the Y chromosome sequencing depth and the copy ratio of each read from the multiple pregnant women in the gene region and / or regulatory region of the reference genome into the machine learning model for training, thereby obtaining the trained machine learning model.

[0051] In one implementation, in the sequencing data acquisition module, the free nucleic acid fragments are derived from the pregnant woman's peripheral plasma, liver, and / or placenta.

[0052] In one embodiment, the gene region and / or regulatory region originates from an autosome, preferably an autosome other than chromosomes 13, 18, and 21.

[0053] In one embodiment, the length of the gene region is 7–2473538 bp, and the length of the regulatory region is 199–43798 bp.

[0054] In one embodiment, the number of the gene regions and / or regulatory regions is more than 10,000, preferably more than 50,000, more preferably more than 100,000, and most preferably more than 200,000.

[0055] In one implementation, the control region is a promoter sub-region.

[0056] In one implementation, the machine learning model is a machine learning regression model.

[0057] In one implementation, the machine learning regression model includes a linear regression model and a nonlinear regression model.

[0058] In one implementation, the machine learning regression model is a ridge regression model, a lasso regression model, a least squares linear regression model, a regression model based on a random forest algorithm, or a regression model based on a deep neural network.

[0059] In one implementation, the gene region and / or regulatory region are derived from ENSEMBLE.

[0060] In one implementation, the fetal nucleic acid concentration (Fraction) is calculated based on Y chromosome depth. fetal )for:

[0061]

[0062] Depth Y The average coverage depth of the Y chromosome, Depth autosomes The average coverage depth of sequencing data on autosomes is used, and preferably, the autosomes do not include chromosomes 13, 18, and 21.

[0063] In one implementation, the copy ratio of the reads compared to the reference genome in gene regions and / or regulatory regions is:

[0064]

[0065] Where copy_ratio ip The number of reads represents the copy ratio of region p of sample i. ip The length is the number of read segments in region p of sample i. p Given the total length of region p, reads-number i The length is the number of read segments for sample i. ref The total length of the reference genome is n, and there are m regions in total.

[0066] The method of this invention does not require high-depth sequencing, PE sequencing, methylation sequencing, or additional sequencing of the fetal parents. The method of this invention can estimate fetal nucleic acid concentration using only NIPT data. Attached Figure Description

[0067] Figure 1 Using data from 2400 pregnant women with male fetuses as training data, the model was used to calculate the predicted fetal nucleic acid concentration for 600 pregnant women. Detailed Implementation

[0068] In this invention, the statistical analysis of biologically functional gene regions and regulatory regions, compared to a simple window segmentation scheme of a certain length, reduces the likelihood that a single biologically functional region unit will simultaneously contain both a region with a high proportion of enriched data from pregnant women and a region with a high proportion of enriched data from fetuses, thus resulting in relatively higher resolution. Preferably, the gene regions and regulatory regions include regions that are extended or expanded from genes or regulators.

[0069] In this invention, a machine learning model is trained using data from multiple pregnant women carrying male fetuses. The training involves inputting the fetal nucleic acid concentration calculated at Y chromosome depth and the copy ratio of each read from the multiple pregnant women in the gene regions and / or regulatory regions of the reference genome into the machine learning model. For the cell-free nucleic acid of the pregnant woman to be tested, the copy ratio of the read aligned to the reference genome in the gene regions and / or regulatory regions is input into the trained machine learning model to obtain the fetal nucleic acid concentration of the pregnant woman to be tested. The same gene regions and / or regulatory regions are preferably used for both the training samples and the test samples. The regulatory region is preferably a promoter region. Because the method of this invention is based on computational methods developed from data of biologically functional genes and regulatory regions, it has higher resolution and can be applied to samples with fetal nucleic acid concentrations below 5%.

[0070] In this invention, the machine learning model is a machine learning regression model, including linear regression and nonlinear regression models. The machine learning regression model can be a ridge regression model, a lasso regression model, a regression model based on the random forest algorithm, or a regression model based on a deep neural network. Ridge regression is a multiple linear regression model, which essentially fits a linear function such that the target variable y is a linear combination of the independent variables x (also called features). Ridge regression further reduces the risk of overfitting by penalizing the coefficients of the independent variables (i.e., feature coefficients) of the linear function (i.e., performing L2 regularization).

[0071] The method and system of the present invention will be described by way of example below with reference to the specific embodiments.

[0072] Example

[0073] Samples from multiple pregnant women known to have carried male fetuses were used as training samples. For both training and test samples, circulating DNA (cfDNA) samples were extracted from the pregnant women's plasma and sequenced to obtain sequencing reads for each sample, which were then used for subsequent analysis.

[0074] The first step involves aligning all raw data (fq format) from samples used for model training and prediction to the human reference chromosome hg38 using the samse mode in BWA after quality control. Picard is used to remove duplicate reads from the alignment results and the duplication rate is calculated. The base quality score correction (BQSR) function in variant detection algorithms such as GATK is used to perform local correction of the alignment results. Prepare non-invasive prenatal testing (NIPT) data for male fetuses, such as alignment files for SE35 data in bam / cram format.

[0075] The second step is to download the gene region file and regulatory region file of hg19 / hg38 from ENSEMBLE, and filter out the sex chromosomes and the three chromosomes 13, 18 and 21 that may cause trisomy.

[0076] The third step is to train the model, specifically for male fetuses, by calculating the average coverage depth of sequencing data on each autosome. autosomes ) and the average coverage depth of the Y chromosome (Depth) Y Then the concentration of male fetal nucleic acid (Fraction) can be obtained. fetal The formula for calculating ) is:

[0077]

[0078] Download the gene region and regulatory region files for hg38 from ENSEMBLE, filtering out sex chromosomes and chromosomes 13, 18, and 21, which may cause trisomy. A total of 54,119 gene regions and 160,209 promoter regions are ultimately found. Calculate the copy ratio of all genes and promoter regions using the following formula:

[0079]

[0080] Where copy_ratio ip The number of reads represents the copy ratio of region p of sample i. ip The length is the number of reads in region p of sample i that meet the quality control (MAPQ>30). p For the total length of region p, reads_number iLet lengt be the number of reads in sample i that meet the quality control (MAPQ>30). xref To use the total length of the reference genome, there are a total of n samples and m regions. Here, m represents 54,119 gene regions and 160,209 promoter regions, or a combination of both (214,328). The length of the gene regions ranges from 7 to 2,473,538 bp; the length of the promoter regions ranges from 199 to 43,798 bp.

[0081] The fourth step involves training a machine learning model, such as Ridge Regression, using the fetal nucleic acid concentration estimated from the Y chromosome data as the Y value and the copy ratios of all genes and promoters as the X values. The Ridge Regression model is implemented using the `LinearRegression` module in the `sklearn` Python package. Data from 2400 pregnant women with male fetuses were used for training. L2 regularization and cross-validation were used during training to obtain weight estimates for each region, and the trained Ridge Regression model was saved. In this model, the feature weights are equivalent to the weight coefficients β in the linear regression model, i.e.:

[0082] y i =β0+β1x i1 +β2x i2 +…+β n x ip i = 1, ..., p

[0083] Among them, y i Let {β0...} be the nucleic acid concentration of the male fetus calculated from the Y chromosome depth corresponding to sample i, and {β1...x} be the coefficients of this feature in the model. p} represents the copy ratio of all p genes and promoter regions in sample i.

[0084] In one embodiment of this invention, other machine learning or deep learning algorithms can also be used, such as least squares linear regression or lasso linear regression. The loss function for linear regression is the mean squared error, which is minimized during training. Lasso regression and ridge regression are essentially modified versions of standard linear regression, with L1 and L2 regularization added respectively. Regularization can be used to address the overfitting problem in linear regression. Furthermore, the regularization coefficients can be determined through cross-validation during model building to improve model accuracy and robustness.

[0085] In one embodiment of the present invention, the method of the present invention can introduce phenotypic data such as gestational age and age as features other than gene and promoter regions, and apply the same method to increase the accuracy of calculation.

[0086] The fifth step is to calculate the fetal nucleic acid concentration. For the maternal plasma cfDNA sample from which the fetal nucleic acid concentration is to be calculated, the copy ratio of the reads in the gene and promoter of the sample is calculated. The gene and promoter include the gene and promoter used to train the machine learning model. The fetal nucleic acid concentration is predicted and calculated using the model obtained in the fourth step.

[0087] Results: Using data from 2400 pregnant women with male fetuses as training data, the model was used to calculate the predicted fetal nucleic acid concentration in 600 pregnant women. The comparison with the fetal nucleic acid concentration calculated using Y chromosome depth showed... Figure 1 As shown in the figure, the correlation (R²) between the fetal nucleic acid concentration (x-axis) calculated based on Y chromosome depth and the fetal nucleic acid concentration (y-axis) calculated by the method of this invention is 0.83. Furthermore, the inventors performed calculations using individual gene regions, and the correlation R² with the results calculated using the Y chromosome-based method was 0.83; the accuracy value for calculations using individual promoter regions was 0.73.

Claims

1. A method for estimating fetal nucleic acid concentration in non-invasive prenatal genetic testing data, the method comprising: (1) Sequencing the cell-free nucleic acid of the pregnant woman to be tested to obtain sequencing data, wherein the sequencing data includes several reads; Preferably, the free nucleic acid is free DNA; (2) Align the reads to the reference genome and calculate the copy ratio of the reads in the gene regions and / or regulatory regions of the reference genome. (3) Input the copy ratio into the trained machine learning model to obtain the fetal nucleic acid concentration of the pregnant woman to be tested. The trained machine learning model is trained using data from multiple pregnant women carrying male fetuses. The data from these multiple pregnant women includes fetal nucleic acid concentrations calculated based on Y chromosome sequencing depth and the copy ratio of each read from each of the multiple pregnant women in the gene regions and / or regulatory regions of the reference genome.

2. The method according to claim 1, wherein the method for constructing the trained machine learning model in (3) comprises: (a) Obtaining sequencing data of cell-free nucleic acids from multiple pregnant women carrying male fetuses, wherein the sequencing data includes several reads; (b) Align the reads to a reference genome, and for each of the plurality of pregnant women, calculate the fetal nucleic acid concentration based on the Y chromosome sequencing depth, and calculate the copy ratio of the reads to be aligned to the reference genome in the gene region and / or regulatory region. (c) The fetal nucleic acid concentration calculated based on the Y chromosome sequencing depth and the copy ratio are input into the machine learning model for training to obtain the trained machine learning model.

3. The method according to claim 1 or 2, wherein in (1), the free nucleic acid is derived from the pregnant woman's peripheral blood plasma, the pregnant woman's liver, and / or placenta.

4. The method according to any one of claims 1-3, wherein in (1), the gene region and / or regulatory region originates from an autosome, preferably an autosome other than chromosomes 13, 18 and 21.

5. The method according to any one of claims 1-4, wherein the number of gene regions and / or regulatory regions is more than 10,000, preferably more than 50,000, more preferably more than 100,000, and most preferably more than 200,000.

6. The method according to any one of claims 1-4, wherein the control region is a promoter region.

7. The method according to any one of claims 1-6, wherein the machine learning model is a machine learning regression model, including linear regression models and nonlinear regression models, preferably ridge regression models, lasso regression models, least squares linear regression models, regression models based on random forest algorithms, or regression models based on deep neural networks.

8. The method of any one of claims 1-7, the Y chromosome depth calculated fetal nucleic acid concentration (Fraction fetal ) is: Depth Y is the average coverage depth for the Y chromosome, Depth autosomes is the average coverage depth for sequencing data on autosomes, preferably the autosomes do not include chromosomes 13, 18 and 21.

9. The method according to any one of claims 1-7, wherein the copy ratio of the reads compared to the reference genome in the gene region and / or regulatory region is: where copy_ratio ip is the copy ratio of region p of sample i, reads_number ip is the number of reads of region p of sample i, length p is the total length of region p, reads_number i is the number of reads of sample i, length ref is the total length of the reference genome, total number of samples n, total number of regions m.

10. A machine learning model constructed by the method according to any one of claims 2-9.

11. A system for estimating fetal nucleic acid concentration in non-invasive prenatal genetic testing data, the system comprising: The sequencing data acquisition module is used to sequence the cell-free nucleic acid of the pregnant woman to be tested and obtain sequencing data, wherein the sequencing data includes several reads; preferably, the cell-free nucleic acid is cell-free DNA; The copy ratio calculation module is used to align the read to a reference genome and calculate the copy ratio of the read aligned to the reference genome in gene regions and / or regulatory regions. The prediction module is used to input the copy ratio into a trained machine learning model to obtain the fetal nucleic acid concentration of the pregnant woman to be tested. The trained machine learning model is trained using data from multiple pregnant women carrying male fetuses. The training includes inputting the fetal nucleic acid concentration calculated based on the Y chromosome sequencing depth and the copy ratio of each read from the multiple pregnant women in the gene region and / or regulatory region of the reference genome into the machine learning model for training, thereby obtaining the trained machine learning model.

12. The system according to claim 11, wherein in the sequencing data acquisition module, the free nucleic acid is derived from the pregnant woman's peripheral plasma, the pregnant woman's liver, and / or placenta.

13. The system according to claim 11 or 12, wherein the gene region and / or regulatory region originates from an autosome, preferably an autosome other than chromosomes 13, 18 and 21.

14. The system according to any one of claims 11-13, wherein the number of gene regions and / or regulatory regions is more than 10,000, preferably more than 50,000, more preferably more than 100,000, and most preferably more than 200,000.

15. The system according to any one of claims 11-14, wherein the control region is a startup sub-region.

16. The system according to any one of claims 11-15, wherein the machine learning model is a machine learning regression model, including linear regression models and nonlinear regression models, preferably ridge regression models, lasso regression models, least squares linear regression models, regression models based on random forest algorithms, or regression models based on deep neural networks.

17. The system of any one of claims 11-16, the fetal nucleic acid concentration FractionY chromosome depth calculated as: fetal FractionY chromosome depth = 2 * (Y chromosome depth - 0.5) Depth Y is the average coverage depth for the Y chromosome, Depth autosomes is the average coverage depth for sequencing data on autosomes, preferably the autosomes do not include chromosomes 13, 18 and 21.

18. The system according to any one of claims 11-17, wherein the copy ratio of the reads compared to the reference genome in the gene region and / or regulatory region is: copy_ratio i p is the copy ratio of region p of sample i, reads_number ip reads_number is the number of reads of region p of sample i, length p length is the total length of region p, reads_number i reads_number is the number of reads of sample i, length ref length is the total length of the reference genome, sample total n, region total m.