Simulation method, system and equipment for free nucleic acid data of plasma of pregnant woman and medium

By performing data analysis and polynomial fitting on mother-child paired sample files, the fetal free nucleic acid concentration in pregnant women is simulated, which solves the problem of difficulty in effectively simulating the concentration of specific fetal free nucleic acid in the prior art, improving detection accuracy and reducing costs.

CN119943152AActive Publication Date: 2025-05-06CAPITALBIO GENOMICS
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411886581.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-06
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

Existing non-invasive prenatal screening technology is difficult to effectively simulate plasma samples of specific fetal free nucleic acid concentrations, resulting in high detection accuracy and cost.

Method used

By obtaining the family sample files of mother-child pairing, data volume analysis is performed based on the preset fetal free nucleic acid concentration, the sampling ratio is obtained, and the sampling ratio is corrected by polynomial fit to simulate the plasma free nucleic acid data of pregnant women.

Benefits of technology

It reduces sampling and experiment time and cost, improves the accuracy of plasma simulates the proportion of fetal free nucleic acid concentrations, and enhances the flexibility of obtaining positive samples in non-invasive prenatal screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943152A_ABST
    Figure CN119943152A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a simulation method, system and equipment for free nucleic acid data of plasma of a pregnant woman and a medium, and belongs to the technical field of noninvasive prenatal screening. Acquiring a parent-child paired family sample file; performing data volume analysis on the family sample file according to a preset fetal free nucleic acid concentration to obtain a first sampling proportion; sampling and merging the family sample files according to a first sampling proportion to obtain simulated first pregnant woman plasma free nucleic acid data; performing fetal free nucleic acid concentration prediction on the first pregnant woman plasma free nucleic acid data to obtain a prediction result; performing polynomial fitting according to the prediction result and the first sampling proportion to obtain fitting parameters; taking a preset fetal free nucleic acid concentration as input of a polynomial, and determining a second sampling proportion according to the fitting parameter; sampling and merging the family sample files again according to the second sampling proportion to obtain simulated second pregnant woman plasma free nucleic acid data, so that the cost can be saved, and the plasma simulation accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of non-invasive prenatal screening, and in particular to a method, system, device and medium for simulating free nucleic acid data in the plasma of pregnant women. Background Art

[0002] Chromosomal abnormalities and genetic diseases are important causes of birth defects. Prenatal screening can reduce birth defects and improve the quality of the newborn population. Non-invasive prenatal screening performs non-invasive fetal genome testing by analyzing whether the fetal free nucleic acid in the plasma of pregnant women carries single gene diseases or chromosomal abnormality signals. Generally, there is a positive correlation between the concentration of fetal free nucleic acid in the plasma of pregnant women and the detection value of positive samples. The higher the concentration of fetal free nucleic acid, the higher the accuracy of non-invasive fetal genome testing. Therefore, non-invasive prenatal screening products often need to collect positive samples with different fetal concentration gradients during the research and development and verification stages in order to study the sensitivity and specificity of the test kit and detection algorithm at different concentrations.

[0003] However, the incidence of most monogenic and chromosomal diseases is low in the population. In this case, it is difficult to obtain clinically true positive plasma samples with specific fetal free nucleic acid concentrations. At present, the existing technology often obtains standard products that simulate specific fetal free nucleic acid concentration ratios by preparing simulated plasma samples. However, this simulation method not only requires the collection of a large number of positive samples in advance, but also when there is an error in the concentration of the standard preparation, it can only be discovered in the quantification or sequencing stage after the preparation is completed, affecting the accuracy of the fetal free nucleic acid concentration of the simulated plasma sample. Summary of the invention

[0004] The main purpose of the embodiments of the present application is to propose a method, system, equipment and medium for simulating free nucleic acid data in pregnant women's plasma, aiming to reduce the time and preparation costs spent on sampling and experiments, and to improve the accuracy of the ratio of plasma simulated fetal free nucleic acid concentration.

[0005] To achieve the above purpose, one aspect of the embodiment of the present application proposes a method for simulating plasma free nucleic acid data of pregnant women, the method comprising:

[0006] Obtain pedigree sample files for mother-child pairings;

[0007] Performing data volume analysis on the family sample file according to a preset fetal free nucleic acid concentration to obtain a first sampling ratio;

[0008] Sampling and merging the family sample files according to the first sampling ratio to obtain simulated plasma free nucleic acid data of the first pregnant woman;

[0009] Predicting the fetal free nucleic acid concentration based on the first pregnant woman's plasma free nucleic acid data to obtain a prediction result;

[0010] Perform polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters;

[0011] Using the preset fetal free nucleic acid concentration as the input of the polynomial, and determining the second sampling ratio according to the fitting parameters;

[0012] The family sample files are resampled and merged according to the second sampling ratio to obtain simulated plasma free nucleic acid data of the second pregnant woman.

[0013] In some embodiments, the family sample file includes a mother sample file and a fetal sample file, the first sampling ratio includes a first mother ratio and a first fetal ratio, and the data volume analysis of the family sample file according to the preset fetal free nucleic acid concentration to obtain the first sampling ratio includes the following steps:

[0014] Performing data volume statistics on the mother sample file and the fetus sample file respectively to obtain the mother data volume and the fetus data volume;

[0015] Determining a simulated mixed data volume according to a preset fetal free nucleic acid concentration, wherein the product of the mixed data volume and the preset fetal free nucleic acid concentration does not exceed the fetal data volume;

[0016] Determining the maternal sampling amount and the fetal sampling amount according to the preset fetal free nucleic acid concentration and the mixed data amount;

[0017] The first mother ratio is obtained according to the ratio of the mother sampling amount to the mother data amount, and the first fetus ratio is obtained according to the ratio of the fetus sampling amount to the fetus data amount.

[0018] In some embodiments, the sampling and merging of the family sample files according to the first sampling ratio to obtain simulated first pregnant woman plasma free nucleic acid data includes the following steps:

[0019] Randomly sampling the mother sample file according to the first mother ratio to obtain a first mother sampling file;

[0020] Randomly sampling the fetal sample file according to the first fetal ratio to obtain a first fetal sampling file;

[0021] The first mother sampling file and the first fetus sampling file are mixed to obtain simulated first pregnant woman plasma free nucleic acid data.

[0022] In some embodiments, the step of predicting the fetal free nucleic acid concentration based on the first pregnant woman's plasma free nucleic acid data to obtain a prediction result comprises the following steps:

[0023] Determining the frequency range of genotype occurrence frequency according to the preset concentration range of fetal free nucleic acid concentration;

[0024] According to the preset high-frequency heterozygous single nucleotide polymorphism sites, single nucleotide polymorphism detection is performed on the plasma free nucleic acid data of the first pregnant woman, and the single nucleotide polymorphism sites whose genotype occurrence frequency is within the frequency range are collected to obtain a frequency site set;

[0025] According to the frequency site set, the probability of each frequency value within the frequency range is calculated in a gradient manner by a statistical method, and the frequency value with the largest probability is determined as the actual fetal free nucleic acid concentration to obtain a prediction result.

[0026] In some embodiments, performing polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters comprises the following steps:

[0027] According to the prediction results corresponding to different preset fetal free nucleic acid concentrations and the first mother ratio, fitting the relationship between the prediction results and the first mother ratio through a polynomial model to obtain a mother fitting parameter;

[0028] According to the predicted results corresponding to different preset fetal free nucleic acid concentrations and the first fetal proportion, a polynomial model is used to fit the relationship between the predicted results and the first fetal proportion to obtain fetal fitting parameters.

[0029] In some embodiments, the second sampling ratio includes a second maternal ratio and a second fetal ratio, and the step of using the preset fetal free nucleic acid concentration as an input of the polynomial and determining the second sampling ratio according to the fitting parameters comprises the following steps:

[0030] Using the preset fetal free nucleic acid concentration as an input of a polynomial, determining a second mother proportion for extracting data from the mother sample file according to the mother fitting parameter;

[0031] The preset fetal free nucleic acid concentration is used as an input of a polynomial, and the second fetal proportion of data extracted from the fetal sample file is determined according to the fetal fitting parameters.

[0032] In some embodiments, the method for simulating plasma free nucleic acid data of pregnant women further comprises the following steps:

[0033] Enriching the free nucleic acid data of the plasma free nucleic acid data of the second pregnant woman whose inserted fragment length is less than a preset screening threshold value to obtain short fragment data;

[0034] The fetal free nucleic acid concentration is predicted for the short fragment data to obtain an enriched prediction result.

[0035] To achieve the above purpose, another aspect of the embodiment of the present application provides a simulation system for plasma free nucleic acid data of pregnant women, the system comprising:

[0036] The first module is used to obtain the pedigree sample files of mother-child pairing;

[0037] The second module is used to perform data volume analysis on the family sample file according to a preset fetal free nucleic acid concentration to obtain a first sampling ratio;

[0038] The third module is used to sample and merge the family sample files according to the first sampling ratio to obtain simulated plasma free nucleic acid data of the first pregnant woman;

[0039] The fourth module is used to predict the fetal free nucleic acid concentration based on the first pregnant woman's plasma free nucleic acid data to obtain a prediction result;

[0040] A fifth module is used to perform polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters;

[0041] A sixth module is used to use the preset fetal free nucleic acid concentration as an input of a polynomial and determine a second sampling ratio according to the fitting parameters;

[0042] The seventh module is used to resample and merge the family sample files according to the second sampling ratio to obtain simulated plasma free nucleic acid data of the second pregnant woman.

[0043] To achieve the above objective, another aspect of an embodiment of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above method when executing the computer program.

[0044] To achieve the above objective, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above method when executed by a processor.

[0045] The embodiments of the present application include at least the following beneficial effects: The present application provides a method, system, device and medium for simulating plasma free nucleic acid data of pregnant women. The scheme obtains a pedigree sample file of a mother-child pair; performs a data volume analysis on the pedigree sample file according to a preset fetal free nucleic acid concentration to obtain a first sampling ratio; samples and merges the pedigree sample file according to the first sampling ratio to obtain simulated first plasma free nucleic acid data of pregnant women; predicts the fetal free nucleic acid concentration of the first plasma free nucleic acid data of pregnant women to obtain a prediction result; performs polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters; uses the preset fetal free nucleic acid concentration as the input of the polynomial, and determines the second sampling ratio according to the fitting parameters; samples and merges the pedigree sample file again according to the second sampling ratio to obtain simulated second plasma free nucleic acid data of pregnant women, which can reduce the time spent on sampling and experiments and the preparation cost, reduce the complexity of experiments, improve the accuracy of the plasma simulated fetal free nucleic acid concentration ratio, and can batch simulate pregnant women's plasma data with any fetal free nucleic acid concentration ratio in a short time, which helps to improve the flexibility of non-invasive prenatal screening to obtain positive samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a flow chart of a method for simulating plasma free nucleic acid data of pregnant women provided in an embodiment of the present application;

[0047] Figure 2 This is a schematic diagram of the genotype combination in the peripheral blood of pregnant women provided in the embodiments of the present application;

[0048] Figure 3 This is a scatter plot of the first simulated fetal free nucleic acid concentration provided in the embodiment of the present application;

[0049] Figure 4 is a scatter plot of the second simulation of fetal free nucleic acid concentration provided in the embodiment of the present application;

[0050] Figure 5 is a flow chart of a method for simulating plasma free nucleic acid data of pregnant women provided in another embodiment of the present application;

[0051] Figure 6 This is a result schematic diagram of the amount of mother-child pairing sample data and the length of the inserted fragment provided in the embodiment of the present application;

[0052] Figure 7 is a schematic diagram of the result of the first sampling ratio provided in an embodiment of the present application;

[0053] Figure 8 is a schematic diagram of the comparison between the actual fetal free nucleic acid concentration and the preset fetal free nucleic acid concentration in the first simulation provided by the embodiment of the present application;

[0054] Fig. 9It is a scatter plot of the fetal free DNA concentration and the offspring sampling ratio of different families provided in the examples of the present application;

[0055] Fig.10 is a schematic diagram of the results of different family fitting parameters provided in the examples of the present application;

[0056] Fig.11 is a schematic diagram of the result of the second sampling ratio provided in the embodiment of the present application;

[0057] Fig.12 is a schematic diagram of the comparison between the actual fetal free nucleic acid concentration and the preset fetal free nucleic acid concentration in the second simulation provided by the embodiment of the present application;

[0058] Fig.13 is a distribution diagram of the length of the insert fragment of the first family provided in the examples of the present application;

[0059] Fig.14 This is a graph showing the experimental results of the concentration of fetal free nucleic acid before and after enrichment of short fragments provided in the examples of the present application;

[0060] Fig.15 is a schematic diagram of the structure of a simulation system for pregnant women's plasma free nucleic acid data provided in an embodiment of the present application;

[0061] Fig.16 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.

[0063] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".

[0064] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0066] Before describing the embodiments of the present application in detail, the terms involved in the embodiments of the present application and some related technologies are described as follows.

[0067] (1) Read length: Reads refer to the base sequence obtained by a single sequencing by a sequencer. Due to the limitations of current sequencing capabilities, the genome needs to be broken into DNA fragments before sequencing. Different sequencing instruments have different read lengths.

[0068] (2) BAM (Binary Alignment / Map format): BAM is the most common alignment data storage format in genetic data analysis. It is used to store binary files of large-scale sequencing data. BAM files are usually used to store alignment information of sequencing data, such as the alignment position of DNA sequencing reads with the reference genome and the corresponding quality value.

[0069] (3) Single Nucleotide Polymorphism (SNP): Single nucleotide polymorphism mainly refers to DNA sequence polymorphism caused by the variation of a single nucleotide at the genome level. It is the most common type of heritable variation in humans, accounting for more than 90% of all known polymorphisms. SNP is widely present in the human genome, with an average of 1 in every 300 base pairs. When a base pair is converted or transposed, inserted or deleted, it is a SNP site.

[0070] Relevant studies have shown that the total incidence of birth defects in my country is about 5.6%. Based on the annual birth number of 16 million in China, the number of new birth defects reaches 900,000 cases each year, of which 60% are related to chromosomal abnormalities and genetic diseases. Therefore, prenatal screening is an important means to reduce birth defects and improve the quality of the newborn population. However, although traditional prenatal screening methods such as amniocentesis, chorionic villus sampling, and umbilical vein puncture are highly accurate, they are accompanied by a 0.2% to 0.5% risk of miscarriage and infection, which makes non-invasive prenatal screening technology a hot topic of research.

[0071] In 1997, scientists discovered the presence of free fetal nucleic acid (DNA) in pregnant women’s plasma, which laid the theoretical foundation for non-invasive prenatal screening of fetal genomic abnormalities using maternal plasma.

[0072] Subsequently, the application of high-throughput sequencing technology further promoted the development of non-invasive fetal chromosomal aneuploidy genetic testing (NIPT), which can screen fetal chromosomal aneuploidy abnormalities by detecting free DNA in maternal peripheral blood.

[0073] Since then, maternal peripheral blood free DNA has also been shown to be useful for detecting fetal microdeletion and microduplication syndromes as well as fetal monogenic diseases.

[0074] In 2022, a team of scientists pioneered a "three-in-one" comprehensive non-invasive prenatal screening technology that can simultaneously screen for chromosomal aneuploidy, chromosomal microdeletion syndrome and monogenic dominant genetic diseases. It uses specific molecular label (UMI) double-end library capture sequencing technology to accurately calculate the number of DNA molecules in fetal free DNA. Combined with fetal concentration information, low-level fetal mutation sites can be screened out.

[0075] The detection accuracy of the above prenatal technologies is mainly affected by the concentration of fetal free DNA. Therefore, accurate assessment of the proportion of fetal free DNA concentration is very critical for non-invasive prenatal screening.

[0076] The free DNA in the plasma of pregnant women includes both the free DNA of the pregnant women themselves and the free DNA of the fetus, of which the concentration of the free DNA of the fetus is generally within 3% to 30%, with an average of about 13%. Under normal circumstances, the concentration of the free DNA of the fetus in the plasma of pregnant women is positively correlated with the detection value of the positive sample, that is, the higher the concentration of the free DNA of the fetus, the higher the accuracy of the non-invasive fetal genomic disease detection. Taking NIPT as an example, a too low concentration of the free DNA of the fetus will increase the risk of missing positive samples. The detection limit of NIPT is 3% to 4%, but due to individual differences in pregnant women, sample quality, experimental errors, etc., in the prior art, when the concentration of the free DNA of the fetus is lower than 5%, the NIPT test result may be a false negative.

[0077] However, the incidence of most single gene diseases and chromosomal diseases is low. For example, among the three common aneuploidy abnormalities, trisomy 21 (Down syndrome), trisomy 18 (Edwards syndrome) and trisomy 13 (Patau syndrome) are the three most common autosomal aneuploidy diseases, with incidences of 1 / (600-800), 1 / (3500-8000) and 1 / (7000-20000) in newborns, respectively. In this case, it is difficult to obtain clinically real positive plasma samples with specific fetal concentrations to fully study the detection performance.

[0078] At present, in order to solve the problem of difficulty in obtaining positive samples, most studies usually mix a certain proportion of positive or negative DNA into plasma samples to prepare simulated plasma samples. However, this simulation method has the following limitations:

[0079] (1) The experimental operation is highly complicated and requires the collection of a large number of positive DNA and negative plasma samples in advance for preparation.

[0080] (2) The problem of concentration error in the preparation of standard products, that is, there is an error between the actual fetal concentration measured after preparation and the expected preparation concentration. These errors usually need to be discovered in the quantification or sequencing stage after the preparation is completed, which may result in the need to re-prepare the entire batch of standard products.

[0081] (3) The preparation cost is high, and the cost increases proportionally with the density of the fetal concentration gradient. It is difficult to obtain the results of a high-density fetal concentration gradient (such as 1%) for detailed analysis.

[0082] In view of this, a method, system, device and medium for simulating plasma free nucleic acid data of pregnant women are provided in an embodiment of the present application. The scheme obtains a pedigree sample file of a mother-child pair; performs a data volume analysis on the pedigree sample file according to a preset fetal free nucleic acid concentration to obtain a first sampling ratio; samples and merges the pedigree sample file according to the first sampling ratio to obtain simulated first plasma free nucleic acid data of pregnant women; predicts the fetal free nucleic acid concentration of the first plasma free nucleic acid data of pregnant women to obtain a prediction result; performs polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters; uses the preset fetal free nucleic acid concentration as the input of the polynomial, and determines the second sampling ratio according to the fitting parameters; samples and merges the pedigree sample file again according to the second sampling ratio to obtain simulated second plasma free nucleic acid data of pregnant women, which can reduce the time and preparation cost spent on sampling and experiments, reduce the complexity of experiments, and improve the accuracy of the plasma simulated fetal free nucleic acid concentration ratio. Pregnant plasma data of pregnant women with any fetal free nucleic acid concentration ratio can be simulated in batches in a short time, which helps to improve the flexibility of obtaining positive samples in non-invasive prenatal screening.

[0083] The simulation method of pregnant women's plasma free nucleic acid data provided in the embodiment of the present application relates to the field of non-invasive prenatal screening technology. The simulation method of pregnant women's plasma free nucleic acid data provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or it can be configured as a server cluster or distributed system composed of multiple physical servers, and can also be configured to provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms and other basic cloud computing services. The cloud server, the server can also be a node server in the blockchain network; the software can be an application of the simulation method of pregnant women's plasma free nucleic acid data, etc., but is not limited to the above form.

[0084] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0085] Figure 1 is an optional flow chart of the simulation method of pregnant women's plasma free nucleic acid data provided in the embodiment of the present application, Figure 1 The method may include but is not limited to steps S101 to S107.

[0086] Step S101, obtaining a pedigree sample file of mother-child pairing.

[0087] Step S102, performing data volume analysis on the family sample file according to the preset fetal free nucleic acid concentration to obtain a first sampling ratio.

[0088] Step S103, sampling and merging the family sample files according to the first sampling ratio to obtain simulated plasma free nucleic acid data of the first pregnant woman.

[0089] Step S104, predicting the fetal free nucleic acid concentration based on the first pregnant woman's plasma free nucleic acid data to obtain a prediction result.

[0090] Step S105, performing polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters.

[0091] Step S106, using the preset fetal free nucleic acid concentration as the input of the polynomial, and determining the second sampling ratio according to the fitting parameters.

[0092] Step S107, re-sampling and merging the family sample files according to the second sampling ratio to obtain simulated plasma free nucleic acid data of the second pregnant woman.

[0093] Specifically, first, obtain the family sample file of the mother-child pairing, process the DNA samples of the mother and fetus in the same family through high-throughput sequencing and bioinformatics process analysis, align each DNA sample to the reference genome, record the sequencing reads data to generate the corresponding BAM file, and obtain the family sample file.

[0094] Next, the data volume of the family sample file is analyzed according to the preset fetal free nucleic acid concentration to obtain the first sampling ratio. According to the preset fetal free nucleic acid concentration, the amount of data to be extracted from the mother and fetus BAM files is calculated, and the first sampling ratio is obtained based on the proportion of the extracted data volume in the corresponding BAM file.

[0095] After obtaining the first sampling ratio, the family sample files are sampled and merged according to the first sampling ratio to obtain the simulated first maternal plasma free nucleic acid data. According to the first sampling ratio, the first data sampling is performed on the BAM files of the mother and the fetus respectively, and the BAM files obtained by each sampling are merged to generate the first computer simulated mixed maternal plasma BAM file to obtain the first maternal plasma free nucleic acid data.

[0096] Then, the fetal free nucleic acid concentration is predicted for the first pregnant woman's plasma free nucleic acid data to obtain a prediction result. The simulated first pregnant woman's plasma free nucleic acid data is subjected to SNP detection to detect all possible SNP sites in the data, and the fetal free nucleic acid concentration is predicted using the SNP sites, and the actual fetal free nucleic acid concentration in the simulated data is calculated to obtain a prediction result.

[0097] Furthermore, after obtaining the prediction result, the actual fetal free nucleic acid concentration is compared with the preset fetal free nucleic acid concentration, and it is found that the actual simulated fetal free nucleic acid concentration is larger than the preset fetal free nucleic acid concentration. In order to correct the deviation of the fetal free nucleic acid concentration introduced in the simulation process, a polynomial fitting is performed according to the prediction result and the first sampling ratio to obtain the fitting parameters. The calculated actual fetal free nucleic acid concentration (denoted as cff) is combined with the first sampling ratio (denoted as R) to perform polynomial fitting. According to multiple data points (cff, R) corresponding to different preset fetal free nucleic acid concentrations, the coefficients of each order of the polynomial are solved using the least squares method, and the optimal fitting parameters are found by minimizing the residual sum of squares between the data points and the fitting function.

[0098] The preset fetal free nucleic acid concentration is used as the input of the polynomial, and the second sampling ratio is determined according to the fitting parameters. The obtained fitting parameters are used to establish a polynomial to restore the variation trend between the sampling ratio and the actual fetal free nucleic acid concentration, and the preset fetal free nucleic acid concentration is used as the input of the polynomial, and the calibrated second sampling ratio is calculated according to the polynomial.

[0099] The family sample files are resampled and merged according to the second sampling ratio to obtain the simulated second pregnant woman plasma free nucleic acid data. The mother and fetus BAM files are sampled again according to the second sampling ratio, and finally the separately sampled BAM files are merged to obtain the simulated second pregnant woman plasma free nucleic acid data. By predicting the fetal concentration of the second pregnant woman plasma free nucleic acid data obtained by the second simulation, it can be found that the predicted result is close to the preset fetal free nucleic acid concentration. At this time, the simulation of the pregnant woman plasma free nucleic acid data is successful.

[0100] In this embodiment, only the sequencing data of the mother and the fetus need to be obtained to simulate the plasma of pregnant women, which can reduce the time and preparation cost spent on sampling and experiments and reduce the complexity of experiments. By fitting the variation trend between the sampling ratio and the actual fetal free nucleic acid concentration and determining the secondary sampling ratio, the accuracy of the plasma simulation of the fetal free nucleic acid concentration ratio can be improved, and the plasma data of pregnant women with any fetal free nucleic acid concentration ratio can be simulated in batches in a short time, which helps to improve the flexibility of non-invasive prenatal screening to obtain positive samples.

[0101] In some embodiments, the pedigree sample file includes a mother sample file and a fetus sample file, the first sampling ratio includes a first mother ratio and a first fetus ratio, and step S102 may include but is not limited to steps S201 to S204.

[0102] Step S201, respectively perform data volume statistics on the mother sample file and the fetus sample file to obtain the mother data volume and the fetus data volume.

[0103] Step S202, determining the simulated mixed data volume according to the preset fetal free nucleic acid concentration, wherein the product of the mixed data volume and the preset fetal free nucleic acid concentration does not exceed the fetal data volume.

[0104] Step S203, determining the maternal sampling amount and the fetal sampling amount according to the preset fetal free nucleic acid concentration and the mixed data amount.

[0105] Step S204, obtaining a first mother ratio according to the ratio of the mother sampling amount to the mother data amount, and obtaining a first fetus ratio according to the ratio of the fetus sampling amount to the fetus data amount.

[0106] Specifically, the data volume statistics of the mother sample file and the fetus sample file are respectively performed to obtain the mother data volume and the fetus data volume. The data volume statistics of the reads data contained in the mother sample file are performed to obtain the mother data volume (denoted as Mm), and at the same time, the data volume statistics of the reads data contained in the fetus sample file are performed to obtain the fetus data volume (denoted as Mf).

[0107] The amount of mixed data to be simulated is determined according to the preset fetal free nucleic acid concentration, wherein the product of the mixed data amount and the preset fetal free nucleic acid concentration does not exceed the amount of fetal data. In order to simulate the plasma data of pregnant women under different fetal free nucleic acid concentrations, the fetal free nucleic acid concentration is preset (denoted as ff) to generate simulated data of any desired fetal free nucleic acid concentration. In order to ensure that during the simulation process, enough data can be extracted from the fetal sample file to reach the preset fetal free nucleic acid concentration, it is necessary to determine the amount of mixed data to be simulated (denoted as Ms) according to the preset fetal free nucleic acid concentration.

[0108] It should be noted that the setting of the mixed data volume needs to meet the condition that the product of the mixed data volume and the preset fetal free nucleic acid concentration does not exceed the fetal data volume. If the fetal data volume required for the preset fetal free nucleic acid concentration exceeds the actual fetal data volume of the fetal sample file, then effective data sampling cannot be performed, and it is impossible to simulate the maternal plasma data that meets the requirements. Therefore, the value obtained by dividing the fetal data volume by the preset fetal free nucleic acid concentration is determined as the maximum data volume of the simulation, and then the simulated mixed data volume is set not to exceed the maximum data volume.

[0109] Next, the maternal sampling amount and the fetal sampling amount are determined according to the preset fetal free nucleic acid concentration and the mixed data amount. The maternal sampling amount is the amount of data to be extracted from the maternal sample file, and the fetal sampling amount is the amount of data to be extracted from the fetal sample file.

[0110] Exemplarily, the maternal sampling amount extracted from the maternal sample file is determined to be Ms×(1-ff) based on the preset fetal free nucleic acid concentration and the mixed data volume, and the fetal sampling amount extracted from the fetal sample file is determined to be Ms×ff based on the preset fetal free nucleic acid concentration and the mixed data volume.

[0111] Finally, the first mother ratio is obtained according to the ratio of the mother sampling amount to the mother data amount, and the first fetal ratio is obtained according to the ratio of the fetal sampling amount to the fetal data amount. The mother sampling amount and the fetal sampling amount are respectively compared with the corresponding BAM file data amount. According to the obtained mother sampling amount, the sampling proportion of the mother data amount in the mother sample file is calculated to obtain the first mother ratio Rm=Ms×(1-ff) / Mm. At the same time, according to the obtained fetal sampling amount, the sampling proportion of the fetal data amount in the fetal sample file is calculated to obtain the first fetal ratio Rf=Ms×ff / Mf.

[0112] In some embodiments, step S103 may include but is not limited to steps S301 to S303.

[0113] Step S301, randomly sampling the mother sample file according to the first mother ratio to obtain the first mother sampling file.

[0114] Step S302: Randomly sample the fetal sample file according to the first fetal ratio to obtain the first fetal sampling file.

[0115] Step S303, mixing the first mother sampling file and the first fetus sampling file to obtain simulated first pregnant woman plasma free nucleic acid data.

[0116] Specifically, the mother sample file is randomly sampled according to the first mother ratio to obtain the first mother sampling file, and the fetal sample file is randomly sampled according to the first fetal ratio to obtain the first fetal sampling file. The mother sample file is sampled according to the first mother ratio obtained in step S204, and a corresponding number of reads data are randomly selected from the mother sample file to obtain the first mother sampling file. At the same time, the fetal sample file is sampled according to the first fetal ratio obtained in step S204, and a corresponding number of reads data are randomly selected from the fetal sample file to obtain the first fetal sampling file.

[0117] Exemplarily, the DownsampleSam module in the bioinformatics tool Picard is used to perform data sampling on the family sample file, and the calculated first mother ratio and first fetus ratio are input as parameters into the DownsampleSam module of Picard to perform data sampling on the BAM files of the mother and fetus respectively, to generate corresponding sampling files.

[0118] Furthermore, the first mother sampling file and the first fetus sampling file are mixed to obtain simulated first pregnant woman plasma free nucleic acid data. The first mother sampling file and the first fetus sampling file obtained by each sampling are mixed, and the two sampling files are combined to obtain the first pregnant woman plasma free nucleic acid data for simulating the concentration of fetal free nucleic acid in real pregnant woman plasma.

[0119] In some embodiments, step S104 may include but is not limited to steps S401 to S403.

[0120] Step S401, determining the frequency range of genotype occurrence frequency according to a preset concentration range of fetal free nucleic acid concentration.

[0121] Step S402, performing single nucleotide polymorphism detection on the plasma free nucleic acid data of the first pregnant woman according to preset high-frequency heterozygous single nucleotide polymorphism sites, collecting single nucleotide polymorphism sites whose genotype occurrence frequency is within the frequency range, and obtaining a frequency site set.

[0122] Step S403, based on the frequency site set, the probability of each frequency value within the frequency range is calculated by a statistical method in a gradient manner, and the frequency value with the largest probability is determined as the predicted value of the fetal free nucleic acid concentration to obtain a predicted result.

[0123] Specifically, the frequency range of the genotype occurrence frequency is determined according to the preset concentration range of the fetal free nucleic acid concentration. The genotype in the peripheral blood of pregnant women is mainly composed of the genotypes of the mother and the fetus. Figure 2 , genotype A and genotype B represent the most frequent allele and the second most frequent allele in a single SNP locus, respectively. When the SNP genotype consists of two genotype A, it is a homozygous genotype, and when the SNP genotype consists of one genotype A and one genotype B, it is a heterozygous genotype.

[0124] Since the concentration of fetal free nucleic acid in pregnant women's plasma is usually between 3% and 30%, it is necessary to determine a reasonable frequency range of genotype occurrence based on the preset concentration range of fetal free nucleic acid concentration to ensure that the SNP sites screened out subsequently have a strong correlation with the fetal free nucleic acid concentration.

[0125] Furthermore, based on the preset high-frequency heterozygous single nucleotide polymorphism sites, single nucleotide polymorphism detection is performed on the plasma free nucleic acid data of the first pregnant woman, and single nucleotide polymorphism sites with genotype occurrence frequencies within the frequency range are collected to obtain a frequency site set.

[0126] First, high-frequency heterozygous SNP sites are selected from the SNP sites in the population, and these sites are preset as SNP sites for detection.

[0127] Next, the SNP detection is performed on the plasma free nucleic acid data of the first pregnant woman using a bioinformatics tool (such as Samtools software), and the SNP sites whose genotype occurrence frequencies are within the frequency range are screened using the frequency range determined in step S401 to obtain a frequency site set.

[0128] Finally, according to the frequency site set, the probability of each frequency value in the frequency range is calculated by gradient through statistical methods, and the frequency value with the largest probability is determined as the predicted value of fetal free nucleic acid concentration to obtain the predicted result.

[0129] Optionally, after collecting the frequency site set, the binomial distribution probability calculation and the statistical method of maximum likelihood estimation can be combined to calculate the probability of each frequency value within the frequency range according to a certain gradient, and the frequency value with the highest probability is determined as the predicted value of the fetal free nucleic acid concentration to obtain a predicted result.

[0130] Exemplarily, the simulated plasma free nucleic acid data of the first pregnant woman are subjected to SNP detection using the mpileup module of the Samtools software, and the fetal concentration is predicted using high-frequency heterozygous SNP sites in 820 populations to obtain prediction results.

[0131] Combining the maternal and fetal SNP genotypes can produce four types of maternal peripheral blood. Figure 2 Taking the second type of SNP genotype combination as an example (i.e., the mother is a homozygous genotype and the fetus is a heterozygous genotype), the fetal free nucleic acid concentration is predicted. In the second case, the gene frequency of genotype B in the fetus is positively correlated with the fetal free nucleic acid concentration. Since the mean frequency of genotype B in the fetus is half of the fetal free nucleic acid concentration, when the preset fetal free nucleic acid concentration range is set to 3% to 30%, it can be seen that the genotype B frequency range to be collected is 1.5% to 15%.

[0132] SNP sites with a genotype B frequency of 1.5% to 15% were collected, and the probability values ​​were calculated using statistical methods within the concentration range of 1.5% to 15% with a gradient of 0.01%. When the probability value reached the maximum, the corresponding frequency value was the actual fetal free nucleic acid concentration predicted based on the plasma free nucleic acid data of the first pregnant woman.

[0133] In some embodiments, step S105 may include but is not limited to step S501 and step S502.

[0134] Step S501, according to the prediction results and the first mother ratio corresponding to different preset fetal free nucleic acid concentrations, the relationship between the prediction results and the first mother ratio is fitted by a polynomial model to obtain the mother fitting parameters.

[0135] Step S502, according to the prediction results and the first fetal proportion corresponding to different preset fetal free nucleic acid concentrations, a relationship between the prediction results and the first fetal proportion is fitted by a polynomial model to obtain fetal fitting parameters.

[0136] Specifically, refer to Figure 3 Since the actual fetal free nucleic acid concentration predicted based on the first pregnant woman’s plasma free nucleic acid data is larger than the theoretical preset free nucleic acid concentration, in order to correct the deviation of the fetal free nucleic acid concentration introduced during the simulation process, it is necessary to use the calculated actual fetal free nucleic acid concentration data combined with the sampling ratios of the mother and the fetus to perform polynomial fitting respectively.

[0137] In statistics and data analysis, polynomial fitting is often used to explore the relationship between variables. According to the prediction results and the first mother ratio corresponding to different preset fetal free nucleic acid concentrations, the relationship between the prediction results and the first mother ratio is fitted by a polynomial model to obtain the mother fitting parameters; according to the prediction results and the first fetal ratio corresponding to different preset fetal free nucleic acid concentrations, the relationship between the prediction results and the first fetal ratio is fitted by a polynomial model to obtain the fetal fitting parameters.

[0138] The actual predicted fetal free nucleic acid concentration (cff) under each preset fetal nucleic acid concentration is matched with the first mother ratio (Rm) used in the first simulation, and the relationship between the predicted result and the first mother ratio is fitted by a polynomial model to obtain the mother fitting parameters.

[0139] The actual predicted fetal free nucleic acid concentration (cff) under each preset fetal nucleic acid concentration is matched with the first fetal ratio (Rf) used in the first simulation, and the relationship between the predicted result and the first fetal ratio is fitted by a polynomial model to obtain the fetal fitting parameters.

[0140] Exemplarily, a polynomial quadratic model is established, and the expression of the polynomial quadratic model is:

[0141] R = cff 2 ×a+cff×b+c (1),

[0142] Among them, R represents the first sampling ratio, cff represents the actual fetal free nucleic acid concentration (ie, the predicted result), a represents the quadratic term coefficient, b represents the linear term coefficient, and c represents the constant term coefficient.

[0143] Using the polynomial quadratic model, according to the prediction results and the first mother ratio corresponding to different preset fetal free nucleic acid concentrations, the relationship between the prediction results and the first mother ratio was fitted to obtain the quadratic term coefficient Ma, the linear term coefficient Mb and the constant term coefficient Mc in the mother fitting parameters.

[0144] Using the polynomial quadratic model, according to the predicted results and the first fetal proportion corresponding to different preset fetal free nucleic acid concentrations, the relationship between the predicted results and the first fetal proportion was fitted to obtain the quadratic term coefficient Fa, the linear term coefficient Fb and the constant term coefficient Fc in the fetal fitting parameters.

[0145] It can be understood that polynomial fitting is a method of approximating the actual data relationship by constructing a polynomial function. By selecting the appropriate degree of the polynomial (such as first, second, third, etc.), different degrees of fitting effects can be obtained. The higher the degree of the polynomial, the higher the accuracy of the fitting. The polynomial quadratic model is only exemplary. The degree of the fitted polynomial can be adjusted according to the experimental accuracy requirements, and this application does not impose specific restrictions.

[0146] In some embodiments, the second sampling ratio includes a second mother ratio and a second fetus ratio, and step S106 may include but is not limited to step S601 and step S602.

[0147] Step S601, using the preset fetal free nucleic acid concentration as the input of the polynomial, and determining the second mother ratio for extracting data from the mother sample file according to the mother fitting parameters.

[0148] Step S602, using the preset fetal free nucleic acid concentration as the input of the polynomial, and determining the second fetal proportion of data extracted from the fetal sample file according to the fetal fitting parameters.

[0149] In step S601 of some embodiments, the preset fetal free nucleic acid concentration is used as the input of the polynomial, and the second mother ratio of data extracted from the mother sample file is determined according to the mother fitting parameters. The mother fitting parameters (Ma, Mb, Mc) obtained in step S501 are substituted into formula (1), and the corresponding polynomial quadratic power model is established, and the sampling ratio of the mother sample file is corrected. The preset fetal free nucleic acid concentration (ff) is used as the input of the polynomial, and the second mother ratio Rm' of data extracted from the mother sample file is calculated by the model.

[0150] In step S602 of some embodiments, the preset fetal free nucleic acid concentration is used as the input of the polynomial, and the second fetal ratio of data extracted from the fetal sample file is determined according to the fetal fitting parameters. The fetal fitting parameters (Fa, Fb, Fc) obtained in step S5502 are substituted into formula (1), and the corresponding polynomial quadratic model is established to correct the sampling ratio of the fetal sample file, and the preset fetal free nucleic acid concentration (ff) is used as the input of the polynomial, and the second fetal ratio Rf' of data extracted from the fetal sample file is calculated by the model.

[0151] In some implementation instances, after re-sampling and merging the family sample files according to the second sampling ratio to obtain simulated second pregnant woman plasma free nucleic acid data, the fetal free nucleic acid concentration can also be predicted for the second pregnant woman plasma free nucleic acid data to verify whether the predicted result from the second simulated BAM file is close to the preset fetal free nucleic acid concentration.

[0152] Reference Figure 4 After sampling the second maternal plasma free nucleic acid data obtained from step S601 and the second fetal ratio obtained from step S602, the fetal free nucleic acid concentration is predicted for the second maternal plasma free nucleic acid data, and a scatter plot is established with the preset fetal free nucleic acid concentration as the horizontal axis and the fetal free nucleic acid concentration predicted by the simulated data as the vertical axis. Figure 4 It can be seen that the predicted results are close to the preset fetal free nucleic acid concentration, so it can be considered that the simulation of pregnant women's plasma data is successful.

[0153] In some embodiments, the simulation method of pregnant women's plasma free nucleic acid data may also include but is not limited to steps S701 to S702.

[0154] Step S701, enriching the free nucleic acid data of the second pregnant woman's plasma free nucleic acid data whose inserted fragment length is less than a preset screening threshold to obtain short fragment data.

[0155] Step S702, predicting the fetal free nucleic acid concentration of the short fragment data to obtain an enriched prediction result.

[0156] Specifically, the free nucleic acid data with an inserted fragment length less than a preset slice screening threshold in the free nucleic acid data of the second pregnant woman's plasma are enriched to obtain short fragment data. Enrichment is an operation step of collecting a specific type of molecule (such as DNA, RNA or protein) to be measured from a large amount of maternal material to a smaller volume, thereby increasing its proportion in the sample. Related studies have shown that enriching free nucleic acids less than 160bp in pregnant woman's plasma by experimental methods can effectively enrich fetal nucleic acids, thereby effectively increasing the concentration of fetal free nucleic acids. Therefore, for the simulated second pregnant woman's plasma free nucleic acid data, the inserted fragment length information in the BAM file can be used for fragment screening, and the upper limit of fragment screening is determined by the preset slice screening threshold, and the free nucleic acid data corresponding to the inserted fragment length less than the preset slice screening threshold are collected to obtain short fragment data.

[0157] Optionally, the preset screening threshold is 160dp, and the plasma free nucleic acid data of the second pregnant woman is screened according to the preset screening threshold, and the reads data corresponding to the inserted fragment information less than 160dp are enriched to obtain short fragment data.

[0158] It is understandable that the preset film screening threshold can be set according to experimental requirements, and 160dp is only an example. As long as the short segment data that meets the experimental requirements can be screened out from the data according to the preset film screening threshold, the embodiment of the present application does not make any specific restrictions.

[0159] After obtaining the short fragment data, the fetal free nucleic acid concentration is predicted for the short fragment data to obtain the enriched prediction result. The enriched short fragment data can be used for subsequent high-throughput sequencing or other detection methods to predict the fetal free nucleic acid concentration. The short fragment data is used to achieve the effect of increasing the fetal free nucleic acid concentration and obtain the enriched prediction result.

[0160] The following is a detailed introduction and description of the solution of the embodiment of the present invention in conjunction with a specific application example.

[0161] In the embodiment of the present application, a program for simulation process control and model fitting is written based on Perl and Python programming languages, and the DownsampleSam module in the bioinformatics tool Picard is used in combination to sample data from the Bam file, and the mpileup module in the bioinformatics tool Samtools is used to detect SNP sites for subsequent fetal concentration prediction.

[0162] Reference Figure 5 , Figure 5 This is a flow chart of a method for simulating free nucleic acid data in pregnant women's plasma provided in another embodiment of the present application, wherein the red dotted line represents the process of the first mixing according to the target mixing ratio, the orange solid line represents the process of polynomial fitting and correction, and the blue solid line represents the process of the second mixing using the corrected sampling ratio.

[0163] First, the family information was obtained, and the plasma free DNA data of pregnant women were simulated using three pairs of mother-child paired DNA samples. Among them, the offspring NA19978 was paired with the mother NA19976 to form the first family, the offspring NA07415 was paired with the mother NA07422 to form the second family, and the offspring NA09376 was paired with the mother NA09375 to form the third family.

[0164] Optionally, during the experimental library construction process, the time of the disrupting enzyme treatment for the offspring DNA samples is extended, and the screening coefficient is increased, so that the offspring DNA fragments screened out are shorter, which can simulate the effect of shorter fetal free DNA fragments.

[0165] Subsequently, the single samples of offspring and mothers were analyzed through high-throughput sequencing and bioinformatics process analysis. The reads_number was obtained through sequencing technology to generate a BAM file, and the corresponding data volume and insert fragment length information were recorded.

[0166] Reference Figure 6 , the amount of data for each offspring (Mf) and the amount of data for the mother (Mm) as well as the length of the inserted fragment were counted respectively, and the maximum amount of data that could be simulated for mixed data was estimated using the statistically obtained data amount, and the amount of mixed data Ms used for simulation was determined to be 50M (1M = 1,000,000 reads).

[0167] Next, the simulation range of the fetal free DNA concentration is set to 1% to 30%, and two lower concentrations of 0.5% and 0.8% are set. According to the amount of mixed data and the preset fetal free DNA concentration, the amount of mixed offspring data that needs to be extracted from the offspring BAM file is calculated to be Ms×ff, and the amount of mother data extracted from the mother BAM file for mixing is Ms×(1-ff). According to the calculated amount of data for mixing, the sampling proportion for mixing is calculated respectively, and the sampling proportion of the offspring is calculated to be Rf=Ms×ff / Mf, and the sampling proportion of the mother is Rm=Ms×(1-ff) / Mm.

[0168] For example, refer to Figure 6 The data volume Mf of the offspring NA19978 is 64353174, and the data volume Ms of the mother NA19976 is 67348599. When the total data volume after mixing is set to 50M, the sampling ratio can be calculated according to the preset fetal free DNA concentration (ff).

[0169] Taking ff as 2% as an example, the amount of offspring data extracted from the BAM file of offspring NA19978 after mixing is calculated to be 50M×2%=1000000, and the amount of mother data extracted from the BAM file of mother NA19976 for mixing is 50M×(1-2%)=49000000. According to the target mixing ratio of mixed sampling, the sampling ratio of offspring NA19978-Rf=1000000 / 64353174≈0.0155 (retain four decimal places), and the sampling ratio of mother NA19976-Rm=49000000 / 67348599≈0.7276 (retain four decimal places) are calculated respectively.

[0170] According to the above calculation method, the sampling ratios of the first mixing simulation corresponding to different preset fetal free DNA concentrations are calculated for the first family, the second family, and the third family, as follows: Figure 7 shown.

[0171] The calculated sampling ratio is used as a parameter to input Picard-DownsampleSam through the bioinformatics software for random sampling and merging to simulate the plasma samples of pregnant women. After obtaining the plasma free nucleic acid data of the first pregnant woman with different mixing ratios, the high-frequency heterozygous SNP sites in the 820 populations are used to predict the fetal free DNA concentration, and the actual fetal free DNA concentration result cff obtained based on the plasma free nucleic acid data of the first pregnant woman is calculated, as shown in Figure 8 shown.

[0172] right Figure 8 A preliminary analysis of the data in the results showed that the fetal free DNA concentrations of the three families after the first simulated mixing were quite different from the preset fetal free DNA concentrations. Therefore, it was necessary to perform polynomial fitting based on the sampling ratios of offspring and mothers.

[0173] Reference Fig. 9 , with the actual predicted fetal free DNA concentration as the horizontal axis and the offspring sampling ratio as the vertical axis, a scatter plot was drawn based on the data of offspring from different families, and a curve reflecting the changing trend between the offspring sampling ratio and the actual predicted fetal free DNA concentration after mixing was obtained through polynomial fitting.

[0174] It should be noted that the fitting parameters are not universal among different families, and the fitting parameters of each family need to be fitted separately. By establishing a polynomial quadratic model, polynomial fitting is performed on the offspring and mother of each family respectively to obtain the fitting parameters of the offspring and mother of each family.

[0175] Reference Fig.10 For ease of presentation, the fitting parameters are retained to four decimal places. Substituting the calculated fitting parameters into the polynomial quadratic model, the curves of each family are obtained as follows Fig. 9 As shown, the blue curve represents the offspring of the first family, the orange curve represents the offspring of the second family, and the yellow curve represents the offspring of the third family.

[0176] The preset fetal free DNA concentration was used as the input parameter and entered into the polynomial quadratic model corresponding to each family. The corrected sampling ratios of offspring and mothers were calculated respectively. For ease of display, the calculated sampling ratio results were retained to eight decimal places, such as Fig.11 shown.

[0177] Taking the first family as an example, taking ff=2% as input, the corrected offspring sampling ratio NA19978-Rf'≈0.00877808 and the corrected mother sampling ratio NA19976-Rm'≈0.73406132 were calculated.

[0178] The corrected sampling ratio was used for sampling and mixing to re-simulate the maternal plasma sample. After obtaining the second mixed second maternal plasma free nucleic acid data, the second fetal free DNA concentration prediction was performed using the high-frequency heterozygous SNP sites in the 820 populations, and the second mixed fetal free DNA concentration result cff' was obtained to observe whether it was consistent with the preset fetal free DNA concentration.

[0179] Reference Fig.12 After the second mixing, the concentration of fetal free DNA was roughly consistent with the preset concentration, and the error between the measured fetal free DNA concentration ratio and the expected fetal free DNA concentration ratio was within ±1%. Therefore, the subsequent research and development of non-invasive prenatal screening can be carried out based on the plasma free nucleic acid data of the second pregnant woman obtained by the second simulation.

[0180] In an embodiment of the present application, by using two comparison data files of mother-child pairings, combined with bioinformatics tools and model fitting methods, mixed data that meets the expected fetal concentration can be output to simulate maternal plasma data, thereby enabling research and development work related to prenatal screening technology to be carried out, greatly reducing the time cost spent on sampling and experiments, and being able to batch generate maternal plasma simulation data with different fetal concentrations under different data volumes in a short period of time, which helps to improve the flexibility of non-invasive prenatal screening in obtaining positive samples.

[0181] In some embodiments, the fragment length distribution after mixing and the fragment length distribution of the mother and child single samples can be observed later. Fig.13 Taking the first family as an example, the peak value of the fragment length distribution (green line) of the offspring NA19978 is lower than that of the mother NA19976 (blue line), while the fragment length distribution (red line) of the simulated data after mixing is also biased towards the short fragment range due to the influence of the shorter fragments of the offspring. Therefore, the concentration of fetal free DNA can be increased by enriching short fragments.

[0182] For example, refer to Fig.14 , using the inserted fragment length information in the BAM file, the corresponding reads data of inserted fragments less than 160bp were collected and then the fetal concentration was predicted. The comparison results showed that after enriching short fragments, the fetal free DNA concentration could be increased to about 1.5 times the original.

[0183] Reference Fig.15 The embodiment of the present application also provides a simulation system for pregnant women's plasma free nucleic acid data, which can implement the simulation method for pregnant women's plasma free nucleic acid data, and the system includes:

[0184] The first module is used to obtain the family sample file of mother-child pairing.

[0185] The second module is used to perform data volume analysis on the family sample file according to the preset fetal free nucleic acid concentration to obtain the first sampling ratio.

[0186] The third module is used to sample and merge the family sample files according to the first sampling ratio to obtain the simulated plasma free nucleic acid data of the first pregnant woman.

[0187] The fourth module is used to predict the fetal free nucleic acid concentration based on the plasma free nucleic acid data of the first pregnant woman to obtain the prediction result.

[0188] The fifth module is used to perform polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters.

[0189] The sixth module is used to use the preset fetal free nucleic acid concentration as the input of the polynomial and determine the second sampling ratio according to the fitting parameters.

[0190] The seventh module is used to resample and merge the family sample files according to the second sampling ratio to obtain the simulated plasma free nucleic acid data of the second pregnant woman.

[0191] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0192] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned method for simulating free nucleic acid data in plasma of pregnant women when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.

[0193] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0194] Reference Fig.16 , Fig.16 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:

[0195] The processor 901 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0196] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 902, and the processor 901 calls and executes the simulation method of the plasma free nucleic acid data of pregnant women in the embodiment of this application.

[0197] The input / output interface 903 is used to implement information input and output.

[0198] The communication interface 904 is used to realize the communication interaction between this device and other devices. The communication can be realized through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0199] The bus 905 transmits information between various components of the device (eg, the processor 901 , the memory 902 , the input / output interface 903 , and the communication interface 904 ).

[0200] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .

[0201] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned simulation method of pregnant women's plasma free nucleic acid data.

[0202] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0203] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0204] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0205] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0206] The system embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.

[0207] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0208] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0209] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. A method for simulating plasma free nucleic acid data of pregnant women, characterized in that: The method comprises the following steps: Obtain pedigree sample files for mother-child pairings; Performing data volume analysis on the family sample file according to a preset fetal free nucleic acid concentration to obtain a first sampling ratio; Sampling and merging the family sample files according to the first sampling ratio to obtain simulated plasma free nucleic acid data of the first pregnant woman; Predicting the fetal free nucleic acid concentration based on the first pregnant woman's plasma free nucleic acid data to obtain a prediction result; Perform polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters; Using the preset fetal free nucleic acid concentration as the input of the polynomial, and determining the second sampling ratio according to the fitting parameters; The family sample files are resampled and merged according to the second sampling ratio to obtain simulated plasma free nucleic acid data of the second pregnant woman.

2. The method according to claim 1, characterized in that: The family sample file includes a mother sample file and a fetus sample file, the first sampling ratio includes a first mother ratio and a first fetus ratio, and the data volume analysis of the family sample file according to the preset fetal free nucleic acid concentration to obtain the first sampling ratio includes the following steps: Performing data volume statistics on the mother sample file and the fetus sample file respectively to obtain the mother data volume and the fetus data volume; Determining a simulated mixed data volume according to a preset fetal free nucleic acid concentration, wherein the product of the mixed data volume and the preset fetal free nucleic acid concentration does not exceed the fetal data volume; Determine the maternal sampling amount and the fetal sampling amount according to the preset fetal free nucleic acid concentration and the mixed data amount; The first mother ratio is obtained according to the ratio of the mother sampling amount to the mother data amount, and the first fetus ratio is obtained according to the ratio of the fetus sampling amount to the fetus data amount.

3. The method according to claim 2, characterized in that The step of sampling and merging the family sample files according to the first sampling ratio to obtain simulated plasma free nucleic acid data of the first pregnant woman comprises the following steps: Randomly sampling the mother sample file according to the first mother ratio to obtain a first mother sampling file; Randomly sampling the fetal sample file according to the first fetal ratio to obtain a first fetal sampling file; The first mother sampling file and the first fetus sampling file are mixed to obtain simulated first pregnant woman plasma free nucleic acid data.

4. The method according to claim 1, characterized in that: The step of predicting the fetal free nucleic acid concentration based on the first pregnant woman's plasma free nucleic acid data to obtain a prediction result comprises the following steps: Determining the frequency range of genotype occurrence frequency according to the preset concentration range of fetal free nucleic acid concentration; According to the preset high-frequency heterozygous single nucleotide polymorphism sites, single nucleotide polymorphism detection is performed on the plasma free nucleic acid data of the first pregnant woman, and the single nucleotide polymorphism sites whose genotype occurrence frequency is within the frequency range are collected to obtain a frequency site set; According to the frequency site set, the probability of each frequency value within the frequency range is calculated in a gradient manner by a statistical method, and the frequency value with the largest probability is determined as the actual fetal free nucleic acid concentration to obtain a prediction result.

5. The method according to claim 2, characterized in that: The step of performing polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters comprises the following steps: According to the prediction results corresponding to different preset fetal free nucleic acid concentrations and the first mother ratio, a relationship between the prediction results and the first mother ratio is fitted by a polynomial model to obtain a mother fitting parameter; According to the predicted results corresponding to different preset fetal free nucleic acid concentrations and the first fetal proportion, a polynomial model is used to fit the relationship between the predicted results and the first fetal proportion to obtain fetal fitting parameters.

6. The method according to claim 5, characterized in that The second sampling ratio includes a second mother ratio and a second fetus ratio, and the method of using the preset fetal free nucleic acid concentration as the input of the polynomial and determining the second sampling ratio according to the fitting parameters includes the following steps: Using the preset fetal free nucleic acid concentration as an input of a polynomial, determining a second mother proportion for extracting data from the mother sample file according to the mother fitting parameter; The preset fetal free nucleic acid concentration is used as an input of a polynomial, and the second fetal proportion of data extracted from the fetal sample file is determined according to the fetal fitting parameters.

7. The method according to claim 1, characterized in that The simulation method of the pregnant woman's plasma free nucleic acid data also includes the following steps: Enriching the free nucleic acid data of the plasma free nucleic acid data of the second pregnant woman whose inserted fragment length is less than a preset screening threshold value to obtain short fragment data; The fetal free nucleic acid concentration is predicted for the short fragment data to obtain an enriched prediction result.

8. A simulation system for plasma free nucleic acid data of pregnant women, characterized in that: The system comprises: The first module is used to obtain the pedigree sample files of mother-child pairing; The second module is used to perform data volume analysis on the family sample file according to a preset fetal free nucleic acid concentration to obtain a first sampling ratio; The third module is used to sample and merge the family sample files according to the first sampling ratio to obtain simulated plasma free nucleic acid data of the first pregnant woman; The fourth module is used to predict the fetal free nucleic acid concentration based on the first pregnant woman's plasma free nucleic acid data to obtain a prediction result; A fifth module is used to perform polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters; A sixth module is used to use the preset fetal free nucleic acid concentration as an input of a polynomial and determine a second sampling ratio according to the fitting parameters; The seventh module is used to resample and merge the family sample files according to the second sampling ratio to obtain simulated plasma free nucleic acid data of the second pregnant woman.

9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Determination method of fetal DNA content in maternal plasma, based on single-nucleotide polymorphic loci

    CN103215350A

  • Method for improving proportion of fetal free DNA in maternal plasma free DNA sequencing library

    CN105926043A

  • Method for constructing target gene library, detection device and application thereof

    CN112996926A

  • Simulated plasma circulating tumor DNA standard substance as well as preparation method and application thereof

    CN113930512A

  • Method for simulating high-depth sequencing TSS features

    CN116805512A