Method, system, device and medium for simulating pregnant women's plasma free nucleic acid data
Through data analysis and polynomial fitting of mother-child paired sample files, the problems of high complexity and large concentration error in simulating maternal plasma samples in the existing technology were solved, efficient and accurate simulation of fetal free nucleic acid concentration was achieved, and the flexibility of non-invasive prenatal screening was improved.
Patent Information
- Application Number
- CN202411886581.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-20
AI Technical Summary
The existing technology requires a large number of positive samples when preparing simulated plasma samples of pregnant women. The operation is complicated and costly, and the preparation concentration error is large. It is difficult to accurately simulate the concentration of fetal free nucleic acid, which affects the accuracy of non-invasive prenatal screening.
By obtaining family sample files of mother-child pairs, performing data volume analysis, obtaining the sampling ratio, performing polynomial fitting, calibrating the sampling ratio, merging sample files, and generating simulated maternal plasma free nucleic acid data, the concentration accuracy is improved.
It reduces sampling and experimental time and cost, reduces experimental complexity, improves the accuracy of plasma simulated fetal free nucleic acid concentration, and enhances the flexibility of non-invasive prenatal screening.
Smart Images

Figure CN119943152B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of non-invasive prenatal screening, and in particular to a method, system, device and medium for simulating plasma free nucleic acid data of pregnant women. Background Art
[0002] Chromosomal abnormalities and genetic diseases are important causes of birth defects. Prenatal screening can reduce birth defects and improve the quality of the newborn population. Non-invasive prenatal screening performs non-invasive fetal genome testing by analyzing whether the fetal free nucleic acid in the plasma of pregnant women carries single-gene diseases or chromosomal abnormality signals. Generally, the concentration of fetal free nucleic acid in maternal plasma and the detection value of positive samples are positively correlated. The higher the concentration of fetal free nucleic acid, the higher the accuracy of non-invasive fetal genome testing. Therefore, during the research and development and verification stages of non-invasive prenatal screening products, it is often necessary to collect positive samples with different fetal concentration gradients in order to study the sensitivity and specificity of the test kit and detection algorithm at different concentrations.
[0003] However, the prevalence of most monogenic and chromosomal diseases is low, making it difficult to obtain clinically accurate positive plasma samples with specific fetal free nucleic acid concentrations. Currently, existing technologies often use simulated plasma samples to obtain standards that simulate specific fetal free nucleic acid concentration ratios. However, this simulation method not only requires the collection of a large number of positive samples in advance, but also requires that errors in the prepared standard concentrations be detected only after the preparation is complete, either in the quantitative or sequencing phase, affecting the accuracy of the fetal free nucleic acid concentrations in the simulated plasma samples. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose a method, system, equipment and medium for simulating free nucleic acid data in pregnant women's plasma, aiming to reduce the time and preparation costs spent on sampling and experiments, and improve the accuracy of the ratio of plasma simulated fetal free nucleic acid concentration.
[0005] To achieve the above objectives, one aspect of the present invention provides a method for simulating plasma cell-free nucleic acid data of pregnant women, the method comprising:
[0006] Obtain family sample files for mother-child pairing;
[0007] Performing data volume analysis on the family sample file according to a preset fetal free nucleic acid concentration to obtain a first sampling ratio;
[0008] Sampling and merging the family sample files according to the first sampling ratio to obtain simulated plasma cell-free nucleic acid data of the first pregnant woman;
[0009] Predicting the fetal free nucleic acid concentration based on the first pregnant woman's plasma free nucleic acid data to obtain a prediction result;
[0010] performing polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters;
[0011] taking the preset fetal cell-free nucleic acid concentration as an input of the polynomial, and determining a second sampling ratio according to the fitting parameters;
[0012] re-sampling and merging the family sample file according to the second sampling ratio to obtain simulated second pregnant woman plasma cell-free nucleic acid data.
[0013] In some embodiments, the family sample file includes a mother sample file and a fetal sample file, the first sampling ratio includes a first mother ratio and a first fetal ratio, and the data volume analysis of the family sample file according to the preset fetal cell-free nucleic acid concentration to obtain the first sampling ratio includes the following steps:
[0014] respectively performing data volume statistics on the mother sample file and the fetal sample file to obtain a mother data volume and a fetal data volume;
[0015] determining a simulated mixed data volume according to the preset fetal cell-free nucleic acid concentration, wherein the product of the mixed data volume and the preset fetal cell-free nucleic acid concentration does not exceed the fetal data volume;
[0016] determining a mother sampling volume and a fetal sampling volume according to the preset fetal cell-free nucleic acid concentration and the mixed data volume;
[0017] obtaining the first mother ratio according to the ratio of the mother sampling volume to the mother data volume, and obtaining the first fetal ratio according to the ratio of the fetal sampling volume to the fetal data volume.
[0018] In some embodiments, the sampling and merging of the family sample file according to the first sampling ratio to obtain simulated first pregnant woman plasma cell-free nucleic acid data includes the following steps:
[0019] randomly sampling the mother sample file according to the first mother ratio to obtain a first mother sampling file;
[0020] randomly sampling the fetal sample file according to the first fetal ratio to obtain a first fetal sampling file;
[0021] performing data mixing on the first mother sampling file and the first fetal sampling file to obtain simulated first pregnant woman plasma cell-free nucleic acid data.
[0022] In some embodiments, the fetal cell-free nucleic acid concentration prediction of the first pregnant woman plasma cell-free nucleic acid data to obtain a prediction result includes the following steps:
[0023] Determining the frequency range of genotype occurrence frequency according to the preset concentration range of fetal free nucleic acid concentration;
[0024] performing single nucleotide polymorphism detection on the plasma cell-free nucleic acid data of the first pregnant woman according to the preset high-frequency heterozygous single nucleotide polymorphism sites, collecting single nucleotide polymorphism sites whose genotype occurrence frequency is within the frequency range, and obtaining a frequency site set;
[0025] According to the frequency site set, the probability of each frequency value within the frequency range is calculated in a gradient manner using a statistical method, and the frequency value with the largest probability is determined as the actual fetal free nucleic acid concentration to obtain a prediction result.
[0026] In some embodiments, performing polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters includes the following steps:
[0027] According to the prediction results corresponding to different preset fetal free nucleic acid concentrations and the first maternal ratio, a relationship between the prediction results and the first maternal ratio is fitted using a polynomial model to obtain a maternal fitting parameter;
[0028] According to the predicted results and the first fetal ratio corresponding to different preset fetal free nucleic acid concentrations, a polynomial model is used to fit the relationship between the predicted results and the first fetal ratio to obtain fetal fitting parameters.
[0029] In some embodiments, the second sampling ratio includes a second maternal ratio and a second fetal ratio, and the step of using the predetermined fetal free nucleic acid concentration as an input of a polynomial and determining the second sampling ratio according to the fitting parameters comprises the following steps:
[0030] Using the preset fetal free nucleic acid concentration as an input to a polynomial, determining a second mother proportion of data extracted from the mother sample file based on the mother fitting parameters;
[0031] The preset fetal free nucleic acid concentration is used as an input of a polynomial, and a second fetal proportion of data extracted from the fetal sample file is determined according to the fetal fitting parameters.
[0032] In some embodiments, the method for simulating plasma cell-free nucleic acid data of pregnant women further comprises the following steps:
[0033] Enriching the free nucleic acid data of the plasma free nucleic acid data of the second pregnant woman with an inserted fragment length less than a preset screening threshold to obtain short fragment data;
[0034] The fetal free nucleic acid concentration is predicted for the short fragment data to obtain an enriched prediction result.
[0035] To achieve the above objectives, another aspect of the present invention provides a system for simulating plasma cell-free nucleic acid data of pregnant women, the system comprising:
[0036] The first module is used to obtain family sample files for mother-child pairing;
[0037] The second module is configured to perform data volume analysis on the family sample file according to a preset fetal free nucleic acid concentration to obtain a first sampling ratio;
[0038] A third module is configured to sample and merge the family sample files according to the first sampling ratio to obtain simulated plasma cell-free nucleic acid data of the first pregnant woman;
[0039] A fourth module is used to predict the fetal free nucleic acid concentration based on the first pregnant woman's plasma free nucleic acid data to obtain a prediction result;
[0040] A fifth module is configured to perform polynomial fitting based on the prediction result and the first sampling ratio to obtain fitting parameters;
[0041] A sixth module is configured to use the preset fetal free nucleic acid concentration as an input of a polynomial and determine a second sampling ratio according to the fitting parameters;
[0042] The seventh module is used to resample and merge the family sample files according to the second sampling ratio to obtain simulated plasma free nucleic acid data of the second pregnant woman.
[0043] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned method when executing the computer program.
[0044] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above-mentioned method when executed by a processor.
[0045] Embodiments of the present application include at least the following beneficial effects: The present application provides a method, system, device, and medium for simulating plasma free nucleic acid data of pregnant women. The solution obtains family sample files of mother-child pairs; performs data volume analysis on the family sample files based on a preset fetal free nucleic acid concentration to obtain a first sampling ratio; samples and merges the family sample files based on the first sampling ratio to obtain simulated first plasma free nucleic acid data of pregnant women; predicts the fetal free nucleic acid concentration of the first plasma free nucleic acid data to obtain a prediction result; performs polynomial fitting based on the prediction result and the first sampling ratio to obtain fitting parameters; uses the preset fetal free nucleic acid concentration as the input of the polynomial, and determines a second sampling ratio based on the fitting parameters; and re-samples and merges the family sample files based on the second sampling ratio to obtain simulated second plasma free nucleic acid data of pregnant women. This solution can reduce the time and preparation costs spent on sampling and experiments, reduce experimental complexity, and improve the accuracy of the simulated fetal free nucleic acid concentration ratio of plasma. It can simulate pregnant women's plasma data with any fetal free nucleic acid concentration ratio in a batch in a short time, which helps to improve the flexibility of obtaining positive samples in non-invasive prenatal screening. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a flow chart of a method for simulating plasma free nucleic acid data of pregnant women provided in an embodiment of the present application;
[0047] Figure 2 This is a schematic diagram of the genotype combination pattern in the peripheral blood of pregnant women provided in the examples of the present application;
[0048] Figure 3 This is a scatter plot of the first simulated fetal free nucleic acid concentration provided in the examples of the present application;
[0049] Figure 4 This is a scatter plot of the second simulated fetal free nucleic acid concentration provided in the examples of the present application;
[0050] Figure 5 This is a flow chart of a method for simulating plasma free nucleic acid data of pregnant women provided in another embodiment of the present application;
[0051] Figure 6 This is a schematic diagram of the results of the mother-child pairing sample data volume and insert fragment length provided in the embodiments of the present application;
[0052] Figure 7 This is a schematic diagram of the results of the first sampling ratio provided in the embodiment of the present application;
[0053] Figure 8 This is a schematic diagram showing the comparison between the actual fetal free nucleic acid concentration and the preset fetal free nucleic acid concentration in the first simulation provided by the embodiment of the present application;
[0054] Figure 9This is a scatter plot of the fetal free DNA concentration and offspring sampling ratio of different families provided in the examples of this application;
[0055] Figure 10 This is a schematic diagram of the results of different family fitting parameters provided in the examples of this application;
[0056] Figure 11 This is a schematic diagram of the results of the second sampling ratio provided in the embodiment of the present application;
[0057] Figure 12 This is a schematic diagram showing the comparison between the actual fetal free nucleic acid concentration and the preset fetal free nucleic acid concentration in the second simulation provided by the embodiment of the present application;
[0058] Figure 13 This is a distribution diagram of the length of the insert fragment of the first family provided in the examples of this application;
[0059] Figure 14 This is a graph showing the experimental results of the fetal free nucleic acid concentration before and after enrichment of short fragments provided in the examples of this application;
[0060] Figure 15 This is a schematic diagram of the structure of a simulation system for pregnant women's plasma free nucleic acid data provided in an embodiment of the present application;
[0061] Figure 16 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0063] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0064] As used herein, the terms "at least one", "multiple", "each", "any of" and the like, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding plurality, and any of refers to any one of the plurality.
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to be limiting of this application.
[0066] Before the embodiments of the present application are described in detail, first, the terms involved in the embodiments of the present application and part of the related technology are described as follows.
[0067] (1) reads: reads refer to the base sequence obtained by a single sequencing of a sequencer. Due to the limitation of the current sequencing level, the genome needs to be broken into DNA fragments before genome sequencing. Different sequencing instruments have different read lengths.
[0068] (2) BAM (Binary Alignment / Map format): BAM is the most commonly used alignment data storage format in current genetic data analysis, which is a binary file used to store large-scale sequencing data. The BAM file is usually used to store the alignment information of the sequencing data, such as the alignment position of the DNA sequencing reads with the reference genome and the corresponding quality value.
[0069] (3) SNP (Single Nucleotide Polymorphism): SNP mainly refers to the DNA sequence polymorphism caused by the variation of a single nucleotide at the genome level. It is the most common type of human heritable variation, accounting for more than 90% of all known polymorphisms. SNP exists widely in the human genome, with an average of 1 SNP per 300 base pairs. When a base pair is converted or transposed, inserted or deleted, it is an SNP site.
[0070] Related studies show that the total incidence of birth defects in China is about 5.6%, and with 160 million births per year in the country, there are 900,000 new cases of birth defects each year, of which 60% are related to chromosomal abnormalities and genetic diseases. Therefore, prenatal screening is an important means to reduce birth defects and improve the quality of the population. However, traditional prenatal screening methods such as amniocentesis, chorionic villus biopsy, and umbilical vein puncture, although accurate, are associated with a 0.2% to 0.5% risk of miscarriage and infection, which makes non-invasive prenatal screening technology a research hotspot.
[0071] In 1997, scientists discovered the presence of fetal free nucleic acid (DNA) in maternal plasma, which laid the theoretical foundation for non-invasive prenatal screening of fetal genomic abnormalities through maternal plasma.
[0072] Subsequently, the application of high-throughput sequencing technology further promoted the development of non-invasive fetal chromosomal aneuploidy genetic testing (NIPT), which can screen fetal chromosomal aneuploidy abnormalities by detecting free DNA in maternal peripheral blood.
[0073] Since then, maternal peripheral blood free DNA has also been shown to be useful for detecting fetal microdeletion and microduplication syndromes as well as fetal monogenic diseases.
[0074] In 2022, a team of scientists pioneered a "three-in-one" comprehensive non-invasive prenatal screening technology that can simultaneously screen for chromosomal aneuploidy, chromosomal microdeletion syndrome and single-gene dominant genetic diseases. It uses specific molecular index (UMI) double-end library capture sequencing technology to accurately calculate the number of DNA molecules in fetal free DNA. Combined with fetal concentration information, low-level fetal mutation sites can be screened out.
[0075] The detection accuracy of the above prenatal technologies is mainly affected by the concentration of fetal free DNA. Therefore, accurate assessment of the ratio of fetal free DNA concentration is very critical for non-invasive prenatal screening.
[0076] The free DNA in maternal plasma includes both the maternal free DNA and the fetal free DNA, of which the concentration of fetal free DNA is generally within 3% to 30%, with an average of about 13%. Under normal circumstances, the concentration of fetal free DNA in maternal plasma is positively correlated with the detection value of positive samples, that is, the higher the concentration of fetal free DNA, the higher the accuracy of non-invasive fetal genomic disease detection. Taking NIPT as an example, a too low concentration of fetal free DNA will increase the risk of missing positive samples. The detection limit of NIPT is 3% to 4%, but due to individual differences in pregnant women, sample quality, and experimental errors, in existing technologies, when the concentration of fetal free DNA is lower than 5%, the NIPT test result may be a false negative.
[0077] However, the incidence of most single-gene and chromosomal diseases is low. For example, among the three common aneuploidy abnormalities, trisomy 21 (Down syndrome), trisomy 18 (Edwards syndrome), and trisomy 13 (Patau syndrome) are the three most common autosomal aneuploidies, with incidences of 1 / 600-800, 1 / 3500-8000, and 1 / 7000-20000 in newborns, respectively. In such cases, it is difficult to obtain clinically positive plasma samples with specific fetal concentrations to fully study the performance of the test.
[0078] Currently, to address the difficulty in obtaining positive samples, most studies typically mix a certain proportion of positive or negative DNA into plasma samples to prepare simulated plasma samples. However, this simulation method has the following limitations:
[0079] (1) The experimental operation is highly complex and requires the collection of a large number of positive DNA and negative plasma samples in advance for preparation.
[0080] (2) The problem of concentration error in the preparation of standard products, that is, there is an error between the actual fetal concentration measured after preparation and the expected preparation concentration. These errors usually need to be discovered in the quantification or sequencing stage after the preparation is completed, which may result in the need to re-prepare the entire batch of standard products.
[0081] (3) The preparation cost is high, and the cost increases proportionally with the density of the fetal concentration gradient. It is difficult to obtain the results of a high-density fetal concentration gradient (such as 1%) for detailed analysis.
[0082] In view of this, embodiments of the present application provide a method, system, device, and medium for simulating maternal plasma free nucleic acid data. The method obtains a family sample file of a mother-child pair; performs a data volume analysis on the family sample file based on a preset fetal free nucleic acid concentration to obtain a first sampling ratio; samples and merges the family sample files based on the first sampling ratio to obtain simulated first maternal plasma free nucleic acid data; predicts the fetal free nucleic acid concentration of the first maternal plasma free nucleic acid data to obtain a prediction result; performs polynomial fitting based on the prediction result and the first sampling ratio to obtain fitting parameters; uses the preset fetal free nucleic acid concentration as the input of the polynomial, and determines a second sampling ratio based on the fitting parameters; and re-samples and merges the family sample files based on the second sampling ratio to obtain simulated second maternal plasma free nucleic acid data. This method can reduce the time and preparation costs spent on sampling and experiments, reduce experimental complexity, and improve the accuracy of the simulated fetal free nucleic acid concentration ratio of plasma. It can simulate maternal plasma data with any fetal free nucleic acid concentration ratio in a short time, which helps to improve the flexibility of obtaining positive samples in non-invasive prenatal screening.
[0083] The simulation method of the plasma free nucleic acid data of pregnant women provided in the embodiment of the present application relates to the technical field of non-invasive prenatal screening. The simulation method of the plasma free nucleic acid data of pregnant women provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the simulation method of the plasma free nucleic acid data of pregnant women, etc., but is not limited to the above forms.
[0084] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0085] Figure 1 This is an optional flow chart of the simulation method for pregnant women's plasma free nucleic acid data provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S107.
[0086] Step S101: Obtain a family sample file for mother-child pairing.
[0087] Step S102 : performing data volume analysis on the family sample file according to the preset fetal free nucleic acid concentration to obtain a first sampling ratio.
[0088] Step S103: sampling and merging the family sample files according to the first sampling ratio to obtain simulated plasma cell-free nucleic acid data of the first pregnant woman.
[0089] Step S104: predicting the fetal free nucleic acid concentration based on the plasma free nucleic acid data of the first pregnant woman to obtain a prediction result.
[0090] Step S105 , performing polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters.
[0091] Step S106: Using the preset fetal free nucleic acid concentration as the input of the polynomial, and determining the second sampling ratio according to the fitting parameters.
[0092] Step S107: resample and merge the family sample files according to the second sampling ratio to obtain simulated plasma cell-free nucleic acid data of the second pregnant woman.
[0093] Specifically, first, obtain the family sample file of the mother-child pairing, process the DNA samples of the mother and fetus in the same family through high-throughput sequencing and bioinformatics process analysis, align each DNA sample to the reference genome, record the sequencing reads data to generate the corresponding BAM file, and obtain the family sample file.
[0094] Next, the data volume of the family sample files is analyzed based on the preset fetal free nucleic acid concentration to determine the first sampling ratio. Based on the preset fetal free nucleic acid concentration, the amount of data to be extracted from the maternal and fetal BAM files is calculated, and the first sampling ratio is determined based on the proportion of the extracted data in the corresponding BAM files.
[0095] After obtaining the first sampling ratio, the family sample files were sampled and merged according to the first sampling ratio to obtain the simulated first maternal plasma free nucleic acid data. The first data sampling was performed on the maternal and fetal BAM files respectively according to the first sampling ratio. The BAM files obtained from the individual sampling were merged to generate the first computer-simulated mixed maternal plasma BAM file, obtaining the first maternal plasma free nucleic acid data.
[0096] Then, the fetal free nucleic acid concentration is predicted based on the plasma free nucleic acid data of the first pregnant woman to obtain a prediction result. SNP detection is performed on the simulated plasma free nucleic acid data of the first pregnant woman to detect all possible SNP sites in the data. The SNP sites are used to predict the fetal free nucleic acid concentration, and the actual fetal free nucleic acid concentration in the simulated data is calculated to obtain a prediction result.
[0097] Furthermore, after obtaining the predicted results, the actual fetal free nucleic acid concentration was compared with the preset fetal free nucleic acid concentration. It was found that the actual simulated fetal free nucleic acid concentration was slightly higher than the preset fetal free nucleic acid concentration. To correct the deviation of the fetal free nucleic acid concentration introduced during the simulation process, a polynomial fit was performed based on the predicted results and the first sampling ratio to obtain fitting parameters. The polynomial fit was performed using the calculated actual fetal free nucleic acid concentration (denoted as cff) in combination with the first sampling ratio (denoted as R). Based on multiple data points (cff, R) corresponding to different preset fetal free nucleic acid concentrations, the coefficients of each order of the polynomial were solved using the least squares method. The optimal fitting parameters were found by minimizing the sum of squares of the residuals between the data points and the fitting function.
[0098] The preset fetal free nucleic acid concentration is used as the input of the polynomial, and the second sampling ratio is determined based on the fitting parameters. The obtained fitting parameters are used to establish a polynomial to restore the changing trend between the sampling ratio and the actual fetal free nucleic acid concentration. The preset fetal free nucleic acid concentration is used as the input of the polynomial, and the calibrated second sampling ratio is calculated based on the polynomial.
[0099] The family sample files were resampled and merged according to the second sampling ratio to obtain simulated plasma free nucleic acid data for the second maternal pregnancy. The maternal and fetal BAM files were sampled again according to the second sampling ratio, and finally the individually sampled BAM files were merged to obtain simulated plasma free nucleic acid data for the second maternal pregnancy. By re-predicting the fetal concentration of the second maternal plasma free nucleic acid data obtained from the second simulation, it was found that the predicted result was close to the preset fetal free nucleic acid concentration. At this point, the simulation of the maternal plasma free nucleic acid data was successful.
[0100] In this example, only maternal and fetal sequencing data are needed to simulate maternal plasma, reducing sampling and experimental time and preparation costs, and lowering experimental complexity. By fitting the trend between the sampling ratio and the actual fetal free nucleic acid concentration and determining the secondary sampling ratio, the accuracy of the plasma-simulated fetal free nucleic acid concentration ratio can be improved. Furthermore, maternal plasma data with any fetal free nucleic acid concentration ratio can be simulated in batches in a short period of time, helping to increase the flexibility of obtaining positive samples for non-invasive prenatal screening.
[0101] In some embodiments, the family sample file includes a maternal sample file and a fetal sample file, the first sampling ratio includes a first maternal ratio and a first fetal ratio, and step S102 may include but is not limited to steps S201 to S204.
[0102] Step S201 , performing data volume statistics on the mother sample file and the fetus sample file respectively to obtain the mother data volume and the fetus data volume.
[0103] Step S202: determining the simulated mixed data volume according to the preset fetal free nucleic acid concentration, wherein the product of the mixed data volume and the preset fetal free nucleic acid concentration does not exceed the fetal data volume.
[0104] Step S203, determining the maternal sampling amount and the fetal sampling amount according to the preset fetal free nucleic acid concentration and the mixed data amount.
[0105] Step S204: obtaining a first maternal ratio according to the ratio of the maternal sample size to the maternal data size, and obtaining a first fetal ratio according to the ratio of the fetal sample size to the fetal data size.
[0106] Specifically, data volume statistics are performed on the maternal sample file and the fetal sample file, respectively, to obtain the maternal data volume and the fetal data volume. Data volume statistics are performed on the reads data contained in the maternal sample file to obtain the maternal data volume (denoted as Mm). At the same time, data volume statistics are performed on the reads data contained in the fetal sample file to obtain the fetal data volume (denoted as Mf).
[0107] The amount of simulated mixed data is determined based on the preset fetal free nucleic acid concentration, wherein the product of the mixed data amount and the preset fetal free nucleic acid concentration does not exceed the fetal data amount. To simulate maternal plasma data at different fetal free nucleic acid concentrations, a preset fetal free nucleic acid concentration (denoted as ff) is used to generate simulated data for any desired fetal free nucleic acid concentration. To ensure that a sufficient amount of data can be extracted from the fetal sample file to reach the preset fetal free nucleic acid concentration during the simulation, the amount of simulated mixed data (denoted as Ms) needs to be determined based on the preset fetal free nucleic acid concentration.
[0108] It should be noted that the setting of the mixed data volume must satisfy the condition that the product of the mixed data volume and the preset fetal free nucleic acid concentration does not exceed the fetal data volume. If the fetal data volume required for the preset fetal free nucleic acid concentration exceeds the actual fetal data volume in the fetal sample file, effective data sampling cannot be performed, and it is impossible to simulate maternal plasma data that meets the requirements. Therefore, the value obtained by dividing the fetal data volume by the preset fetal free nucleic acid concentration is determined as the maximum data volume for simulation, and the simulated mixed data volume is then set to not exceed the maximum data volume.
[0109] Next, the maternal sampling volume and the fetal sampling volume are determined based on the preset fetal free nucleic acid concentration and the mixed data volume. The maternal sampling volume is the amount of data to be extracted from the maternal sample file, and the fetal sampling volume is the amount of data to be extracted from the fetal sample file.
[0110] Exemplarily, the maternal sampling amount extracted from the maternal sample file is determined to be Ms×(1-ff) based on the preset fetal free nucleic acid concentration and the mixed data volume, and the fetal sampling amount extracted from the fetal sample file is determined to be Ms×ff based on the preset fetal free nucleic acid concentration and the mixed data volume.
[0111] Finally, the first maternal ratio is calculated based on the ratio of the maternal sample size to the maternal data volume, and the first fetal ratio is calculated based on the ratio of the fetal sample size to the fetal data volume. The maternal and fetal sample sizes are each ratioed to the corresponding BAM file data volume. Based on the obtained maternal sample size, the sampling percentage of the maternal data volume in the maternal sample file is calculated to obtain the first maternal ratio Rm = Ms × (1-ff) / Mm. At the same time, based on the obtained fetal sample size, the sampling percentage of the fetal data volume in the fetal sample file is calculated to obtain the first fetal ratio Rf = Ms × ff / Mf.
[0112] In some embodiments, step S103 may include but is not limited to steps S301 to S303.
[0113] Step S301: Randomly sample the mother sample file according to the first mother ratio to obtain the first mother sampling file.
[0114] Step S302: Randomly sample the fetal sample file according to the first fetal ratio to obtain a first fetal sampling file.
[0115] Step S303: Mix the first mother sampling file and the first fetus sampling file to obtain simulated first pregnant woman plasma free nucleic acid data.
[0116] Specifically, the maternal sample file is randomly sampled according to the first maternal ratio to obtain a first maternal sampling file, and the fetal sample file is randomly sampled according to the first fetal ratio to obtain a first fetal sampling file. The maternal sample file is sampled according to the first maternal ratio obtained in step S204, and a corresponding number of reads are randomly selected from the maternal sample file to obtain a first maternal sampling file. Simultaneously, the fetal sample file is sampled according to the first fetal ratio obtained in step S204, and a corresponding number of reads are randomly selected from the fetal sample file to obtain a first fetal sampling file.
[0117] Exemplarily, the DownsampleSam module in the bioinformatics tool Picard is used to perform data sampling on the family sample file. The calculated first maternal ratio and first fetal ratio are input as parameters into the DownsampleSam module of Picard to perform data sampling on the BAM files of the mother and fetus respectively to generate corresponding sampling files.
[0118] Furthermore, the first maternal sampling file and the first fetal sampling file are data-mixed to obtain simulated first maternal plasma free nucleic acid data. The first maternal sampling file and the first fetal sampling file, each obtained by sampling, are data-mixed and combined to obtain first maternal plasma free nucleic acid data that simulates the concentration of fetal free nucleic acid in actual maternal plasma.
[0119] In some embodiments, step S104 may include but is not limited to steps S401 to S403 .
[0120] Step S401 : determining the frequency range of genotype occurrence frequency according to a preset concentration range of fetal free nucleic acid concentration.
[0121] Step S402 : performing single nucleotide polymorphism detection on the plasma cell-free nucleic acid data of the first pregnant woman according to preset high-frequency heterozygous single nucleotide polymorphism sites, collecting single nucleotide polymorphism sites whose genotype occurrence frequencies are within the frequency range, and obtaining a frequency site set.
[0122] Step S403, based on the frequency site set, the probability of each frequency value in the frequency range is calculated by a statistical method in a gradient manner, and the frequency value with the largest probability is determined as the predicted value of the fetal free nucleic acid concentration to obtain a prediction result.
[0123] Specifically, the frequency range of genotype occurrence is determined based on the preset concentration range of fetal free nucleic acid concentration. The genotype in the peripheral blood of pregnant women is mainly composed of the genotypes of the mother and the fetus, with reference to Figure 2 Genotype A and genotype B represent the most frequent allele and the second most frequent allele, respectively, at a single SNP. A homozygous genotype is defined by two genotype A variants, while a heterozygous genotype is defined by one genotype A and one genotype B variant.
[0124] Since the concentration of fetal free nucleic acid in maternal plasma is usually between 3% and 30%, it is necessary to determine a reasonable frequency range of genotype occurrence based on the preset concentration range of fetal free nucleic acid concentration to ensure that the SNP sites subsequently screened have a strong correlation with the fetal free nucleic acid concentration.
[0125] Furthermore, based on the preset high-frequency heterozygous single nucleotide polymorphism sites, single nucleotide polymorphism detection is performed on the plasma free nucleic acid data of the first pregnant woman, and single nucleotide polymorphism sites with genotype occurrence frequencies within the frequency range are collected to obtain a frequency site set.
[0126] First, high-frequency heterozygous SNP sites are selected from the SNP sites in the population, and these sites are preset as SNP sites for detection.
[0127] Next, SNP detection is performed on the plasma free nucleic acid data of the first pregnant woman using a bioinformatics tool (such as Samtools software), and the frequency range determined in step S401 is used to screen out SNP sites whose genotype occurrence frequency is within the frequency range to obtain a frequency site set.
[0128] Finally, based on the frequency site set, the probability of each frequency value within the frequency range is calculated gradiently by statistical methods, and the frequency value with the largest probability is determined as the predicted value of fetal free nucleic acid concentration to obtain the prediction result.
[0129] Optionally, after collecting the frequency site set, the binomial distribution probability calculation and the statistical method of maximum likelihood estimation can be combined to calculate the probability of each frequency value within the frequency range according to a certain gradient, and the frequency value with the highest probability is determined as the predicted value of the fetal free nucleic acid concentration to obtain the predicted result.
[0130] Exemplarily, the simulated plasma free nucleic acid data of the first pregnant woman are subjected to SNP detection using the mpileup module of the Samtools software, and the fetal concentration is predicted using the high-frequency heterozygous SNP sites in 820 populations to obtain the prediction results.
[0131] Combining the maternal and fetal SNP genotypes can produce four types of maternal peripheral blood. Figure 2 Taking the second type of SNP genotype combination as an example (i.e., the mother is a homozygous genotype and the fetus is a heterozygous genotype), the fetal free nucleic acid concentration is predicted. In the second case, the gene frequency of genotype B in the fetus is positively correlated with the fetal free nucleic acid concentration. Since the average frequency of genotype B in the fetus is half of the fetal free nucleic acid concentration, when the preset fetal free nucleic acid concentration range is set to 3% to 30%, it can be seen that the genotype B frequency range to be collected is 1.5% to 15%.
[0132] SNP sites with a genotype B frequency of 1.5% to 15% were collected, and the probability values were calculated using statistical methods from the concentration range of 1.5% to 15% with a gradient of 0.01%. When the probability value reached the maximum, the corresponding frequency value was the actual fetal free nucleic acid concentration predicted based on the plasma free nucleic acid data of the first pregnant woman.
[0133] In some embodiments, step S105 may include but is not limited to step S501 and step S502.
[0134] Step S501 , according to the prediction results corresponding to different preset fetal free nucleic acid concentrations and the first maternal ratio, a polynomial model is used to fit the relationship between the prediction results and the first maternal ratio to obtain maternal fitting parameters.
[0135] Step S502 , according to the prediction results corresponding to different preset fetal free nucleic acid concentrations and the first fetal ratio, a polynomial model is used to fit the relationship between the prediction results and the first fetal ratio to obtain fetal fitting parameters.
[0136] Specifically, refer to Figure 3 Since the actual fetal free nucleic acid concentration predicted based on the first pregnant woman’s plasma free nucleic acid data is larger than the theoretical preset free nucleic acid concentration, in order to correct the deviation of the fetal free nucleic acid concentration introduced during the simulation process, it is necessary to use the calculated actual fetal free nucleic acid concentration data combined with the sampling ratio of mother and fetus to perform polynomial fitting respectively.
[0137] In statistics and data analysis, polynomial fitting is often used to explore relationships between variables. Based on the predicted results and the first maternal ratio corresponding to different preset fetal free nucleic acid concentrations, a polynomial model is used to fit the relationship between the predicted results and the first maternal ratio to obtain the maternal fitting parameters. Based on the predicted results and the first fetal ratio corresponding to different preset fetal free nucleic acid concentrations, a polynomial model is used to fit the relationship between the predicted results and the first fetal ratio to obtain the fetal fitting parameters.
[0138] The actual predicted fetal free nucleic acid concentration (cff) under each preset fetal nucleic acid concentration was matched with the first maternal ratio (Rm) used in the first simulation, and the relationship between the predicted result and the first maternal ratio was fitted by a polynomial model to obtain the maternal fitting parameters.
[0139] The actual predicted fetal free nucleic acid concentration (cff) under each preset fetal nucleic acid concentration was matched with the first fetal ratio (Rf) used in the first simulation, and the relationship between the predicted result and the first fetal ratio was fitted by a polynomial model to obtain the fetal fitting parameters.
[0140] For example, a polynomial quadratic model is established, and the expression of the polynomial quadratic model is:
[0141] R=cff 2 ×a+cff×b+c (1),
[0142] Wherein, R represents the first sampling ratio, cff represents the actual fetal free nucleic acid concentration (ie, the predicted result), a represents the quadratic term coefficient, b represents the linear term coefficient, and c represents the constant term coefficient.
[0143] Using the polynomial quadratic model, the relationship between the predicted results and the first maternal ratio corresponding to different preset fetal free nucleic acid concentrations was fitted to obtain the quadratic term coefficient Ma, the linear term coefficient Mb and the constant term coefficient Mc in the mother fitting parameters.
[0144] Using a polynomial quadratic model, the relationship between the predicted results and the first fetal proportion corresponding to different preset fetal free nucleic acid concentrations was fitted to obtain the quadratic term coefficient Fa, the linear term coefficient Fb, and the constant term coefficient Fc in the fetal fitting parameters.
[0145] It can be understood that polynomial fitting is a method of approximating the actual data relationship by constructing a polynomial function. By selecting the appropriate polynomial degree (such as first, second, third, etc.), different degrees of fitting effects can be obtained. The higher the polynomial degree, the higher the fitting accuracy. The polynomial quadratic model is only exemplary. The fitting polynomial degree can be adjusted according to the experimental accuracy requirements, and this application does not impose specific restrictions.
[0146] In some embodiments, the second sampling ratio includes a second mother ratio and a second fetus ratio, and step S106 may include but is not limited to step S601 and step S602.
[0147] Step S601: Using the preset fetal free nucleic acid concentration as the input of the polynomial, the second mother ratio for extracting data from the mother sample file is determined according to the mother fitting parameters.
[0148] Step S602: Using the preset fetal free nucleic acid concentration as the input of the polynomial, the second fetal ratio of data extracted from the fetal sample file is determined according to the fetal fitting parameters.
[0149] In step S601 of some embodiments, a preset fetal free nucleic acid concentration is used as an input to a polynomial, and a second maternal ratio of data extracted from the maternal sample file is determined based on the maternal fitting parameters. The maternal fitting parameters (Ma, Mb, Mc) obtained in step S501 are substituted into formula (1), and a corresponding polynomial quadratic power model is established. The sampling ratio of the maternal sample file is corrected, and the preset fetal free nucleic acid concentration (ff) is used as an input to the polynomial, and the second maternal ratio Rm' of data extracted from the maternal sample file is calculated using the model.
[0150] In step S602 of some embodiments, a preset fetal free nucleic acid concentration is used as an input to a polynomial, and a second fetal ratio of data extracted from the fetal sample file is determined based on the fetal fitting parameters. The fetal fitting parameters (Fa, Fb, Fc) obtained in step S5502 are substituted into formula (1), and a corresponding polynomial quadratic power model is established. The sampling ratio of the fetal sample file is corrected, and the preset fetal free nucleic acid concentration (ff) is used as an input to the polynomial, and the second fetal ratio Rf' of data extracted from the fetal sample file is calculated using the model.
[0151] In some implementation examples, after re-sampling and merging the family sample files according to the second sampling ratio to obtain simulated second pregnant woman plasma free nucleic acid data, the fetal free nucleic acid concentration can also be predicted for the second pregnant woman plasma free nucleic acid data to verify whether the predicted result of the second simulated BAM file is close to the preset fetal free nucleic acid concentration.
[0152] Reference Figure 4 After sampling the second maternal plasma free nucleic acid data obtained from the second maternal ratio obtained in step S601 and the second fetal ratio obtained in step S602, the fetal free nucleic acid concentration is predicted for the second maternal plasma free nucleic acid data, and a scatter plot is created with the preset fetal free nucleic acid concentration as the horizontal axis and the fetal free nucleic acid concentration obtained by the actual prediction of the simulated data as the vertical axis. Figure 4 It can be seen that the predicted results are close to the preset fetal free nucleic acid concentration, so it can be considered that the simulation of maternal plasma data is successful.
[0153] In some embodiments, the simulation method of pregnant women's plasma free nucleic acid data may also include but is not limited to steps S701 to S702.
[0154] Step S701: enriching the free nucleic acid data of the plasma free nucleic acid data of the second pregnant woman with an inserted fragment length less than a preset screening threshold to obtain short fragment data.
[0155] Step S702: predict the fetal free nucleic acid concentration of the short fragment data to obtain an enriched prediction result.
[0156] Specifically, the free nucleic acid data of the second pregnant woman’s plasma free nucleic acid data with an insert length less than a preset sieving threshold is enriched to obtain short fragment data. Enrichment is an operational step of collecting a specific type of molecule to be measured (such as DNA, RNA or protein) from a large amount of maternal material to a smaller volume, thereby increasing its proportion in the sample. Related studies have shown that enriching free nucleic acids less than 160bp in pregnant woman’s plasma through experimental methods can effectively enrich fetal nucleic acids, thereby effectively increasing the concentration of fetal free nucleic acids. Therefore, for the simulated second pregnant woman’s plasma free nucleic acid data, the insert length information in the BAM file can be used to perform fragment screening, and the upper limit of the fragment screening is determined by the preset sieving threshold, and the free nucleic acid data corresponding to the insert length less than the preset sieving threshold are collected to obtain short fragment data.
[0157] Optionally, the preset screening threshold is 160dp, and the plasma free nucleic acid data of the second pregnant woman is screened according to the preset screening threshold, and the reads data corresponding to the inserted fragment information less than 160dp are enriched to obtain short fragment data.
[0158] It is understandable that the preset screening threshold can be set according to experimental requirements. 160dp is only an example. As long as short fragments of data that meet the experimental requirements can be screened out from the data according to the preset screening threshold, the embodiment of the present application does not impose any specific restrictions.
[0159] After obtaining the short fragment data, the fetal cell-free nucleic acid concentration is predicted based on the short fragment data to obtain an enriched prediction result. The enriched short fragment data can be used in subsequent high-throughput sequencing or other detection methods to predict the fetal cell-free nucleic acid concentration. Using the short fragment data can increase the fetal cell-free nucleic acid concentration and obtain an enriched prediction result.
[0160] The following describes and explains the solution of the embodiment of the present invention in detail with reference to specific application examples.
[0161] In the embodiment of the present application, a program for simulation process control and model fitting is written based on the Perl and Python programming languages, and the DownsampleSam module in the bioinformatics tool Picard is used in combination with the mpileup module of the bioinformatics tool Samtools to perform data sampling on the Bam file, and SNP site detection is used for subsequent fetal concentration prediction.
[0162] Reference Figure 5 , Figure 5 This is a flow chart of a simulation method for plasma free nucleic acid data of pregnant women provided in another embodiment of the present application, wherein the red dotted line represents the process of the first mixing according to the target mixing ratio, the orange solid line represents the process of polynomial fitting and correction, and the blue solid line represents the process of the second mixing using the corrected sampling ratio.
[0163] First, family information was obtained, and three pairs of mother-offspring paired DNA samples were used to simulate the plasma free DNA data of pregnant women. Among them, the offspring NA19978 was paired with the mother NA19976 to form the first family, the offspring NA07415 was paired with the mother NA07422 to form the second family, and the offspring NA09376 was paired with the mother NA09375 to form the third family.
[0164] Optionally, during the experimental library construction process, the time of the DNA sample of the offspring is extended by treating it with the disrupting enzyme, and the screening coefficient is increased at the same time, so that the offspring DNA fragments screened out are shorter, which can simulate the effect of shorter fetal free DNA fragments.
[0165] Subsequently, the offspring single sample and the mother single sample were analyzed through high-throughput sequencing and bioinformatics process analysis. The reads_number was obtained through sequencing technology to generate a BAM file, and the corresponding data volume and insert fragment length information were recorded.
[0166] Reference Figure 6 , the data volume of each offspring (Mf) and the mother data volume (Mm) as well as the length of the inserted fragment were counted respectively. The statistically obtained data volume was used to estimate the maximum data volume of the simulated mixed data, and the mixed data volume Ms used for simulation was determined to be 50M (1M=1000000 reads).
[0167] Next, the simulation range of fetal cell-free DNA concentration was set from 1% to 30%, with two lower concentrations of 0.5% and 0.8% set. Based on the amount of mixed data and the preset fetal cell-free DNA concentration, the amount of mixed offspring data to be extracted from the offspring BAM file was calculated as Ms×ff, and the amount of maternal data to be extracted from the maternal BAM file was calculated as Ms×(1-ff). Based on the calculated amount of mixed data, the sampling ratio for mixing was calculated, resulting in the offspring sampling ratio Rf = Ms×ff / Mf and the maternal sampling ratio Rm = Ms×(1-ff) / Mm.
[0168] For example, referring to Figure 6 The data volume Mf of the offspring NA19978 is 64353174, and the data volume Ms of the mother NA19976 is 67348599. When the total data volume after mixing is set to 50M, the sampling ratio can be calculated according to the preset fetal free DNA concentration (ff).
[0169] Taking ff as 2% as an example, the amount of offspring data extracted from the BAM file of offspring NA19978 after mixing is calculated to be 50M × 2% = 1,000,000, and the amount of maternal data extracted from the BAM file of mother NA19976 for mixing is calculated to be 50M × (1-2%) = 49,000,000. Based on the target mixing ratio for mixed sampling, the offspring sampling ratio NA19978-Rf = 1,000,000 / 64,353,174 ≈ 0.0155 (rounded to four decimal places), and the mother sampling ratio NA19976-Rm = 49,000,000 / 67,348,599 ≈ 0.7276 (rounded to four decimal places).
[0170] According to the above calculation method, the sampling ratios of the first mixing simulation corresponding to different preset fetal free DNA concentrations are calculated for the first family, the second family, and the third family, as follows: Figure 7 shown.
[0171] The calculated sampling ratio is input as a parameter into Picard-DownsampleSam through the bioinformatics software to perform random sampling and merging to simulate maternal plasma samples. After obtaining the first maternal plasma free nucleic acid data with different mixing ratios, the high-frequency heterozygous SNP sites in the 820 populations are used to predict the fetal free DNA concentration. The actual fetal free DNA concentration result cff obtained based on the first maternal plasma free nucleic acid data is calculated, as shown in the figure below: Figure 8 shown.
[0172] right Figure 8 A preliminary analysis of the data in the , it can be observed that the fetal free DNA concentrations of the three groups of families after the first simulated mixing were quite different from the preset fetal free DNA concentrations. Therefore, it is necessary to perform polynomial fitting based on the sampling ratios of offspring and mothers.
[0173] Reference Figure 9 , with the actual predicted fetal free DNA concentration as the horizontal axis and the offspring sampling ratio as the vertical axis, a scatter plot was drawn based on the data of offspring from different families, and a curve reflecting the changing trend between the offspring sampling ratio and the actual predicted fetal free DNA concentration after mixing was obtained through polynomial fitting.
[0174] It should be noted that the fitting parameters are not universal across different families and must be individually fitted for each family. By establishing a quadratic polynomial model and performing polynomial fitting on the offspring and mother of each family, the fitting parameters for the offspring and mother of each family were obtained.
[0175] Reference Figure 10 For ease of presentation, the fitting parameters are retained to four decimal places. Substituting the calculated fitting parameters into the polynomial quadratic model, the curves of each family are obtained as follows Figure 9 As shown, the blue curve represents the offspring of the first family, the orange curve represents the offspring of the second family, and the yellow curve represents the offspring of the third family.
[0176] The preset fetal free DNA concentration was used as the input parameter and input into the polynomial quadratic model corresponding to each family. The corrected sampling ratios of offspring and mother were calculated respectively. For ease of display, the calculated sampling ratio results were rounded to eight decimal places, as shown in the following example: Figure 11 shown.
[0177] Taking the first family as an example, with ff=2% as input, the corrected offspring sampling ratio NA19978-Rf'≈0.00877808 and the corrected mother sampling ratio NA19976-Rm'≈0.73406132 are calculated.
[0178] The corrected sampling ratio was used to perform sample mixing and re-simulate the maternal plasma sample. After obtaining the second maternal plasma free nucleic acid data from the second mixing, the 820 high-frequency heterozygous SNP sites in the population were used to perform a second fetal free DNA concentration prediction. The second fetal free DNA concentration result, cff', was obtained to verify its consistency with the preset fetal free DNA concentration.
[0179] Reference Figure 12 After the second mixing, the fetal free DNA concentration was roughly consistent with the preset concentration, and the error between the measured fetal free DNA concentration ratio and the expected fetal free DNA concentration ratio was within ±1%. Therefore, the plasma free nucleic acid data of the second pregnant woman obtained in the second simulation can be used to carry out subsequent research and development of non-invasive prenatal screening.
[0180] In an embodiment of the present application, by using two comparison data files of mother-child pairings, combined with bioinformatics tools and model fitting methods, mixed data that meets the expected fetal concentration can be output to simulate maternal plasma data, thereby enabling research and development work related to prenatal screening technology to be carried out, greatly reducing the time cost spent on sampling and experiments, and being able to batch generate maternal plasma simulation data with different fetal concentrations under different data volumes in a short period of time, which helps to improve the flexibility of non-invasive prenatal screening in obtaining positive samples.
[0181] In some embodiments, the fragment length distribution after mixing and the fragment length distribution of the mother and child samples can be observed later. Figure 13 Taking the first family as an example, the peak of the fragment length distribution of offspring NA19978 (green line) is lower than that of the mother NA19976 (blue line). The fragment length distribution of the simulated data after mixing (red line) is also biased toward the short fragment range due to the influence of the shorter fragments in the offspring. Therefore, it is possible to increase the concentration of free fetal DNA by enriching for short fragments.
[0182] For example, referring to Figure 14 Using the inserted fragment length information in the BAM file, the corresponding reads data of inserted fragments less than 160bp were collected and then the fetal concentration was predicted. The comparison results showed that after enriching short fragments, the fetal free DNA concentration could be increased to about 1.5 times the original.
[0183] Reference Figure 15 The present application also provides a system for simulating cell-free nucleic acid data in pregnant women's plasma, which can implement the above-mentioned method for simulating cell-free nucleic acid data in pregnant women's plasma. The system includes:
[0184] The first module is used to obtain family sample files for mother-child pairing.
[0185] The second module is used to perform data volume analysis on the family sample file according to the preset fetal free nucleic acid concentration to obtain a first sampling ratio.
[0186] The third module is used to sample and merge the family sample files according to the first sampling ratio to obtain the simulated plasma free nucleic acid data of the first pregnant woman.
[0187] The fourth module is used to predict the fetal free nucleic acid concentration based on the plasma free nucleic acid data of the first pregnant woman to obtain a prediction result.
[0188] The fifth module is used to perform polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters.
[0189] The sixth module is used to use the preset fetal free nucleic acid concentration as the input of the polynomial and determine the second sampling ratio according to the fitting parameters.
[0190] The seventh module is used to resample and merge the family sample files according to the second sampling ratio to obtain the simulated plasma free nucleic acid data of the second pregnant woman.
[0191] It can be understood that the contents of the above method embodiments are applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0192] The present application also provides an electronic device comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method for simulating cell-free nucleic acid data in maternal plasma. The electronic device can be any smart terminal, including a tablet computer and an in-vehicle computer.
[0193] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0194] Reference Figure 16 , Figure 16 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0195] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0196] The memory 902 can be implemented in the form of a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM), etc. The memory 902 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present specification are implemented by software or firmware, the related program codes are stored in the memory 902 and are called and executed by the processor 901 to perform the simulation method of the pregnant woman's plasma free nucleic acid data provided by the embodiments of the present application.
[0197] The input / output interface 903 is configured to realize information input and output.
[0198] The communication interface 904 is configured to realize the communication interaction between the device and other devices. The communication can be realized by a wired manner (for example, a USB, a network cable, etc.) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, etc.).
[0199] The bus 905 is configured to transmit information between various components (for example, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904) of the device.
[0200] The processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are connected to each other through the bus 905 to realize the communication connection between the device.
[0201] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the simulation method of the pregnant woman's plasma free nucleic acid data.
[0202] It can be understood that the contents in the above method embodiments are applicable to the present storage medium embodiments. The functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved by the present storage medium embodiments are also the same as those achieved by the above method embodiments.
[0203] The memory is a non-transitory computer readable storage medium, which can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory can include a high-speed random access memory and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0204] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0205] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0206] The system embodiment described above is merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0207] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0208] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0209] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A method for simulating plasma free nucleic acid data of pregnant women, characterized in that: The method comprises the following steps: Obtaining a pedigree sample file for mother-child pairing, wherein the pedigree sample file includes a mother sample file and a fetus sample file; Performing data volume analysis on the family sample file according to a preset fetal free nucleic acid concentration to obtain a first sampling ratio, wherein the first sampling ratio includes a first maternal ratio and a first fetal ratio; Sampling and merging the family sample files according to the first sampling ratio to obtain simulated plasma cell-free nucleic acid data of the first pregnant woman; Predicting the fetal free nucleic acid concentration based on the first pregnant woman's plasma free nucleic acid data to obtain a prediction result; Performing polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters; Using the preset fetal free nucleic acid concentration as an input of a polynomial, and determining a second sampling ratio according to the fitting parameters; Re-sampling and merging the family sample files according to the second sampling ratio to obtain simulated plasma cell-free nucleic acid data of the second pregnant woman; The step of performing data volume analysis on the family sample file according to the preset fetal free nucleic acid concentration to obtain a first sampling ratio includes the following steps: Performing data volume statistics on the maternal sample file and the fetal sample file respectively to obtain maternal data volume and fetal data volume; Determining a simulated mixed data volume according to a preset fetal free nucleic acid concentration, wherein a product of the mixed data volume and the preset fetal free nucleic acid concentration does not exceed the fetal data volume; Determining the maternal sampling amount and the fetal sampling amount according to the preset fetal free nucleic acid concentration and the mixed data amount; The first maternal ratio is obtained according to the ratio of the maternal sample size to the maternal data size, and the first fetal ratio is obtained according to the ratio of the fetal sample size to the fetal data size.
2. The method according to claim 1, characterized in that The step of sampling and merging the family sample files according to the first sampling ratio to obtain simulated plasma cell-free nucleic acid data of the first pregnant woman comprises the following steps: Randomly sampling the mother sample file according to the first mother ratio to obtain a first mother sampling file; Randomly sampling the fetal sample file according to the first fetal ratio to obtain a first fetal sampling file; The first mother sampling file and the first fetus sampling file are mixed to obtain simulated first pregnant woman plasma free nucleic acid data.
3. The method according to claim 1, characterized in that The step of predicting the fetal free nucleic acid concentration based on the first pregnant woman's plasma free nucleic acid data to obtain a prediction result comprises the following steps: Determining the frequency range of genotype occurrence frequency according to the preset concentration range of fetal free nucleic acid concentration; performing single nucleotide polymorphism detection on the plasma cell-free nucleic acid data of the first pregnant woman according to the preset high-frequency heterozygous single nucleotide polymorphism sites, collecting single nucleotide polymorphism sites whose genotype occurrence frequency is within the frequency range, and obtaining a frequency site set; According to the frequency site set, the probability of each frequency value within the frequency range is calculated in a gradient manner using a statistical method, and the frequency value with the largest probability is determined as the actual fetal free nucleic acid concentration to obtain a prediction result.
4. The method according to claim 1, wherein The step of performing polynomial fitting according to the prediction result and the first sampling ratio to obtain fitting parameters comprises the following steps: According to the prediction results corresponding to different preset fetal free nucleic acid concentrations and the first maternal ratio, a relationship between the prediction results and the first maternal ratio is fitted using a polynomial model to obtain a maternal fitting parameter; According to the predicted results and the first fetal ratio corresponding to different preset fetal free nucleic acid concentrations, a polynomial model is used to fit the relationship between the predicted results and the first fetal ratio to obtain fetal fitting parameters.
5. The method according to claim 4, characterized in that The second sampling ratio includes a second maternal ratio and a second fetal ratio, and the method of using the preset fetal free nucleic acid concentration as an input of the polynomial and determining the second sampling ratio according to the fitting parameters comprises the following steps: Using the preset fetal free nucleic acid concentration as an input to a polynomial, determining a second mother proportion of data extracted from the mother sample file based on the mother fitting parameters; The preset fetal free nucleic acid concentration is used as an input of a polynomial, and a second fetal proportion of data extracted from the fetal sample file is determined according to the fetal fitting parameters.
6. The method according to claim 1, characterized in that The simulation method of the plasma free nucleic acid data of pregnant women also includes the following steps: Enriching the free nucleic acid data of the plasma free nucleic acid data of the second pregnant woman in which the length of the inserted fragment is less than a preset screening threshold to obtain short fragment data; The fetal free nucleic acid concentration is predicted for the short fragment data to obtain an enriched prediction result.
7. A simulation system for plasma free nucleic acid data of pregnant women, characterized in that: The system comprises: The first module is used to obtain a family sample file for mother-child pairing, wherein the family sample file includes a mother sample file and a fetus sample file; The second module is configured to perform data volume analysis on the family sample file according to a preset fetal free nucleic acid concentration to obtain a first sampling ratio, wherein the first sampling ratio includes a first maternal ratio and a first fetal ratio; A third module is configured to sample and merge the family sample files according to the first sampling ratio to obtain simulated plasma cell-free nucleic acid data of the first pregnant woman; A fourth module is used to predict the fetal free nucleic acid concentration based on the first pregnant woman's plasma free nucleic acid data to obtain a prediction result; A fifth module is configured to perform polynomial fitting based on the prediction result and the first sampling ratio to obtain fitting parameters; A sixth module is configured to use the preset fetal free nucleic acid concentration as an input of a polynomial and determine a second sampling ratio according to the fitting parameters; A seventh module is configured to resample and merge the family sample files according to the second sampling ratio to obtain simulated plasma cell-free nucleic acid data of a second pregnant woman; The step of performing data volume analysis on the family sample file according to the preset fetal free nucleic acid concentration to obtain a first sampling ratio includes the following steps: Performing data volume statistics on the maternal sample file and the fetal sample file respectively to obtain maternal data volume and fetal data volume; Determining a simulated mixed data volume according to a preset fetal free nucleic acid concentration, wherein a product of the mixed data volume and the preset fetal free nucleic acid concentration does not exceed the fetal data volume; Determining the maternal sampling amount and the fetal sampling amount according to the preset fetal free nucleic acid concentration and the mixed data amount; The first maternal ratio is obtained according to the ratio of the maternal sample size to the maternal data size, and the first fetal ratio is obtained according to the ratio of the fetal sample size to the fetal data size.
8. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Determination method of fetal DNA content in maternal plasma, based on single-nucleotide polymorphic loci
CN103215350A
Method for improving proportion of fetal free DNA in maternal plasma free DNA sequencing library
CN105926043A