Noninvasive prenatal fetal chromosome detection method, device and kit
Through magnetic bead purification and sliding window technology combined with dynamic copy number correction, the non-invasive prenatal fetal chromosome detection method is solved, and the fetal DNA concentration is insufficient and the repeat sequence error is achieved, efficient and accurate chromosome detection, especially the detection of rare chromosome abnormalities, improving the detection efficiency and accuracy.
Patent Information
- Application Number
- CN202510455507.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing non-invasive prenatal fetal chromosome detection technology has problems such as insufficient fetal DNA concentration, repeat sequence error, high false positive rate, and insufficient rare chromosome detection, resulting in low detection accuracy and efficiency.
The method of enhancing fetal DNA by purifying magnetic beads, combining sliding window technology and dynamic copy number correction, a plasma free DNA sequencing library was constructed, and a negative and positive Z value was used for joint analysis, genome duplication regions were identified and chromosome copy number correction was performed, and a sex chromosome detection decision tree model was established.
It significantly improves fetal DNA concentration, reduces detection failure and false negative rates, improves the accuracy and sensitivity of chromosome detection, covers 23 aneuploid abnormal detection of chromosomes, and enhances the technical support for prenatal screening.
Smart Images

Figure CN120505400A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of chromosome detection, and in particular to a method, device and kit for non-invasive prenatal fetal chromosome detection. Background Art
[0002] With the continuous advancement of high-throughput sequencing technology (NGS) and the gradual reduction in costs, NIPT technology has been widely used in clinical practice and has become the preferred method for screening the three common chromosomal aneuploidies (T13, T18, and T21). Currently, non-invasive prenatal screening technology (NIPT) for the three common chromosomal aneuploidies (T13, T18, and T21) is relatively mature and has high accuracy. However, it still has certain technical limitations, as follows:
[0003] 1) Insufficient fetal DNA in maternal plasma. The proportion of free fetal DNA in plasma varies greatly between different maternal plasmas, usually ranging from 3% to 30%. Samples with low fetal concentrations often lead to test failure or false negative results.
[0004] 2) Currently available sequencing data preprocessing and noise reduction methods fail to fully consider the interference of positive chromosomes on the copy number calculation of other chromosomes and the repetitive sequence errors introduced by repetitive regions. The chromosome copy number is the ratio of the reads of the target chromosome to the total reads of the autosomes. If there is an aneuploid abnormality in the chromosome, the reads of the chromosome will be significantly higher than the expected value of the negative sample, resulting in a significant deviation in the total read value of the autosomes, affecting the copy number calculation of other chromosomes. The traditional Picard-based method for removing repetitive sequences fails to effectively remove repetitive regions in the genomic region and will introduce a large number of false positive results.
[0005] 3) Currently, conventional variant detection algorithms are usually based on the Z-test method calculated on the negative sample set. This method tests a single dimension, and the mean and variance of the Z-score calculation formula are established based on the negative sample set data. It only considers the distribution of negative samples and ignores the distribution of test samples in the positive sample set. Therefore, it cannot take into account both sensitivity and specificity.
[0006] 4) NIPT testing generally suffers from insufficient positive predictive value (PPV). Results from multiple prospective clinical studies of NIPT have shown that the PPV for T13 is approximately 10-40%, for T18 is approximately 60-90%, and for T21 is approximately 80-90%. Furthermore, false-positive results can lead to unnecessary biopsy, causing psychological stress for pregnant women. Because fetal DNA in maternal plasma primarily originates from placental trophoblast cells rather than fetal cells, inconsistencies in the karyotypes of placental trophoblast cells and fetal cells can result in false-positive NIPT results. The most common type of this is restricted placental mosaicism, where the placental trophoblast cells are abnormal but the fetus is normal. Patents 202210825534.2 and 202311069138.2 propose methods for calculating the placental mosaicism ratio using fetal concentration and copy number values and using the mosaicism ratio to reduce false positives in NIPT detection. However, such methods have obvious technical limitations, specifically: First, the calculation of the mosaicism ratio depends on the fetal DNA concentration. The accuracy of the current prediction of fetal concentration for female fetuses still has large errors, which affects the accuracy of the mosaicism ratio calculation; Second, placental mosaicism includes multiple classifications, such as fetal mosaicism 4 (TFM4) and fetal mosaicism 6 (TFM6), in which fetal cells are abnormal; The above method simply corrects samples with positive placental mosaic signals to negative, which can reduce the false positive rate to a certain extent, but will introduce a significant risk of missed detection.
[0007] 5) Research on rare chromosomal aneuploidies (RATs) other than chromosomes 13, 18, and 21 is insufficient. Rare chromosome-positive samples are rare, and there is a lack of sufficient samples for reference set establishment, algorithm modeling, and evaluation, which affects the reliability and accuracy of the test results.
[0008] 6) Non-invasive prenatal screening (NIPT) can detect fetal sex chromosome aneuploidy (SCA) abnormalities, including monosomy X (45,X), Klinefelter syndrome (47,XXY), superfetogenic syndrome (47,XXX), and superandrogenic syndrome (47,XYY). These abnormalities have a combined prevalence of 1:500 and are more common than the more common trisomy syndromes. Currently, NIPT primarily detects autosomal abnormalities. Rare sex chromosome aneuploidies, such as XOXY and XXXY, are often not included in testing models due to their low clinical prevalence, leading to missed detection or misdiagnosis.
[0009] Due to the above technical shortcomings, the efficiency and accuracy of non-invasive prenatal fetal chromosomal aneuploidy detection have decreased. Summary of the Invention
[0010] The main purpose of the embodiments of the present application is to propose a method, device and kit for non-invasive prenatal fetal chromosome detection, which can improve the efficiency and accuracy of non-invasive prenatal fetal chromosome detection.
[0011] In one aspect, the present invention provides a method for non-invasive prenatal fetal chromosome detection, comprising the following steps:
[0012] Based on the magnetic bead purification and enrichment kit for fetal DNA and the library construction system, a plasma free DNA sequencing library corresponding to maternal peripheral blood was constructed;
[0013] Obtaining genome sequencing data from the plasma cell-free DNA sequencing library;
[0014] Aligning the genome sequencing data to a human reference genome, dividing the human reference genome into a plurality of unit windows according to a preset unit window width, and calculating the number of sequencing sequences corresponding to the genome sequencing data aligned to each of the unit windows;
[0015] Screening out a plurality of target candidate windows from the plurality of unit windows according to the number of sequencing sequences corresponding to each of the unit windows, wherein the number of sequencing sequences corresponding to the target candidate windows exceeds a given average level threshold;
[0016] Determining, based on each of the target candidate windows, a plurality of genomic duplication regions in the genomic sequencing data, wherein the genomic duplication regions are genomic regions in which the target candidate windows appear continuously and the number of occurrences is greater than a given threshold;
[0017] Obtaining multiple plasma free DNA samples to be tested in the current test batch; the plasma free DNA samples to be tested are obtained by enriching fetal free DNA in the plasma of the pregnant woman to be tested by a magnetic bead purification method;
[0018] For each chromosome to be detected in each of the plasma free DNA samples to be detected, performing regional screening on the chromosome to be detected based on the genome sequencing data and each of the genome repetitive regions to determine the number of gene sequences corresponding to the chromosome to be detected;
[0019] For each of the plasma free DNA samples to be tested, copy number statistics are performed based on the number of gene sequences corresponding to each of the chromosomes to be tested to determine the copy number of each of the chromosomes to be tested;
[0020] For each of the plasma cell-free DNA samples to be tested, obtaining a negative sample data set corresponding to each of the chromosomes to be tested, and determining a negative Z value corresponding to each of the chromosomes to be tested based on the copy number corresponding to each of the chromosomes to be tested and the negative sample data set;
[0021] For each of the plasma free DNA samples to be tested, dynamically correcting the copy number of each of the chromosomes to be tested according to the negative Z value corresponding to each of the chromosomes to be tested and a given quantity index threshold, and determining a copy number correction value corresponding to each of the chromosomes to be tested;
[0022] According to the copy number correction value corresponding to each of the chromosomes to be detected in each of the plasma free DNA samples to be detected, performing chromosome copy number error correction for the current detection batch, and determining a copy number batch correction value corresponding to each of the chromosomes to be detected in each of the plasma free DNA samples to be detected;
[0023] For each of the plasma free DNA samples to be tested, a corresponding positive sample data set is obtained, and a negative Z value and a positive Z value are determined based on the negative sample data set and the positive sample data set. Chromosome detection is performed based on the copy number batch correction value, the negative Z value, and the positive Z value corresponding to each of the chromosomes to be detected in each of the plasma free DNA samples to be tested, and the chromosome detection result of each of the plasma free DNA samples to be tested is determined, wherein the positive sample data set is simulated and established based on the negative sample data set.
[0024] In some embodiments, for each chromosome to be detected in each plasma free DNA sample to be detected, regional screening of the chromosome to be detected is performed based on the genome sequencing data and each genome duplication region to determine the number of gene sequences corresponding to the chromosome to be detected, specifically comprising:
[0025] Determine, based on the genome sequencing data, multiple gene sequencing regions in the chromosome to be detected and the number of regions and region length corresponding to each gene sequencing region, where the region length is the sequence length corresponding to the gene sequencing region, and the number of regions is the frequency of occurrence of the gene sequencing region in the chromosome to be detected;
[0026] Dividing the gene sequencing regions into multiple gene sequencing groups according to the region lengths corresponding to the gene sequencing regions, each gene sequencing group including multiple gene sequencing regions with the same region length;
[0027] According to each of the genomic duplication regions, screening the chromosome duplication region corresponding to each of the genomic duplication regions from the multiple gene sequencing regions, and determining the target gene sequencing group where the chromosome duplication region is located from the multiple gene sequencing groups;
[0028] For each of the chromosome duplication regions, updating the region number corresponding to the chromosome duplication region to the average region number corresponding to the target gene sequencing group, where the average region number is the average of the region numbers of multiple gene sequencing regions in the gene sequencing group;
[0029] The number of gene sequences corresponding to the chromosome to be detected is determined based on the number of regions corresponding to each of the gene sequencing regions except the multiple chromosome duplication regions and the number of regions after update of each of the chromosome duplication regions.
[0030] In some embodiments, for each of the plasma free DNA samples to be tested, copy number statistics are performed based on the number of gene sequences corresponding to each of the chromosomes to be tested to determine the copy number of each of the chromosomes to be tested, specifically comprising:
[0031] Determining the total number of gene sequences corresponding to the plasma cell-free DNA sample to be tested based on the number of gene sequences corresponding to each of the chromosomes to be tested;
[0032] The copy number corresponding to each chromosome to be detected is determined based on the number of gene sequences corresponding to each chromosome to be detected and the total number of gene sequences. The copy number is calculated by the following formula:
[0033] RR ij =RN ij / AN i ;
[0034]
[0035] Among them, RR ij is the copy number of the jth chromosome to be detected in the i-th plasma free DNA sample to be detected, RN ij is the number of gene sequences corresponding to the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, AN i is the total number of gene sequences corresponding to the i-th plasma free DNA sample to be tested, and N is the number of chromosomes to be tested in the i-th plasma free DNA sample to be tested.
[0036] In some embodiments, for each of the plasma free DNA samples to be tested, obtaining a negative sample data set corresponding to each of the chromosomes to be tested, and determining a negative Z value corresponding to each of the chromosomes to be tested based on the copy number corresponding to each of the chromosomes to be tested and the negative sample data set, specifically includes:
[0037] Determine the corresponding negative sample copy number mean and negative sample copy number standard deviation according to the negative sample data set corresponding to each of the chromosomes to be detected;
[0038] According to the copy number corresponding to the chromosome to be detected, the mean copy number of the negative sample and the standard deviation of the copy number of the negative sample, the negative Z value corresponding to the chromosome to be detected is determined, and the negative Z value is calculated by the following formula:
[0039] Z_neg ij =(RR ij -μneg j ) / σneg j ;
[0040] Among them, Z_neg ij is the negative Z value corresponding to the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, RR ij is the copy number of the jth chromosome to be detected in the i-th plasma free DNA sample to be detected, μneg j is the mean number of negative sample copies corresponding to the jth chromosome to be tested in the negative sample data set, σneg j is the standard deviation of the negative sample copy number corresponding to the jth chromosome to be detected in the negative sample data set.
[0041] In some embodiments, for each of the plasma free DNA samples to be tested, dynamically correcting the copy number of each of the chromosomes to be tested according to the negative Z value corresponding to each of the chromosomes to be tested and a given quantity index threshold, and determining the copy number correction value corresponding to each of the chromosomes to be tested, specifically includes:
[0042] Determining a plurality of first abnormal chromosomes from the plurality of chromosomes to be detected according to the negative Z value corresponding to each of the chromosomes to be detected and the quantitative index threshold, wherein the negative Z value corresponding to the first abnormal chromosome exceeds the quantitative index threshold;
[0043] obtaining any of the first abnormal chromosomes as a current corrected chromosome;
[0044] Obtaining the total number of gene sequences corresponding to the plasma cell-free DNA sample to be tested and the mean number of negative sample copies corresponding to the currently corrected chromosome;
[0045] According to the total number of gene sequences corresponding to the plasma free DNA sample to be tested and the mean number of negative sample copies corresponding to the current corrected chromosome, the number of gene sequences corresponding to the current corrected chromosome is updated, and the updated number of gene sequences is calculated by the following formula:
[0046]
[0047] in, is the updated number of gene sequences corresponding to the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, AN i is the total number of gene sequences corresponding to the i-th plasma free DNA sample to be tested, μneg j is the mean copy number of negative samples corresponding to the jth chromosome to be tested;
[0048] updating the total number of gene sequences according to the number of updated gene sequences corresponding to the currently corrected chromosome and the number of gene sequences corresponding to the remaining chromosomes to be tested;
[0049] Recalculating the chromosome copy number based on the total number of updated gene sequences, the number of gene sequences of the remaining chromosomes to be detected, and the updated number of gene sequences corresponding to the currently corrected chromosome, to determine the corrected copy number value corresponding to the currently corrected chromosome and the corrected copy number values corresponding to the remaining chromosomes to be detected;
[0050] When there is any first abnormal chromosome whose gene sequence quantity has not been updated, the first abnormal chromosome whose gene sequence quantity has not been updated is obtained as the current corrected chromosome, and then the total number of gene sequences corresponding to the plasma free DNA sample to be tested and the mean copy number of the negative samples corresponding to the current corrected chromosome are obtained again until the gene sequence quantity of all the first abnormal chromosomes has been updated.
[0051] In some embodiments, performing chromosome copy number error correction for the current test batch based on the copy number correction value corresponding to each of the chromosomes to be tested in each of the plasma free DNA samples to be tested, and determining the copy number batch correction value corresponding to each of the chromosomes to be tested in each of the plasma free DNA samples to be tested, specifically includes:
[0052] Acquire any of the plasma free DNA samples to be detected as a target detection sample, acquire any of the chromosomes to be detected in the target detection sample as a current target detection chromosome, and acquire the copy number correction value corresponding to the current target detection chromosome as the current copy number;
[0053] Obtaining the negative sample copy number mean, the negative sample copy number standard deviation, and the current batch chromosome copy number mean corresponding to the current target detection chromosome, wherein the current batch chromosome copy number mean is the copy number mean corresponding to the current target detection chromosome in all the plasma free DNA samples to be detected;
[0054] Determine the normalized copy number corresponding to the current target detection chromosome according to the current copy number, the mean copy number of the negative samples and the mean copy number of the chromosomes in the current batch;
[0055] Update the negative Z value corresponding to the current target detection chromosome according to the normalized copy number, the negative sample copy number mean and the negative sample copy number standard deviation;
[0056] Determining whether the current target detection chromosome is the second abnormal chromosome according to the updated negative Z value corresponding to the current target detection chromosome and the quantity index threshold;
[0057] When the updated negative Z value corresponding to the current target detection chromosome is greater than the quantity index threshold, the current target detection chromosome is determined to be the second abnormal chromosome, the negative sample copy number mean is determined as the current copy number, and the current batch chromosome copy number mean is re-determined based on the current copy number, and then the normalized copy number corresponding to the current target detection chromosome is determined based on the current copy number, the negative sample copy number mean, and the current batch chromosome copy number mean, until the updated negative Z value corresponding to the current target detection chromosome is less than the quantity index threshold;
[0058] When the updated negative Z value corresponding to the current target detection chromosome is less than the quantity index threshold, determining the normalized copy number as the copy number batch correction value corresponding to the current target detection chromosome;
[0059] When there is any chromosome to be detected in the target detection sample that has not undergone chromosome copy number error correction, any remaining chromosome to be detected in the target detection sample that has not undergone chromosome copy number error correction is obtained as the current target detection chromosome, the copy number correction value corresponding to the current target detection chromosome is obtained as the current copy number, and then the negative sample copy number mean, the negative sample copy number standard deviation, and the current batch chromosome copy number mean corresponding to the current target detection chromosome are obtained, until all the chromosomes to be detected in the target detection sample have undergone chromosome copy number error correction;
[0060] The normalized copy number was calculated by the following formula:
[0061]
[0062] Among them, the current detection chromosome is the jth chromosome to be detected in the i-th plasma free DNA sample to be detected, The normalized copy number corresponding to the current target chromosome detection, Correction value for the copy number of the current target chromosome, μb j is the mean number of chromosome copies in the current batch corresponding to the current target detection chromosome in the current detection batch, μneg j The mean copy number of negative samples corresponding to the current target chromosome.
[0063] In some embodiments, for each of the plasma free DNA samples to be tested, a corresponding positive sample data set is obtained, a negative Z value and a positive Z value are determined based on the negative sample data set and the positive sample data set, and chromosome detection is performed based on the copy number batch correction value, the negative Z value, and the positive Z value corresponding to each of the chromosomes to be tested in each of the plasma free DNA samples to be tested, to determine the chromosome detection result of each of the plasma free DNA samples to be tested, specifically comprising:
[0064] Obtaining the negative sample data set corresponding to the chromosome to be detected;
[0065] Determining the expected copy number value of the positive sample data set according to the fetal concentration distribution of the negative sample data set, and generating the positive sample data set according to the expected copy number value and random error;
[0066] Performing copy number distribution statistics on the negative sample data set to determine the negative Z value judgment threshold;
[0067] Update the negative Z value corresponding to the chromosome to be detected according to the copy number batch correction value corresponding to the chromosome to be detected, the negative sample copy number mean and the negative sample copy number standard deviation;
[0068] When the chromosome to be detected is an autosome, obtaining a positive sample data set corresponding to the chromosome to be detected, performing copy number distribution statistics on the positive sample data set, and determining the positive Z value judgment threshold, the positive sample copy number mean, and the positive sample copy number standard deviation;
[0069] The positive Z value corresponding to the chromosome to be detected is determined according to the copy number batch correction value corresponding to the chromosome to be detected, the mean copy number of the positive samples, and the standard deviation of the copy number of the positive samples. The positive Z value is calculated by the following formula:
[0070]
[0071] Among them, Z_pos ij is the positive Z value corresponding to the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, μpos j is the mean sample copy number corresponding to the jth chromosome to be tested in the positive sample data set, is the batch correction value of the copy number of the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, σpos j is the standard deviation of the sample copy number corresponding to the jth chromosome to be tested in the positive sample data set;
[0072] Performing a joint test based on the negative Z value judgment threshold, the negative Z value, the positive Z value judgment threshold, and the positive Z value to determine the chromosome detection result;
[0073] When the negative Z value is greater than the negative Z value judgment threshold, and the positive Z value is greater than the positive Z value judgment threshold, determining that the chromosome detection result is positive for the chromosome to be detected in the plasma free DNA sample to be detected;
[0074] When the negative Z value is less than or equal to the negative Z value judgment threshold, and the positive Z value is less than or equal to the positive Z value judgment threshold, determining that the chromosome detection result is negative for the chromosome to be detected in the plasma free DNA sample to be detected;
[0075] When the chromosome to be detected is a sex chromosome, a sex chromosome detection decision tree model is obtained, and the chromosome detection result is determined using the sex chromosome detection decision tree model according to the negative Z value judgment threshold and the negative Z value.
[0076] In some embodiments, when the chromosome to be detected is a sex chromosome, obtaining a sex chromosome detection decision tree model, and using the sex chromosome detection decision tree model to determine the chromosome detection result according to the negative Z value judgment threshold and the negative Z value, specifically includes:
[0077] Obtain the first negative Z value corresponding to the X chromosome and the second negative Z value corresponding to the Y chromosome and the Z value judgment threshold;
[0078] The first negative Z value, the second negative Z value and the Z value judgment threshold are input into the sex chromosome detection decision tree model, and the sex chromosome detection decision tree model is used to output the fetal sex and the corresponding sex chromosome detection result.
[0079] In some embodiments, the kit for fetal DNA enrichment based on magnetic beads and the library construction system are used to construct a plasma cell-free DNA sequencing library corresponding to maternal peripheral blood, specifically comprising:
[0080] Extraction of plasma free DNA from maternal plasma;
[0081] performing fragment screening on the plasma free DNA to obtain fragment-screened plasma free DNA;
[0082] The plasma free DNA after the fragment screening is filled with A and phosphorylated to obtain the corresponding end repair product;
[0083] Performing adapter ligation on the end-repair product to obtain a corresponding adapter ligation product;
[0084] Purifying the adapter ligation product with magnetic beads, and performing library PCR amplification on the adapter ligation product after magnetic bead purification to obtain corresponding PCR products;
[0085] The PCR products are purified by magnetic beads, and the library is quantified and sequenced to generate the plasma free DNA sequencing library.
[0086] On the other hand, an embodiment of the present application provides a non-invasive prenatal fetal chromosome detection device, the device comprising:
[0087] A data preprocessing module is used to construct a plasma free DNA sequencing library corresponding to the peripheral blood of pregnant women based on a magnetic bead purification and enrichment fetal DNA kit and a library construction system; obtain genome sequencing data from the plasma free DNA sequencing library; align the genome sequencing data to a human reference genome, divide the human reference genome into multiple unit windows according to a preset unit window width, and calculate the number of sequencing sequences corresponding to the genome sequencing data aligned to each unit window; based on the number of sequencing sequences corresponding to each unit window, screen out a number of target candidate windows from the multiple unit windows, wherein the number of sequencing sequences corresponding to the target candidate windows exceeds a given average level threshold; based on each target candidate window, determine a number of genomic duplication regions in the genomic sequencing data, wherein the genomic duplication regions are genomic regions in which the target candidate windows appear continuously and the number of occurrences is greater than a given threshold;
[0088] The chromosome copy number dynamic correction module is used to obtain multiple plasma free DNA samples to be tested in the current test batch; the plasma free DNA samples to be tested are obtained by enriching fetal free DNA in the plasma of the pregnant woman to be tested by a magnetic bead purification method; for each chromosome to be tested in each plasma free DNA sample to be tested, according to the genome sequencing data and each genome duplication region, the chromosome to be tested is regionally screened to determine the number of gene sequences corresponding to the chromosome to be tested; for each plasma free DNA sample to be tested, according to the number of gene sequences corresponding to each chromosome to be tested, copy number statistics are performed to determine the copy number of each chromosome to be tested; for each plasma free DNA sample to be tested, the copy number of each chromosome to be tested is obtained. The negative sample data set corresponding to the chromosome to be detected, determining the negative Z value corresponding to each chromosome to be detected according to the copy number corresponding to each chromosome to be detected and the negative sample data set; for each plasma free DNA sample to be detected, dynamically correcting the copy number of each chromosome to be detected according to the negative Z value corresponding to each chromosome to be detected and a given quantity index threshold, and determining the copy number correction value corresponding to each chromosome to be detected; performing chromosome copy number error correction for the current detection batch according to the copy number correction value corresponding to each chromosome to be detected in each plasma free DNA sample to be detected, and determining the copy number batch correction value corresponding to each chromosome to be detected in each plasma free DNA sample to be detected;
[0089] A chromosome detection module is used to obtain a corresponding positive sample data set for each of the plasma free DNA samples to be tested, determine a negative Z value and a positive Z value based on the negative sample data set and the positive sample data set, perform chromosome detection based on the copy number batch correction value, the negative Z value, and the positive Z value corresponding to each of the chromosomes to be tested in each of the plasma free DNA samples to be tested, and determine the chromosome detection result of each of the plasma free DNA samples to be tested, wherein the positive sample data set is simulated and established based on the negative sample data set.
[0090] On the other hand, an embodiment of the present application proposes a detection kit, which is used to achieve the construction of the plasma free DNA sequencing library described above.
[0091] Embodiments of the present application include at least the following beneficial effects: The present application provides a method, apparatus, and kit for non-invasive prenatal fetal chromosome detection, which can construct a DNA sequencing library enriched with fetal DNA concentration, effectively increase the proportion of fetal DNA in the DNA sequencing library, thereby reducing the probability of detection failure and false negatives caused by low fetal DNA concentration and significantly improving the reliability of chromosome detection. The kit is used to construct a DNA sequencing library, improve the stability of the library construction results, achieve sliding window-based elimination of genomic repeat regions and dynamic correction of chromosome copy number, effectively identify abnormal reads caused by repetitive sequences, and reduce the impact of aneuploid abnormal chromosomes on the copy number calculation of other chromosomes, thereby improving the accuracy of chromosome copy number calculation. Based on the negative and positive Z values of the chromosome to be detected, combined with the Z values in the negative reference set and the positive reference set, a joint analysis is performed to achieve chromosome detection, which better meets the expected clinical detection sensitivity and specificity requirements. At the same time, based on the negative sample data set, a positive sample data set is simulated and established, effectively solving the problems of insufficient positive samples for rare chromosomal aneuploidy abnormalities and the difficulty in obtaining positive sample data sets, and has strong application value. In summary, this application can effectively improve the accuracy and comprehensiveness of fetal chromosome testing, cover the detection of aneuploidy abnormalities of 23 pairs of chromosomes, improve the sensitivity and specificity of chromosome testing, and provide more powerful technical support for prenatal screening. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0093] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0094] Figure 1 This is a flow chart of a non-invasive prenatal fetal chromosome detection method provided in an embodiment of the present application;
[0095] Figure 2 1 is a comparative diagram of the chromosome copy number CV values of the conventional Picard deduplication method and the sliding window-based duplication region elimination method in the embodiments of the present application;
[0096] Figure 3 This is a schematic structural diagram of a non-invasive prenatal fetal chromosome detection device provided in an embodiment of the present application;
[0097] Figure 4This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0098] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0099] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0100] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.
[0101] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0102] Reference Figure 1 , Figure 1 This is an optional flow chart of a non-invasive prenatal fetal chromosome detection method provided in an embodiment of the present application. The method may include but is not limited to steps S102 to S112:
[0103] Step S101, constructing a plasma cell-free DNA sequencing library corresponding to the maternal peripheral blood based on a magnetic bead purification and enrichment kit for fetal DNA and a library construction system;
[0104] Step S102, obtaining genome sequencing data from the plasma cell-free DNA sequencing library;
[0105] Step S103, aligning the genome sequencing data to the human reference genome, dividing the human reference genome into multiple unit windows according to a preset unit window width, and calculating the number of sequencing sequences corresponding to the genome sequencing data aligned to each unit window;
[0106] Step S104, based on the number of sequencing sequences corresponding to each unit window, a number of target candidate windows are screened out from the multiple unit windows, where the number of sequencing sequences corresponding to the target candidate windows exceeds a given average level threshold;
[0107] Step S105, determining a number of genomic duplication regions in the genome sequencing data based on each target candidate window, where the genomic duplication region is a genomic region in which the target candidate window appears continuously and the number of occurrences is greater than a given threshold;
[0108] Step S106, obtaining a plurality of plasma free DNA samples to be tested in the current test batch; the plasma free DNA samples to be tested are obtained by enriching fetal free DNA in the plasma of the pregnant woman to be tested by a magnetic bead purification method;
[0109] Step S107, for each chromosome to be detected in each plasma free DNA sample to be detected, regional screening is performed on the chromosome to be detected based on the genome sequencing data and the repeated regions of each genome to determine the number of gene sequences corresponding to the chromosome to be detected;
[0110] Step S108, for each plasma free DNA sample to be tested, performing copy number statistics based on the number of gene sequences corresponding to each chromosome to be tested, and determining the copy number of each chromosome to be tested;
[0111] Step S109, for each plasma cell-free DNA sample to be tested, obtaining a negative sample data set corresponding to each chromosome to be tested, and determining a negative Z value corresponding to each chromosome to be tested based on the copy number corresponding to each chromosome to be tested and the negative sample data set;
[0112] Step S110, for each plasma free DNA sample to be tested, dynamically correcting the copy number of each chromosome to be tested according to the negative Z value corresponding to each chromosome to be tested and a given quantity index threshold, and determining the copy number correction value corresponding to each chromosome to be tested;
[0113] Step S111, performing chromosome copy number error correction for the current testing batch based on the copy number correction value corresponding to each chromosome to be tested in each plasma free DNA sample to be tested, and determining the copy number batch correction value corresponding to each chromosome to be tested in each plasma free DNA sample to be tested;
[0114] Step S112: For each plasma free DNA sample to be tested, a corresponding positive sample data set is obtained, and a negative Z value and a positive Z value are determined based on the negative sample data set and the positive sample data set. Chromosome testing is performed based on the copy number batch correction value, negative Z value, and positive Z value corresponding to each chromosome to be tested in each plasma free DNA sample to be tested, and the chromosome detection result of each plasma free DNA sample to be tested is determined, wherein the positive sample data set is simulated and established based on the negative sample data set.
[0115] In some embodiments, step S101 may include but is not limited to steps S201 to S206:
[0116] Step S201, extracting plasma free DNA from pregnant woman's plasma;
[0117] Step S202, performing fragment screening on plasma free DNA to obtain fragment screened plasma free DNA;
[0118] Step S203, the plasma free DNA after fragment screening is filled with A and phosphorylated to obtain the corresponding end-repair products;
[0119] Step S204, performing adapter ligation on the end-repair product to obtain a corresponding adapter ligation product;
[0120] Step S205, performing magnetic bead purification on the adapter ligation product, and performing library PCR amplification on the adapter ligation product after magnetic bead purification to obtain the corresponding PCR product;
[0121] Step S206 , performing magnetic bead purification on the PCR products, performing library quantification and sequencing on the PCR products after magnetic bead purification, and generating a plasma free DNA sequencing library.
[0122] In steps S201 to S206 of some embodiments, specifically, the construction of the plasma free DNA sequencing library includes the following steps:
[0123] The first step is to take a certain amount (e.g., 400 μL) of maternal plasma, extract plasma DNA using a nucleic acid extraction or purification kit, and then elute the DNA with a corresponding amount (e.g., 50 μL) of elution buffer;
[0124] The second step is sheet screening purification: Sheet screening refers to taking the extracted plasma free DNA and screening the fragments to increase the fetal DNA concentration. The specific steps are as follows:
[0125] a) Add 1.3-1.7 times the volume of purified magnetic beads (using solid-phase reversible immobilized magnetic beads for enrichment, including but not limited to Beckman Agencourt AMPure XP magnetic beads and TransGen Magnetic DNA beads) to the extracted plasma free DNA, mix thoroughly, let stand for 5 minutes, and place on a magnetic rack until clarified;
[0126] b) Pipette the supernatant into 1-2 times the volume of magnetic beads, mix the supernatant and magnetic beads thoroughly, let it stand for 5 minutes, and place it on a magnetic rack until it becomes clear;
[0127] c) aspirating the supernatant, washing twice with 75% to 80% ethanol, and then dissolving and eluting with 16.5 μL of TE solution to obtain the screened plasma free DNA;
[0128] Step 3: Filling-in and phosphorylation of plasma free DNA: Prepare a reaction mixture according to Table 1 from the sieved free DNA. Perform PCR on the reaction mixture using a PCR instrument to obtain end-repair products. The PCR instrument settings include: heated lid at 105°C, 30°C for 30 min, 72°C for 30 min, and 4°C hold. Table 1 shows the following:
[0129] Table 1
[0130] Reagent name Volume (μL) End repair buffer 2 End repair enzyme mixture 1 DNA 15 Total volume 18
[0131] Step 4: Adapter ligation: Use the end-repair products to prepare the reaction mixture shown in Table 2. Use a PCR instrument to perform a PCR reaction on the reaction mixture to obtain adapter ligation products. The PCR instrument setting parameters include: hot cover off, 20°C for 15 minutes, and 4°C hold. Table 2 shows the following:
[0132] Table 2
[0133] Reagent name Volume (μL) DNA after final modification and addition of A 18 connector 3 Rapid Ligation Buffer 9 Fast ligase 3 Total volume 33
[0134] Step 5: Magnetic bead purification: Use 30 μL of purification magnetic beads to purify the adapter ligation product and finally dissolve it back in 19.5 μL TE buffer;
[0135] Step 6: PCR amplification of the library: The adapter-ligated products after magnetic bead purification were used to prepare the reaction mixture shown in Table 3. The mixture was subjected to PCR reaction using a PCR instrument to obtain PCR products. The PCR instrument settings included: heated lid at 105°C, 98°C for 45 seconds, (98°C for 15 seconds, 60°C for 30 seconds, 72°C for 30 seconds)*11, 72°C for 1 minute, and a hold time of 4°C. Table 3 shows the following:
[0136] Table 3
[0137]
[0138]
[0139] Step 7: Purify the PCR product: Use 44 μL of purification magnetic beads to purify the PCR product and then dissolve it in 25 μL of TE buffer.
[0140] Step 8: Library quantification: Use Qubit to measure the concentration of the PCR product after magnetic bead purification to determine the corresponding library concentration;
[0141] Step 9: Library sequencing: Use an Illumina platform-related sequencer to sequence the PCR products after magnetic bead purification, with a data volume of 5M for each library.
[0142] In some embodiments, the above-mentioned method for constructing a plasma free DNA sequencing library first uses a magnetic bead purification method to enrich short-fragment free DNA, thereby increasing the concentration of fetal free DNA in the sample; at the same time, by optimizing the operating parameters of experimental steps such as end repair, linker connection, and PCR, the complexity of library construction operations and experimental time are reduced, thereby improving the success rate of library construction.
[0143] For example, six maternal plasma samples (L1 to L6) were used for DNA sequencing library construction and sequencing. Two DNA samples were extracted from each sample, one of which was processed using the method for constructing a maternal plasma cell-free DNA library with increased fetal DNA concentration proposed in the present invention; the other DNA sample was used as a comparison method, in which the fragment screening step was omitted during library construction, as a comparison of the library construction effect without fetal concentration enrichment. The specific steps for DNA sequencing library construction are as follows:
[0144] 1) Plasma DNA extraction: 400 μL of maternal plasma was collected and plasma DNA was extracted using the above kit. Two DNA aliquots were extracted from each sample, and DNA samples with the same ID were mixed together for future use.
[0145] 2) The extracted DNA was divided into two equal parts. The first DNA sample served as a control group and was not subjected to fragment screening. The second DNA sample was subjected to fragment screening and purification to increase the fetal DNA concentration. The specific steps are as follows:
[0146] a) Add 1.3-1.7 times the volume of purified magnetic beads (using solid-phase reversible immobilized magnetic beads for enrichment, including but not limited to Beckman Agencourt AMPure XP magnetic beads and TransGen Magnetic DNA beads) to the extracted plasma free DNA, mix thoroughly, let stand for 5 minutes, and place on a magnetic rack until clarified;
[0147] b) Pipette the supernatant into 1-2 times the volume of magnetic beads, mix the supernatant and magnetic beads thoroughly, let it stand for 5 minutes, and place it on a magnetic rack until it becomes clear;
[0148] c) aspirating the supernatant, washing twice with 75% to 80% ethanol, and then dissolving and eluting with 16.5 μL of TE solution to obtain the screened plasma free DNA;
[0149] 3) Filling and phosphorylation of plasma free DNA: Prepare the reaction mixture according to Table 1 using the free DNA after sieving. Table 4 is as follows:
[0150] Table 4
[0151] Reagent name Volume (μL) End repair buffer 2 End repair enzyme mixture 1 DNA 15 Total volume 18
[0152] Reaction on PCR instrument: heated cover 105°C; 30°C for 30 min, 72°C for 30 min, 4°C hold.
[0153] 4) Connector connection: Prepare the reaction mixture shown in Table 5, which is as follows:
[0154] Table 5
[0155] Reagent name Volume (μL) DNA after final modification and addition of A 18 connector 3 Rapid Ligation Buffer 9 Fast ligase 3 Total volume 33
[0156] Reaction on PCR instrument: hot cover off; 20℃ for 15 min, 4℃ hold;
[0157] 5) Magnetic bead purification: Purify the adapter ligation product using 30 μL of purification magnetic beads and finally dissolve it back in 19.5 μL of TE buffer;
[0158] 6) Library PCR amplification: Prepare the reaction mixture shown in Table 6, which is as follows:
[0159] Table 6
[0160] Reagent name Volume (μL) Adapter-ligated DNA 18.5 PCR amplification primers 1.5 PCR amplification enzyme mix 20 Total reaction volume 40
[0161] The reaction was carried out on a PCR instrument: heated lid at 105°C; 98°C for 45 seconds, (98°C for 15 seconds, 60°C for 30 seconds, 72°C for 30 seconds)*11; 72°C for 1 minute; 4°C hold;
[0162] 7) PCR product purification: Purify using 44 μL of purification magnetic beads and then dissolve in 25 μL of TE buffer;
[0163] 8) Library detection: The library concentration was determined using Qubit. The library fragment size was detected using Agilent 2100.
[0164] 9) Library sequencing: Sequencing was performed using an Illumina platform-related sequencer to obtain sequencing reads of maternal plasma cell-free DNA.
[0165] 10) Determination of the effect of sheet-screening purification on fetal DNA concentration: To compare the effect of sheet-screening on fetal DNA concentration, we first determined whether short fragments of free DNA <158 bp were successfully enriched based on sequencing read length. Second, we determined the multiple increase in fetal DNA concentration based on the Z-score of the Y chromosome. The results in Table 7 show that the sheet-screening purification method reduced the concentration of the DNA sequencing library to 1 / 3 compared to the traditional method without sheet-screening, shortened the DNA length by approximately 30 bp, and increased the fetal DNA concentration by 2.12 times that of the method without sheet-screening. This indicates that the sheet-screening purification method can effectively enrich short fragments of free fetal DNA and increase the fetal concentration. The specific results in Table 7 are shown below:
[0166] Table 7 Comparison of results between the sieve purification method and the non-sieve purification method
[0167]
[0168] In some embodiments, the genomic sequencing data (reads) corresponding to the plasma cell-free DNA sequencing library are aligned to the human reference genome using alignment software (including but not limited to BWA, Bowtie, etc.), and duplicate sequences and low-quality alignment results are removed. For example, duplicate sequences can be aligned using, but not limited to, the Picard MarkDuplicates tool, in combination with the alignment software's alignment flag and alignment quality value setting threshold combination to remove low-quality alignment results.
[0169] In steps S102 to S105 of some embodiments, a sliding window-based repetitive region removal method (SW-RR-Removal) is proposed to effectively identify and remove repetitive regions in the genome, especially those where the repetition rate is normal at a single site but a large number of repetitive sequences appear in continuous segments, causing abnormal readings.
[0170] The sliding window-based repeated region elimination method (SW-RR-Removal) is as follows:
[0171] First, the human reference genome is divided into multiple unit windows with a preset unit window width (windows can be 10bp, 20bp, etc.). The number of sequencing sequences aligned to each window (represented by the read number, hereinafter referred to as reads) is calculated. From the multiple unit windows, several target candidate windows are screened for those whose window reads are significantly higher than the average level threshold (the significance level is selected but not limited to 1‰);
[0172] Then, the clustering of each target candidate window in the genome is analyzed, and for each target candidate window, the genomic region where the target candidate window appears continuously in the genome sequencing data and is greater than a given threshold is set as the genomic duplication region.
[0173] Optionally, refer to Figure 2 , Figure 2 This is an optional comparative schematic diagram of the CV (Coefficient of Variation, CV) values of chromosome copy numbers of the conventional Picard deduplication method and the sliding window-based genomic duplication region elimination method in the embodiment of the present application. In order to verify the effectiveness of the sliding window genomic duplication region elimination method (SW-RR-Removal), the sequencing results of 5 negative pregnant women's plasma samples were selected, and the differences between the conventional Picard deduplication and the chromosome sliding window duplication region elimination method (SW-RR-Removal) were compared. The CV value of the chromosome copy number, the smaller the CV value, the better the noise reduction effect. According to the sequencing results of the 5 negative pregnant women's plasma samples, it is fully demonstrated that the sliding window-based duplication region elimination method (SW-RR-Removal) can more effectively remove duplicate regions than the conventional Picard deduplication method, reduce data noise, and improve the accuracy of chromosome copy number.
[0174] In some embodiments, all plasma free DNA samples to be tested in the current test batch are from the plasma of the same pregnant woman. Optionally, a magnetic bead purification method is used to enrich short fragments of free DNA in multiple tubes of plasma from the same pregnant woman to obtain corresponding multiple plasma free DNA samples to be tested.
[0175] In some embodiments, step S107 may include but is not limited to steps S301 to S305:
[0176] Step S301, based on the genome sequencing data, determining multiple gene sequencing regions in the chromosome to be detected and the number of regions and region length corresponding to each gene sequencing region, where the region length is the sequence length corresponding to the gene sequencing region, and the number of regions is the frequency of occurrence of the gene sequencing region in the chromosome to be detected;
[0177] Step S302, dividing the gene sequencing regions into multiple gene sequencing groups according to the region lengths corresponding to the gene sequencing regions, wherein each gene sequencing group includes multiple gene sequencing regions with the same region length;
[0178] Step S303, based on each genomic duplication region, screening the chromosome duplication region corresponding to each genomic duplication region from multiple gene sequencing regions, and determining the target gene sequencing group where the chromosome duplication region is located from the multiple gene sequencing groups;
[0179] Step S304: for each chromosome duplication region, the region number corresponding to the chromosome duplication region is updated to the average region number corresponding to the target gene sequencing group, where the average region number is the average of the region numbers of multiple gene sequencing regions in the gene sequencing group;
[0180] Step S305 , determining the number of gene sequences corresponding to the chromosome to be detected based on the number of regions corresponding to each gene sequencing region except for the multiple chromosome duplication regions and the updated number of regions in each chromosome duplication region.
[0181] In some embodiments, multiple gene sequencing groups are divided according to the region length of each gene sequencing region in the chromosome to be detected. For example, all gene sequencing regions with a region length of 20kb in the chromosome to be detected are divided into the same gene sequencing group, all gene sequencing regions with a region length of 30kb in the chromosome to be detected are divided into another gene testing group, and so on.
[0182] In some embodiments, when calculating the number of gene sequences of the chromosome to be detected, if there are genomic duplication regions, the region number of the genomic duplication region on the chromosome is replaced by an expected value (such as the above-mentioned average region number). For example, a genomic duplication region with a region length of 20 kb is detected in the chromosome to be detected, and the region number of the genomic duplication region is 2000. However, the average region number on the chromosome to be detected corresponding to all gene sequencing regions with a region length of 20 kb in the chromosome to be detected is 100. Then, 100 is used instead of 2000 to update the region number of the genomic duplication region on the chromosome to be detected, effectively reducing the deviation introduced into the statistics of the number of gene sequences of the chromosome by the sequencing results of a large number of duplications in the genomic duplication region. Finally, the number of gene sequences corresponding to the chromosome to be detected is re-determined based on the region number corresponding to each gene sequencing region excluding multiple chromosomal duplication regions and the updated region number of each chromosomal duplication region.
[0183] In some embodiments, step S108 may include but is not limited to steps S401 to S402:
[0184] Step S401, determining the total number of gene sequences corresponding to the plasma cell-free DNA sample to be tested based on the number of gene sequences corresponding to each chromosome to be tested;
[0185] Step S402: Determine the copy number of each chromosome to be detected based on the number of gene sequences corresponding to each chromosome to be detected and the total number of gene sequences. The copy number is calculated using the following formula:
[0186] RR ij =RN ij / AN i ;
[0187]
[0188] Among them, RR ij is the copy number of the jth chromosome to be detected in the i-th plasma free DNA sample to be detected, RN ij is the number of gene sequences corresponding to the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, AN i is the total number of gene sequences corresponding to the i-th plasma free DNA sample to be tested, and N is the number of chromosomes to be tested in the i-th plasma free DNA sample to be tested.
[0189] In some embodiments, the number of gene sequences RN of chromosome j to be detected in plasma cell-free DNA sample i to be detected is counted. ij First, calculate the total number of gene sequences AN of all chromosomes in the plasma free DNA sample i to be tested i Finally, the copy number RR of the chromosome j to be tested in the plasma free DNA sample i to be tested is counted ij .
[0190] In some embodiments, step S109 may include but is not limited to steps S501 to S502:
[0191] Step S501, determining the corresponding negative sample copy number mean and negative sample copy number standard deviation according to the negative sample data set corresponding to each chromosome to be detected;
[0192] Step S502: Determine the negative Z value corresponding to the chromosome to be detected based on the copy number corresponding to the chromosome to be detected, the mean copy number of negative samples, and the standard deviation of the copy number of negative samples. The negative Z value is calculated by the following formula:
[0193] Z_neg ij =(RR ij -μneg j ) / σneg j ;
[0194] Among them, Z_neg ij is the negative Z value corresponding to the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, RR ij is the copy number of the jth chromosome to be detected in the i-th plasma free DNA sample to be detected, μneg j is the mean number of negative sample copies corresponding to the jth chromosome to be tested in the negative sample data set, σneg j is the standard deviation of the negative sample copy number corresponding to the jth chromosome to be detected in the negative sample data set.
[0195] In some embodiments, step S110 may include but is not limited to steps S601 to S607:
[0196] Step S601, determining a plurality of first abnormal chromosomes from a plurality of chromosomes to be detected based on the negative Z value and the quantitative index threshold corresponding to each chromosome to be detected, wherein the negative Z value corresponding to the first abnormal chromosome exceeds the quantitative index threshold;
[0197] Step S602, obtaining any first abnormal chromosome as the current corrected chromosome;
[0198] Step S603, obtaining the total number of gene sequences corresponding to the plasma cell-free DNA sample to be tested and the mean number of negative sample copies corresponding to the current corrected chromosome;
[0199] Step S604: Update the number of gene sequences corresponding to the current corrected chromosome based on the total number of gene sequences corresponding to the plasma cell-free DNA sample to be tested and the mean number of negative sample copies corresponding to the current corrected chromosome. The updated number of gene sequences is calculated using the following formula:
[0200]
[0201] in, is the updated number of gene sequences corresponding to the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, AN i is the total number of gene sequences corresponding to the i-th plasma free DNA sample to be tested, μneg j is the mean copy number of negative samples corresponding to the jth chromosome to be tested;
[0202] Step S605, updating the total number of gene sequences based on the updated number of gene sequences corresponding to the currently corrected chromosome and the number of gene sequences corresponding to the remaining chromosomes to be tested;
[0203] Step S606, recalculating the chromosome copy number based on the total number of updated gene sequences, the number of gene sequences of the remaining chromosomes to be tested, and the updated number of gene sequences corresponding to the currently corrected chromosome, to determine a corrected copy number value corresponding to the currently corrected chromosome and a corrected copy number value corresponding to the remaining chromosomes to be tested;
[0204] Step S607: When there is any first abnormal chromosome whose gene sequence number has not been updated, obtain any first abnormal chromosome whose gene sequence number has not been updated as the current corrected chromosome, and then return to obtain the total number of gene sequences corresponding to the plasma free DNA sample to be tested and the mean copy number of negative samples corresponding to the current corrected chromosome until the gene sequence numbers of all first abnormal chromosomes have been updated.
[0205] In some embodiments, step S111 may include but is not limited to steps S701 to S708:
[0206] Step S701, obtaining any plasma free DNA sample to be detected as a target detection sample, obtaining any chromosome to be detected in the target detection sample as a current target detection chromosome, and obtaining a copy number correction value corresponding to the current target detection chromosome as the current copy number;
[0207] Step S702, obtaining the mean copy number of negative samples corresponding to the current target detection chromosome, the standard deviation of the negative sample copy number, and the mean copy number of chromosomes in the current batch, wherein the mean copy number of chromosomes in the current batch is the mean copy number corresponding to the current target detection chromosome in all plasma free DNA samples to be detected;
[0208] Step S703, determining the normalized copy number corresponding to the current target detection chromosome based on the current copy number, the mean copy number of negative samples, and the mean copy number of chromosomes in the current batch;
[0209] Step S704, updating the negative Z value corresponding to the current target detection chromosome according to the normalized copy number, the mean copy number of negative samples, and the standard deviation of the copy number of negative samples;
[0210] Step S705, determining whether the current target detection chromosome is the second abnormal chromosome based on the updated negative Z value and quantity index threshold corresponding to the current target detection chromosome;
[0211] Step S706: When the updated negative Z value corresponding to the current target detection chromosome is greater than the quantity index threshold, the current target detection chromosome is determined to be the second abnormal chromosome, the mean copy number of the negative samples is determined as the current copy number, and the mean copy number of the chromosomes in the current batch is re-determined based on the current copy number. Then, the normalized copy number corresponding to the current target detection chromosome is determined based on the current copy number, the mean copy number of the negative samples, and the mean copy number of the chromosomes in the current batch, until the updated negative Z value corresponding to the current target detection chromosome is less than the quantity index threshold;
[0212] Step S707 , when the updated negative Z value corresponding to the current target detection chromosome is less than the quantity index threshold, the normalized copy number is determined as the copy number batch correction value corresponding to the current target detection chromosome;
[0213] Step S708: When there is any chromosome to be detected in the target detection sample that has not undergone chromosome copy number error correction, any other chromosome to be detected in the target detection sample that has not undergone chromosome copy number error correction is obtained as the current target detection chromosome, the copy number correction value corresponding to the current target detection chromosome is obtained as the current copy number, and then the process returns to obtain the negative sample copy number mean, negative sample copy number standard deviation, and current batch chromosome copy number mean corresponding to the current target detection chromosome, until all chromosomes to be detected in the target detection sample have undergone chromosome copy number error correction.
[0214] In some embodiments, the normalized copy number is calculated by the following formula:
[0215] Among them, the current detection chromosome is the jth chromosome to be detected in the i-th plasma free DNA sample to be detected, RR′ ij The normalized copy number corresponding to the current target chromosome detection, Correction value for the copy number of the current target chromosome, μb j is the mean number of chromosome copies in the current batch corresponding to the current target detection chromosome in the current detection batch, μneg j The mean copy number of negative samples corresponding to the current target chromosome.
[0216] In some embodiments, the chromosome copy number is the ratio of the target chromosome reading to the total chromosome reading. If there is an aneuploid abnormality in the chromosome, the reading of the chromosome will be significantly higher than the expected value of the negative sample, resulting in a significant deviation in the total chromosome reading value, affecting the copy number calculation of other chromosomes. To reduce this effect, a dynamic correction method for chromosome copy number is proposed. In each round of iteration, the copy number of the target chromosome and the negative Z value are calculated to identify chromosomes that may have aneuploid abnormalities. The interference of these abnormal chromosomes on the overall copy number calculation is dynamically corrected to improve the accuracy of chromosome copy number quantification. The dynamic correction method for chromosome copy number includes copy number correction within the current sample and copy number correction within the current batch. The specific method for dynamic correction of chromosome copy number is as follows:
[0217] The first step is to correct the copy number in the current sample. The specific steps are as follows:
[0218] 1) Obtain the copy number RR of the chromosome j to be tested in the plasma free DNA sample i to be tested ij ;
[0219] 2) Using the negative sample data set corresponding to the chromosome j to be tested, calculate the mean negative sample copy number μneg of the chromosome j to be tested j and the standard deviation of negative sample copy number σneg j, calculate the negative Z value Z_neg of the chromosome j to be tested ij , where Z_neg ij =(RR ij -μneg j ) / σneg j ;
[0220] 3) Set the number index threshold of the chromosome j to be detected to determine whether the chromosome j to be detected is a chromosome with chromosomal aneuploidy, such as Z_neg of the chromosome j to be detected ij If it is 6, which exceeds the set quantity index threshold, the quantity index threshold is marked as the first abnormal chromosome;
[0221] 4) Assuming that the chromosome j to be tested is the first abnormal chromosome, the chromosome j to be tested in the plasma free DNA sample i to be tested is obtained as the current corrected chromosome, and the expected number of gene sequences of the first abnormal chromosome is calculated: Replace the number of gene sequences of chromosome j to be detected with the expected number of gene sequences Complete the update of the number of gene sequences, recalculate the number of gene sequences of the remaining chromosomes to be tested in the plasma free DNA sample i to be tested, and update the total number of gene sequences of the plasma free DNA sample i to be tested. The updated total number of gene sequences is AN′ i ;
[0222] 5) According to the updated total number of gene sequences is AN' i , recalculate the copy number correction values of the remaining chromosomes to be detected and the copy number correction value of the chromosome j to be detected: Then, the next first abnormal chromosome is obtained as the current corrected chromosome, and the process returns to step 4). Finally, the process iterates steps 4) to 5) for k1 times, and the iteration is stopped. The corrected copy number values of each chromosome to be tested in the plasma free DNA sample i to be tested after multiple rounds of iterative correction are output. Wherein, k1 is the number of the first abnormal chromosome in the plasma cell-free DNA sample i to be tested;
[0223] The second step is to perform copy number correction within the current batch. The specific steps are as follows:
[0224] 1) Obtain the plasma free DNA sample i to be tested as the target detection sample, obtain the chromosome j to be detected in the plasma free DNA sample i to be tested as the current target detection chromosome, and obtain the copy number correction value of the chromosome j to be detected in the plasma free DNA sample i to be tested after the sample copy number correction is the current copy number;
[0225] 2) Calculate the mean copy number μb of the chromosome j to be tested in all plasma free DNA samples to be tested in the current test batch j ;
[0226] 3) Obtain the mean copy number μneg of the negative samples corresponding to the chromosome j to be tested j and the standard deviation of negative sample copy number σneg j , perform batch error correction to obtain the normalized copy number of the chromosome j to be tested in the plasma free DNA sample i to be tested in,
[0227] 4) Update the negative Z value of the chromosome j to be detected according to the normalized copy number, the mean copy number of the negative sample and the standard deviation of the copy number of the negative sample. The updated negative Z value is
[0228] 5) determining whether the chromosome j to be detected is a second abnormal chromosome that may have a chromosomal aneuploidy abnormality based on a number index threshold of the chromosome j to be detected;
[0229] 6) Assuming that the chromosome j to be detected is the second abnormal chromosome, the copy number batch correction value of the second abnormal chromosome is determined as the negative sample copy number mean μneg j , recalculate the mean copy number μb′ of the chromosome j to be tested in all plasma free DNA samples to be tested in the current test batch j , looping steps 3) to 6) until the negative Z value of the chromosome j to be tested is less than the quantitative index threshold;
[0230] 7) When the negative Z value of the chromosome j to be tested is less than the quantitative index threshold, the normalized copy number Determine the copy number batch correction value of the chromosome j to be detected;
[0231] 8) Determine whether any chromosome to be detected that has not undergone chromosome copy number error correction exists in the plasma free DNA sample i to be detected. If any chromosome to be detected that has not undergone chromosome copy number error correction exists, obtain any remaining chromosome to be detected in the target detection sample that has not undergone chromosome copy number error correction as the current target detection chromosome, obtain the copy number correction value corresponding to the current target detection chromosome as the current copy number, and repeat the above steps 3) to 8) until all the chromosomes to be detected in the plasma free DNA sample i to be detected have undergone chromosome copy number error correction;
[0232] 9) When all the chromosomes to be detected in the plasma free DNA sample i to be detected are subjected to chromosome copy number error correction, the copy number batch correction value of each chromosome to be detected in the plasma free DNA sample i to be detected is output.
[0233] In some embodiments, illustratively, the sequencing results of 18 aneuploidy-positive pregnant women's plasma samples (including 2 T13 (numbered S1-S2), 5 T18 (numbered S3-S7), and 11 T21 (numbered S8-S18)) were selected to compare the performance difference between the conventional method without dynamic copy number correction and the above-mentioned chromosome copy number dynamic correction method, wherein Table 8 is a result table of chromosomes 13, 18, and 21 Z values of each aneuploidy-positive pregnant women's plasma sample, and Table 8 is specifically shown as follows:
[0234] Table 8 Chromosome Z-score results of plasma samples from aneuploidy-positive pregnant women
[0235]
[0236]
[0237] The results in Table 8 show that the use of the dynamic correction method for chromosome copy number can effectively improve the detection value of positive fetal samples compared with the conventional method without dynamic correction of copy number, with reference to the Z value of chromosome 13 of the T13 sample, the Z value of chromosome 18 of the T18 sample, and the Z value of chromosome 21 of the T21 sample, without causing abnormal values of other chromosomes.
[0238] In some embodiments, step S112 may include but is not limited to steps S801 to S810:
[0239] Step S801, obtaining a negative sample data set corresponding to the chromosome to be detected;
[0240] Step S802, determining the expected copy number value of the positive sample data set based on the fetal concentration distribution of the negative sample data set, and generating the positive sample data set based on the expected copy number value and the random error;
[0241] Step S803, performing copy number distribution statistics on the negative sample data set to determine the negative Z value judgment threshold;
[0242] Step S804, updating the negative Z value corresponding to the chromosome to be detected according to the copy number batch correction value corresponding to the chromosome to be detected, the negative sample copy number mean and the negative sample copy number standard deviation;
[0243] Step S805: When the chromosome to be detected is an autosome, a positive sample data set corresponding to the chromosome to be detected is obtained, and copy number distribution statistics of the positive sample data set are performed to determine the positive Z value judgment threshold, the positive sample copy number mean, and the positive sample copy number standard deviation;
[0244] Step S806: Determine the positive Z value corresponding to the chromosome to be detected based on the batch correction value of the copy number corresponding to the chromosome to be detected, the mean copy number of the positive samples, and the standard deviation of the copy number of the positive samples. The positive Z value is calculated by the following formula:
[0245]
[0246] Among them, Z_pos ij is the positive Z value corresponding to the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, μpos j is the mean sample copy number corresponding to the jth chromosome to be tested in the positive sample data set, is the batch correction value of the copy number of the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, σpos j is the standard deviation of the sample copy number corresponding to the jth chromosome to be tested in the positive sample data set;
[0247] Step S807, performing a joint test based on the negative Z value judgment threshold, the negative Z value, the positive Z value judgment threshold, and the positive Z value to determine the chromosome detection result;
[0248] Step S808, when the negative Z value is greater than the negative Z value judgment threshold, and the positive Z value is greater than the positive Z value judgment threshold, the chromosome detection result is determined to be positive for the chromosome to be detected in the plasma free DNA sample to be detected;
[0249] Step S809, when the negative Z value is less than or equal to the negative Z value judgment threshold, and the positive Z value is less than or equal to the positive Z value judgment threshold, determining that the chromosome detection result is negative for the chromosome to be detected in the plasma free DNA sample to be detected;
[0250] Step S810: When the chromosome to be detected is a sex chromosome, a sex chromosome detection decision tree model is obtained, and the sex chromosome detection decision tree model is used to determine the chromosome detection result according to the negative Z value judgment threshold and the negative Z value.
[0251] In some embodiments, a negative sample dataset corresponding to an autosome is established: the negative sample dataset is composed of sequencing data of plasma free DNA of more than 10,000 negative pregnant women. The negative sample dataset can be used to obtain the copy number distribution of negative samples on 22 autosomes, and the mean copy number (μneg) of the chromosome j to be detected in the negative sample is calculated. j) and standard deviation (σneg j ).
[0252] When the chromosome to be detected is an autosome, a positive sample data set corresponding to the autosome is established: the data source of the positive sample data set includes the collected real positive plasma samples diagnosed with fetal aneuploidy abnormalities; the positive sample data set can be used to obtain the copy number distribution of the positive samples in the chromosome, and the mean value (μpos j ) and standard deviation (σpos j ).
[0253] Optionally, for aneuploidy types with fewer than 100 positive plasma samples, a theoretical distribution of an aneuploidy reference set was established using the fetal DNA concentration distribution. Taking the construction of the negative and positive sample datasets of T21 as an example, the specific steps are as follows:
[0254] 1) Establish T21 negative sample data set: The negative sample data set consists of sequencing data of plasma free DNA of more than 10,000 negative pregnant women. The negative sample data set can be used to obtain the copy number distribution of negative samples in 22 autosomes, and calculate the mean copy number of chromosome j in negative samples (μneg j ) and standard deviation (σneg j )
[0255] 2) Establish a T21 positive sample dataset, including:
[0256] a) Collect fetal concentration values of 10,000 negative samples;
[0257] b) Calculate the expected copy number of 10,000 positive samples: Assume that the fetal concentration of plasma free DNA in 10,000 negative samples is sample fc i , when the chromosome j to be tested in the sample is trisomic, the expected copy number is
[0258] c) The expected copy number value is added with random error to generate the copy number distribution value of the 10,000 positive sample data set. Specifically, the mean value of the negative sample (μneg j ) and standard deviation (σneg j ), using the R language rnorm function to generate a mean of 0 and a standard deviation of σneg j Ten thousand random errors ε ij , and the random error ε ij and copy number expectations Add together to obtain the copy number distribution value T_RR of the simulated positive sample data set ij, the calculation formula is
[0259] The positive sample data set is used to obtain the copy number distribution of the positive sample on the chromosome j to be tested, and the mean value of chromosome j in the positive sample data set (μpos j ) and standard deviation (σpos j ).
[0260] In some embodiments, optionally, according to the sensitivity of the kit, the percentile value (1-sensitivity) of the positive sample data set corresponding to the Zneg value is determined, and it is used as the judgment threshold of Zneg according to the Cutoff Zneg_sensitivity (Such as the negative Z value judgment threshold mentioned above), for example, if the sensitivity of the kit is 99%, the Zneg corresponding to the 1% percentile of the positive sample is set as Cutoff Zneg_sensitivity , positive samples meet Zneg>Cutoff Zneg_sensitivity According to the specificity of the kit, find the percentile value (specificity) of the negative reference set corresponding to Zneg as the judgment threshold of Zneg based on Cutoff Zneg_specificity Assuming that the specificity of the kit is not less than 99.95%, the Zneg corresponding to the 99.95% quantile of the negative sample is set as Cutoff Zneg_specificity , positive samples meet Zneg>Cutoff Zneg_specificity ; The above two conditions Zneg>Cutoff Zneg_specificity and Zneg>Cutoff Zneg_sensitivity All must be met, set the negative Z value judgment threshold Zneg=max(Cutoff Zneg_sensitivity , Cutoff Zneg_specificity ).
[0261] In some embodiments, according to the sensitivity of the kit, the percentile value (1-sensitivity) corresponding to the positive sample data set is found and used as the judgment threshold of Zpos according to the Cutoff value. Zpos_sensitivity (Such as the positive Z value judgment threshold mentioned above). For example, if the sensitivity of the kit is 99%, the Zneg corresponding to the 1% percentile of the positive sample is set as Cutoff Zpos_sensitivity , positive samples meet Zpos>Cutoff Zpos_sensitivity According to the specificity of the test kit, find the percentile value (specificity) of the negative reference set corresponding to Zpos, which is used as the judgment threshold of Zpos according to Cutoff Zpos_specificity Assuming that the specificity of the kit is not less than 99.95%, the Zneg corresponding to the 99.95% quantile of the negative sample is set as Cutoff Zneg_specificity, positive samples meet Zneg>Cutoff Zneg_specificity ; Above Zpos>Cutoff Zpos_sensitivity and Zneg>Cutoff Zneg_specificity Both conditions must be met, set the positive Z value threshold Zpos=max(Cutoff Zpossensitivity , Cutoff Zposspecificity ).
[0262] In some embodiments, assuming that the negative Z value of the chromosome to be detected is Zneg, the positive Z value is Zpos, the negative Z value judgment threshold is Cutoff_Zneg, and the positive Z value judgment threshold is Cutoff_Zpos, the joint test rules are shown in Table 9 below:
[0263] Table 9 Joint inspection rules
[0264]
[0265] 1) If Zneg>Cutof_Zneg and Zpos>Cutof_Zpos, the chromosome to be tested is judged to be positive; if Zneg≤Cutoff_Zneg and Zpos≤Cutoff_Zpos, the chromosome to be tested is judged to be negative; if any of the above conditions are not met, the result is unclear and the test is judged to be retested. Cutoff_Zneg and Cutoff_Zpos are determined based on the percentile index values of Zneg and Zpos of T21 negative and positive samples and the sensitivity of the kit. The percentile index values of Zneg and Zpos of T21 negative and positive samples are specifically shown in Table 10. Table 10 is as follows:
[0266] Table 10 Percentile index values of Zneg and Zpos for negative and positive samples of T21
[0267]
[0268]
[0269] 2) According to the results in Table 10, select the judgment threshold value of Zneg, Cutoff_Zneg: Assuming that the sensitivity of the test kit is not less than 99%, that is, the percentile of the corresponding positive sample is 0.10%, and Zneg>4.79; assuming that the specificity of the test kit is not less than 99.95%, that is, the percentile of the corresponding negative sample is 99.95%, and Zneg>3.68, both of the above conditions must be met. Therefore, determine Cutoff_Zneg=4.79;
[0270] 3) Based on the results in Table 10, select the judgment threshold value of Zpos, Cutoff_Zpos: Assuming that the sensitivity of the test kit is not less than 99%, that is, the percentile of the corresponding positive sample is 0.10%, Zpos>-2.48; assuming that the specificity of the test kit is not less than 99.95%, that is, the percentile of the corresponding negative sample is 99.95%, Zpos>-2.63. Both of the above conditions must be met. Therefore, determine Cutoff_Zpos=-2.63;
[0271] 4) Determine the Z value joint inspection rule. Assuming that the determination threshold of Zneg is Cutoff_Zneg and the determination threshold of Zpos is Cutoff_Zpos, the Z value joint inspection rule is as shown in Table 9 above.
[0272] 5) Calculate the two-dimensional Z values (Zneg and Zpos) of chromosome 21 of the plasma cell-free DNA sample to be tested: Calculate the negative Z value (Zneg): Calculate the positive Z value (Zpos):
[0273] 6) Combined with Table 9, a two-dimensional Z value (Zneg and Zpos) joint test was performed to determine whether the sample under test had T21 abnormalities: if Zneg>4.79 and Zpos>-2.63, it was judged as positive; if Zneg<4.79 and Zpos<-2.63, it was judged as negative; samples that did not meet any of the above conditions had unclear results and were judged to be retested.
[0274] In some embodiments, a test data set was established, including 1012 gold standard T21 positive samples and 98,988 follow-up negative samples, totaling 100,000 samples for testing. In order to verify the effectiveness of the above-mentioned Z-score combined test method, each sample in the test data set was tested using the conventional Z-score test method, the conventional Z-score test + placental mosaicism method, and the above-mentioned Z-score combined test method. The specific results are as follows:
[0275] The above Z-value joint test method was used to test the 1012 positive samples in the test data set. The results are shown in Table 11. The details of Table 11 are as follows:
[0276] Table 11 Positive sample test results
[0277]
[0278] According to Table 11, there were 1012 T21-positive samples in the test data set, 14 of which were judged as uncertain, with a retest rate of 1.38%; among the 998 samples that were successfully tested, 988 were judged as positive and 10 were judged as negative, and the detection sensitivity of samples with successful quality control was 98.99%.
[0279] The above Z-value joint test method was used to test the 98,988 negative samples in the test data set. The results are shown in Table 12. The details of Table 12 are as follows:
[0280]
[0281] According to Table 12, among the 98,988 T21-negative samples in the test data set, 83 were judged as uncertain, with a retest rate of 0.08‰; among the 98,895 samples that were successfully tested, 96 were judged as positive and 98,809 were judged as negative. The detection specificity of samples with successful quality control was 99.91%, which was close to the sensitivity level intended by the kit.
[0282] In summary, the above-mentioned Z-value joint test method was used to test 100,000 samples in the test data set, and the final retest rate was 0.1% (97 / 100,000), the sensitivity was 98.99% (988 / 998), the specificity was 99.91% (98809 / 98895), and the positive predictive value was 91.14% (988 / 1084).
[0283] In comparison, conventional Z-value test methods such as NIPT use the commonly used Zneg=3 in NIPT as the single-dimensional detection threshold, and use NIPT to test each sample in the test data set. The test results are as follows: 1007 of 1012 T21-positive samples were judged as positive, and 5 were judged as negative, with a sensitivity of 99.5% (1007 / 1012); of 98988 T21-negative samples, 98538 were judged as negative, and 450 were judged as positive, with a specificity of 99.5% (98538 / 98988); the positive predictive value was 69.11% (1007 / 1457).
[0284] The conventional Z-value test + placental mosaicism method uses Zneg combined with fetal DNA concentration to identify the accuracy of false positive samples of placental mosaicism. The copy number expectation of positive samples can be calculated based on the fetal DNA concentration. Assume that the fetal concentration of sample i is sample fc i , when the chromosome j of the sample is trisomic, the expected copy number is Then the placental mosaic ratio is Mosaic ij The following calculations can be performed based on the measured copy number and the expected value: The judgment rule for positive samples is: Zneg>3 and Mosaic ij <Set the mosaic ratio threshold; other samples that do not meet the positive sample rules are judged as negative. For comparison, the conventional Z-score test + placental mosaicism method is used to test the test data set. The specific test results are as follows:
[0285] 1) Test 1: Mosaicij 152 samples with a sensitivity below 80% were labeled as placental mosaicism. The results of Zneg combined with placental mosaicism information for T21 positivity were as follows: 969 of 1012 T21-positive samples were judged as positive and 43 as negative, with a sensitivity of 95.75% (969 / 1012); of 98,988 T21-negative samples, 98,954 were judged as negative and 34 were judged as positive, with a specificity of 99.96% (98,954 / 98,988); the positive predictive value was 96.61% (969 / 1058);
[0286] 2) Test 2: Mosaic ij 136 samples with a sensitivity below 60% were labeled as placental mosaicism. The results of Zneg combined with placental mosaicism information for T21 positivity were as follows: 981 of 1012 T21-positive samples were judged as positive and 31 as negative, with a sensitivity of 96.93% (981 / 1012); 98,988 T21-negative samples, 98,912 as negative and 76 as positive, with a specificity of 99.92% (98,538 / 98,988); and a positive predictive value of 92.80% (871 / 1321).
[0287] 3) Test 3: Mosaic ij 89 samples with a sensitivity below 30% were labeled as placental mosaicism. Zneg combined with placental mosaicism information for T21 positivity was as follows: 988 of 1012 T21-positive samples were judged positive and 25 were negative, with a sensitivity of 97.62% (981 / 1012). Of 98,988 T21-negative samples, 98,899 were judged negative and 89 were judged positive, with a specificity of 99.91% (98,538 / 98,988). The positive predictive value was 91.73% (871 / 1368).
[0288] Finally, the detection performance of the three methods is summarized and compared, as shown in Table 13. The details of Table 13 are as follows:
[0289] Table 13 Comparison of detection performance of three methods
[0290]
[0291]
[0292] According to Table 13, the combined Z-score test method has the best overall performance. Under the premise of sensitivity of 98.99%, the specificity reaches 99.91%, and the comprehensive PPV is 91.14%. The sensitivity of the conventional Z-score test method is 99.50%, the specificity is 99.50%, and the PPV is 69.11%. However, the performance of the conventional Z-score test plus placental mosaicism method is related to the set mosaicism ratio. When the mosaicism ratio is set to 30%, the sensitivity is only 97.62%, and the specificity (99.91%) and PPV (91.73%) are close to those of the combined Z-score test method. Therefore, the combined Z-score test method has the best overall performance, which can meet the requirements of sensitivity and specificity at the same time, with a sensitivity close to 99%, a specificity of 99.9%, and a PPV>90%. The conventional Z-score test method has the highest sensitivity, but the specificity and PPV are the lowest among all methods. The conventional Z-score test plus placental mosaicism method has insufficient sensitivity and an excessively high missed detection rate.
[0293] In some embodiments, step S809 may include but is not limited to steps S901 to S902:
[0294] Step S901, obtaining a first negative Z value corresponding to the X chromosome, a second negative Z value corresponding to the Y chromosome, and a Z value judgment threshold;
[0295] Step S902: Input the first negative Z value, the second negative Z value, and the Z value judgment threshold into the sex chromosome detection decision tree model, and use the sex chromosome detection decision tree model to output the fetal sex and the corresponding sex chromosome detection result.
[0296] In some embodiments, for the detection of fetal sex chromosome aneuploidy, it is not necessary to construct a positive sample dataset, only a negative sample dataset of sex chromosomes needs to be constructed. For the detection of aneuploidy of the other 23 pairs of autosomes of the fetus, it is necessary to construct a positive sample dataset and a negative sample dataset. Taking the detection of fetal sex chromosome aneuploidy as an example, the above-mentioned non-invasive prenatal fetal chromosome detection method is applied. It should be noted that by constructing a negative reference set for female fetuses and a negative reference set for male fetuses, the Z values (Z values) of the X chromosome and Y chromosome of the sample to be tested are systematically calculated. X and Z Y ), and combined with the decision tree model and Z X and Z Y The combined testing method can detect and classify chromosomal aneuploidy abnormalities. The specific steps are as follows:
[0297] 1. Constructing a negative reference set for female and male fetuses: Before calculating the Z value, a negative reference set must be constructed. The female negative reference set consists of sequencing data from plasma cell-free DNA from more than 2,000 pregnant women whose gold standard indicates a female fetus. The female negative reference set can be used to obtain the copy number distribution of the X and Y chromosomes in negative female fetal samples, and the mean copy number of the X chromosome in negative female fetal samples, μ X and standard deviation σ X ; Mean copy number of Y chromosome in negative female fetal samples μ Y and standard deviation σ Y Male fetus negative reference set: It consists of sequencing data of plasma free DNA of more than 2000 pregnant women whose gold standard showed that they were female fetuses. The male fetus negative reference set is mainly used to construct Z X and Z Y The linear relationship between the male fetus sample Z X and Z Y Information about the values that make up the reference set;
[0298] 2. Calculate the Z value of the sample X chromosome and Y chromosome:
[0299] Calculate the Z value of chromosome X (Z X ):
[0300]
[0301] Among them, Z X is the negative Z value of chromosome X, RR′ iX is the batch correction value of the copy number of chromosome X to be detected;
[0302] Calculate the Y chromosome Z value (Z Y ):
[0303]
[0304] Among them, Z Y is the negative Z value of the Y chromosome, RR′ iY is the batch correction value of the copy number of the Y chromosome to be detected;
[0305] 3. 2D Z value (Z X and Z Y ) Joint test: After calculating the first negative Z value (Z X ) and the second negative Z value corresponding to the Y chromosome (Z Y ) and then perform a joint test on these two Z values to identify various sex chromosome aneuploidy abnormalities and predict the Z value range of the Y chromosome of the negative male fetus reference sample with a 99% confidence interval. ref_X and Zref_y The linear relationship is established by combining the Z value of the fetal sample to be tested. X The maximum value Z of the 99% confidence interval of the fetal sample to be tested under negative conditions is obtained. Y_max and the minimum value Z Y_min , according to the first reference Z of the male fetus negative reference set Y Value and the second reference Z of the female fetus negative reference set Y value, establish Z Y The reference threshold value interval is from Z Y The threshold Z of the Y chromosome Z value used to distinguish male and female fetuses is determined in the reference threshold interval Y_boy , optionally, Z X 、Z Y and Z Y_boy Input into the trained sex chromosome detection decision tree model. First, the sex chromosome detection decision tree model is used to determine the sex of the fetus. Assuming that sample Z Y >Z Y_boy , it is judged as a male fetus sample, otherwise it is judged as a female fetus sample; then, for the female fetus sample, the sex chromosome detection decision tree model is used to determine the sex chromosomes according to the Z X To determine the abnormality of sex chromosomes in female fetus samples, if Z X >3, judged as XXX, if Z X <-3, judged as XO, if Z X If the value is between -3 and 3, it is judged as XX. For male fetus samples, the sex chromosome detection decision tree model is used according to Z Y 、Z X 、Z Y_max and Z Y_min , to determine the abnormality of sex chromosomes in male fetus samples, if Z X Between [-3, 3] and Z Y >Z Y_max , judged as XXY, if Z X <-3 and Z Y >Z Y_max , judged as XYY, if Z X <-3 and Z Y <Z Y_max , judged as XOXY, if Z X <-3 and Z Y In [Z Y_min , Z Y_max ] is judged as XY.
[0306] For example, the process and results of determining the sex chromosome copy number variation are shown in Table 14. The details of Table 14 are as follows:
[0307] Table 14 Sex chromosome copy number variation judgment results
[0308]
[0309]
[0310] The present application provides a detection kit for non-invasive prenatal screening, which is used in combination with a library construction system to implement the above-mentioned step S101 and complete the construction of a plasma free DNA sequencing library. The detection kit constructs a plasma free DNA sequencing library by optimizing steps such as end repair, adapter ligation, PCR library expansion, library quantification, and sequencing, thereby improving the success rate of library construction. Optionally, the detection kit is a magnetic bead purification and enrichment kit for fetal DNA. The main components of the detection kit are shown in Table 15 below. Table 15 is as follows:
[0311] Table 15 Main components of the detection kit
[0312]
[0313]
[0314] Reference Figure 3 , Figure 3 This is an optional structural diagram of a non-invasive prenatal fetal chromosome detection device provided in an embodiment of the present application. The device is used to implement the above-mentioned non-invasive prenatal fetal chromosome detection method. The device may include:
[0315] A data preprocessing module is used to construct a plasma free DNA sequencing library corresponding to the peripheral blood of pregnant women based on a magnetic bead purification and enrichment fetal DNA kit and a library construction system; obtain genomic sequencing data from the plasma free DNA sequencing library; align the genomic sequencing data to the human reference genome, divide the human reference genome into multiple unit windows according to a preset unit window width, and calculate the number of sequencing sequences corresponding to the genomic sequencing data aligned to each unit window; based on the number of sequencing sequences corresponding to each unit window, screen out a number of target candidate windows from the multiple unit windows, where the number of sequencing sequences corresponding to the target candidate windows exceeds a given average level threshold; based on each target candidate window, determine a number of genomic duplication regions in the genomic sequencing data, where the genomic duplication region is a genomic region in which the target candidate window appears continuously and the number of times the target candidate window appears is greater than a given threshold;
[0316] The chromosome copy number dynamic correction module is used to obtain multiple plasma free DNA samples to be tested in the current test batch; the plasma free DNA samples to be tested are obtained by enriching the fetal free DNA in the plasma of the pregnant woman to be tested by the magnetic bead purification method; for each chromosome to be tested in each plasma free DNA sample to be tested, the chromosome to be tested is regionally screened according to the genome sequencing data and the repeated regions of each genome, and the number of gene sequences corresponding to the chromosome to be tested is determined; for each plasma free DNA sample to be tested, the copy number statistics are performed according to the number of gene sequences corresponding to each chromosome to be tested, and the copy number of each chromosome to be tested is determined; for each plasma free DNA sample to be tested, Obtain a negative sample data set corresponding to each chromosome to be detected, and determine the negative Z value corresponding to each chromosome to be detected based on the copy number corresponding to each chromosome to be detected and the negative sample data set; for each plasma free DNA sample to be detected, dynamically correct the copy number of each chromosome to be detected based on the negative Z value corresponding to each chromosome to be detected and a given quantity index threshold, and determine the copy number correction value corresponding to each chromosome to be detected; based on the copy number correction value corresponding to each chromosome to be detected in each plasma free DNA sample to be detected, perform chromosome copy number error correction for the current detection batch, and determine the copy number batch correction value corresponding to each chromosome to be detected in each plasma free DNA sample to be detected;
[0317] The chromosome detection module is used to obtain the corresponding positive sample data set for each plasma free DNA sample to be tested, determine the negative Z value and the positive Z value based on the negative sample data set and the positive sample data set, perform chromosome detection based on the copy number batch correction value, negative Z value and positive Z value corresponding to each chromosome to be tested in each plasma free DNA sample to be tested, and determine the chromosome detection result of each plasma free DNA sample to be tested, wherein the positive sample data set is simulated and established based on the negative sample data set.
[0318] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0319] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described non-invasive prenatal fetal chromosome testing method. The electronic device can be any smart terminal, such as a tablet computer.
[0320] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0321] See also Figure 4 , Figure 4 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0322] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0323] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called by the processor 901 to execute the non-invasive prenatal fetal chromosome detection method of the embodiments of this application;
[0324] Input / output interface 903, used to implement information input and output;
[0325] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0326] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );
[0327] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0328] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned non-invasive prenatal fetal chromosome detection method.
[0329] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0330] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0331] The embodiments of the present application provide a method, device, and kit for non-invasive prenatal fetal chromosome detection, which enriches the fetal free DNA in the plasma of the pregnant woman to be detected through a magnetic bead purification method to obtain a plasma free DNA sample to be detected, thereby increasing the proportion of the fetal free DNA concentration in the sample to be detected, and can realize automatic detection of the plasma free DNA sample to be detected, reduce the occurrence of detection failure or false negative results, and perform chromosome detection based on the negative Z value and positive Z value corresponding to the chromosome to be detected, thereby effectively improving the efficiency and accuracy of non-invasive prenatal fetal chromosome detection.
[0332] The embodiments of the present application provide a method, device, and kit for non-invasive prenatal fetal chromosome detection, and propose a library construction method that can enrich the concentration of fetal DNA. By adding a DNA short fragmentation enrichment step before DNA end repair, the proportion of fetal DNA in the library is effectively increased, thereby reducing the probability of detection failure and false negatives caused by low fetal DNA concentration and significantly improving the reliability of the detection. A plasma free DNA library construction kit that can enrich the concentration of fetal DNA is proposed. The kit includes a complete library construction step, including DNA fragment screening, end repair, and ligation PCR amplification. The components of the kit are optimized in each link to improve the stability of the library construction results. A sliding window-based repetitive region removal method (SW-RR-Removal) and a chromosome copy number dynamic correction method are proposed, which can effectively identify abnormal reads caused by repetitive sequences and reduce the impact of aneuploid abnormal chromosomes on the copy number calculation of other chromosomes, thereby improving the accuracy of chromosome copy number calculation. A negative and positive Z value combined test method is proposed, which combines the Z values in the negative reference set and the positive reference set for joint analysis. The combined test takes into account the requirements of sensitivity and specificity of the kit at the same time, while the traditional single The Z test method only considers the specificity requirement. This method can flexibly set differentiated negative Z value and positive Z value threshold combinations for different chromosomes, so as to better meet the clinical expected detection sensitivity and specificity requirements and reduce the missed detection rate. The traditional Z test + identification of placental mosaicism method is not sensitive enough and the missed detection rate is too high. At the same time, the simulation establishes a negative reference data set, which effectively solves the problem of insufficient positive samples of rare chromosomal aneuploidy abnormalities and difficulty in obtaining positive sample data sets. The copy number distribution value of the positive sample is deduced by the fetal concentration distribution of the negative sample, which realizes the reasonable expansion of the aneuploidy positive reference set. The mean and standard deviation of the sex samples generate random errors, making the simulated data closer to the real distribution characteristics and improving the credibility of the positive sample data set. The established positive sample reference data can be extended to different types of chromosome copy number variation detection scenarios, such as non-invasive chromosome microdeletion and microduplication detection, and has strong application value. By constructing a sex chromosome detection decision tree model and combining the differences between male and female fetal samples, the negative Z value of the X chromosome and the negative Z value of the Y chromosome are calculated, and the expected value of the Y chromosome is calculated using the regression curve to detect the significance of its deviation, so as to accurately identify the type of fetal sex chromosome aneuploidy abnormality.
[0333] In summary, the non-invasive prenatal fetal chromosome detection method, device and kit provided in the embodiments of the present application can effectively improve the accuracy and comprehensiveness of fetal chromosome detection, cover the detection of aneuploidy abnormalities of 23 pairs of chromosomes, improve the sensitivity and specificity of chromosome detection, and provide more powerful technical support for prenatal screening.
[0334] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0335] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0336] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A non-invasive prenatal fetal chromosome detection method, characterized in that: The method comprises the following steps: Based on the magnetic bead purification and enrichment kit for fetal DNA and the library construction system, a plasma free DNA sequencing library corresponding to maternal peripheral blood was constructed; Obtaining genome sequencing data from the plasma cell-free DNA sequencing library; Aligning the genome sequencing data to a human reference genome, dividing the human reference genome into a plurality of unit windows according to a preset unit window width, and calculating the number of sequencing sequences corresponding to the genome sequencing data aligned to each of the unit windows; Screening out a plurality of target candidate windows from the plurality of unit windows according to the number of sequencing sequences corresponding to each of the unit windows, wherein the number of sequencing sequences corresponding to the target candidate windows exceeds a given average level threshold; Determining, based on each of the target candidate windows, a plurality of genomic duplication regions in the genomic sequencing data, wherein the genomic duplication regions are genomic regions in which the target candidate windows appear continuously and the number of occurrences is greater than a given threshold; Obtaining multiple plasma free DNA samples to be tested in the current test batch; the plasma free DNA samples to be tested are obtained by enriching fetal free DNA in the plasma of the pregnant woman to be tested by a magnetic bead purification method; For each chromosome to be detected in each of the plasma free DNA samples to be detected, performing regional screening on the chromosome to be detected based on the genome sequencing data and each of the genome repetitive regions to determine the number of gene sequences corresponding to the chromosome to be detected; For each of the plasma free DNA samples to be tested, copy number statistics are performed based on the number of gene sequences corresponding to each of the chromosomes to be tested to determine the copy number of each of the chromosomes to be tested; For each of the plasma cell-free DNA samples to be tested, obtaining a negative sample data set corresponding to each of the chromosomes to be tested, and determining a negative Z value corresponding to each of the chromosomes to be tested based on the copy number corresponding to each of the chromosomes to be tested and the negative sample data set; For each of the plasma free DNA samples to be tested, dynamically correcting the copy number of each of the chromosomes to be tested according to the negative Z value corresponding to each of the chromosomes to be tested and a given quantity index threshold, and determining a copy number correction value corresponding to each of the chromosomes to be tested; According to the copy number correction value corresponding to each of the chromosomes to be detected in each of the plasma free DNA samples to be detected, performing chromosome copy number error correction for the current detection batch, and determining a copy number batch correction value corresponding to each of the chromosomes to be detected in each of the plasma free DNA samples to be detected; For each of the plasma free DNA samples to be tested, a corresponding positive sample data set is obtained, and a negative Z value and a positive Z value are determined based on the negative sample data set and the positive sample data set. Chromosome detection is performed based on the copy number batch correction value, the negative Z value, and the positive Z value corresponding to each of the chromosomes to be detected in each of the plasma free DNA samples to be tested, and the chromosome detection result of each of the plasma free DNA samples to be tested is determined, wherein the positive sample data set is simulated and established based on the negative sample data set.
2. The non-invasive prenatal fetal chromosome detection method according to claim 1, characterized in that: For each chromosome to be detected in each plasma free DNA sample to be detected, performing regional screening on the chromosome to be detected based on the genome sequencing data and each genome duplication region to determine the number of gene sequences corresponding to the chromosome to be detected, specifically comprising: Determine, based on the genome sequencing data, multiple gene sequencing regions in the chromosome to be detected and the number of regions and region length corresponding to each gene sequencing region, where the region length is the sequence length corresponding to the gene sequencing region, and the number of regions is the frequency of occurrence of the gene sequencing region in the chromosome to be detected; Dividing the gene sequencing regions into multiple gene sequencing groups according to the region lengths corresponding to the gene sequencing regions, each gene sequencing group including multiple gene sequencing regions with the same region length; According to each of the genomic duplication regions, screening the chromosome duplication region corresponding to each of the genomic duplication regions from the multiple gene sequencing regions, and determining the target gene sequencing group where the chromosome duplication region is located from the multiple gene sequencing groups; For each of the chromosome duplication regions, updating the region number corresponding to the chromosome duplication region to the average region number corresponding to the target gene sequencing group, where the average region number is the average of the region numbers of multiple gene sequencing regions in the gene sequencing group; The number of gene sequences corresponding to the chromosome to be detected is determined based on the number of regions corresponding to each of the gene sequencing regions except the multiple chromosome duplication regions and the number of regions after update of each of the chromosome duplication regions.
3. The non-invasive prenatal fetal chromosome detection method according to claim 1, characterized in that: For each of the plasma free DNA samples to be detected, performing copy number statistics according to the number of the gene sequences corresponding to each of the chromosomes to be detected to determine the copy number of each of the chromosomes to be detected, specifically includes: Determining the total number of gene sequences corresponding to the plasma cell-free DNA sample to be tested based on the number of gene sequences corresponding to each of the chromosomes to be tested; The copy number corresponding to each chromosome to be detected is determined based on the number of gene sequences corresponding to each chromosome to be detected and the total number of gene sequences. The copy number is calculated by the following formula: RR ij =RN ij / AN i ; Among them, RR ij is the copy number of the jth chromosome to be detected in the i-th plasma free DNA sample to be detected, RN ij is the number of gene sequences corresponding to the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, AN i is the total number of gene sequences corresponding to the i-th plasma free DNA sample to be tested, and N is the number of chromosomes to be tested in the i-th plasma free DNA sample to be tested.
4. The non-invasive prenatal fetal chromosome detection method according to claim 1, characterized in that: The method further comprises: obtaining a negative sample data set corresponding to each chromosome to be detected for each plasma free DNA sample to be detected, and determining a negative Z value corresponding to each chromosome to be detected based on the copy number corresponding to each chromosome to be detected and the negative sample data set. Determine the corresponding negative sample copy number mean and negative sample copy number standard deviation according to the negative sample data set corresponding to each of the chromosomes to be detected; According to the copy number corresponding to the chromosome to be detected, the mean copy number of the negative sample and the standard deviation of the copy number of the negative sample, the negative Z value corresponding to the chromosome to be detected is determined, and the negative Z value is calculated by the following formula: Z_neg ij =(RR ij -μneg j ) / σneg j ; Among them, Z_neg ij is the negative Z value corresponding to the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, RR ij is the copy number of the jth chromosome to be detected in the i-th plasma free DNA sample to be detected, μneg j is the mean number of negative sample copies corresponding to the jth chromosome to be tested in the negative sample data set, σneg j is the standard deviation of the negative sample copy number corresponding to the jth chromosome to be detected in the negative sample data set.
5. The non-invasive prenatal fetal chromosome detection method according to claim 4, characterized in that: For each of the plasma free DNA samples to be detected, dynamically correcting the copy number of each of the chromosomes to be detected according to the negative Z value corresponding to each of the chromosomes to be detected and a given quantity index threshold, and determining the copy number correction value corresponding to each of the chromosomes to be detected specifically includes: Determining a plurality of first abnormal chromosomes from the plurality of chromosomes to be detected according to the negative Z value corresponding to each of the chromosomes to be detected and the quantitative index threshold, wherein the negative Z value corresponding to the first abnormal chromosome exceeds the quantitative index threshold; obtaining any of the first abnormal chromosomes as a current corrected chromosome; Obtaining the total number of gene sequences corresponding to the plasma cell-free DNA sample to be tested and the mean number of negative sample copies corresponding to the currently corrected chromosome; According to the total number of gene sequences corresponding to the plasma free DNA sample to be tested and the mean number of negative sample copies corresponding to the current corrected chromosome, the number of gene sequences corresponding to the current corrected chromosome is updated, and the updated number of gene sequences is calculated by the following formula: in, is the updated number of gene sequences corresponding to the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, AN i is the total number of gene sequences corresponding to the i-th plasma free DNA sample to be tested, μneg j is the mean copy number of negative samples corresponding to the jth chromosome to be tested; updating the total number of gene sequences according to the number of updated gene sequences corresponding to the currently corrected chromosome and the number of gene sequences corresponding to the remaining chromosomes to be tested; Recalculating the chromosome copy number based on the total number of updated gene sequences, the number of gene sequences of the remaining chromosomes to be detected, and the updated number of gene sequences corresponding to the currently corrected chromosome, to determine the corrected copy number value corresponding to the currently corrected chromosome and the corrected copy number values corresponding to the remaining chromosomes to be detected; When there is any first abnormal chromosome whose gene sequence quantity has not been updated, the first abnormal chromosome whose gene sequence quantity has not been updated is obtained as the current corrected chromosome, and then the total number of gene sequences corresponding to the plasma free DNA sample to be tested and the mean copy number of the negative samples corresponding to the current corrected chromosome are obtained again until the gene sequence quantity of all the first abnormal chromosomes has been updated.
6. The non-invasive prenatal fetal chromosome detection method according to claim 4, characterized in that: The step of performing chromosome copy number error correction for the current detection batch based on the copy number correction value corresponding to each chromosome to be detected in each plasma free DNA sample to be detected, and determining the copy number batch correction value corresponding to each chromosome to be detected in each plasma free DNA sample to be detected, specifically includes: Acquire any of the plasma free DNA samples to be detected as a target detection sample, acquire any of the chromosomes to be detected in the target detection sample as a current target detection chromosome, and acquire the copy number correction value corresponding to the current target detection chromosome as the current copy number; Obtaining the negative sample copy number mean, the negative sample copy number standard deviation, and the current batch chromosome copy number mean corresponding to the current target detection chromosome, wherein the current batch chromosome copy number mean is the copy number mean corresponding to the current target detection chromosome in all the plasma free DNA samples to be detected; Determine the normalized copy number corresponding to the current target detection chromosome according to the current copy number, the mean copy number of the negative samples and the mean copy number of the chromosomes in the current batch; Update the negative Z value corresponding to the current target detection chromosome according to the normalized copy number, the negative sample copy number mean and the negative sample copy number standard deviation; Determining whether the current target detection chromosome is the second abnormal chromosome according to the updated negative Z value corresponding to the current target detection chromosome and the quantity index threshold; When the updated negative Z value corresponding to the current target detection chromosome is greater than the quantity index threshold, the current target detection chromosome is determined to be the second abnormal chromosome, the negative sample copy number mean is determined as the current copy number, and the current batch chromosome copy number mean is re-determined based on the current copy number, and then the normalized copy number corresponding to the current target detection chromosome is determined based on the current copy number, the negative sample copy number mean, and the current batch chromosome copy number mean, until the updated negative Z value corresponding to the current target detection chromosome is less than the quantity index threshold; When the updated negative Z value corresponding to the current target detection chromosome is less than the quantity index threshold, determining the normalized copy number as the copy number batch correction value corresponding to the current target detection chromosome; When there is any chromosome to be detected in the target detection sample that has not undergone chromosome copy number error correction, any remaining chromosome to be detected in the target detection sample that has not undergone chromosome copy number error correction is obtained as the current target detection chromosome, the copy number correction value corresponding to the current target detection chromosome is obtained as the current copy number, and then the negative sample copy number mean, the negative sample copy number standard deviation, and the current batch chromosome copy number mean corresponding to the current target detection chromosome are obtained, until all the chromosomes to be detected in the target detection sample have undergone chromosome copy number error correction; The normalized copy number was calculated by the following formula: Among them, the current detection chromosome is the jth chromosome to be detected in the i-th plasma free DNA sample to be detected, The normalized copy number corresponding to the current target chromosome detection, Correction value for the copy number of the current target chromosome, μb j is the mean number of chromosome copies in the current batch corresponding to the current target detection chromosome in the current detection batch, μneg j The mean copy number of negative samples corresponding to the current target chromosome.
7. The non-invasive prenatal fetal chromosome detection method according to claim 4, characterized in that: The method of obtaining a corresponding positive sample data set for each of the plasma free DNA samples to be tested, determining a negative Z value and a positive Z value based on the negative sample data set and the positive sample data set, performing chromosome testing based on the copy number batch correction value, the negative Z value, and the positive Z value corresponding to each of the chromosomes to be tested in each of the plasma free DNA samples to be tested, and determining a chromosome testing result for each of the plasma free DNA samples to be tested, specifically includes: Obtaining the negative sample data set corresponding to the chromosome to be detected; Determining the expected copy number value of the positive sample data set according to the fetal concentration distribution of the negative sample data set, and generating the positive sample data set according to the expected copy number value and random error; Performing copy number distribution statistics on the negative sample data set to determine the negative Z value judgment threshold; Update the negative Z value corresponding to the chromosome to be detected according to the copy number batch correction value corresponding to the chromosome to be detected, the negative sample copy number mean and the negative sample copy number standard deviation; When the chromosome to be detected is an autosome, obtaining a positive sample data set corresponding to the chromosome to be detected, performing copy number distribution statistics on the positive sample data set, and determining the positive Z value judgment threshold, the positive sample copy number mean, and the positive sample copy number standard deviation; The positive Z value corresponding to the chromosome to be detected is determined according to the copy number batch correction value corresponding to the chromosome to be detected, the mean copy number of the positive samples, and the standard deviation of the copy number of the positive samples. The positive Z value is calculated by the following formula: Among them, Z_pos ij is the positive Z value corresponding to the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, μpos j is the mean sample copy number corresponding to the jth chromosome to be tested in the positive sample data set, is the batch correction value of the copy number of the jth chromosome to be tested in the i-th plasma free DNA sample to be tested, σpos j is the standard deviation of the sample copy number corresponding to the jth chromosome to be tested in the positive sample data set; Performing a joint test based on the negative Z value judgment threshold, the negative Z value, the positive Z value judgment threshold, and the positive Z value to determine the chromosome detection result; When the negative Z value is greater than the negative Z value judgment threshold, and the positive Z value is greater than the positive Z value judgment threshold, determining that the chromosome detection result is positive for the chromosome to be detected in the plasma free DNA sample to be detected; When the negative Z value is less than or equal to the negative Z value judgment threshold, and the positive Z value is less than or equal to the positive Z value judgment threshold, determining that the chromosome detection result is negative for the chromosome to be detected in the plasma free DNA sample to be detected; When the chromosome to be detected is a sex chromosome, a sex chromosome detection decision tree model is obtained, and the chromosome detection result is determined using the sex chromosome detection decision tree model according to the negative Z value judgment threshold and the negative Z value.
8. The non-invasive prenatal fetal sex chromosome detection method according to claim 7, characterized in that: When the chromosome to be detected is a sex chromosome, obtaining a sex chromosome detection decision tree model, and using the sex chromosome detection decision tree model to determine the chromosome detection result according to the negative Z value judgment threshold and the negative Z value, specifically includes: Obtain the first negative Z value corresponding to the X chromosome and the second negative Z value corresponding to the Y chromosome and the Z value judgment threshold; The first negative Z value, the second negative Z value and the Z value judgment threshold are input into the sex chromosome detection decision tree model, and the sex chromosome detection decision tree model is used to output the fetal sex and the corresponding sex chromosome detection result.
9. The non-invasive prenatal fetal chromosome detection method according to claim 1, characterized in that: The fetal DNA enrichment kit based on magnetic beads and the library construction system are used to construct a plasma free DNA sequencing library corresponding to the peripheral blood of pregnant women, specifically including: Extraction of plasma free DNA from maternal plasma; performing fragment screening on the plasma free DNA to obtain fragment-screened plasma free DNA; The plasma free DNA after the fragment screening is filled with A and phosphorylated to obtain the corresponding end repair product; Performing adapter ligation on the end-repair product to obtain a corresponding adapter ligation product; Purifying the adapter ligation product with magnetic beads, and performing library PCR amplification on the adapter ligation product after magnetic bead purification to obtain corresponding PCR products; The PCR products are purified by magnetic beads, and the library is quantified and sequenced to generate the plasma free DNA sequencing library.
10. A non-invasive prenatal fetal chromosome detection device, characterized in that: The device comprises: A data preprocessing module is used to construct a plasma free DNA sequencing library corresponding to the peripheral blood of pregnant women based on a magnetic bead purification and enrichment fetal DNA kit and a library construction system; obtain genome sequencing data from the plasma free DNA sequencing library; align the genome sequencing data to a human reference genome, divide the human reference genome into multiple unit windows according to a preset unit window width, and calculate the number of sequencing sequences corresponding to the genome sequencing data aligned to each unit window; based on the number of sequencing sequences corresponding to each unit window, screen out a number of target candidate windows from the multiple unit windows, wherein the number of sequencing sequences corresponding to the target candidate windows exceeds a given average level threshold; based on each target candidate window, determine a number of genomic duplication regions in the genomic sequencing data, wherein the genomic duplication regions are genomic regions in which the target candidate windows appear continuously and the number of occurrences is greater than a given threshold; The chromosome copy number dynamic correction module is used to obtain multiple plasma free DNA samples to be tested in the current test batch; the plasma free DNA samples to be tested are obtained by enriching fetal free DNA in the plasma of the pregnant woman to be tested by a magnetic bead purification method; for each chromosome to be tested in each plasma free DNA sample to be tested, according to the genome sequencing data and each genome duplication region, the chromosome to be tested is regionally screened to determine the number of gene sequences corresponding to the chromosome to be tested; for each plasma free DNA sample to be tested, according to the number of gene sequences corresponding to each chromosome to be tested, copy number statistics are performed to determine the copy number of each chromosome to be tested; for each plasma free DNA sample to be tested, the copy number of each chromosome to be tested is obtained. A negative sample data set corresponding to the chromosome to be detected, determining the negative Z value corresponding to each chromosome to be detected according to the copy number corresponding to each chromosome to be detected and the negative sample data set; for each plasma free DNA sample to be detected, dynamically correcting the copy number of each chromosome to be detected according to the negative Z value corresponding to each chromosome to be detected and a given quantity index threshold, and determining the copy number correction value corresponding to each chromosome to be detected; according to the copy number correction value corresponding to each chromosome to be detected in each plasma free DNA sample to be detected, performing chromosome copy number error correction for the current detection batch, and determining the copy number batch correction value corresponding to each chromosome to be detected in each plasma free DNA sample to be detected; A chromosome detection module is used to obtain a corresponding positive sample data set for each of the plasma free DNA samples to be tested, determine a negative Z value and a positive Z value based on the negative sample data set and the positive sample data set, perform chromosome detection based on the copy number batch correction value, the negative Z value, and the positive Z value corresponding to each of the chromosomes to be tested in each of the plasma free DNA samples to be tested, and determine the chromosome detection result of each of the plasma free DNA samples to be tested, wherein the positive sample data set is simulated and established based on the negative sample data set.
11. A detection kit, characterized in that: The detection kit is used to implement the method described in claim 9.
Citation Information
Patent Citations
Method and device for detecting fetal aneuploid and copy number variation in free DNA of pregnant woman plasma and application
CN117095745A
Kit, device and method for improving concentration of fetal free DNA in maternal peripheral blood
CN107541561A
Device for noninvasive prenatal detection of chromosome abnormality
CN112522387A
Method and device for detecting aneuploid abnormality of fetal chromosome and storage medium
CN115223654A
Optimization method for noninvasive prenatal detection of fetal chromosome copy number abnormality
CN117230165A