Methods and devices for detecting fetal aneuploidy and copy number variations in maternal plasma cell-free DNA and applications

By employing non-targeted whole-genome sequencing and machine learning methods, the false-positive problem in detecting fetal aneuploidy and copy number variations in cell-free DNA from maternal plasma has been solved in existing technologies, achieving high-precision and low-cost detection results.

CN117095745BActive Publication Date: 2025-12-12ANNOROAD GENE TECHNOLOGY (BEIJING) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311069138.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-12-28
Filing Date
2023-08-23
Publication Date
2025-12-12
Estimated Expiration
2043-08-23

AI Technical Summary

Technical Problem

Existing technologies for detecting fetal aneuploidy and copy number variations in cell-free DNA in maternal plasma suffer from high false-positive rates, complex procedures, lack of scientific rigor and inter-platform compatibility, making it difficult to achieve rapid and accurate detection.

Method used

By employing non-targeted whole-genome sequencing combined with machine learning methods, a machine learning device is constructed by extracting deep information and feature vectors from sequencing data to accurately distinguish between fetal and maternal variations, and classifier models such as logistic regression and random forest are used for detection.

Benefits of technology

It achieves accurate differentiation of fetal aneuploidy and copy number variation, with strong compatibility, wide applicability, low cost, and high-precision detection with only ultra-low sequencing depth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004410684520000141
    Figure BDA0004410684520000141
Patent Text Reader

Abstract

The application relates to a method and device for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA and application. The method comprises the following steps: obtaining sequencing data based on a conventional non-targeted whole genome sequencing mode, so as to group sequencing templates and generate analyzable sub-libraries; extracting sequencing depth information of different genomic intervals in the sequencing data; grouping according to the feature vector / matrix extracted from the sequencing data, and counting and extracting the depth information of the genomic intervals from different groups; and constructing a classifier through machine learning to further accurately distinguish the variations carried by the fetus and the variations carried by the mother. The method and device can accurately distinguish the variations carried by the fetus and the variations carried by the mother, and the method and device have strong compatibility, wide applicability and low cost, and can realize accurate distinction and detection of fetal free DNA genome aneuploidy and copy number variation only by using ultra-low sequencing depth.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of gene detection, and particularly relates to a method and device for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA and application. BACKGROUND

[0002] Trisomies 21, 18 and 13 (T21, T18 and T13) in fetal chromosome aneuploidies are the most common chromosomal aneuploidy diseases in clinic. They correspond to 21-trisomy syndrome (also known as Down syndrome, congenital idiocy or Down syndrome), 18-trisomy syndrome (also known as Edwards syndrome) and 13-trisomy syndrome (also known as Patau syndrome), respectively, with an incidence of about 1 / 700, 1 / 6000 and 1 / 10000, respectively. Most of the children suffer from severe mental retardation and organ malformations, and cannot take care of themselves, which not only affects the life and health of children and the quality of life, but also affects the healthy and sustainable development of the economy and society. Copy number variation (CNV) is caused by rearrangement of the genome, generally refers to the increase or decrease of the copy number of a large fragment of the genome with a length of more than 1 kb, and mainly manifests as submicroscopic deletion and duplication. CNV is an important part of structural variation (SV) of the genome. The mutation rate of CNV sites is much higher than that of SNP (Single nucleotide polymorphism), and CNV is one of the important pathogenic factors of human diseases.

[0003] At present, the method commonly used for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA is low-depth whole genome sequencing of peripheral blood free DNA. This method has false positive results caused by placental chimerism, maternal copy number variation and other interferences in the above detection. The method commonly used for identifying placental chimerism at present is to judge the similarity between the variation detection value and the fetal concentration by experience, which lacks scientificity and universality between platforms, and it is difficult to quickly and accurately judge by a fixed threshold, and the interpretability is low.

[0004] Therefore, in view of the low accuracy and complex operation of the detection products on the market, it is urgent to design a method and device for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA, which can specifically improve the detection accuracy of fetal aneuploidy and copy number variation, and can be compatible with different library construction sequencing methods, and has high universality. SUMMARY

[0005] To address the aforementioned problems, this invention provides a method and apparatus for detecting fetal aneuploidy and copy number variations in cell-free DNA from maternal plasma. Using this method and apparatus, it is possible to accurately distinguish between fetal-carried variations and maternally-carried variations. Furthermore, the method provided by this invention is highly compatible, widely applicable, low-cost, and requires no probe design; it only requires ultra-low sequencing depth to achieve accurate differentiation and detection of aneuploidy and copy number variations in the fetal cell-free DNA genome.

[0006] Specifically, the present invention relates to methods and apparatus for detecting fetal aneuploidy and copy number variations in cell-free DNA in maternal plasma, and their applications.

[0007] 1. A method for detecting fetal aneuploidy and copy number variations in cell-free DNA in maternal plasma, comprising the following steps:

[0008] Step 1: Based on non-targeted whole-genome sequencing, obtain sequencing data in order to group sequencing templates and generate analyzable sub-libraries;

[0009] Step 2: Extract sequencing depth information D from different genomic regions in the sequencing data. i , where D i This represents the i-th unit counting window on the genome;

[0010] Step 3: For each sequencing template, group them according to the feature vectors / matrices extracted from the sequencing data. j , j∈(1,2,3......,N); and statistically analyze and extract the depth information D of genomic regions from different grouped sub-libraries. i,j D i,j For S j Genome depth information in the i-th counting window under grouping;

[0011] Step 4: Detect aneuploidy and copy number variation across the entire genome for different groupings, and output the detection values ​​Z corresponding to aneuploidy and copy number variation for different groupings. t,j t represents different detection targets;

[0012] Step 5: Using machine learning methods and samples of the known real results, construct a feature-grouped detection value {Z}. t,j The training set of the machine learning device is used to obtain the learning device;

[0013] Step 6: For the sample to be tested, perform statistical analysis based on the same Z-squared value. t,j and D i,jThe feature grouping detection value feature depth vector is input into the learning device constructed in step 5, and the detection result is finely distinguished according to the predicted label.

[0014] 2. The method according to the above, wherein the non-targeted whole genome sequencing method is selected from at least one of methylation sequencing, paired-end short sequence sequencing and single-end full-length sequencing.

[0015] 3. The method according to the above, wherein the depth information is selected from at least one of Reads, Unique reads, Mapability, Genomic GC, Reads GC, Unique reads GC.

[0016] 4. The method according to the above, wherein the feature vector / matrix is selected from insert size, sequence end base distribution frequency.

[0017] 5. The method according to the above, wherein the method for constructing the training set of the machine learning device based on the feature grouping detection value {Z t,j} includes inputting data, processing the input data by constructing a classifier model to obtain a judgment result of detecting a target, and then performing label classification to obtain a known label.

[0018] 6. The method according to the above, wherein the input data is selected from D i , S j , D i,j , Z t,j .

[0019] 7. The method according to the above, wherein the classifier model is selected from at least one of logistic regression, random forest, support vector, linear regression, decision tree and neural network.

[0020] 8. The method according to the above, wherein the type of label classification is selected from negative and positive; preferably, the type of label classification is selected from negative, fetal positive and maternal positive; more preferably, the type of label classification is selected from negative, fetal positive, maternal positive and chimera.

[0021] 9. A device for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA, comprising a data acquisition unit, a window division unit, a grouping unit, an aneuploidy and copy number variation detection unit, a modeling unit and a to-be-tested sample result output unit; wherein,

[0022] The data acquisition unit is used to obtain sequencing data based on non-targeted whole genome sequencing, so as to group the sequencing templates and produce analyzable sub-libraries;

[0023] The dividing window unit is configured to extract sequencing depth information D of different genomic intervals in the sequencing data i , wherein D i is the i-th unit count window on the genome

[0024] The grouping unit is configured to group each sequencing template according to the feature vector / matrix extracted from the sequencing data, to obtain different groups S j , j∈(1,2,3......,N); and to count and extract the depth information D of the genomic interval from different group sub-libraries i,j , D i,j is the genomic depth information of the i-th count window under the S j grouping

[0025] The aneuploidy and copy number variation detection unit is configured to detect aneuploidy and copy number variation in the whole genome range by grouping different groups, to output the detection value Z corresponding to the aneuploidy and copy number variation of different groups t,j , t represents different detection targets

[0026] The modeling unit is configured to use the known true result sample to construct a training set of a machine learning device based on the feature grouping detection value {Z t,j} by a machine learning method, to obtain a learning device

[0027] The to-be-tested sample result output unit is configured to, for a to-be-tested sample, count the feature grouping detection value feature depth vector based on the same format of Z t,j and D i,j , and import it into the learning device constructed by the modeling unit, and then finely distinguish the detection result according to the predicted label.

[0028] 10. The application of the above method for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA or the above device for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA in the field of genetic testing. DETAILED DESCRIPTION

[0029] In order to enable persons skilled in the art to better understand the present application, the technical solutions involved in the present application will be described clearly and completely below. Obviously, the specific implementation described is only a part of the embodiments of the present application, but not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor should belong to the scope of protection of the present application.

[0030] It should be noted that the terms "first", "second", and the like in the description and claims of the present application are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] Term explanation:

[0032] (1) Aneuploid, is the lack of or additional increase of one or several chromosomes in euploid chromosomes, generally in meiosis, a pair of homologous chromosomes do not separate or separate early to form n-1 or 721 gametes. Its composition is different from the structure of the usual polyploid, chromosome or chromosome fragment or multiple loss. The number of individual chromosomes is not multiplied or reduced, but added or reduced by single or several. The formation mechanism of aneuploidy is caused by chromosome nonseparation, loss during cell division.

[0033] (2) Sequencing depth refers to the ratio of the total amount of bases (bp) obtained by sequencing to the size of the genome (Genome), which is one of the indicators for evaluating the amount of sequencing. Ultra-low sequencing depth may be 0.1x, for example.

[0034] (3) Single-end sequencing (Single-End sequencing) refers to first fragmenting the DNA sample to form 200-500bp fragments, connecting primer sequences to one end of the DNA fragments, then adding adapters to the end, fixing the fragments on the flowcell to generate DNA clusters, and sequencing single-end reads.

[0035] (4) Paired-end sequencing (Paired-end sequencing) refers to adding sequencing primer binding sites to the adapters on both ends when constructing the DNA library to be tested. After the first round of sequencing is completed, the template strand of the first round of sequencing is removed, and the complementary strand is guided to the original position by the paired-end module (Paired-End Module) to achieve the template amount used for the second round of sequencing, and the second round of complementary strand synthesis sequencing is performed.

[0036] (5) Reads: the plural of read, a short sequencing fragment sequence generated by a high-throughput sequencing platform.

[0037] (6) Unique reads: refers to reads uniquely aligned to the genome. In the sequencing process, some reads can be aligned to multiple positions in the genome at the same time. Unique reads are filtered from all non-dup reads to remove reads aligned to multiple positions, and the remaining unique reads are obtained.

[0038] (7) Mapability: For some windows, short sequences have low uniqueness, which may be due to repetitive sequences from large heterochromatin or more complex biological reasons. Therefore, the efficiency of each window is calculated using the parameter of mapability.

[0039] (8) Genomic GC: This parameter represents the GC of the genome corresponding to each window.

[0040] (9) Reads GC: The GC of all reads in each window.

[0041] (10) Unique reads GC: represents the GC of unique reads in each window.

[0042] The first aspect of the present application provides a method for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA, comprising the following steps:

[0043] Step 1: Based on non-targeted whole genome sequencing, sequencing data is obtained to group sequencing templates and produce analyzable sub-libraries;

[0044] Step 2: Extract the sequencing depth information D of different genomic intervals in the sequencing data i , wherein D i is the i-th unit count window on the genome;

[0045] Step 3: For each sequencing template, group S j according to the feature vector / matrix extracted from the sequencing data, j∈(1,2,3......,N); and count and extract the depth information D i,j of the genomic interval from different grouped sub-libraries, D i,j is the genomic depth information in the i-th count window under S j grouping;

[0046] Step 4: Detect aneuploidy and copy number variation in the whole genome range by different groups to output the detection value Z t,j corresponding to the aneuploidy and copy number variation of different groups, t represents different detection targets;

[0047] Step 5, using the known true result samples, a training set of machine learning device based on feature group detection value {Z t,j} is constructed, and a learning device is obtained;

[0048] Step 6, for the sample to be tested, the feature group detection value feature depth vector based on the same format Z t,j and D i,j is counted, and it is introduced into the learning device constructed in step 5, and the detection result is finely distinguished according to the predicted label.

[0049] In the present application, "fine distinction" means that the sample to be tested can be distinguished into negative and positive; preferably, the sample to be tested can be distinguished into negative, fetal positive and maternal positive; more preferably, the sample to be tested can be distinguished into negative, fetal positive, maternal positive and chimera.

[0050] According to the method of the present application, the method of non-targeted whole genome sequencing can be selected from at least one of methylation sequencing, double-end short sequence sequencing and single-end full-length sequencing.

[0051] According to the method of the present application, the depth information is, for example but not limited to, at least one selected from Reads, Unique reads, Mapability, Genomic GC, Reads GC and Unique reads GC.

[0052] According to the method of the present application, the feature vector / matrix is selected from the group consisting of insert length, sequence end base distribution frequency. Alternatively, it can also include a feature vector / matrix with epigenetic characteristics. In the present application, the "feature vector / matrix" can be a feature vector, a matrix or a numerical value.

[0053] According to the method of the present application, there is no requirement for the library construction method and the sequencing platform. Specifically, the library construction method is, for example but not limited to, using the NIPT library construction kit (such as the State Food and Drug Administration Registration No. 20173400331) for library construction. The sequencing platform is, for example but not limited to, using the NextSeq 550AR gene sequencer PE75 read length mode. The analysis method is, for example but not limited to, the BMMPA registration certificate No. 20192210692.

[0054] According to the method of the present application, the number of counting windows, the sequencing depth and the variation size to be distinguished have a logical relationship. At least, the number of stable templates in a unit interval should be sufficient to ensure the stability and accuracy of the method. Since the content of fetal free DNA in plasma is relatively low, the number of sequencing templates in the unit interval should not be less than the average fetal concentration in theory.

[0055] According to the method of the present application, for each sequencing template, the feature vectors / matrices extracted from the sequencing data are grouped according to different groups S j , j ∈ (1, 2, 3, …, N). Wherein, the grouping method can be determined according to the analysis needs, for example, based on typical eigenvalues or based on limited clustering grouping of unsupervised classifiers. In the present application, there can be an intersection between different groups, for example, but not limited to, S1=1~128, S2=109~166, S3=140~223.

[0056] According to the method of the present application, the depth information D i,j of the genomic interval from different groups is counted and extracted i,j , D j is the genomic depth information of the i-th counting window under the S 1,1 group, for example, D 1,2 , D 1,3 , etc.

[0057] According to the method of the present application, the detection of aneuploidy and copy number variation in the whole genome range is carried out by different groups to output the detection value Z t,j corresponding to the aneuploidy and copy number variation, t represents different detection targets. The detection target can also be understood as the detection object, for example, but not limited to, trisomy of chromosome 13, trisomy of chromosome 18, trisomy of chromosome 21, any copy number variation syndrome. Correspondingly, Z t,j can be expressed as Z 13,j , Z 18,j , Z 21,j , Z CNV,j , j ∈ (1, 2, 3, …, N).

[0058] According to the method of the present application, the method for constructing the training set of the machine learning device based on the feature grouping detection value {Z t,j} includes: inputting data, constructing a classifier model to process the input data to obtain the determination result of the detection target, then performing label classification to obtain the known label. In subsequent analysis, the known label is used to classify and evaluate the detection result of the sample to be tested.

[0059] According to the method of the present application, the type of the label classification is selected from negative and positive; preferably, the type of the label classification is selected from negative, fetal positive and maternal positive; more preferably, the type of the label classification is selected from negative, fetal positive, maternal positive and chimera. For example, but not limited to, fetal positive, maternal positive, chimera or fetal negative. In the present application, negative can include fetal negative. For example, after the sample is verified by clinical amniocentesis, the fetus is not trisomy 13 syndrome, but the corresponding placenta of the sample exists trisomy 13 chimera, so the sample is given the classification label of 'chimera or fetal negative' for model training. The present application also includes 'fetal positive' and'maternal positive' classification labels. Fetal positive refers to judging the corresponding target as positive by the detection value and verifying that the fetus exists the abnormality of the target by puncture; maternal positive refers to judging the corresponding target as positive by the detection value, but verifying that the fetus and placenta do not exist the abnormality of the target by puncture, but the mother's somatic cells exist the abnormality of the target.

[0060] According to the method of the present application, the input data is selected from D i , S j , D i,j , Z t,j .

[0061] According to the method of the present application, the classifier model is selected from at least one of logistic regression, random forest, support vector, linear regression, decision tree and neural network. According to the information of a certain number of known samples of fetal positive and fetal negative (false positive), the multi-dimensional target detection value vector (D i,j , Z t,j ) of the above samples is used, and other statistical quantities other than the detection value and fetal concentration described above which have qualitative prediction ability for clinical targets can also be included as input vectors, and a machine learning training step is performed. The preset type number can be 0 or 1 of binary method corresponding to known positive and negative. The present device does not limit the selection of the classifier, and common classifiers that can process numerical variables such as logistic regression, random forest, support vector machine can be used. The optimal training model is selected for evaluation of the to-be-tested sample by a method similar to cross-validation. At the same time, the internal structure and weight of the classifier model can be selected or constructed for visualization of the intermediate variable.

[0062] According to the method of the present application, steps 1 to 4 are repeated with a certain number of samples with known true results, and then the multi-dimensional target detection value vector (D i,j , Z t,j) the pre-trained model in step 5 for classification prediction, and finally obtain the learning device. In the present application, the number can be determined according to the needs, and the increase of the number has a positive correlation with the accuracy of the results, but at the same time, factors such as cost and economy also need to be considered.

[0063] In the present application, step 6, for the sample to be tested, the Z t,j and D i,j feature grouping detection value feature depth vector is obtained based on the same format, and is introduced into the learning device constructed in step 5, and the detection result is finely distinguished according to the predicted label. Among them, the obtaining method of the Z t,j and D i,j feature grouping detection value feature depth vector can be obtained according to the methods of steps 1 to 4.

[0064] According to the specific embodiment of the present application, a method for detecting fetal aneuploidy and copy number variation in pregnant women's plasma free DNA includes the following steps:

[0065] 1. Based on the conventional non-targeted whole genome sequencing method, the characteristic vector / matrix carrying known differences biologically existing in different tissue cell sources is obtained, including but not limited to insert length and sequence end base distribution frequency. This method has no requirements for library construction method and sequencing platform, and the extraction method of characteristic vector / matrix includes but is not limited to methylation sequencing, double-end short sequence sequencing and single-end full-length sequencing, etc., in order to obtain the characteristic vector / matrix carrying specific traceability as the input data of the device;

[0066] 2. The sequencing depth information D i (the i-th single counting window on the genome) in different genomic intervals in the sequencing data is extracted as input data. The number of counting windows has a certain logical relationship with the sequencing depth and the variation size to be distinguished, and at least the number of stable templates in the unit interval should be sufficient to ensure the stability and accuracy of the method, and since the content of fetal free DNA in plasma is relatively low, the number of sequencing templates in the unit interval should not be less than the inverse of the average fetal concentration in theory;

[0067] 3. For each sequencing template, the characteristic vector / matrix extracted from the sequencing data is grouped, and the grouping method includes but is not limited to known typical characteristic values or limited clustering grouping based on unsupervised classifier. Further, the depth information D ij , j∈(1,21......N), S j is different template groups, there can be an intersection between different groups, and D ij is S jthe genomic depth information in the i-th counting window under the S

[0068] 4. By different D j grouping, the detection of aneuploidy and copy number variation in the whole genome range can use conventional analysis methods to output the detection values corresponding to various common aneuploidy and copy number variation, and the detection value Z t,j As the input data of the device, t represents different detection targets such as trisomy of chromosome 13 or a certain copy number variation syndrome;

[0069] 5. By the method of machine learning, a training set of a machine learning device based on the feature grouping detection value {Z t,j} is constructed using a certain number of samples with known true results, and the diagnosis results of known detection targets are used as known labels, including but not limited to fetal / mother / chimeric / negative, etc., for classification and evaluation of the detection results of the to-be-detected sample;

[0070] 6. For a new to-be-detected sample, the feature grouping detection value feature depth vector based on the same format Z t,j and D ij is counted, and is introduced into the constructed learning device, and the detection result is further finely distinguished according to the predicted label.

[0071] The second aspect of the present application provides a device for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA, comprising a data acquisition unit, a window division unit, a grouping unit, an aneuploidy and copy number variation detection unit, a modeling unit and a to-be-detected sample result output unit; wherein,

[0072] The data acquisition unit obtains sequencing data based on non-targeted whole genome sequencing, so as to group the sequencing templates and generate analyzable sub-libraries;

[0073] The window division unit is used for extracting the sequencing depth information D i of different genomic intervals in the sequencing data, wherein D i is the i-th unit counting window on the genome;

[0074] The grouping unit is used for grouping different groups S j for each sequencing template according to the feature vector / matrix extracted from the sequencing data, j ∈ (1, 2, 3, …, N), and counting and extracting the depth information D i,j of the genomic interval from the different grouping sub-libraries, D i,j is the genomic depth information in the i-th counting window under the S j grouping;

[0075] The aneuploidy and copy number variation detection unit is configured to output detection values Z corresponding to aneuploidy and copy number variations of different groups by detecting aneuploidy and copy number variations in a whole genome range of different groups. t,j , t represents different detection targets;

[0076] The modeling unit is configured to use the samples with known true results to construct a training set of a machine learning device based on feature group detection values {Z t,j} by a machine learning method, and obtain a learning device.

[0077] The sample result output unit is configured to, for a sample to be tested, statistically obtain a feature group detection value feature depth vector based on the same format of Z t,j and D i,j , and import the feature group detection value feature depth vector into the learning device constructed by the modeling unit, and finely distinguish the detection results according to the predicted labels.

[0078] In the present application, the feature vector / matrix is selected from the group consisting of an insertion fragment length and a sequence end base distribution frequency.

[0079] In the present application, the method of non-targeted whole genome sequencing is at least one selected from the group consisting of methylation sequencing, double-end short sequence sequencing and single-end full-length sequencing.

[0080] In the present application, there is no requirement for the library construction method and the sequencing platform.

[0081] In the present application, the number of counting windows, the sequencing depth and the variation size to be distinguished have a logical relationship, and at least the number of stable templates in a unit interval should be sufficient to ensure the stability and accuracy of the method.

[0082] In the present application, for each sequencing template, the feature vector / matrix extracted from the sequencing data is grouped S j , j ∈ (1, 2, 3,..., N). The grouping method can be determined according to the analysis needs, for example, a typical feature value-based grouping or a limited clustering grouping based on an unsupervised classifier. In the present application, there can be an intersection between different groups, for example, but not limited to, S1 = 1-128, S2 = 109-166, and S3 = 140-223.

[0083] In the present application, the depth information D i,j of the genomic interval from different groups is statistically obtained and extracted. i,j D jThe genomic depth information of the i-th counting window under the grouping, for example, D 1,1 , D 1,2 , D 1,3 , etc.

[0084] In the present application, the depth information is selected from at least one of Reads, Unique reads, Mapability, Genomic GC, Reads GC, Unique reads GC.

[0085] In the present application, the detection of aneuploidy and copy number variation in the whole genome range is carried out by different groupings to output the detection value Z t,j corresponding to the aneuploidy and copy number variation, where t represents different detection targets. The detection target can also be understood as a detection object, for example, but not limited to, trisomy 13, trisomy 18, trisomy 21, any copy number variation syndrome. Correspondingly, Z t,j may be expressed as Z 13,j , Z 18,j , Z 21,j , Z CNV,j , etc.

[0086] In the present application, the method for constructing the training set of the machine learning device based on the feature grouping detection value {Z t,j} includes inputting data, processing the input data by constructing a classifier model to obtain the determination result of the detection target, then performing label classification to obtain the known label.

[0087] In the present application, the input data is selected from D i , S j , D i,j , Z t,j .

[0088] In the present application, the classifier model is selected from at least one of logistic regression, random forest, support vector, linear regression, decision tree and neural network. According to the information of a certain number of known samples of fetal origin positive and fetal origin negative (false positive), the multi-dimensional target detection value vector (D i,j , Z t,j), while other statistical quantities that have qualitative predictive power on clinical targets other than the above-mentioned detection values and fetal concentrations can also be included as input vectors, and the training step of machine learning is performed. The preset type number can be a dichotomy corresponding to 0 or 1 of known positive and negative. The device does not limit the selection of the classifier, and common classifiers that can process numerical variables such as logistic regression, random forest, and support vector machine can be used. The optimal training model is selected for evaluation of the to-be-tested sample by a method similar to cross-validation. Meanwhile, the internal structure and weights of the classifier model can be selected or constructed for visualization of intermediate variables.

[0089] In the present application, the type of label classification is selected from negative and positive; preferably, the type of label classification is selected from negative, fetal positive and maternal positive; more preferably, the type of label classification is selected from negative, fetal positive, maternal positive and chimera. For example, after verification by clinical amniocentesis, the fetus is not trisomy 13 syndrome, but the corresponding placenta of the sample has trisomy 13 chimera, so the sample is given a classification label of 'chimera or fetal negative' for model training. In the present application, 'fetal positive' and'maternal positive' classification labels are also included. Fetal positive refers to judging the corresponding target as positive by the detection value and verifying that the fetus has the abnormality of the target by puncture; maternal positive refers to judging the corresponding target as positive by the detection value, but verifying that neither the fetus nor the placenta has the abnormality of the target but the mother's somatic cells have the abnormality of the target.

[0090] In the present application, a certain number of samples with known true results are repeatedly subjected to steps 1 to 4, and then the multidimensional target detection value vectors (D i,j , Z t,j ) of the samples are grouped according to the inserted fragments and introduced into the pre-trained model in step 5 for classification prediction, and finally a learning device is obtained. In the present application, the number can be determined as needed. The increase in the number has a positive correlation with the accuracy of the results, but at the same time, factors such as cost and economy also need to be considered.

[0091] The third aspect of the present application provides the use of the above-mentioned method for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA or the above-mentioned device for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA in the field of genetic testing.

[0092] The beneficial effects of the present application are:

[0093] (1) The method and device of the present application can accurately distinguish between fetal variations and maternal variations.

[0094] (2) The method and device provided by the application have strong compatibility, wide applicability, low cost, do not need to design probes, and only need ultra-low sequencing depth to realize accurate distinction and detection of fetal free DNA genome aneuploidy and copy number variation.

[0095] The application will be described below with reference to specific examples. It should be noted that the examples are merely illustrative and should not be construed as limiting the application.

[0096] Example 1

[0097] A method for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA:

[0098] Step 1: Collect prenatal peripheral blood samples with known true results, extract free DNA using 5 mL of peripheral blood to obtain sample ALB73W04375. Use the NIPT library building kit (National Medical Product Registration No. 20173400331) to build a library. The sequencing platform uses a NextSeq 550AR gene sequencer in PE75 mode to generate about 30M of sequencing data.

[0099] Step 2: Extract the sequencing depth information D of different genomic intervals in the sequencing data i , wherein D i is the i-th unit count window on the genome. The depth information includes Unique reads.

[0100] Step 3: Estimate the insert size of each sequencing template by aligning the paired-end sequences, and use the length of the insert as the feature vector of the template grouping S j . According to the grouping with certain biological significance, wherein the templates derived from the fetus are shorter, each template is divided into S1=1-128, S2=109-166, and S3=140-223 three different groups with intersection according to the feature vector of the insert. The depth information D i,j of the genomic interval from different groupings is counted and extracted, D i,j is the genomic depth information in the i-th count window under S j grouping, and the Unique reads of the total library of sample ALB73W04375 in the first unit interval on chr1 and the Unique reads of the sub-libraries of the corresponding three different fragment groups in the same unit interval are shown in Table 1.

[0101] Table 1

[0102] Chromosome Unique reads [D1] chr1 1021

[00011] D 1,1 ]] chr1 125 <![CDATA[D 1,2 ]]> chr1 100 D 1,3 ]]> chr1 49

[0103] Step 4, for each group of sequencing templates generated in step 3, the total amount D of sequencing templates in the genomic unit interval is determined using the conventional analysis method BMMPA registration certificate number 20192210692 i,j Modeling is performed to generate a detection value Z of fetal chromosomal aneuploidy t,j , and other and detection-related fetal quantity characteristics (fetal concentration). The Z of sample ALB73W04375 13,j and Z fc,j As shown in the following table, according to the instructions of the library construction kit (national equipment registration number 20173400331), the Z of the S1 group 13 is greater than 4, which belongs to a positive result, but the corresponding detection values of the other two subsets are all below the gray zone 3, which do not belong to a positive result, as shown in the following table 2.

[0104] Table 2

[0105] t\j [S1 = 1 ~ 128] [S2 = 109-166] [S3 = 140-223] Z 13,j ]]> 4.079794437 2.971799521 1.302150172 Z fc,j ]]> 0.32948574 0.278243513 0.066748496

[0106] wherein Z 13,j represents the detection value of the jth group of trisomy 13, Z fc,j represents the detection value of the jth group of fetal concentration.

[0107] Step 5, according to the information of known sample fetal positive and fetal negative (false positive), using the multi-dimensional target detection value vector (D i,j , Z t,j ) of sample ALB73W04375, a training step of machine learning is performed. The preset type number can be a binary method of 0 or 1 corresponding to known positive and negative, and a classifier for processing numerical variables using logistic regression. The optimal training model is selected for evaluation of the to-be-tested sample by cross-validation. At the same time, the internal structure and weight of the classifier model can be selected or constructed for visualization of intermediate variables. Sample ALB73W04375 is verified after clinical amniocentesis, and the fetus is not trisomy 13 syndrome, but the corresponding placenta of this sample exists trisomy 13 chimera, so this sample is given a classification label of 'chimera or fetal negative' for model training.

[0108] In addition, according to the information of a certain number of known sample fetal positive and fetal negative (false positive), using the multi-dimensional target detection value vector (D i,j , Z t,j ) of the above sample, two classification labels of 'fetal positive' and'maternal positive' are obtained, fetal positive refers to judging the corresponding target positive by the detection value and verifying the existence of the target abnormality of the fetus by puncture; maternal positive refers to judging the corresponding target positive by the detection value, but verifying that neither the fetus nor the placenta exists the target abnormality but the mother's somatic cell exists the target abnormality by puncture.

[0109] The three classification labels of 'chimeric or fetal negative', 'fetal positive' and'maternal positive' are established, and finally the learning device is obtained.

[0110] Step 6, the 12 samples to be tested are grouped by statistics based on the same format Z t,j and D i,j characteristic grouping detection value feature depth vector, and are imported into the learning device constructed in step 5, and the detection results are finely distinguished according to the predicted labels, specifically, the 12 samples to be tested repeat the steps of steps 1-4 above, and then the multi-dimensional target detection value vector (D i,j , Z t,j ) of the inserted fragment grouping is imported into the pre-trained model (learning device) in step 5 for classification prediction. The results are shown in Table 3 below.

[0111] Table 3

[0112] Predicted\True value Fetal positive Maternal positive Mosaic or fetal negative Fetal positive 5 0 0 Maternal positive 0 1 0 Mosaic or fetal negative 0 0 6

[0113] Then the results of the 12 samples to be tested are compared with the clinical diagnosis results, and the consistency is 100%.

[0114] Comparative Example 1

[0115] A method for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA:

[0116] Step 1, collect prenatal peripheral blood samples with known true results, extract free DNA from 5ml peripheral blood to obtain sample ALB73W04375. Use the NIPT library construction kit (National Medical Device Registration No. 20173400331) to construct the library. The sequencing platform uses NextSeq 550AR gene sequencer PE75 mode to generate about 30M sequencing templates.

[0117] Step 2, by the conventional analysis method BMMPA registration certificate No. 20192210692, the total amount of sequencing templates in the genomic unit interval D i is modeled to generate the detection value Z t of fetal chromosome aneuploidy. The total library Z 13 value of sample ALB73W04375 is greater than 4 (see Table 4), which belongs to a positive result according to the instructions of the library construction kit (National Medical Device Registration No. 20173400331).

[0118] Table 4

[0119] Z 13 ]]> Z fc ]]> ALB73W04375 9.10801616 0.461583297

[0120] Wherein, Z 13 represents the detection value of trisomy 13, Zfc a detection value representing a fetal concentration.

[0121] Comparative Example 2

[0122] A method for detecting fetal aneuploidy and copy number variation in maternal plasma free DNA:

[0123] Step 1, for the 12 prenatal peripheral blood samples to be tested in Example 1, 5ml of peripheral blood was used to extract free DNA. The NIPT library construction kit (National Medical Product Registration No. 20173400331) was used to construct the library. The sequencing platform uses the NextSeq 550AR gene sequencer PE75 mode to generate about 30M sequencing templates.

[0124] Step 2, by the conventional analysis method BMMPA registration certificate No. 20192210692, the total amount of sequencing templates in the genomic unit interval D i Modeling was performed to generate a detection value Z t of fetal chromosomal aneuploidy. The results are shown in Table 5 below.

[0125] Table 5

[0126]

[0127] From the results of Example 1 and Comparative Example 1, it can be seen that ALB73W04375 is detected by the conventional method, and the result is positive; while using the method of the present application, it is classified as a chimera or fetal negative, not a positive result.

[0128] From the results of Example 1 and Comparative Example 2, it can be seen that the results of samples reported as positive (fetal) by the traditional method (Comparative Example 2) are evaluated for the 12 samples to be tested, and the application method can further accurately subdivide the positive results, and a part of the known false positive results are further classified as placental chimeras (chimeras or fetal negative) and maternal positive subtypes, which are two non-fetal positive subtypes, which are consistent with the clinical diagnosis results 100%.

[0129] The basic principles of the present application are described above in conjunction with specific embodiments, but it should be noted that the advantages, advantages, effects, etc. mentioned in the present application are only examples and not limitations, and these advantages, advantages, effects, etc. cannot be considered as each embodiment of the present application must have. In addition, the above specific details disclosed are only for the purpose of example and for the purpose of understanding, and are not limited.

Claims

1. A method for detecting fetal aneuploidy and copy number variation in maternal plasma cell-free DNA, comprising the following steps: Step 1. Based on non-targeted whole genome sequencing, obtaining sequencing data in order to group sequencing templates and produce analyzable sub-libraries; Step 2, extracting sequencing depth information D of different genomic intervals in the sequencing data i where D i is the i-th unit count window on the genome; Step 3, for each sequencing template, grouping according to the feature vector / matrix extracted from the sequencing data, obtaining different groups S j , j e (1, 2, 3,..., N); and counting and extracting the depth information D of the genomic interval from different group sub-libraries i,j , D i,j is the genomic depth information of the i-th counting window under S j group Step 4, outputting the detection value Z corresponding to the aneuploidy and copy number variation of different groups by detecting the aneuploidy and copy number variation in the whole genome range of different groups t,j , t represents different detection targets; Step 5, using the samples with known true results, construct a training set for the machine learning device based on the feature grouping detection values {Z t,j} to obtain a learning device; Step 6, for the sample to be tested, the feature depth vector of the feature group detection value based on the same format Z t,j and D i,j is detected and introduced into the learning device constructed in step 5, and the detection result is finely distinguished according to the predicted label.

2. The method of claim 1, wherein, The method of non-targeted whole genome sequencing is selected from at least one of methylation sequencing, paired-end short sequence sequencing and single-end full-length sequencing.

3. The method of claim 1, wherein, The depth information is selected from at least one of Reads, Unique reads, Mapability, Genomic GC, Reads GC, Unique reads GC.

4. The method of claim 1, wherein, The feature vector / matrix is selected from at least one of insert size, sequence end base distribution frequency.

5. The method of claim 1, wherein, The construction is based on feature grouping detection value {Z t,j The method for constructing a training set of a machine learning device of a feature grouping detection value {Z includes: inputting data, processing the input data by a classifier model to obtain a determination result of detecting a target, and then performing label classification to obtain a known label.

6. The method of claim 5, wherein, The input data is selected from D i , S j , D i,j , Z t,j .

7. The method of claim 5, wherein, The classifier model is selected from at least one of logistic regression, random forest, support vector, linear regression, decision tree and neural network.

8. The method of claim 5, wherein, The type of label classification is selected from negative and positive.

9. The method of claim 8, wherein, The type of label classification is selected from negative, fetal positive and maternal positive.

10. The method of claim 8, wherein, The type of label classification is selected from negative, fetal positive, maternal positive and chimera.

11. An apparatus for detecting fetal aneuploidy and copy number variation in maternal plasma cell-free DNA, comprising a data acquisition unit, a window division unit, a grouping unit, an aneuploidy and copy number variation detection unit, a modeling unit and a to-be-tested sample result output unit; Wherein, The data acquisition unit, based on non-targeted whole genome sequencing, is used to obtain sequencing data in order to group sequencing templates and produce analyzable sub-libraries; The division window unit is configured to extract sequencing depth information D of different genomic intervals in the sequencing data i wherein D i is the i-th unit count window on the genome; The grouping unit is configured to group, for each sequencing template, according to a feature vector / matrix extracted from the sequencing data, to obtain different groups S j , j e (1, 2, 3,..., N); and count and extract depth information D of the genomic interval from different group sub-libraries i,j , D i,j is the genomic depth information of the i-th counting window under the S j grouping The aneuploidy and copy number variation detection unit is configured to output detection values Z corresponding to aneuploidy and copy number variations of different groups by detecting aneuploidy and copy number variations in a whole genome range of the different groups t,j , and t represents different detection targets. The modeling unit is configured to use a sample with a known real result to construct a training set of a machine learning device based on a feature grouping detection value {Z t,j} to obtain a learning device. The to-be-tested sample result output unit is configured to, for a to-be-tested sample, detect a feature depth vector based on a same-format Z t,j and D i,j characteristic grouping detection value, and import the feature depth vector into a learning device constructed by a modeling unit, and then finely distinguish the detection result according to a predicted label.

12. The use of the method for detecting fetal aneuploidy and copy number variation in maternal plasma cell-free DNA according to any one of claims 1-10 or the apparatus for detecting fetal aneuploidy and copy number variation in maternal plasma cell-free DNA according to claim 11 in the field of genetic testing.

Citation Information

Patent Citations

  • CNV testing device

    CN109979529A

  • Method and device for detecting chromosomal variations

    CN110268044A