A method and system for detecting fetal chromosomal abnormalities
Through the deep mixed model, the free nucleic acid fragment sequencing data and clinical feature data of pregnant women are combined with the characteristics matrix and machine learning detection, which solves the problems of low accuracy and insufficient specificity of fetal chromosomal abnormality detection in the prior art, and improves the detection accuracy of 18-trisomal syndrome and 13-trisomal syndrome.
Patent Information
- Application Number
- CN202080107528.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-11-27
AI Technical Summary
The existing fetal chromosomal abnormality detection methods have problems such as low detection accuracy, high false positive rate, high false negative rate and insufficient detection accuracy for some syndromes, especially poor detection accuracy for 18-Trisomy syndrome and 13-Trisomy syndrome.
Using a deep mixed model, by obtaining free nucleic acid fragment sequencing data and clinical phenotypic feature data of pregnant women, window division is performed to generate a sequence feature matrix, and sequence feature vectors are extracted in combination with machine learning model, and combined with pregnant women's phenotypic feature vectors to form a combined feature vector, and input the classification detection model for fetal chromosome abnormality detection.
The accuracy of fetal chromosomal abnormality detection was improved, and the false positive and false negative rates were reduced, especially the detection performance of 18-trisomy syndrome and 13-trisomy syndrome was significantly improved.
Smart Images

Figure CN116648752B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and more particularly to a method and system for detecting fetal chromosomal abnormalities. Background Art
[0002] Chromosomal aneuploidies are serious genetic disorders in which the number of individual chromosomes in the fetus increases or decreases, thereby affecting normal gene expression. These disorders primarily include trisomy 21, trisomy 18, trisomy 13, and 5p-syndrome. Chromosomal aneuploidies carry a high risk of death and disability, and there are no effective treatments. Currently, prenatal screening and diagnosis are the primary means of reducing the birth rate of children with chromosomal aneuploidies.
[0003] Traditional non-global chromosomal testing primarily includes non-invasive prenatal screening based on ultrasound or serological screening, and invasive prenatal diagnosis using sampling. Ultrasound-based prenatal screening measures the thickness of the nuchal translucency (NT) at 10-14 weeks of gestation to determine whether the fetus has chromosomal abnormalities. An NT greater than 3 mm is generally considered to indicate a higher risk of fetal chromosomal abnormalities. Serological-based prenatal screening measures maternal serum alpha-fetoprotein (AFP) and human chorionic gonadotropin (HCG) concentrations at 13-16 weeks of gestation, calculating the risk of fetal chromosomal abnormalities based on the maternal due date, age, and gestational age at the time of blood sampling. Invasive prenatal diagnosis typically involves obtaining fetal samples from amniocentesis, cord blood puncture, or chorionic villus sampling at 16-24 weeks of gestation to detect chromosomal abnormalities. Combined screening based on ultrasound and serology does not directly measure fetal chromosomes but rather estimates fetal risk. Its accuracy is 50% to 95%, with a false-positive rate as high as 3% to 7% [1, 2]. Invasive sampling methods can directly and accurately diagnose fetal aneuploidy and are the "gold standard" for detecting and diagnosing fetal chromosomal abnormalities. However, this method results in a certain miscarriage rate (0.5% to 2%), and pregnant women with infectious diseases such as hepatitis B are not suitable for invasive sampling (such as amniocentesis) due to the risk of fetal infection. In addition, amniocentesis requires ultrasound guidance, takes a long time, and requires high technical skills from the operator.
[0004] With the discovery of fetal cell-free DNA (cfDNA) in maternal peripheral blood, the maturity of next-generation high-throughput sequencing (NGS) technology, the significant reduction in sequencing costs, and the development of information analysis technology, non-invasive prenatal testing (NIPT) based on NGS technology is becoming the most widely used prenatal screening method for fetal chromosomal aneuploidy. NIPT technology uses maternal peripheral blood and sequences the free DNA (including fetal cell-free DNA) in maternal peripheral plasma using NGS technology. Combined with bioinformatics analysis, it obtains fetal genetic information, thereby detecting whether the fetus has chromosomal abnormalities such as trisomy 21 (Down syndrome), trisomy 18 (Edwards syndrome), and trisomy 13 (Patau syndrome).
[0005] NIPT technology has high sensitivity and specificity (sensitivities of T21, T18, and T13 are all above 99%) and a low false-positive rate (<0.1%), and is now widely used in clinical practice [3-5]. NIPT technology can reduce the false-positive rate of serological screening and avoid the risks of intrauterine infection and miscarriage associated with invasive prenatal diagnostic procedures (such as amniocentesis and chorionic villus sampling). It is a highly safe, non-invasive prenatal screening technology for the first and second trimesters.
[0006] Conventional NIPT based on NGS technology detects fetal chromosomal abnormalities by calculating the read count of sequencing and using the Baseline Z-test [6]. Its principle is as follows: First, collect maternal peripheral blood samples at 12 - 22 weeks of gestation, use NGS technology to sequence the cell-free DNA in the peripheral blood samples and align the obtained sequencing reads with the human reference genome sequence (while correcting the read count for GC content); then count the number of uniquely aligned reads for each chromosome and calculate the proportion of the number of uniquely aligned reads for each chromosome in the sample to the total number of uniquely aligned reads for all chromosomes in the sample; next, subtract the mean of the proportion of uniquely aligned reads for the corresponding chromosome in the control group sample (i.e., normal sample) from the proportion of uniquely aligned reads for the corresponding chromosome in the test sample, and then divide by the standard deviation of the proportion of uniquely aligned reads for the corresponding chromosome in the control group sample to obtain the Z value of the chromosome of the test sample; finally, compare the Z value with a given threshold. If it exceeds the threshold, it is judged that there is a high risk of trisomy syndrome, otherwise it is judged as a low risk. Here, the mean of the number of uniquely aligned reads for each chromosome in the normal samples of the control group is the Baseline Value. Thus, the more normal samples in the control group, the more accurate the mean and standard deviation of the proportion of uniquely aligned reads, and the more accurate the obtained Z value. The given threshold for the Z value is generally 3, which is defined statistically, that is, 99.9% deviation from the normal expectation.
[0007] Different statistical hypothesis tests can be selected according to different reference values. For example, the correlation analysis and T-test adopted in reference [7] use the median of the read counts of each chromosome in a fixed-size window in the sample to represent the read count of that chromosome, and the median of the read counts of all chromosomes in the sample to represent the read count of the sample. Then, the read count of each chromosome is divided by the read count of the sample to obtain the normalized read count of the corresponding chromosome. Finally, the confidence interval is calculated using the normalized read counts of each chromosome of all samples in the control group. When the value of the sample to be tested is not within this confidence interval, it is an abnormal sample. Another example is that in reference [8], for samples with known karyotypes, a reference chromosome with a similar GC content is found for the chromosome of interest (such as chromosome 21), and after performing a Z-test using the read count of this reference chromosome as the reference value, the accuracy of detecting abnormalities in the chromosome of interest in samples with known karyotypes reaches the maximum. This reference chromosome used as the reference value is the so-called internal chromosome. Yet another example is that reference [9] proposed the Noninvasive Fetal Trisomy (NIFTY) method. In addition to comparing the read counts of chromosomes with those of normal samples in the control group, this method also considers the proportion of fetal free DNA. This method uses binary hypothesis testing, log-likelihood ratio, and the FCAPS binary segmentation algorithm to determine the test results. NIFTY is a genome-wide method. This method has been verified by a large population and has high accuracy, but the process is relatively complex. The above-mentioned statistical hypothesis testing methods based on read counts (Z-test or T-test) are the core of current NIPT analysis.
[0008] The above-mentioned analysis methods of statistical hypothesis testing based on the number of reads (such as Z-test) are the current mainstream NIPT analysis methods. However, these analysis methods have obvious limitations: (1) The current NIPT analysis methods will cause deviations in the sequencing read distribution of individual samples, resulting in fluctuations in different situations in the calculation of Z values, thus affecting the final result determination and related performance indicators; (2) The current NIPT analysis methods highly rely on the proportion of fetal free DNA in maternal peripheral plasma. Due to the large individual differences among pregnant women, too low a proportion of fetal free DNA (<4%) will increase the risk of false negative error detection; (3) The current NIPT analysis methods perform well in the detection of trisomy 21 syndrome. However, due to individual differences among pregnant women and the deviation of GC content on different chromosomes, their accuracy in the detection of trisomy 18 syndrome and trisomy 13 syndrome is poor; (4) The current NIPT analysis methods mainly detect common trisomy syndromes represented by Down syndrome, and have limited clinical effects on the detection of chromosomal microdeletion and microduplication syndromes with relatively high combined incidence, such as DiGeorge Syndrome, Prader-Willi Syndrome, etc.
[14] .
[0009] In addition, some people have proposed a new technology based on machine learning models to detect chromosomal abnormalities using NIPT sequencing results. For example, reference
[10] proposed a method using support vector machines (SVM) to assist NIPT judgment. This method obtains 6 different Z-value results by calculating different baseline values, and also adds the clinical indications of the sample to train the SVM model for chromosomal abnormality judgment. For another example, reference
[11] designed a Bayesian method for chromosomal abnormality judgment. This method uses the prior information of the fetal free DNA ratio, uses the Hidden Markov Model (HMM) to eliminate the interference of the population level and maternal CNV, and performs GC content correction. The Bayes factor is calculated by combining the Z test likelihood value and the prior value of the fetal free DNA ratio inferred from the sex chromosome content. At the same time, multiple risk factors such as maternal age are included in the prior probability to correct the Bayes factor. The z-value and Bayes factor are combined to evaluate whether the chromosome is abnormal. For another example, patent disclosure
[12] proposed using NIPT sequencing results to train a simple convolutional neural network model to detect chromosome copy number variation and chromosome aneuploidy abnormalities. For example, the patent disclosure
[13] proposes to first separate fetal free DNA and maternal free DNA from a peripheral blood sample, amplify multiple single nucleotide variation (SNV) loci from the separated free DNA, sequence the amplified products to determine the genetic sequencing data or genetic array data of the multiple SNV loci, and then train an artificial neural network model based on the genetic sequencing data or genetic array data of these loci to detect the ploidy level state (Ploidy State) of individual chromosomes, tissue canceration (Cancer State) or organ transplant rejection (Transplatation Rejection State).
[0010] The aforementioned machine learning models based on NIPT sequencing results for detecting chromosomal abnormalities also have the following limitations: most of these methods are based on the number of reads in the sequencing data to calculate the features required for model training; most of them rely on the calculation of Z values; either the calculation is too complicated (such as reference
[11] ), or the model design is too simple (such as patent disclosure
[12] ), or genetic sequencing data or genetic array data based on SNV loci are required (such as patent disclosure
[13] ), and their clinical application prospects, model scalability and detection accuracy are limited; the detection accuracy needs to be improved. Summary of the Invention
[0011] In view of the problems in the existing technology for detecting chromosomal abnormalities, especially aneuploidy, in order to more effectively detect chromosomal abnormalities, the purpose of the present invention is at least to further improve the detection accuracy of chromosomal abnormalities based on a deep hybrid model.
[0012] Thus, in a first aspect, the present invention provides a method for detecting fetal chromosomal abnormalities, the method comprising:
[0013] (1) Obtaining sequencing data and clinical phenotypic characteristic data of cell-free nucleic acid fragments of a pregnant woman to be tested, wherein the sequencing data includes a plurality of reads, and the clinical phenotypic characteristic data of the pregnant woman to be tested forms a pregnant woman phenotypic characteristic vector;
[0014] (2) Dividing at least a part of a reference genomic chromosome into windows to obtain a plurality of sliding windows, counting the reads falling within the sliding windows, and generating a sequence characteristic matrix of the chromosomal sequence;
[0015] (3) Inputting the sequence characteristic matrix into a trained machine learning model to extract a sequence characteristic vector of the chromosomal sequence;
[0016] (4) Combining the sequence characteristic vector and the pregnant woman phenotypic characteristic vector to form a combined characteristic vector, inputting the combined characteristic vector into a classification detection model, and obtaining the fetal chromosomal abnormality condition of the pregnant woman to be tested.
[0017] In one embodiment, in (1), the cell-free nucleic acid fragments are from the peripheral plasma of the pregnant woman, the liver of the pregnant woman, and / or the placenta.
[0018] In one embodiment, in (1), the cell-free nucleic acid fragments are cell-free DNA.
[0019] In one embodiment, in (1), the sequencing data is from ultra-low depth sequencing; preferably, the sequencing depth of the ultra-low depth sequencing is 1×, 0.1× or 0.01×.
[0020] In one embodiment, in (1), the reads are aligned with a reference genome to obtain uniquely aligned reads (preferably with GC content correction); preferably, the subsequent steps are performed using the uniquely aligned reads (preferably the reads corrected for GC content).
[0021] In one embodiment, the GC content correction is performed according to the following steps:
[0022] a. First, randomly sample m fragments of length from a certain chromosome of the human reference genome;
[0023] b. Calculate the number of fragments Ni with GC content of i:
[0024]
[0025] where f(k) is the GC content of fragment k, and i represents the GC content (i = 0%, 1%,..., 100%);
[0026] c. Calculate the number of sequencing reads Fi with GC content i:
[0027]
[0028] where represents the GC content of fragment k, and F i represents the number of sequencing reads with GC content i and the starting position of the sequencing reads being the same as the starting position of this fragment;
[0029] d. Calculate the observed-expected ratio of GC content λ i :
[0030]
[0031] where r is the global scaling factor, and its definition is:
[0032]
[0033] e. Correct the number of sequencing reads:
[0034]
[0035] where R i represents the expected number of sequencing reads with GC content i after correction.
[0036] In one embodiment, in (1), the phenotypic characteristic data of the pregnant woman is selected from the combination of one or more of the following: age, gestational week, height, weight, BMI, results of antenatal biochemical tests, results of ultrasound examinations, and concentration of cell-free fetal DNA in plasma.
[0037] In one embodiment, in (1), the phenotypic characteristic data of the pregnant woman is subjected to outlier processing, missing value processing, and / or null value processing.
[0038] In one embodiment, in (1), the phenotypic data of the pregnant woman sample will be determined as outliers and these outliers will be set as null values if the following records occur:
[0039] a. x age < 10 or x age > 80;
[0040] b. x GW < 5 or x GW > 50;
[0041] c. x height < 40 or xheight > 300;
[0042] d.x weight < 10 or x weight > 200.
[0043] In one embodiment, missing values and null values are filled using the missForest algorithm.
[0044] In one embodiment, in (2), the chromosome is chromosome 21, chromosome 18, chromosome 13, and / or sex chromosome.
[0045] In one embodiment, (2) includes: (2.1) using a window of length b with a step size of t to perform overlapping sliding on a chromosome sequence of length L on the reference genome to obtain a number of sliding windows, where b is a positive integer and b = [10000, 10000000], t is any positive integer, L is a positive integer, and L ≥ b; (2.2) counting the reads falling into each of the sliding windows to generate a sequence feature matrix of the chromosome sequence.
[0046] In one embodiment, in (2), the sequence feature matrix includes the number of reads in the sliding window, base quality, and alignment quality.
[0047] In one embodiment, the base quality includes the mean, standard deviation, skewness, and / or kurtosis of the base quality.
[0048] In one embodiment, the alignment quality includes the mean, standard deviation, skewness, and / or kurtosis of the alignment quality.
[0049] In one embodiment, in (2), the sequence feature matrix is:
[0050] X = (x ij ) h×w
[0051] where h represents the number of times the window slides, w represents the number of sequence features in a single sliding window, and x ij represents the jth sequence feature value in the ith sliding window.
[0052] In one embodiment, in (3), the sequence feature matrix is normalized.
[0053] In one embodiment, in (3), the sequence feature matrix is normalized using formula (I):
[0054]
[0055] where, is the sequence feature matrix of sample k after standardization, represents the jth sequence eigenvalue in the i-th sliding window of sample k, μ i,j and σ i,j They represent the mean and standard deviation of the j-th sequence feature value in the i-th sliding window of all samples respectively.
[0056] In one embodiment, in (3), the trained machine learning model is a neural network model or an AutoEncoder model; preferably, the neural network model is a deep neural network model; more preferably, the neural network model is a deep neural network model based on 1-dimensional convolution.
[0057] In one embodiment, the structure of the deep neural network model includes:
[0058] An input layer, configured to receive the sequence feature matrix;
[0059] A front module, connected to the input layer, for performing a first convolution and activation operation on the sequence feature matrix from the input layer to obtain a feature map;
[0060] The core module is connected to the front module and is used to further abstract and extract features from the feature map from the front module, thereby enhancing the neural network's expressive power by effectively increasing the depth of the neural network model;
[0061] A post-module, connected to the core module, for performing feature abstraction representation on the feature map from the core module;
[0062] The first global average pooling layer is connected to the post-module and is used to vectorize the feature map represented by the feature abstraction and output a sequence feature vector of the chromosome sequence.
[0063] In one embodiment, the front module includes:
[0064] (I) 1D Convolution;
[0065] (II) Batch Normalization layer, connected to the 1D convolutional layer in (I);
[0066] (III) ReLU activation layer, connected to the batch normalization layer described in (II).
[0067] In one embodiment, the core module is composed of one or more residual submodules with the same structure, wherein the output of each residual module is the input of the next residual module.
[0068] In one embodiment, the residual sub-module includes:
[0069] (A) A pre-module of the core module, each sub-module including a 1D convolutional layer, a Dropout layer connected to the 1D convolutional layer, a batch normalization layer connected to the Dropout layer, and a ReLU activation layer connected to the batch normalization layer;
[0070] (B) A first 1D Average Pooling layer, connected to the pre-module of the core module in (A);
[0071] (C) A Squeeze-Excite module (SE module), and / or a Spatial Squeeze-Excite module (sSE module), connected to the first 1D Average Pooling layer in (B);
[0072] (D) A first Add layer, connected to the SE module and / or sSE module in (C);
[0073] (E) A second 1D Average Pooling layer, connected to the ReLU activation layer in the pre-module;
[0074] (F) A second Add layer, connected to the first Add layer in (D) and the second 1D Average Pooling layer in (E).
[0075] In one embodiment, the SE module includes:
[0076] (a) A second global average pooling layer, connected to the first 1D Average Pooling layer in the residual sub-module;
[0077] (b) A Reshape layer, connected to the second global average pooling layer in (a), the Reshape layer outputting a feature map with a size of 1×f, where f is the number of 1D convolutional kernels;
[0078] (c) A first fully connected layer, connected to the Reshape layer in (b), the first fully connected layer outputting the number of neurons as where f is the number of 1D convolutional kernels, and r SE is the reduction rate of the SE module;
[0079] (d) A second fully connected layer, connected to the first fully connected layer in (c), the second fully connected layer outputting the number of neurons as f, where f is the number of 1D convolutional kernels;
[0080] (e) A Multiply layer, connected to the second fully connected layer in (d) and the first 1D Average Pooling layer in the residual sub-module.
[0081] In one embodiment, the sSE module comprises:
[0082] a. A 1×1 convolutional layer connected to the first 1D average pooling layer (B). The 1×1 convolutional layer uses the sigmoid function as the activation function.
[0083] b. A Multiply layer connected to the first 1-dimensional average pooling layer (B) and the 1×1 1-dimensional convolutional layer described in a.
[0084] In one embodiment, in (4), the sequence feature vector and the maternal phenotype feature vector are concatenated to obtain a combined feature vector.
[0085] In one embodiment, in (4), the combined feature vector x is subjected to a pyramidalization process:
[0086]
[0087] where x′ i is the ith sequence eigenvalue in the standardized combined eigenvector x, x i is the ith sequence eigenvalue in the combined eigenvector x, μ i is the mean of the eigenvalues of the i-th sequence in the combined eigenvector, σ i is the standard deviation of the eigenvalues of the i-th sequence in the combined eigenvector.
[0088] In one embodiment, in (4), the classification detection model is an ensemble learning model.
[0089] In one embodiment, the integrated learning model is an integrated learning model based on stacking or majority voting; preferably, the integrated learning model is one or more of the following: support vector machine model, naive Bayes classifier, random forest classifier, XGBoos and logistic regression.
[0090] In one embodiment, the chromosomal abnormality comprises at least one or more of the following: trisomy 21, trisomy 18, trisomy 13, 5p-syndrome, chromosomal microdeletions, and chromosomal microduplications.
[0091] In a second aspect, the present invention provides a method for constructing a classification detection model for detecting fetal chromosomal abnormalities, the method comprising:
[0092] (1) obtaining sequencing data and clinical phenotypic characteristic data of free nucleic acid fragments of a pregnant woman, wherein the sequencing data includes a plurality of read segments, and the chromosome status of the fetus of the pregnant woman is known, and the clinical phenotypic characteristic data of the pregnant woman form a pregnant woman phenotypic characteristic vector;
[0093] (2) dividing at least a portion of the reference genome chromosome into windows to obtain a plurality of sliding windows, counting the reads falling within the sliding windows, and generating a sequence feature matrix of the chromosome sequence;
[0094] (3) constructing a training data set using the sequence feature matrix and the fetal chromosome status, and training a machine learning model to extract the sequence feature vector of the chromosome sequence;
[0095] (4) Combining the sequence feature vector and the pregnant woman's phenotype feature vector to form a combined feature vector, and using the combined feature vector and the fetal chromosome status to train a classification model to obtain the trained classification detection model.
[0096] In one embodiment, the chromosomal status of the fetus of the pregnant woman is one or more of the following: normal diploidy, chromosomal aneuploidy, partial monosomy syndrome, chromosomal microdeletion and chromosomal microduplication.
[0097] In one embodiment, the chromosomal aneuploidy comprises at least one or more of the following: trisomy 21, trisomy 18, and trisomy 13.
[0098] In one embodiment, the partial monosomy syndrome comprises 5p- syndrome.
[0099] In one embodiment, the number of pregnant women is greater than 10, and the ratio of the number of normal diploid fetuses to the number of fetuses with chromosomal aneuploidy is 1 / 2 to 2.
[0100] In one embodiment, in (3), the training dataset is represented as:
[0101]
[0102]
[0103] Where N represents the number of training samples, and N is an integer ≥ 1; is the sequence feature matrix of training sample k after normalization, and k∈[1, N], i is an integer ≥1, and j is an integer ≥1.
[0104] For the same technical features as the first aspect of the present invention except for the trained machine learning model, the limitations of the first aspect of the present invention in the implementation scheme also apply hereto. In this aspect, the trained machine learning module includes an output layer. For example, the structure of the deep neural network model includes an output layer after the first global average pooling layer, and the output layer is connected to the first global average pooling layer, is a fully connected layer, has 1 output neuron, and is used to output chromosomal abnormalities.
[0105] In a third aspect, the present invention provides a system for detecting fetal chromosomal abnormalities, comprising:
[0106] a data acquisition module, configured to obtain sequencing data and clinical phenotypic characteristic data of free nucleic acid fragments of a pregnant woman sample to be tested, wherein the sequencing data includes a plurality of read segments, and the clinical phenotypic characteristic data of the pregnant woman sample to be tested forms a pregnant woman phenotypic characteristic vector;
[0107] A sequence feature matrix generation module is used to divide at least a portion of the reference genome chromosome into windows to obtain a plurality of sliding windows, count the reads falling within the sliding windows, and generate a sequence feature matrix of the chromosome sequence;
[0108] A sequence feature vector extraction module, configured to input the sequence feature matrix into a trained machine learning model to extract a sequence feature vector of the chromosome sequence;
[0109] The classification detection module is used to combine the sequence feature vector and the pregnant woman's phenotype feature vector to form a combined feature vector, input the combined feature vector into a classification detection model, and obtain the fetal chromosomal abnormality of the pregnant woman to be tested.
[0110] In one embodiment, the system further comprises an alignment module for aligning the reads of the sequencing data with a reference genome to obtain uniquely aligned reads.
[0111] In one embodiment, in the data acquisition module, the free nucleic acid fragments are derived from maternal peripheral plasma, maternal liver and / or placenta.
[0112] In one embodiment, in the data acquisition module, the free nucleic acid fragments are free DNA.
[0113] In one embodiment, in the data acquisition module, the sequencing data comes from ultra-low depth sequencing; preferably, the sequencing depth of the ultra-low depth sequencing is 1×, 0.1× or 0.01×.
[0114] In one embodiment, in the data acquisition module, the reads are aligned with the reference genome to obtain uniquely aligned reads (preferably GC content corrected); preferably, subsequent steps are performed using the uniquely aligned reads (preferably GC content corrected reads).
[0115] In one embodiment, GC content correction is performed as follows:
[0116] a. First, randomly sample m chromosomes of length fragments;
[0117] b. Calculate the number of fragments N with GC content i i :
[0118]
[0119] in f(k) is the GC content of fragment k, i represents the GC content (i = 0%, 1%, ..., 100%);
[0120] c. Calculate the number of sequencing reads F with GC content i i :
[0121]
[0122] in represents the GC content of fragment k, F i Indicates the number of sequencing reads with a GC content of i and the same sequencing read start point as the fragment start point;
[0123] d. Calculate the observed-to-expected ratio λ of GC content i :
[0124]
[0125] Where r is the global scaling factor, which is defined as:
[0126]
[0127] e. Correction of sequencing read numbers:
[0128]
[0129] Where R i represents the expected number of sequencing reads with GC content i after correction.
[0130] In one embodiment, in the data acquisition module, the phenotypic characteristic data of the pregnant woman is selected from the combination of one or more of the following: age, gestational week, height, weight, BMI, prenatal biochemical test results, ultrasound examination diagnosis results, and the concentration of cell-free fetal DNA in plasma.
[0131] In one embodiment, in the data acquisition module, the phenotypic characteristic data of the pregnant woman is processed for outliers, missing values, and / or null values.
[0132] In one embodiment, in the data acquisition module, the phenotypic data of the pregnant woman's sample will be determined as outliers and these outliers will be set as null values if the following records occur:
[0133] a. x age <10 or x age >80;
[0134] b. x GW <5 or x GW >50;
[0135] c. x heigh <40 or x heigh >300;
[0136] d. x weight <10 or x weight >200.
[0137] In one embodiment, the missing values and null values are filled using the missForest algorithm.
[0138] In one embodiment, in the sequence feature matrix generation module, the chromosomes are chromosome 21, chromosome 18, chromosome 13, and / or sex chromosomes.
[0139] In one embodiment, in the sequence feature matrix generation module, the following operations are performed: (2.1) Using a window of length b with a step size of t to perform overlapping sliding on the chromosome sequence of length L on the reference genome to obtain a number of sliding windows, where b is a positive integer and b = [10000, 10000000], t is any positive integer, L is a positive integer, and L ≥ b; (2.2) Counting the reads falling into each of the sliding windows to generate the sequence feature matrix of the chromosome sequence.
[0140] In one embodiment, in the sequence feature matrix generation module, the sequence feature matrix includes the number of reads in the sliding window, base quality, and alignment quality.
[0141] In one embodiment, the base quality includes the mean, standard deviation, skewness, and / or kurtosis of the base quality.
[0142] In one embodiment, the alignment quality includes the mean, standard deviation, skewness, and / or kurtosis of the alignment quality.
[0143] In one embodiment, in the sequence feature matrix generation module, the sequence feature matrix is:
[0144] X = (x ij ) h×w
[0145] where h represents the number of times the window slides, w represents the number of sequence features within a single sliding window, and x ij represents the j-th sequence feature value in the i-th sliding window.
[0146] In one embodiment, in the sequence feature vector extraction module, the sequence feature matrix is normalized.
[0147] In one embodiment, in the sequence feature vector extraction module, the normalization of the sequence feature matrix is performed using formula (1):
[0148]
[0149] where, is the sequence feature matrix of sample k after normalization processing, represents the j-th sequence feature value in the i-th sliding window of sample k, μ i,j and σ i,j respectively represent the mean and standard deviation of the j-th sequence feature values in the i-th sliding window of all samples.
[0150] In one embodiment, in the sequence feature vector extraction module, the trained machine learning model is a neural network model or an AutoEncoder model; preferably, the neural network model is a deep neural network model; more preferably, the neural network model is a deep neural network model based on 1D convolution.
[0151] For the deep neural network model, the limitations in the embodiments of the first aspect of the present invention also apply here.
[0152] In one embodiment, in the classification and detection module, the sequence feature vector is concatenated with the pregnant woman phenotype to obtain a combined feature vector.
[0153] In one embodiment, in the classification and detection module, the combined feature vector x is normalized:
[0154]
[0155] where x′ iis the ith sequence eigenvalue in the standardized combined eigenvector x, x i is the ith sequence eigenvalue in the combined eigenvector x, μ i is the mean of the eigenvalues of the i-th sequence in the combined eigenvector, σ i is the standard deviation of the eigenvalues of the i-th sequence in the combined eigenvector.
[0156] In one embodiment, in the classification detection module, the classification detection model is an integrated learning model.
[0157] In one embodiment, the integrated learning model is an integrated learning model based on stacking or majority voting; preferably, the integrated learning model is one or more of the following: support vector machine model, naive Bayes classifier, random forest classifier, XGBoos and logistic regression.
[0158] In a fourth aspect, the present invention provides a system for constructing a classification detection model for detecting fetal chromosomal abnormalities, the system comprising:
[0159] a data acquisition module, configured to obtain sequencing data of cell-free nucleic acid fragments of a pregnant woman and clinical phenotypic characteristic data of the pregnant woman, wherein the sequencing data includes a plurality of read segments, the chromosome status of the fetus of the pregnant woman is known, and the clinical phenotypic characteristic data of the pregnant woman form a phenotypic characteristic vector of the pregnant woman;
[0160] A sequence feature matrix generation module is used to divide at least a portion of the reference genome chromosome into windows to obtain a plurality of sliding windows, count the reads falling within the sliding windows, and generate a sequence feature matrix of the chromosome sequence;
[0161] A sequence feature vector extraction module is used to construct a training data set using the sequence feature matrix and the fetal chromosome status, and train a machine learning model to extract the sequence feature vector of the chromosome sequence;
[0162] The classification detection model acquisition module is used to combine the sequence feature vector and the pregnant woman's phenotype feature vector to form a combined feature vector, and use the combined feature vector and the fetal chromosome situation to train the classification detection model to obtain the trained classification detection model.
[0163] In one embodiment, the system further comprises an alignment module for aligning the reads of the sequencing data with a reference genome to obtain uniquely aligned reads.
[0164] For the technical features that are the same as those of the third aspect of the present invention except for the trained machine learning model, the limitations in the embodiments of the third aspect of the present invention also apply thereto. In this aspect, the trained machine learning module includes an output layer. For example, the structure of the deep neural network model includes an output layer after the first global average pooling layer. The output layer is connected to the first global average pooling layer and is a fully connected layer with 1 output neuron for outputting chromosomal abnormalities. The method and model of the present invention are based on an innovative algorithm for sequencing data and do not rely on the Z-test, avoiding the clinical problem of being difficult to judge when the result score is in the "gray area" due to relying on a threshold. Moreover, as the number of samples (e.g., sample sequencing data and corresponding pregnant women's phenotype data) increases, the hybrid model proposed by the present invention can be automatically upgraded and optimized to improve the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0165] Figure 1 FIG. shows a flowchart of a method for detecting fetal chromosomal abnormalities based on a deep neural network hybrid model according to an embodiment of the present invention.
[0166] Figure 2 FIG. shows a feature matrix for calculating sequencing data according to an embodiment of the present invention.
[0167] Figure 3 FIG. shows a deep neural network structure according to an embodiment of the present invention.
[0168] Figure 4 FIG. shows a Squeeze-Excite module (SE module) according to an embodiment of the present invention.
[0169] Figure 5 FIG. shows a Spatial Squeeze-Excite module (sSE module) according to an embodiment of the present invention.
[0170] Figure 6 FIG. shows the filling of missing values in a phenotype data set according to an embodiment of the present invention.
[0171] Figure 7 FIG. shows a structure of an ensemble learning model based on Stacking according to an embodiment of the present invention.
[0172] Figure 8 FIG. shows an ROC curve of the training result of a 5-fold cross-validation of an ensemble learning model based on Stacking according to an embodiment of the present invention.
[0173] Figure 9 FIG. shows an ROC curve of the model evaluation based on a test set according to an embodiment of the present invention.
[0174] Figure 10 The Precision-Recall curve of the model evaluation based on the test set according to one embodiment of the present invention is shown.
[0175] Figure 11 FIG. 4 shows a confusion matrix diagram when the decision threshold according to one embodiment of the present invention is at a default value (ie, 0.5).
[0176] Figure 12 The precision and recall as a function of the threshold according to one embodiment of the present invention are shown.
[0177] Figure 13 The confusion matrix diagram is shown when the minimum recall rate is set to 0.95 (ie, limiting the second type error) according to one embodiment of the present invention. DETAILED DESCRIPTION
[0178] In the present invention, the method for detecting fetal chromosomal abnormalities can be implemented by a system for detecting fetal chromosomal abnormalities; the method for constructing a detection model for detecting fetal chromosomal abnormalities can be implemented by a system for constructing a detection model for detecting fetal chromosomal abnormalities.
[0179] In the present invention, the data acquisition module is used to obtain sequencing data of free nucleic acid fragments of pregnant women and clinical phenotypic characteristic data of pregnant women, wherein the sequencing data includes several reads, the chromosome status of the fetus of the pregnant woman is known (training sample) or unknown (sample to be tested), and the clinical phenotypic characteristic data of the pregnant woman form a pregnant woman phenotypic characteristic vector. The data acquisition module may include a data receiving module for receiving the above data. The data acquisition module may also include a sequencer, which obtains sequencing data by inputting the free nucleic acid of the pregnant woman. Sequencing can be high-throughput sequencing, and ultra-low depth sequencing can be performed, and the sequencing depth of ultra-low depth sequencing is 1×, 0.1× or 0.01×. The free nucleic acid of the pregnant woman can be derived from the peripheral plasma, liver and / or placenta of the pregnant woman. The clinical phenotypic characteristics of the pregnant woman and the chromosome status of the fetus of the pregnant woman (training sample) can be obtained through a database, and the chromosome status of the fetus of the pregnant woman can be chromosome aneuploidy, microdeletion and / or microduplication.
[0180] In the present invention, the alignment module is used to align the reads with the reference genome to obtain uniquely aligned reads. Application software for aligning sequences to a reference genome can be obtained from open source developers, such as online from certain websites, or can be developed independently.
[0181] In the present invention, the sequence feature matrix generation module is used to divide at least a part of the reference genome chromosome into windows, obtaining a number of sliding windows, counting the reads falling within the sliding windows, and generating a sequence feature matrix for this segment of chromosome sequence. This can be achieved by sliding a window of fixed length on the chromosome sequence. The fixed length of the window can be 10k, 100k, 1M, 10M, etc. The step size can be of any length. For the convenience of calculation, it is generally set to half of the sliding window length. The length of the chromosome sequence only needs to be greater than the sliding window length, and can be 10k, 100k, 1M, 10M, 100M... up to the length of the entire chromosome. The chromosome can be the target chromosome. For example, for the detection of trisomy 21, it corresponds to chromosome No. 21; for the detection of trisomy 18, it corresponds to chromosome No. 18; for the detection of trisomy 13, it corresponds to chromosome No. 13; for the detection of sex chromosome abnormalities, it corresponds to the XY chromosomes; for the detection of chromosomal microdeletions / microduplications, it corresponds to all autosomes. For each window, its parameters are counted, including the number of reads, base quality (measuring the sequencing accuracy), and alignment quality (a metric for the accuracy of read alignment to the reference genome, the higher the alignment quality, the more unique the position where the read is aligned to the reference genome), etc. This can be accomplished using computer software.
[0182] In the present invention, the sequence feature extraction module is used to extract the sequence features of the chromosome sequence. For the training data set, the sequence feature vector generation module constructs a training data set using the sequence feature matrix and the fetal chromosome abnormality situation of the pregnant woman, and trains a machine learning model to extract the sequence feature vector of this segment of chromosome sequence. For the test data, the sequence feature vector generation module constructs a test data set using the sequence feature matrix, and inputs the trained machine learning model, such as a deep neural network model, to extract the sequence feature vector of this segment of chromosome sequence.
[0183] In the present invention, for the training data set, the classification and detection module, such as the ensemble learning model training module, is used to train the classification and detection model with the combined feature vector formed by the sequence feature vector and the pregnant woman's phenotypic feature vector, as well as the fetal chromosome situation, to obtain the trained classification and detection model.
[0184] For the test data set, the classification and detection module is used to take the combined feature vector formed by combining the sequence feature vector and the pregnant woman's phenotypic feature vector as the input, and use the above-trained classification and detection model to detect the chromosome abnormality situation.
[0185] The present invention proposes a completely innovative method for detecting chromosomal abnormalities, such as aneuploidy, microdeletion or microduplication. Different from traditional methods, the present invention does not directly detect aneuploidy based on the read count and Z - value, without the need for cumbersome data pre - processing and feature extraction and selection work. Instead, it designs a machine - learning model to automatically extract sequence feature vectors from the sequence feature matrix generated in the sequencing data, combines the sequence feature vectors with the clinical phenotypic features of the pregnant woman, and uses a classification detection model for detection, finally obtaining the prediction result of whether there is a genetic abnormality in the fetal chromosome.
[0186] In the present invention, a machine - learning model is used to automatically extract sequence feature vectors from the sequencing data, avoiding the disadvantages of traditional manual extraction of NIPT whole - genome sequence features. The method of the present invention not only fully exploits the information in the sequencing data, but also makes full use of the clinical phenotypic information of the pregnant woman (the phenotypic data information that can be added to the model includes the age of the pregnant woman, gestational age, height, weight, BMI (body mass index), results of antenatal biochemical tests, diagnostic results of ultrasound examinations such as NT value, etc.). By combining the extracted sequence feature vectors with the phenotypic feature vectors of the pregnant woman, the rich feature data information contained in the NIPT sequencing data and the clinical phenotypic results of the pregnant woman is fully exploited, ensuring the high reliability and effectiveness of the detection results. The method of the present invention can be used not only to detect common trisomy syndromes, but also to detect other chromosomal defects, such as chromosomal copy number variations, chromosomal microdeletions, microduplications, etc.
[0187] In the present invention, the extraction of sequence feature vectors can also be performed using deep neural network models such as the Autoencoder network or the Variational Autoencoder network.
[0188] In the present invention, an ensemble learning model based on Stacking or Majority Voting is trained to detect chromosomal abnormalities, making full use of the recognition results of different classifiers for aneuploidy, greatly improving the accuracy of aneuploidy recognition.
[0189] In the present invention, the reference genome refers to a human genome map with normal diploid chromosomes generated through, for example, the Human Genome Project, such as hg38, hg19, etc. The reference genome can be one chromosome or multiple chromosomes, or it can be a part of one chromosome.
[0190] The present invention will be further described below through specific embodiments, but the present invention is not limited by the embodiments.
[0191] Example 1: An example of constructing a detection model
[0192] In an exemplary embodiment, the process and steps of the exemplary model embodiment for constructing the detection model are described in detail as follows.
[0193] 1. Obtain NIPT sequencing data and comparison results
[0194] The training samples were sequenced using the high-throughput sequencing platform BGIseq500 (using SE35 and a sequencing depth of 0.1×), i.e., the free nucleic acid fragments of pregnant women whose fetal chromosomes were known. The sequencing data were aligned with the reference genome and duplicate alignment sequences were filtered to obtain unique alignment sequencing read results.
[0195] 2. Preprocess the sequencing reads obtained in step 1 above and recalibrate the sequence coverage depth of each coverage region of the genome based on the relationship between GC content and sequencing depth. The specific process is as follows (please refer to reference
[15] for detailed process).
[0196] a. First, randomly sample m lengths from a chromosome of the human reference genome (such as chromosome 21). Fragment.
[0197] b. Calculate the number of fragments Ni with GC content i:
[0198]
[0199] in f(k) is the GC content of fragment k, and i represents the GC content (i=0%, 1%, ..., 100%).
[0200] c. Calculate the number of sequencing reads with GC content i:
[0201]
[0202] in represents the GC content of fragment k, and Fi represents the number of sequencing reads with GC content i and the same sequencing read starting point as the fragment starting point.
[0203] d. Calculate the observed-to-expected ratio of GC content:
[0204]
[0205] Where r is the global scaling factor, which is defined as:
[0206]
[0207] e. Correction of sequencing read numbers:
[0208]
[0209] where R i represents the expected number of sequencing reads with a GC content of i after correction.
[0210] 3. Generate a sequence feature matrix
[0211] Calculate the feature matrix based on the result of step 2 above. The calculation process is as follows (as Figure 2 shown):
[0212] Use a sliding window of length b to slide from the starting position to the ending position of the target chromosome of length L with a sliding step of t. Calculate the following features for each region of length b covered by the sliding window:
[0213] a. The number of reads after GC correction in this region;
[0214] b. The mean base quality in this region;
[0215] c. The standard deviation of base quality in this region;
[0216] d. The skewness of base quality in this region;
[0217] e. The kurtosis of base quality in this region;
[0218] f. The mean alignment quality in this region;
[0219] g. The standard deviation of alignment quality in this region;
[0220] h. The skewness of alignment quality in this region;
[0221] i. The kurtosis of alignment quality in this region;
[0222] Thus, a sequence feature matrix is obtained:
[0223] X = (x ij ) h×w
[0224] where h represents the number of times the window slides. For example,
[0225] w represents the number of sequence features within a single sliding window. For example, w = 9 (i.e., each sliding window of length b calculates 9 different features);
[0226] x ij represents the j-th sequence feature value in the i-th sliding window.
[0227] Base Quality quantitatively describes the accuracy of sequencing results; the mean base quality, standard deviation of base quality, skewness of base quality, and kurtosis of base quality respectively refer to the mean, standard deviation, skewness, and kurtosis of all base qualities in the sequencing reads. Map Quality refers to the reliability of aligning a given sequencing read to the reference genome sequence; the mean map quality, standard deviation of map quality, skewness of map quality, and kurtosis of map quality respectively refer to the mean, standard deviation, skewness, and kurtosis of the map quality of a given sequencing read.
[0228] 4. Construct a deep neural network model
[0229] 4.1 Construct a dataset
[0230] Use the results of 3 to construct a training set where N represents the number of samples, and N is an integer ≥ 1; Z (k) is the sequence feature matrix of sample k after standardization (hereinafter referred to as the standardized sequence feature matrix), and k ∈ [1, N], defined as:
[0231] where represents the j-th sequence feature vector in the i-th sliding window of sample k in the training set, μ i,j is the mean of the j-th feature vector in the i-th sliding window in the training set, σ i,j is the standard deviation of the j-th feature vector in the i-th sliding window in the training set, i is an integer ≥ 1, and j is an integer ≥ 1;
[0232]
[0233] 4.2 Construct a deep neural network model
[0234] Construct a deep neural network model, the structure of which is as Figure 3 shown. All convolutional layers involved in the deep neural network model perform 1D convolution operations. Without additional special instructions, the parameters of the 1D convolutional kernel (i.e., 1D filter) are the same, that is, the number of 1D convolutional kernels is f, the size of the 1D convolutional kernel is k, the stride of the 1D convolution operation is s, the 1D convolutional kernel uses L2 regularization and the regularization factor is r L2 , the initialization function of the 1D convolutional kernel is g, set the size of the output feature map (FeatureMap) of the 1D convolution operation to be the same as the input feature map size, the size of the pooling kernel is p, and the pooling stride is p s .
[0235] The dropout rates (Dropout Ratio) used in the Dropout layers involved in the deep neural network model are the same, set to d.
[0236] The deep neural network model structure includes:
[0237] 4.2.1 Input layer
[0238] The input layer is used to receive the sequence feature matrix Z after normalization processing (k) , and its size is h×w.
[0239] 4.2.2 Pre-module
[0240] The pre-module is connected to the input layer and is used to perform the first convolution and activation operations on the input sequence feature matrix to obtain an abstract representation feature map (Feature Map). This module includes: a 1D convolution layer, a batch normalization layer connected thereto, and a ReLU activation layer connected thereto.
[0241] 4.2.3 Core module
[0242] The core module is connected to the pre-module and is used to further abstract and extract features from the feature map, thereby enhancing the neural network expression ability by effectively increasing the depth of the neural network model. This core module is composed of the repeated operation of 3 identical residual modules, where the output of each residual module is the input of the next residual module. Each of the residual modules includes:
[0243] (A) The pre-sub-module of the core module repeated twice, each sub-module having the same structure, including a 1D convolution layer, a Dropout layer connected thereto, a batch normalization layer connected thereto, and a ReLU activation layer connected thereto;
[0244] (B) A first 1D average pooling layer (1D Average Pooling), connected to the second pre-sub-module described in (A);
[0245] (C) A Squeeze-Excite module (i.e., SE module) or a Spaial Squeeze-Excite module (i.e., sSE module), connected to the first 1D average pooling layer described in (B);
[0246] First, set the reduction rate of the SE module to r SE , as Figure 4 shown, the structure of the SE module includes (for detailed description, see reference
[16] ):
[0247] (a) A second global average pooling layer, connected to the first 1D average pooling layer described in (B);
[0248] (b) A Reshape layer connected to the second global average pooling layer described in (a), with the output feature map size of 1×f, where f is the number of 1D convolutional kernels;
[0249] (c) A first fully connected layer connected to the Reshape layer described in (b), with the number of output neurons being where f is the number of 1D convolutional kernels, and r SE is the descent rate of the SE module;
[0250] (d) A second fully connected layer connected to the first fully connected layer described in (c), with the number of output neurons being f, where f is the number of 1D convolutional kernels;
[0251] (e) A Multiply layer, connected to the first 1D average pooling layer of (B) and the second fully connected layer in (d);
[0252] As Figure 5 shown, the sSE module structure includes (for detailed description, see reference 1171):
[0253] a. A 1×1 1D convolutional layer, connected to the first 1D average pooling layer in (B), and the 1×1 1D convolutional layer uses the sigmoid function as the activation function;
[0254] b. A Multiply layer, connected to the first 1D average pooling layer in (B) and the 1×1 1D convolutional layer in a;
[0255] (D) A first addition layer (Add layer), connected to the SE module described in (C) and the sSE module;
[0256] (E) A second 1D average pooling layer, connected to the ReLU activation layer in the pre-module described in 4.2.2;
[0257] (F) A second addition layer (Add layer), connected to the first addition layer described in (D) and the second 1D average pooling layer in (E);
[0258] The above (A)-(D) are the left branch of the residual module, and (E) is the right branch of the residual module.
[0259] 4.2.4 Post-module
[0260] The post-module has the same structure as the pre-module, and the only difference is that the number of 1D convolutional kernels in the post-module is set to n out , which is used for feature abstraction representation before output of the feature map from the core module.
[0261] 4.2.5 First Global Average Pooling Layer
[0262] A first global average pooling layer, connected to the subsequent module, is used to vectorize the feature map of the feature abstraction representation.
[0263] 4.2.6 Output Layer
[0264] The output layer, connected to the first global average pooling layer, is a fully connected layer with 1 output neuron and a sigmoid function as the activation function, and is used to output the chromosomal abnormality situation.
[0265] 5. Calculate Sequence Feature Vector
[0266] Use the training set to train the deep neural network model described in 4, and use the trained deep neural network model to calculate the sequence feature vector of the sample. The process is as follows:
[0267] (1) Calculate the normalized sequence feature vector of each sample according to the description in 4.1 above;
[0268] (2) Input the normalized sequence feature matrix obtained in (1) into the deep neural network model for calculation;
[0269] (3) Save the output of the first global average pooling layer of the deep neural network model described in 4.2.5 as the sequence feature vector seq generated corresponding to the input sample, which is defined as:
[0270]
[0271] where n out is the number of 1D convolutional kernels defined in the subsequent module described in 4.2.4.
[0272] 6. Obtain the Phenotype Result Corresponding to the Pregnant Woman Sample
[0273] Obtain the phenotype result corresponding to the pregnant woman sample and construct the initial phenotype feature vector phe init , which includes 5 features and is defined as:
[0274] phe init = [x age , x GW , x height , x weight , x FF T
[0275] where x age represents the age (years) of the pregnant woman at the time of sampling, and x GW represents the gestational week of the pregnant woman at the time of sampling, xheight Denote the height (cm) of the pregnant woman as \(x\). weiqh Denote the weight (kg) of the pregnant woman as \(y\). FF Denote the concentration of cell-free fetal DNA in the plasma of the pregnant woman as \(z\).
[0276] 7. Preprocessing of phenotypic data
[0277] Preprocess the phenotypic dataset of pregnant women, including outlier handling and missing value or null value handling.
[0278] (1) Outlier handling
[0279] The following records in the phenotypic data of pregnant woman samples will be determined as outliers, and these outliers will be set as null values.
[0280] a. \(x\) age \(< 10\) or \(x\) age \(> 80\);
[0281] b. \(y\) GW \(< 5\) or \(y\) GW \(> 50\);
[0282] c. \(z\) heigh \(< 40\) or \(z\) heig \(> 300\);
[0283] d. \(w\) weight \(< 10\) or \(w\) weight \(> 200\).
[0284] (2) Missing value or null value handling
[0285] Construct a phenotypic data matrix \(P\), which is defined as follows:
[0286]
[0287] where represents the phenotypic feature vector of the \(i\)-th sample in the training set (defined as described in 6), and \(N\) represents the number of samples in the training set. The samples in the training set here are the same as those in the training set described in 4.1. Thus, the phenotypic data matrix \(P\) is a matrix of size \(N\times M\), where \(M\) is the number of phenotypic features, and here \(M = 5\).
[0288] The missing values are filled using the missForest algorithm. The missForest algorithm is a non-parametric missing value filling algorithm based on random forest (for detailed description, see reference
[18] ). Its algorithm is as follows:
[0289]
[0290]
[0291] (3) Calculate BMI
[0292] Calculate the BMI using the phenotypic results after filling in the missing values, and its definition is:
[0293]
[0294] (4) Add the result of (3) to the phenotypic feature vector after filling in the missing values to obtain the final phenotypic feature vector:
[0295] phe = [x age , x GW , x height , x weight , x FF , x BMI T
[0296] 8. Generate the combined feature vector
[0297] Combine the sequence feature vector described in 5 with the final feature vector described in 7 to obtain the combined feature vector:
[0298]
[0299] 9. Standardize the combined feature vector
[0300] Perform standardization processing on the combined feature vector described in 8:
[0301]
[0302] Among them, x′ i is the i-th sequence feature value in the standardized combined feature vector x, x i is the i-th sequence feature value in the combined feature vector x, μ i is the mean of the i-th sequence feature value in the combined feature vector, and σ i is the standard deviation of the i-th sequence feature value in the combined feature vector.
[0303] 10. Construct an ensemble learning model based on Stacking
[0304] Construct a training set using the result in 9
[0305]
[0306] Among them, N represents the number of training samples, and N is an integer ≥ 1; is the sequence feature matrix of the k-th training sample after standardization processing, and k ∈ [1, N], i is an integer ≥ 1, and j is an integer ≥ 1; y = 0 indicates that the fetal chromosome is normal, and y = 1 indicates that the fetal chromosome is abnormal.
[0307] Predict aneuploidy using the Stacking-based ensemble learning algorithm. The algorithm description is as follows (for a detailed description, see reference
[19] ):
[0308]
[0309]
[0310] Example 2. Example of detecting chromosomal abnormalities
[0311] In an exemplary solution, the present invention proposes a method for detecting fetal chromosomal abnormalities. The method uses the nucleic acid sequencing results of non-invasive prenatal testing (NIPT) and the phenotypic data of pregnant women to jointly predict whether there are genetic abnormalities in the fetal chromosomes. In a specific implementation, the process and steps of the method for detecting fetal chromosomal abnormalities are as Figure 1 shown, and the specific process is described as follows.
[0312] 1. Obtain NIPT sequencing data and alignment results
[0313] Sequence the test sample using the high-throughput sequencing platform BGIseq500 (using SE35, sequencing depth 0.1×), align the sequencing data with the reference genome, and filter out the duplicate alignment sequences to obtain the unique alignment sequencing read results.
[0314] 2. Preprocess the sequencing reads obtained in step 1 above. Recalibrate the sequence coverage depth of each covered region of the genome through the relationship between the GC content and the sequencing depth. The specific process refers to Example 1.
[0315] 3. Generate a sequence feature matrix
[0316] Calculate the feature matrix based on the results of step 2 above. The calculation process is as follows (as Figure 2 shown):
[0317] Use a sliding window of length b to slide from the starting point to the ending point of the target chromosome of length L, with a sliding step of t. Calculate the following features for each region of length b covered by the sliding window:
[0318] a. The number of reads corrected by GC in this region;
[0319] b. The mean value (mean) of the base quality in this region;
[0320] c. The standard deviation (std) of the base quality in this region;
[0321] d. The skewness of the base quality in this region;
[0322] e. The kurtosis of the base quality within this region;
[0323] f. The mean of the mapping quality within this region;
[0324] g. The standard deviation (std) of the mapping quality within this region;
[0325] h. The skewness of the mapping quality within this region;
[0326] i. The kurtosis of the mapping quality within this region;
[0327] Thus, a sequence feature matrix is obtained:
[0328] X = (x ij ) h×w
[0329] where h represents the number of sliding times of the window. For example,
[0330] w represents the number of sequence features within a single sliding window. For example, w = 9 (i.e., 9 different features are calculated for each sliding window of length b);
[0331] x ij represents the j-th sequence feature value in the i-th sliding window.
[0332] Base Quality quantitatively describes the accuracy of the sequencing result; the mean of base quality, the standard deviation of base quality, the skewness of base quality, and the kurtosis of base quality respectively refer to the mean, standard deviation, skewness, and kurtosis of all base qualities in the sequencing reads. Map Quality refers to the reliability of a given sequencing read mapped to the reference genome sequence; the mean of map quality, the standard deviation of map quality, the skewness of map quality, and the kurtosis of map quality respectively refer to the mean, standard deviation, skewness, and kurtosis of the map quality of a given sequencing read.
[0333] 4. Calculate the sequence feature vector of the sample using the trained deep neural network model in Example 1. The process is as follows:
[0334] (1) Calculate the standardized sequence feature matrix of the sample according to 4.1 in Example 1;
[0335] (2) Input the standardized sequence feature matrix obtained in (1) into the deep neural network model for calculation;
[0336] (3) Save the output of the first global average pooling layer of the deep neural network model described in 4.2.5 in Example 1 as the sequence feature vector seq generated corresponding to the input sample, defined as:
[0337]
[0338] where n out is the number of 1D convolutional kernels defined in the post-module described in 4.2.4.
[0339] 5. Obtain the phenotypic result corresponding to the pregnant woman sample to be tested
[0340] Obtain the phenotypic result corresponding to the pregnant woman sample to be tested, and construct the initial phenotypic feature vector phe init , which includes 5 features, defined as:
[0341] phe init = [x age , x GW , x height , x weight , x FF T
[0342] where x age represents the age (years) of the pregnant woman at the time of sampling, x GW represents the gestational week of the pregnant woman at the time of sampling, x height represents the height (centimeters) of the pregnant woman, x weight represents the weight (kilograms) of the pregnant woman, x FF represents the concentration of cell-free fetal DNA in the plasma of the pregnant woman.
[0343] 6. Perform outlier processing on the phenotypic data
[0344] The following records in the phenotypic data of the pregnant woman sample to be tested will be determined as outliers, and these outliers will be set to null values.
[0345] a. x age < 10 or x age > 80;
[0346] b. x GW < 5 or x GW > 50;
[0347] c. x height < 40 or x height > 300;
[0348] d. x weight < 10 or x weight > 200.
[0349] 7. Combine the sequence feature vector described in 4 and the final feature vector described in 6 to obtain a combined feature vector:
[0350]
[0351] 8. Standardization of the combined feature vector
[0352] Perform standardization processing on the combined feature vector described in 7:
[0353]
[0354] Among them, x′ i is the i-th sequence eigenvalue in the standardized combined feature vector x, and x i is the i-th sequence eigenvalue in the combined feature vector x, and μ i is the mean value of the i-th sequence eigenvalue in the combined feature vector, and σ i is the standard deviation of the i-th sequence eigenvalue in the combined feature vector.
[0355] 9. Input the combined feature vector into the Stacking-based ensemble learning model constructed in Example 1 to obtain the fetal chromosome condition of the pregnant woman to be tested.
[0356] Example 3. Verification example
[0357] 1. Sample quantity
[0358] In this example, 1205 "trisomy 21 (T21)" are used as positive samples and 1600 normal chromosome (diploid) samples are used as negative samples.
[0359] Table 1 describes the sample quantities of the training samples and test samples
[0360] Total number of samples (N) Number of samples in the training set (90% × N) Number of samples in the test set (10% × N) Positive samples (T21) 1205 1084 121 Negative samples (normal) 1600 1440 160
[0361] 2. Preprocess the sequencing data of all positive and negative samples according to the steps described in 2 of Example 1 above, where the number m of randomly sampled fragments is 50000000, and the fragment
[0362] 3. Generate a sequencing sequence feature matrix for all positive and negative samples according to the steps described in 3 of Example 1 above. The parameter settings are as follows:
[0363] Length of chromosome 21: L = 46709983;
[0364] Sliding window size: b = 1000000;
[0365] Sliding step: t = 500000.
[0366] The resulting sequencing sequence feature matrix for each sample is 9×93 in size, i.e., w = 9, h = 93. Because the initial portion of chromosome 21 has no comparable sequence in the reference genome, in this example, the first eight columns of the sequencing sequence feature matrix are filtered, resulting in a sequencing sequence feature matrix size of 9×85.
[0367] 4. According to the results of 3, use the corresponding sequencing data feature matrix in the training set to train the deep neural network model.
[0368] (1) Standardize the feature matrix of the sequencing data in the training set according to 4.1 of the above-mentioned embodiment 1 and save the standardized model.
[0369] (2) The input tensor for training the deep neural network model is obtained as described in (1), and the tensor size is 2524×85×9.
[0370] (3) Train the deep neural network model according to the above-mentioned 4.2 of Example 1, and set the deep neural network model parameters as follows:
[0371] Number of 1D convolution kernels: f = 32,
[0372] 1D convolution kernel size: k=8,
[0373] 1D convolution operation step size: s = 1,
[0374] 1D convolution kernel 12 regularization factor: r l2 =0.0004,
[0375] The initialization function g of the 1D convolution kernel uses the “He normalization” initialization function described in
[20] .
[0376] The output feature map size of the 1D convolution operation remains unchanged from the input feature map size.
[0377] Pooling kernel size: p=2,
[0378] Pooling step size: p s =2,
[0379] Dropout layer uses the dropout rate: r d =0.5,
[0380] SE module drop rate: r SE =16,
[0381] The number of 1D convolution kernels n in the post-module out =8.
[0382] This embodiment is implemented based on the GPU versions of TensorFlow (version = 1.12.2) and Keras (version = 2.2.4). Table 2 lists the operations of each layer, the size of the output feature map, and the network connections in the deep neural network model according to the said parameters.
[0383]
[0384]
[0385]
[0386] (4) 80% of the samples in the training set are used to train the deep neural network, and 20% of the samples are used for validation to calculate the accuracy rate.
[0387] (5) Set the number of iterations epochs = 100 and the sample batch size mini_batch = 64 for training the deep neural network. The optimization algorithm for gradient descent uses the Adam algorithm (parameters β1 = 0.9, β2 = 0.999), and the initial learning rate (LearningRate) is set to 0.01. During the training process, if the accuracy rate does not improve after two consecutive iterations, the learning rate will be reduced by 2 times (i.e., multiplied by 0.5); if the accuracy rate does not improve after ten consecutive iterations, the training will stop.
[0388] (6) During the training process of the deep neural network model, a class weight factor is introduced (calculate the class weight using the compute_class_weight() function in the machine learning library scikit-learn (version = 0.22.2), and assign this class weight to the samples of the corresponding class).
[0389] (7) Save the trained deep neural network model.
[0390] 5. Calculate the sequence feature vector according to item 5 in the foregoing Embodiment 1:
[0391] (1) Calculate the sequence feature matrix for all samples in the entire data set (including the training set and the test set) according to item 3 in the foregoing Embodiment 1;
[0392] (2) Use the obtained sequence normalization model to perform normalization processing on the sequence feature matrix obtained in (1) above according to 4.1;
[0393] (3) Utilize the deep neural network model obtained in 4, with the input of the model being the result of (2) above, and modify the output layer of the model to a global average pooling layer (i.e., the 65th layer in Table 2);
[0394] (4) Obtain the sequence feature vectors of all samples in the entire data set (including the training set and the test set) according to the process in (3).
[0395] 6. Obtain the phenotypic features of all samples in the entire data set (including the training set and the test set) as described in item 7 of the foregoing Embodiment 1, and process the phenotypic feature outliers.
[0396] 7. Perform missing value filling processing on the phenotypic features in the training set as described in item 7 of the foregoing Embodiment 1, and save the missing value filling model.
[0397] 8. Calculate the BMI for the phenotypic features in the training set after the missing value filling processing as described in item 7 of the foregoing Embodiment 1, as Figure 6 shown.
[0398] 9. Perform a combination operation on the sequence feature vectors in the training set and the phenotypic feature vectors of the corresponding samples as described in item 8 of the foregoing Embodiment 1 to obtain combined feature vectors.
[0399] 10. Standardize the combined feature vectors of each sample in the training set as described in item 9 of the foregoing Embodiment 1 to obtain standardized feature vectors, and save the combined feature vector standardization model.
[0400] 11. According to the process in steps 7-10 above, use the saved missing value filling model to fill in the missing values of the phenotypic features of each sample in the test set, then combine the sequence feature vectors in the test set and the phenotypic feature vectors of the corresponding samples to obtain the combined feature vectors of the test set, and then use the saved combined feature vector standardization model to standardize the combined feature vectors in the test set.
[0401] 12. Train an ensemble learning model based on Stacking using the training set standardized feature vectors obtained in step 10 above, as Figure 7 shown. This embodiment is implemented based on the scikit-learn (version = 0.22.2) machine learning library, where a class weight factor is introduced into each base classifier model and the final meta-classifier model, and default values are used for parameters without special instructions.
[0402] (1) As described in item 10 of the foregoing Embodiment 1, the base classifiers used in this embodiment include:
[0403] · SVC, with parameters C = 0.5, kernel = "rbf"
[0404] · v-SVC, with parameters v = 0.25, kernel = kernel = "rbf"
[0405] · GaussianNB (Gaussian Naive Bayes model)
[0406] ·RandomForestClassifier (Random forest classifier), with parameters n_estimators = 100, criterion = "gini", max_depth = 5, min_samples_leaf = 1, min_samples_split = 2
[0407] ·XGBClassifier (XGBoosting classifier), with parameters n_estimators = 100, min_child_weight = 1, gmnma = 0.1, colsample_bytree = 0.8, subsample = 0.7, reg_alpha = 0.01, max_depth = 5, learning_rate = 0.05,
[0408] ·LogisticRegression (Logistic regression), with parameter C = 0.5
[0409] (2) The final meta-classifier described in item 10 of the foregoing Embodiment 1 is ExtraTreesClassifier (Extreme Random Trees classifier), and the parameters involved in this classifier are respectively set as n_estimators = 110, max_depth = 6, min_samples_split = 3, min_smnples_leaf = 1.
[0410] (3) Perform 5-fold cross-validation training on the Stacking-based ensemble learning model, and the results are as Figure 8 shown. It can be seen that the average AUC of the 5-fold cross-validation training model is 0.96.
[0411] 13. Use the test set to verify the Stacking-based ensemble learning model trained in item 12.
[0412] (1) The ROC curve of the test result is as Figure 9 shown, where AUC = 0.96;
[0413] (2) The Precision-Recall curve of the test result is as Figure 10 shown, where AP = 0.95.
[0414] (3) The confusion matrix when the decision-making threshold is the default value (i.e., 0.5) is as Figure 11 shown. At this time, the recall rate and the precision rate are 0.83 and 0.89 respectively.
[0415] (4) The precision and recall as functions of the threshold are shown as Figure 12 follows.
[0416] (5) Set the minimum recall = 0.95, that is, limit the type II error. The obtained results are shown as Figure 12 follows. At this time, the recall and precision are 0.96 and 0.70 respectively.
[0417] The present invention proposes to extract the sequence feature vectors of NIPT sequencing data by using a machine learning model (such as a deep neural network), and then integrate and combine the sequence feature vectors (the feature items include but are not limited to the number of reads, base quality, and alignment quality) and the pregnant woman phenotype feature vectors (the pregnant woman phenotype features include but are not limited to the pregnant woman's age, gestational week, height, weight, BMI, antenatal biochemical test results, ultrasound examination diagnosis results such as NT value, etc.) in a vector combination manner, and then use a classification model (such as an ensemble learning model based on Stacking) to obtain the final prediction of aneuploidy. In the present invention, the method for extracting the sequence feature vectors is not limited to the method used herein, and other methods including but not limited to an autoencoder network (Autoencoder) or a variational autoencoder network (Variational Autoencoder) can also be used. The model structure proposed by the present invention is a hybrid model, that is, the model includes two stages. In the first stage, a machine learning model (such as a deep neural network) is used to calculate the sequence feature vectors, and in the second stage, a classification model (such as an ensemble learning model based on Stacking) is used to predict aneuploidy by using the integrated and combined sequence feature vectors and phenotype feature vectors. Other ensemble learning models, such as based on majority voting, can also be used.
[0418] Compared with other convolutional neural networks, the network design and architecture characteristics of the verified advanced deep neural network model used in the embodiments of the present invention include: the deep neural network model used in the embodiments of the present invention is designed based on a one-dimensional convolutional model; the deep neural network model used in the embodiments of the present invention is based on a residual network (Residual Net) network model; the deep neural network model used in the embodiments of the present invention introduces the SE module of the Squeeze-Excite network. Based on these designs, the number of layers of the neural network model used in the embodiments of the present invention is more (see Embodiment 3), and the risk of gradient disappearance and overfitting during model training is effectively reduced, and the model stability is improved, thereby effectively improving the accuracy of the model prediction results.
[0419] The present invention can be implemented as a computer-readable storage medium having stored thereon a computer program which, when executed by a processor, causes the steps of the method of the present invention to be performed. In one embodiment, the computer program is distributed across a plurality of network-coupled computer devices or processors such that the computer program is stored, accessed, and executed by one or more computer devices or processors in a distributed manner. A single method step / operation, or two or more method steps / operations, can be performed by a single computer device or processor or by two or more computer devices or processors. One or more method steps / operations can be performed by one or more computer devices or processors, and one or more other method steps / operations can be performed by one or more other computer devices or processors. One or more computer devices or processors can perform a single method step / operation, or perform two or more method steps / operations.
[0420] Those skilled in the art should understand that the division and order of each step in the method for detecting fetal chromosomal abnormalities of the present invention are merely illustrative and not restrictive. Those skilled in the art can make deletions, additions, substitutions, modifications, and changes without departing from the spirit and scope of the present invention as set forth in the appended claims and their equivalent technical solutions. The technical features of the embodiments of the present invention can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as falling within the scope described in this specification.
[0421] Although the present invention has been described with reference to exemplary embodiments, it should be understood that the present invention is not limited to the structures and methods of the above embodiments. On the contrary, the present invention is intended to cover various modifications and equivalent configurations. In addition, although various elements and method steps disclosed in the present invention are shown in various exemplary combinations and configurations, other combinations including more or fewer elements or methods also fall within the scope of the present invention.
[0422] References:
[0423] [1]Evans, Mark I., Stephanie Andriole, and Shara M. Evans. ″Genetics: update on prenatal screening and diagnosis.″ Obstetrics and Gynecology Clinics 42.2 (2015): 193 - 208.
[0424] [2]Norwitz, Errol R., and Brynn Levy. "Noninvasive prenatal testing: the future is now." Reviews in obstetrics and gynecology 6.2 (2013): 48.
[0425] [3]Norton, Mary E., et al. "Cell-free DNA analysis for noninvasive examination of trisomy." New England Journal of Medicine 372.17 (2015): 1589-1597.
[0426] [4]Langlois, Sylvie, et al. "Current status in non-invasive prenatal detection of Down syndrome, trisomy 18, and trisomy 13 using cell-free DNA in maternal plasma." Journal of Obstetrics and Gynaecology Canada 35.2 (2013): 177-181.
[0427] [5]Allyse, Megan, et al. "Non-invasive prenatal testing: a review of international implementation and challenges." International journal of women's health 7 (2015): 113.
[0428] [6]Chiu, Rossa WK, et al. "Noninvasive prenatal diagnosis of fetal chromosomal aneuploidy by massively parallel genomic sequencing of DNA in maternal plasma." Proceedings of the National Academy of Sciences 105.51 (2008): 20458-20463.
[0429] [7]Fan, H. Christina, et al. "Noninvasive diagnosis of fetal aneuploidy by shotgun sequencing DNA from maternal blood." Proceedings of the National Academy of Sciences 105.42 (2008): 16266-16271.
[0430] [8]Lau, Tze Kin, et al. "Noninvasive prenatal diagnosis of common fetal chromosomal aneuploidies by maternal plasma DNA sequencing." The Journal of Maternal-Fetal&Neonatal Medicine 25.8 (2012): 1370-1374.
[0431] [9]Jiang, Fuman, et al. "Noninvasive Fetal Trisomy (NIFTY) test: an advanced noninvasive prenatal diagnosis methodology for fetal autosomal and sex chromosomal aneuploidies." BMC medical genomics 5.1 (2012): 57.
[0432]
[10] Yang, Jianfeng, Xiaofan Ding, and Weidong Zhu. "Improving the calling of non-invasive prenatal testing on 13- / 18- / 21-trisomy by support vector machine discrimination." BioRxiv (2017): 216689.
[0433]
[11] Xu,Hanli,et al.″Informative priors on fetal fraction increasepower of the noninvasive prenatal screen.″Genetics in Medicine 20.8(2018):817-824.
[0434]
[12] Ehrich,Mathias,et al.″Deep learning-based methods,devices,andsystems for prenatal testing″,Publication number:WO2019191319A1,Filing Date:27 March 2019.
[0435]
[13] Egilsson,Agust,et al.″Methods and systems for calling ploidystatus using a neural network″.Publication number:WO2020018522A1,Filing date:16 July 2019.
[0436]
[14] Petersen,Andrea K.,et al.″Positive predictive value estimates forcell-free noninvasive prenatal screening from data of a large referralgenetic diagnostic laboratory.″American journal of obstetrics and gynecology217.6(2017):691-e1.
[0437]
[15] Benjamini,Yuval,and Terence P.Speed.″Summarizing and correctingthe GC content bias in high-throughput sequencing.″Nucleic acids research40.10(2012):e72-e72.
[0438]
[16] Hu, Jie, Li Shen, and Gang Sun. "Squeeze-and-excitation networks." Proceedings of the IEEE conference on computer vision and pattern recognition. 2018.
[0439]
[17] Roy, Abhijit Guha, Nassir Navab, and Christian Wachinger. "Concurrent spatial and channel'squeeze&excitation' in fully convolutional networks." International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, Cham, 2018.
[0440]
[18] Stekhoven, Daniel J., and Peter Bühlmann. "MissForest--non-parametric missing value imputation for mixed-type data." Bioinformatics 28.1 (2012): 112-118.
[0441]
[19] Tang, J., S. Alelyani, and H. Liu. "Data Classification: Algorithms and Applications." Data Mining and Knowledge Discovery Series, CRC Press (2015): pp. 498-500.
[0442]
[20] He, Kaiming, et al. "Delving deep into rectifiers: Surpassing human-level performance on imagenet classification." Proceedings of the IEEE international conference on computer vision. 2015.
Claims
1. A method for detecting fetal chromosomal abnormalities, the method comprising: (1) obtaining sequencing data and clinical phenotypic characteristic data of cell-free nucleic acid fragments of a pregnant woman to be tested, wherein the sequencing data includes a plurality of reads, and the clinical phenotypic characteristic data of the pregnant woman to be tested forms a pregnant woman phenotypic characteristic vector; (2) partitioning at least a part of a reference genome chromosome into windows to obtain a plurality of sliding windows, counting the reads falling within the sliding windows, and generating a sequence characteristic matrix of the chromosome sequence; (3) inputting the sequence characteristic matrix into a trained machine learning model to extract a sequence characteristic vector of the chromosome sequence; (4) combining the sequence characteristic vector and the pregnant woman phenotypic characteristic vector to form a combined characteristic vector, inputting the combined characteristic vector into a classification detection model, and obtaining the fetal chromosomal abnormality condition of the pregnant woman to be tested.
2. The method according to claim 1, in (1), the cell-free nucleic acid fragments are from the peripheral plasma of the pregnant woman, the liver of the pregnant woman, and / or the placenta.
3. The method according to claim 1 or 2, in (1), the cell-free nucleic acid fragments are cell-free DNA.
4. The method according to claim 1 or 2, in (1), the sequencing data is from ultra-low-depth sequencing.
5. The method according to claim 4, the sequencing depth of the ultra-low-depth sequencing is 1×, 0.1× or 0.01×.
6. The method according to any one of claims 1, 2, 5, in (1), aligning the reads with a reference genome to obtain uniquely aligned reads.
7. The method according to claim 6, performing GC content correction on the uniquely aligned reads.
8. The method according to claim 6, the subsequent steps are performed using the uniquely aligned reads.
9. The method according to claim 8, the uniquely aligned reads are reads after GC content correction.
10. The method according to claim 6, the GC content correction is performed according to the following steps: a. First, randomly sample segments of length from a certain chromosome of the human reference genome; b. Calculate the number of segments with a GC content of : Ni wherein , is the fragment 's GC content, indicating the GC content( ); c. Calculate the number of sequencing reads with a GC content of F i: wherein represents the fragment GC content, represents that the GC content is and the number of sequencing reads with the same starting position of the sequencing reads as the starting position of the fragment; d. Calculate the observed-to-expected ratio of GC content : wherein is the global scale factor, which is defined as: e. Correcting the number of sequencing reads: wherein represents the expected number of sequencing reads with a GC content of after correction.
11. The method according to any one of claims 1, 2, 5, 7-10, in (1), the pregnant woman phenotypic characteristic data is selected from a combination of one or more of the following: age, gestational week, height, weight, BMI, prenatal biochemical test results, ultrasound examination diagnosis results, and fetal cell-free DNA concentration in plasma.
12. The method according to any one of claims 1, 2, 5, 7-10, in (1), the pregnant woman phenotypic characteristic data is subjected to outlier processing, missing value processing, and / or null value processing.
13. The method according to claim 12, in (1), the phenotypic data of the pregnant woman sample will be determined as outliers and these outliers will be set as null values if the following records occur: a. or ; b. or ; c. or ; d. or .
14. The method according to claim 12, the missing values and null values are filled using the missForest algorithm.
15. The method according to any one of claims 1, 2, 5, 7-10, 13, 14, in (2), the chromosome is chromosome 21, chromosome 18, chromosome 13, and / or sex chromosome.
16. The method according to any one of claims 1, 2, 5, 7-10, 13, 14, wherein (2) includes (2.1) overlapping and sliding a chromosome sequence of length L on a reference genome using a window of length b with a step size to obtain a plurality of sliding windows, where b is a positive integer and b = [10000, 10000000], t is any positive integer, L is a positive integer, and L ≥ b; (2.2) counting the reads falling into each of the sliding windows to generate a sequence feature matrix of the chromosome sequence.
17. The method according to any one of claims 1, 2, 5, 7 - 10, 13, 14, in (2), the sequence feature matrix includes the number of reads, base quality, and alignment quality within the sliding window.
18. The method according to claim 17, the base quality includes the mean, standard deviation, skewness, and / or kurtosis of the base quality.
19. The method according to claim 17, the alignment quality includes the mean, standard deviation, skewness, and / or kurtosis of the alignment quality.
20. The method according to any one of claims 1, 2, 5, 7 - 10, 13, 14, 18, 19, in (2), the sequence feature matrix is: Among them represents the number of times the window slides, represents the number of sequence features within a single sliding window, represents the j-th sequence feature value in the i-th sliding window.
21. The method according to any one of claims 1, 2, 5, 7 - 10, 13, 14, 18, 19, in (3), the sequence feature matrix is normalized.
22. The method according to claim 21, in (3), formula (I) is used for the normalization of the sequence feature matrix: (I) Among them, is a sample The sequence feature matrix after standardization processing, represents the sample The j-th sequence feature value in the i-th sliding window of, and respectively represent the mean and standard deviation of the j-th sequence feature value in the i-th sliding window of all samples.
23. The method according to any one of claims 1, 2, 5, 7 - 10, 13, 14, 18, 19, 22, in (3), the trained machine learning model is a neural network model or an AutoEncoder model.
24. The method according to claim 23, the neural network model is a deep neural network model.
25. The method according to claim 23, the neural network model is a deep neural network model based on 1 - D convolution.
26. The method according to claim 24 or 25, the structure of the deep neural network model includes: An input layer for receiving the sequence feature matrix; A pre - processing module connected to the input layer for performing the first convolution and activation operations on the sequence feature matrix from the input layer to obtain a feature map; A core module connected to the pre - processing module for further abstracting and feature extraction of the feature map from the pre - processing module, strengthening the neural network expression ability by effectively increasing the depth of the neural network model; A post - processing module connected to the core module for performing feature abstract representation on the feature map from the core module; A first global average pooling layer connected to the post - processing module for vectorizing the feature map of the feature abstract representation and outputting the sequence feature vector of the chromosome sequence.
27. The method according to claim 26, the pre - processing module includes: (I) A 1 - D convolutional layer; (II) A batch normalization layer connected to the 1 - D convolutional layer in (I); (III) A ReLU activation layer connected to the batch normalization layer in (II).
28. The method according to claim 26, the core module is composed of one or more residual sub - modules with the same structure, where the output of each residual module is the input of the next residual module.
29. The method according to claim 28, the residual sub - module includes: (A)The pre-submodule of the core module, where each submodule includes a 1D convolutional layer, a Dropout layer connected to the 1D convolutional layer, a batch normalization layer connected to the Dropout layer, and a ReLU activation layer connected to the batch normalization layer; (B)The first 1D average pooling layer, connected to the pre-submodule of the core module described in (A); (C)The Squeeze-Excite module, and / or the Spatial Squeeze-Excite module, connected to the first 1D average pooling layer described in (B); (D)The first addition layer, connected to the Squeeze-Excite module and / or the Spatial Squeeze-Excite module described in (C); (E)The second 1D average pooling layer, connected to the ReLU activation layer in the pre-module; (F)The second addition layer, connected to the first addition layer described in (D) and the second 1D average pooling layer described in (E).
30. The method according to claim 29, wherein the Squeeze-Excite module includes: (a)The second global average pooling layer, connected to the first 1D average pooling layer in the residual submodule described in (B); (b)Reshape layer, connected to the second global average pooling layer described in (a), and the Reshape layer outputs a feature map with a size of , where is the number of 1D convolutional kernels; (c) The first fully connected layer, connected to the Reshape layer described in (b), the number of output neurons of the first fully connected layer is , where is the number of 1D convolutional kernels, is the descent rate of the Squeeze-Excite module; (d) The second fully connected layer, connected to the first fully connected layer described in (c), and the number of output neurons of the second fully connected layer is , where is the number of 1D convolutional kernels; (e)The Multiply layer, connected to the second fully connected layer described in (d) and the first 1D average pooling layer in the residual submodule described in (B).
31. The method according to claim 29 or 30, wherein the Spatial Squeeze-Excite module includes: a. A 1D convolutional layer, connected to (B) the first 1D average pooling layer, and the 1D convolutional layer uses the sigmoid function as the activation function; b. A Multiply layer, connected thereto are (B) a first 1D average pooling layer and the 1D convolutional layer described in a. The 1D convolutional layer.
32. The method according to any one of claims 1, 2, 5, 7-10, 13, 14, 18, 19, 22, 24, 25, 27-30, in (4), splicing the sequence feature vector and the pregnant woman phenotype feature vector to obtain a combined feature vector.
33. According to the method described in any one of claims 1, 2, 5, 7-10, 13, 14, 18, 19, 22, 24, 25, 27-30, in (4), the combined feature vector is normalized as follows: where is the i-th sequence eigenvalue in the standardized combined feature vector x, is the i-th sequence eigenvalue in the combined feature vector x, is the mean of the i-th sequence eigenvalue in the combined feature vector, is the standard deviation of the i-th sequence eigenvalue in the combined feature vector.
34. The method according to any one of claims 1, 2, 5, 7-10, 13, 14, 18, 19, 22, 24, 25, 27-30, in (4), the classification detection model is an ensemble learning model.
35. The method according to claim 34, wherein the ensemble learning model is an ensemble learning model based on stacking or majority voting.
36. The method according to claim 35, wherein the ensemble learning model is one or more of the following: support vector machine model, naive Bayes classifier, random forest classifier, XGBoos, and logistic regression.
37. The method according to claim 1, wherein the chromosomal abnormality includes at least one or more of the following: trisomy 21, trisomy 18, trisomy 13, 5p-syndrome, chromosomal microdeletion, and chromosomal microduplication.
38. A method for constructing a classification detection model for detecting fetal chromosomal abnormalities, the method comprising: (1)Obtaining sequencing data and clinical phenotype feature data of free nucleic acid fragments of a pregnant woman, wherein the sequencing data includes a plurality of reads, and the fetal chromosomal condition of the pregnant woman is known, and the clinical phenotype feature data of the pregnant woman forms a pregnant woman phenotype feature vector; (2)Partition at least a part of the reference genome chromosome into windows to obtain a number of sliding windows, count the reads falling within the sliding windows, and generate a sequence feature matrix of the chromosome sequence; (3)Construct a training data set using the sequence feature matrix and the fetal chromosome situation, and train a machine learning model to extract the sequence feature vector of the chromosome sequence; (4)Combine the sequence feature vector and the pregnant woman's phenotypic feature vector to form a combined feature vector, and use the combined feature vector and the fetal chromosome situation to train the classification model to obtain the trained classification detection model.
39. The method according to claim 38, wherein the fetal chromosome situation of the pregnant woman is one or more of the following: normal diploid, chromosomal aneuploidy, partial monosomy syndrome, chromosomal microdeletion, and chromosomal microduplication.
40. The method according to claim 39, wherein the chromosomal aneuploidy includes at least one or more of the following: trisomy 21, trisomy 18, and trisomy 13.
41. The method according to claim 39, wherein the partial monosomy syndrome includes 5p- syndrome.
42. The method according to any one of claims 39-41, wherein the number of pregnant women is greater than 10, and the ratio of the number of fetuses with normal diploid to the number of fetuses with chromosomal aneuploidy is ½ to 2.
43. The method according to any one of claims 38-41, wherein in (3), the training data set is represented as: , Among them, represents the number of training samples, and N is an integer greater than or equal to 1; is a training sample is the sequence feature matrix after normalization processing, where k ∈ [1, N], i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1; The chromosomal abnormalities include at least one or more of the following: trisomy 21, trisomy 18, trisomy 13, 5p- syndrome, chromosomal microdeletion, and chromosomal microduplication.
44. A system for detecting fetal chromosomal abnormalities, comprising: A data acquisition module for obtaining sequencing data of free nucleic acid fragments and clinical phenotypic feature data of a pregnant woman to be tested, wherein the sequencing data includes a number of reads, and the clinical phenotypic feature data of the pregnant woman to be tested forms a pregnant woman phenotypic feature vector; A sequence feature matrix generation module for partitioning at least a part of the reference genome chromosome into windows to obtain a number of sliding windows, counting the reads falling within the sliding windows, and generating a sequence feature matrix of the chromosome sequence; A sequence feature vector extraction module for inputting the sequence feature matrix into a trained machine learning model to extract the sequence feature vector of the chromosome sequence; A classification detection module for combining the sequence feature vector and the pregnant woman phenotypic feature vector to form a combined feature vector, inputting the combined feature vector into a classification detection model, and obtaining the fetal chromosomal abnormality situation of the pregnant woman to be tested.
45. The system according to claim 44, further comprising an alignment module for aligning the reads of the sequencing data with the reference genome to obtain uniquely aligned reads.
46. The system according to claim 44 or 45, wherein in the data acquisition module, the free nucleic acid fragments are from the peripheral plasma of the pregnant woman, the liver of the pregnant woman, and / or the placenta.
47. The system according to claim 44 or 45, wherein in the data acquisition module, the free nucleic acid fragment is free DNA.
48. The system according to claim 44 or 45, wherein in the data acquisition module, the sequencing data is from ultra-low depth sequencing.
49. The system according to claim 48, wherein the sequencing depth of the ultra-low depth sequencing is 1×, 0.1× or 0.01×.
50. The system according to any one of claims 44, 45, 49, wherein in the data acquisition module, the reads are aligned with a reference genome to obtain uniquely aligned reads.
51. The system according to claim 50, wherein the uniquely aligned reads are corrected for GC content.
52. The system according to claim 50, wherein the uniquely aligned reads are used in subsequent steps.
53. The system according to claim 52, wherein the uniquely aligned reads are reads corrected for GC content.
54. The system according to any one of claims 44, 45, 49, 51-53, wherein in the data acquisition module, the phenotypic characteristic data of the pregnant woman is selected from the combination of one or more of the following: age, gestational week, height, weight, BMI, prenatal biochemical test results, ultrasound examination diagnosis results, and fetal free DNA concentration in plasma.
55. The system according to any one of claims 44, 45, 49, 51-53, wherein in the data acquisition module, the phenotypic characteristic data of the pregnant woman is processed for outliers, missing values, and / or null values.
56. The system according to any one of claims 44, 45, 49, 51-53, wherein in the data acquisition module, the phenotypic data of the pregnant woman sample will be determined as outliers and these outliers will be set as null values if the following records occur: a. or ; b. or ; c. or ; d. or .
57. The system according to claim 55, wherein the missing values and null values are filled using the missForest algorithm.
58. The system according to any one of claims 44, 45, 49, 51-53, 57, wherein in the sequence feature matrix generation module, the chromosomes are chromosome 21, chromosome 18, chromosome 13, and / or sex chromosomes.
59. The system according to any one of claims 44, 45, 49, 51-53, 57 performs in the sequence feature matrix generation module: (2.1) Using a window of length b with a step size to perform overlapping sliding on a chromosome sequence of length L on the reference genome to obtain a plurality of sliding windows, where b is a positive integer and b = [10000, 10000000], t is any positive integer, L is a positive integer, and L ≥ b; (2.2) Counting the reads falling into each of the sliding windows to generate a sequence feature matrix of the chromosome sequence.
60. The system according to any one of claims 44, 45, 49, 51-53, 57, wherein in the sequence feature matrix generation module, the sequence feature matrix includes the number of reads, base quality, and alignment quality within a sliding window.
61. The system according to claim 60, wherein the base quality includes the mean, standard deviation, skewness, and / or kurtosis of the base quality.
62. The system according to claim 60, wherein the alignment quality includes the mean, standard deviation, skewness, and / or kurtosis of the alignment quality.
63. The system according to any one of claims 44, 45, 49, 51-53, 57, 61, 62, wherein in the sequence feature matrix generation module, the sequence feature matrix is: Among them represents the number of sliding times of the window represents the number of sequence features within a single sliding window represents the j-th sequence feature value in the i-th sliding window 64. The system according to any one of claims 44, 45, 49, 51-53, 57, 61, 62, wherein in the sequence feature vector extraction module, the sequence feature matrix is normalized.
65. The system according to any one of claims 44, 45, 49, 51 - 53, 57, 61, 62, in the sequence feature vector extraction module, the normalization of the sequence feature matrix is performed using formula (I): (I) Among them, is a sample The sequence feature matrix after standardization processing, represents the sample the j-th sequence feature value in the i-th sliding window of, and respectively represent the mean and standard deviation of the j-th sequence feature value in the i-th sliding window of all samples.
66. The system according to any one of claims 44, 45, 49, 51 - 53, 57, 61, 62, in the sequence feature vector extraction module, the trained machine learning model is a neural network model or an AutoEncoder model.
67. The system according to claim 66, wherein the neural network model is a deep neural network model.
68. The system according to claim 66, wherein the neural network model is a deep neural network model based on one - dimensional convolution.
69. The system according to any one of claims 44, 45, 49, 51 - 53, 57, 61, 62, 67, 68, in the classification detection module, the sequence feature vector is spliced with the pregnant woman phenotype to obtain a combined feature vector.
70. The system according to any one of claims 44, 45, 49, 51 - 53, 57, 61, 62, 67, 68, in the classification detection module, the combined feature vector is normalized as follows: wherein is the i-th sequence eigenvalue in the combined feature vector x after standardization, is the i-th sequence eigenvalue in the combined feature vector x, is the mean of the i-th sequence eigenvalue in the combined feature vector, is the standard deviation of the i-th sequence eigenvalue in the combined feature vector.
71. The system according to any one of claims 44, 45, 49, 51 - 53, 57, 61, 62, 67, 68, in the classification detection module, the classification model is an ensemble learning model.
72. The system according to claim 71, wherein the ensemble learning model is an ensemble learning model based on stacking or majority voting.
73. The system according to claim 72, wherein the ensemble learning model is one or more of the following: support vector machine model, naive Bayes classifier, random forest classifier, XGBoos, and logistic regression.
74. A system for constructing a classification detection model for detecting fetal chromosomal abnormalities, the system comprising: A data acquisition module, configured to obtain sequencing data of free nucleic acid fragments of a pregnant woman and clinical phenotype characteristic data of the pregnant woman, wherein the sequencing data includes a plurality of reads, and the fetal chromosomal condition of the pregnant woman is known, and the clinical phenotype characteristic data of the pregnant woman forms a pregnant woman phenotype characteristic vector; A sequence feature matrix generation module, configured to perform window partitioning on at least a part of a reference genome chromosome to obtain a plurality of sliding windows, count the reads falling within the sliding windows, and generate a sequence feature matrix of the chromosomal sequence; A sequence feature vector extraction module, configured to construct a training data set by using the sequence feature matrix and the fetal chromosomal condition, and train a machine learning model to extract a sequence feature vector of the chromosomal sequence; A classification detection model acquisition module, configured to form a combined feature vector by combining the sequence feature vector and the pregnant woman phenotype characteristic vector, and train a classification detection model by using the combined feature vector and the fetal chromosomal condition to obtain the trained classification detection model.
75. The system according to claim 74, further comprising an alignment module, configured to align the reads of the sequencing data with a reference genome to obtain uniquely aligned reads.
Citation Information
Patent Citations
Deep learning-based methods, devices, and systems for prenatal testing
WO2019191319A1
Methods and systems for calling ploidy states using a neural network
WO2020018522A1
Method and device for determining chromosome aneuploidy and constructing classification model
CN111226281A
Method for detecting mutation, electronic equipment and computer storage medium
CN111292802A
Method for screening Down's syndrome and neural tube defect before delivery
CN1600265A