A fetal chromosomal aneuploidy screening model construction method, feature combination and detection method

By analyzing the Z-value, 5' end base sequence frequency, and 15Mb sliding window coverage of pregnant women's plasma cfDNA, and combining this with machine learning algorithms to construct a screening model, the problem of high false positive rate in fetal chromosomal aneuploidy screening was solved, achieving more accurate screening results.

CN122455091APending Publication Date: 2026-07-24THE OBSTETRICS & GYNECOLOGY HOSPITAL OF FUDAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE OBSTETRICS & GYNECOLOGY HOSPITAL OF FUDAN UNIV
Filing Date
2026-04-30
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies have a high false-positive rate in fetal chromosomal aneuploidy screening, leading to unnecessary amniocentesis. Furthermore, they fail to fully exploit cfDNA fragment feature information and machine learning is only used for sex chromosome classification, not for building comprehensive autosomal screening models.

Method used

By analyzing low-depth whole-genome sequencing data of cfDNA from pregnant women's plasma samples, the Z-score was calculated and combined with the 5' end base sequence frequency of cfDNA and the 15Mb sliding window coverage to screen for significantly different features. A screening model was constructed using machine learning algorithms, including the random forest algorithm, and the comprehensive screening model was used to reduce the false positive rate.

Benefits of technology

It effectively reduced the false positive rate, improved the positive predictive value, enhanced the ability to identify negative samples, and reduced unnecessary amniocentesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122455091A_ABST
    Figure CN122455091A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of biomedical detection, and in particular to a method for constructing a screening model for fetal chromosomal aneuploidy, a feature combination and a detection method. The method for constructing the screening model comprises: performing low-depth whole genome sequencing on cfDNA of a maternal plasma sample, analyzing 5' end base motif frequency and 15Mb sliding window coverage of the cfDNA; screening cfDNA features that have significant differences between a true positive group and a healthy control group and have no differences between a false positive group and the healthy control group; and constructing the screening model by machine learning using maternal age, Z value and the screened cfDNA features. The present application also provides a feature combination comprising 5' end base motif frequencies of GAGC, GAGG, TATT and TAGG, cfDNA coverage of regions of chr18:30000000-45000000, chr18:60000000-75000000, chr21:30000000-45000000 and chr21:45000000-46709983, and a Z value and maternal age. The present application can effectively reduce the false positive rate of fetal chromosomal aneuploidy screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical detection technology, specifically to a method for constructing a fetal chromosomal aneuploidy screening model, feature combination, and detection method. Background Technology

[0002] Non-invasive prenatal testing (NIPT) for fetal chromosomal aneuploidy analyzes cell-free DNA (cfDNA) in the pregnant woman's peripheral blood and calculates the deviation (Z-score) of the proportion of sequencing data of the target chromosome (such as chromosomes 21, 18, and 13) from the reference value in the normal population. Based on this, it determines whether the fetus has chromosomal abnormalities. This method has been widely used in clinical screening.

[0003] Patent document CN107133495B discloses a method and system for analyzing aneuploidy bioinformation. This method eliminates sequencing bias through three different GC correction strategies, and reduces the false positive rate by combining fetal concentration prediction, Z-score calculation, and maternal microdeletion / microduplication window filtering. Specifically, this patent calculates the Z-score after dividing the genome into 100kb windows for GC correction. For autosomal aneuploidy, three sets of Z-scores are used for comprehensive judgment; for sex chromosome abnormalities, a prediction model is constructed using a support vector machine (SVM) multi-class classification algorithm; simultaneously, by calculating the Z-score of the test sample in each window, windows containing maternal microdeletions or microduplications are filtered before calculating the Z-score, thereby reducing false positives caused by maternal chromosomal abnormalities.

[0004] However, this technical solution does not involve the analysis of cfDNA fragment end sequence characteristics and genome-specific scale window coverage characteristics, and its machine learning method is only applied to sex chromosome classification and not to the construction of an autosomal comprehensive screening model. Summary of the Invention

[0005] The purpose of this invention is to provide a method for constructing a fetal chromosomal aneuploidy screening model, a feature combination method, and a detection method to address the problem of high false positive rates in existing fetal chromosomal aneuploidy screening, leading to unnecessary amniocentesis. Furthermore, this invention provides a method for constructing a fetal chromosomal aneuploidy screening model that fully utilizes the 5' end base sequence frequency of cfDNA and the 15Mb sliding window coverage as screening features, addressing the problem of insufficient mining of cfDNA fragment feature information in existing technologies; and it provides a method for applying machine learning algorithms to the construction of a comprehensive screening model, addressing the problem that existing technologies only use machine learning for sex chromosome classification and not for constructing comprehensive autosomal screening models.

[0006] The technical solution of the present invention is as follows: Firstly, a method for constructing a fetal chromosomal aneuploidy screening model is provided, comprising the following steps: (1) Feature analysis steps: Obtain low-depth whole-genome sequencing data of cfDNA from pregnant women's plasma samples, compare it with the reference genome, calculate and measure the deviation of the proportion of sequencing data of the tested chromosome from the reference value of the normal population, i.e., the Z value, and analyze the frequency of cfDNA 5' end base motif and the coverage of cfDNA in the 15Mb sliding window on the genome of the same sample; The pregnant women's plasma samples include true positive group samples that are positive for fetal chromosomal aneuploidy by non-invasive prenatal testing and are still positive after amniocentesis, false positive group samples that are positive for fetal chromosomal aneuploidy by non-invasive prenatal testing but are negative after amniocentesis, and healthy control group samples that are negative for fetal chromosomal aneuploidy by non-invasive prenatal testing; (2) Feature screening steps: Based on the analysis results of step (1), statistical screening is performed on cfDNA 5' end base sequence and cfDNA coverage abnormal 15Mb sliding window that show significant differences between the true positive group and the healthy control group but no significant differences between the false positive group and the healthy control group to obtain candidate cfDNA fragment features. (3) Model construction steps: Using the pregnant woman’s age, Z value and candidate cfDNA fragment features obtained in step (2), a screening model is constructed through machine learning algorithm to obtain the optimal feature combination and the corresponding screening model.

[0007] Preferably, the candidate cfDNA fragment features in step (2) include the 5' end base sequence frequencies of GAGC, GAGG, TATT, TAGG, and TCCC cfDNA, and the cfDNA coverage in the regions chr18:30000000-45000000, chr18:60000000-75000000, chr21:30000000-45000000, and chr21:45000000-46709983.

[0008] Preferably, the machine learning algorithm in step (3) is the random forest algorithm.

[0009] Preferably, the above-mentioned method for constructing a fetal chromosomal aneuploidy screening model further includes step (4): based on the model results of step (3), the samples of the true positive group, false positive group and healthy control group are classified and verified in the validation set.

[0010] Secondly, a feature combination for fetal chromosomal aneuploidy screening is provided, including: 5' end base sequence frequency of GAGC, GAGG, TATT, TAGG, TCCC cfDNA, cfDNA coverage in the regions chr18:30000000-45000000, chr18:60000000-75000000, chr21:30000000-45000000, chr21:45000000-46709983, as well as Z-score and maternal age.

[0011] Thirdly, it provides the application of the above-mentioned combination of features in the preparation of detection products for fetal chromosomal aneuploidy screening.

[0012] Fourthly, a method for detecting fetal chromosomal aneuploidy is provided, including: (1) Obtain the characteristic combination data of the pregnant woman to be tested, the characteristic combination including the 5' end base sequence frequency of GAGC, GAGG, TATT, TAGG, TCCC cfDNA, the cfDNA coverage in the regions chr18:30000000-45000000, chr18:60000000-75000000, chr21:30000000-45000000, chr21:45000000-46709983, as well as the Z value and the age of the pregnant woman; (2) Input the feature combination data obtained in step (1) into the screening model for processing to obtain the detection results; The screening model is obtained through any of the construction methods described above.

[0013] Preferably, when the screening model outputs a result of 0, it is determined to be a low risk of fetal chromosomal aneuploidy, and when the output result is 1, it is determined to be a high risk of fetal chromosomal aneuploidy.

[0014] Preferably, the 5' end base sequence frequency of the cfDNA is obtained by calculating the proportion of the 4-base sequence at the 5' end of the effective sequencing sequence to the total number of effective sequencing sequences.

[0015] Preferably, the cfDNA coverage of the 15Mb sliding window is obtained by dividing the genome into 15Mb windows and calculating the number of cfDNA fragments aligned to each window for every 1,000,000 sequencing sequences.

[0016] Compared with the prior art, the advantages of the present invention are: (1) A screening model was constructed by introducing the 5' end base sequence frequency of cfDNA and the 15Mb sliding window coverage as new features, combined with the Z-score and the mother's age. According to the verification results of the examples, 10 false positives occurred in the validation set when the Z-score was used alone. After adding the mother's age and three 5' end base sequence frequencies of cfDNA, 3 out of the 10 false positives were correctly predicted as negative, while ensuring that the false negatives were zero. After adding the mother's age, Tau value, and three 15Mb window features, 6 out of the 10 false positives were correctly predicted as negative. For trisomy 21 screening, after adding the mother's age and two 5' end base sequence frequencies of cfDNA, 2 out of 7 false positives were correctly predicted as negative. After adding the mother's age, Tau value, and two 15Mb windows, 5 out of 7 false positives were correctly predicted as negative. The above data show that the present invention can effectively correct the false positive results generated by using the Z-score alone in the prior art.

[0017] (2) Machine learning (random forest) was applied to the construction of a comprehensive screening model. The frequency of the 5' end base sequence of cfDNA, the coverage of the 15Mb window, the Z-value, and the maternal age were used as multidimensional features input to the model for comprehensive judgment. According to the verification results of the examples, in the three trisomy screenings, the positive predictive value of the model constructed by combining the Z-value, the maternal age, and the three 5' end base sequence frequencies of cfDNA increased from 0.7727 when using the Z-value alone to 0.8293; the positive predictive value of the model constructed by combining the Z-value, the maternal age, the Tau value, and the three 15Mb windows increased from 0.7727 when using the Z-value alone to 0.8947. In trisomy 21 screening, the model constructed by combining Z-score, maternal age, and the frequencies of two cfDNA 5' end bases improved its positive predictive value from 0.7813 to 0.8333; the model constructed by combining Z-score, maternal age, Tau value, and two 15Mb windows improved its positive predictive value from 0.7813 to 0.9259. The improved positive predictive value indicates that the model's ability to identify negative samples (i.e., healthy samples) is enhanced, directly reflecting a reduction in the false positive rate.

[0018] (3) A 15Mb sliding window was used to analyze cfDNA coverage. The window size is 150 times larger than the 100kb window disclosed in CN107133495B. Its purpose is to directly analyze the differences in cfDNA coverage in specific regions of the genome, rather than for GC correction. According to the screening results of the examples, the cfDNA coverage of the regions chr18:30000000-45000000, chr18:60000000-75000000, chr21:30000000-45000000, and chr21:45000000-46709983 showed significant differences between the true positive group and the healthy control group (P<0.05), while there was no significant difference between the false positive group and the healthy control group. This indicates that the cfDNA coverage of these regions can serve as an effective feature to distinguish between true positives and false positives.

[0019] (4) By comparing the true positive group, the false positive group, and the healthy control group, differential analysis was performed to screen out cfDNA features that showed significant differences only between the true positive group and the healthy control group, but not between the false positive group and the healthy control group. This screening strategy ensures that the selected features can distinguish between true positive samples and false positive samples, rather than just between positive samples and negative samples. According to the example, the cfDNA 5' end base sequence (GAGC, GAGG, TATT, TAGG, TCCC) and 15Mb window (chr18:30000000-45000000, chr18:60000000-75000000, chr21:30000000-45000000, chr21:45000000-46709983) screened according to this strategy all met the above differential conditions, and their introduction into the model can effectively correct the misjudgment of false positive samples. Attached Figure Description

[0020] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 The image shows the model performance of screening for three types of fetal chromosomal aneuploidy in the validation set, which combines Z-value, maternal age, and cfDNA 5' end base sequence frequency characteristics in the embodiments of the present invention. Figure 2 The image shows the model performance of trisomy 21 screening in the validation set, which combines Z-score, maternal age, and cfDNA 5' end base sequence frequency characteristics in the embodiments of the present invention. Figure 3 The image shows the model performance of screening for three types of fetal chromosomal aneuploidy in the validation set, combining Z-value, maternal age, cfDNA coverage diversity within 15Mb windows, and specific 15Mb window coverage features in the embodiments of the present invention. Figure 4This is a diagram showing the model performance of trisomy 21 screening in the validation set, combining Z-value, maternal age, cfDNA coverage diversity within the 15Mb window, and specific 15Mb window coverage features in the embodiments of the present invention. Detailed Implementation

[0021] The present invention will be further described in detail below with reference to specific embodiments: I. Origin and characteristics of cfDNA The cell-free DNA (cfDNA) in a pregnant woman's peripheral blood mainly originates from two sources: maternal cell apoptosis and placental trophoblast cell apoptosis. Placental cells originate from the fetus, therefore their DNA carries the fetus's genetic information. These cells naturally die and fragment during normal physiological processes, releasing DNA fragments of approximately 150-200 bp into the bloodstream.

[0022] The generation of cfDNA fragments is not random cutting, but rather regulated by the intracellular chromatin structure and nuclease activity. Chromatin is not uniformly distributed within the cell nucleus; some regions are loosely open (euchromatin), while others are tightly folded (heterochromatin). Nucleases are more likely to cut DNA in open chromatin regions. Therefore, the terminal sequence of cfDNA fragments (i.e., the four bases at the 5' end) is not random, but reflects a DNA sequence preference at that location during cleavage. This preference is related to cell type, physiological state, and the presence of chromosomal abnormalities.

[0023] In cases of fetal chromosomal aneuploidy (such as trisomy 21), fetal cells differ from normal fetal cells in gene expression levels and chromatin openness. These differences affect the fragmentation process of cfDNA, thereby altering the frequency distribution of the 5' terminal 4-base motif. Therefore, analyzing the frequency of the 5' terminal motif in cfDNA can provide additional information about the fetal chromosomal status.

[0024] II. Distribution characteristics of cfDNA in the genome The human genome was divided into several continuous windows, each 15 Mb in length. Since the sequenced cfDNA fragments originated from different locations in the genome, the number of cfDNA fragments contained in each window reflected the DNA abundance of that genomic region ("abundance" refers to the number of DNA fragments in that region).

[0025] In normal samples, the abundance of cfDNA in each window is relatively stable. When the fetus has chromosomal aneuploidy, such as trisomy 21, it means that the copy number of chromosome 21 in the fetal cells is one more than normal. Therefore, the total amount of cfDNA fragments derived from chromosome 21 in the maternal plasma will increase accordingly. This increase is not uniformly distributed across the entire chromosome, but may be more pronounced in specific regions. The 15Mb window size used in this invention is a scale determined through experimental screening: a window that is too small (e.g., 100kb) is greatly affected by random fluctuations, while a window that is too large (e.g., across the entire chromosome) will lose local information. The 15Mb window can reflect the abundance changes in local regions and has sufficient statistical stability.

[0026] By analyzing the cfDNA coverage (i.e., the number of normalized fragments) within each 15Mb window, regions with abnormal abundance in true positive samples can be identified. For example, the chr21:30000000-45000000 and chr21:45000000-46709983 regions screened in this patent showed significantly higher coverage in true positive samples of trisomy 21 than in healthy controls, while showing no significant difference from healthy controls in false positive samples.

[0027] III. Principle Analysis Current clinical screening primarily relies on the Z-score, which measures whether the number of cfDNA fragments on an entire chromosome (such as chromosome 21) deviates from the normal range. However, an elevated Z-score does not necessarily indicate fetal chromosomal aneuploidy. Maternal chromosomal abnormalities (such as microduplications on maternal chromosome 21) and placental mosaicism can also lead to abnormal Z-scores, which are the main sources of false positives.

[0028] This invention supplements the Z-value information by introducing two new types of features: The first type of characteristic is the frequency of the 5' end base sequence of cfDNA. This characteristic reflects the pattern of DNA fragmentation and is directly related to the state of fetal cells, rather than simply reflecting the DNA copy number. When the Z-score is elevated due to maternal abnormalities rather than fetal abnormalities, this characteristic may remain normal.

[0029] The second type of feature is the 15Mb sliding window coverage. This feature reflects the DNA abundance of a specific genomic region. In cases of fetal aneuploidy, certain regions of the abnormal chromosome will exhibit characteristic abundance changes; however, the increased Z-score caused by maternal abnormalities or placental mosaicism may show different abundance change patterns.

[0030] By inputting the two new features mentioned above, along with the Z-value and the pregnant woman's age, into the machine learning model, the model can learn the difference distribution between true positive samples and false positive samples in the multidimensional feature space, thereby more accurately distinguishing between the two and reducing the false positive rate.

[0031] Example

[0032] 1. Sample collection and data acquisition 1.1 Sample inclusion criteria True positive group: Singleton pregnancies where non-invasive prenatal testing shows positive fetal chromosomal aneuploidy (Z-value), and amniocentesis diagnosis is followed by G-banding karyotype analysis which is still positive.

[0033] False positive group: Pregnant women with singleton pregnancies, non-invasive prenatal testing showed positive results for fetal chromosomal aneuploidy (Z-score), but amniocentesis and G-banding karyotype analysis showed negative results.

[0034] Healthy control group: Pregnant women with singleton pregnancies, whose non-invasive prenatal testing results for fetal chromosomal aneuploidy (Z-score) were negative, and whose babies were growing well at birth and had no record of major disease diagnoses within 6 months after birth.

[0035] Exclusion criteria: (1) The pregnant woman has received allogeneic blood transfusion within 1 year or immunotherapy with the introduction of exogenous DNA within 4 weeks; (2) The pregnant woman has undergone transplantation surgery or stem cell therapy; (3) The fetal ultrasound examination indicates structural abnormalities and prenatal diagnosis is required; (4) The pregnant woman herself is a patient with a single gene disease, has given birth to a child with a single gene disease or has a family history of single gene disease; (5) The pregnant woman has a malignant tumor during pregnancy; (6) The pregnant woman is pregnant with twins or multiples, including twins or multiples that have stopped developing or reduced in number.

[0036] 1.2 Sample Collection and Processing Peripheral blood (10 mL) was collected from pregnant women meeting the inclusion criteria and placed in an EDTA anticoagulant tube. Centrifugation was performed within 96 hours of blood collection: first centrifugation at 1600g, 4℃, for 10 minutes to separate the supernatant plasma; second centrifugation at 16000g, room temperature, for 10 minutes to remove residual cellular components. 4 mL of plasma was used for subsequent cfDNA extraction. A commercial DNA extraction kit (magnetic bead method) was used according to the instructions. The extracted cfDNA concentration was determined using a Qubit 3.0 fluorometer (Life Technologies).

[0037] 1.3 Sample Statistics A total of 1,430 samples were selected according to the above criteria, and the specific statistics are shown in Table 1.

[0038] Table 1: Sample Size Statistics

[0039] 1.4 Plasma cell-free DNA sequencing and Z-score calculation Low-depth whole-genome sequencing was performed on the extracted cfDNA using a sequencing platform. Sequencing library construction was performed according to the manufacturer's instructions, with a single-end read length of 36 bp and an average sequencing depth of 6 M reads / sample. The raw sequencing data was converted to FastQ format using software and aligned with the human reference genome GRCh38 using BWA software, with PCR duplicates removed. The number of uniquely mapped reads on each chromosome was counted, and the percentage of unique reads on each chromosome relative to the total number of unique reads on all autosomes was calculated. The Z-score was calculated using the formula: Z i =(x i -μ i ) / σ i Where i is the chromosome number, x i The percentage of unique reads on chromosome i of the sample to be tested relative to the entire genome, μ. i σ represents the mean percentage of unique reads on chromosome i across the entire genome in the healthy control group. i The standard deviation of the percentage of unique reads on chromosome i in the whole genome in the healthy control group.

[0040] 2. Characterization of cell-free DNA fragments in plasma The acquired low-depth whole-genome sequencing data were compared with the reference genome GRCh38, and after removing multiple matching results and PCR replicates, cfDNA fragment feature analysis was performed.

[0041] 2.1 Screening Criteria for Valid Sequencing Sequences After alignment, sequences that meet the following conditions are selected as valid sequencing sequences: (1) alignment quality value ≥ 30; (2) unique alignment to a single position in the reference genome; (3) non-PCR repetitive sequences.

[0042] 2.2 Analysis of 5' terminal base sequence frequency in cfDNA For each valid sequencing sequence, the four bases at its 5' end are extracted to form a 4-base motif. Since DNA is composed of four bases (A, T, C, G), there are 256 possible combinations of 4-base motifs. For each sample, the frequency of each of the 256 4-base motifs is counted and divided by the total number of valid sequencing sequences in that sample to obtain the frequency of each 4-base motif. The frequency calculation formula is: F... k =N k / N total , where F k Let N be the frequency of the k-th motif. k Let N be the number of occurrences of the k-th motif. total This represents the total number of valid sequencing sequences.

[0043] 2.3 Genomic 15Mb sliding window cfDNA coverage analysis The human reference genome GRCh38, chromosomes 1-22, were divided into 15Mb sliding windows with a sliding step size of 15Mb. For each window, the number of cfDNA fragments that were effectively aligned to the sequence was counted. To eliminate the influence of differences in sequencing depth among different samples, the number of cfDNA fragments within each window was standardized: C i =(R i / R total )×1000000, where C i R is the normalized coverage of the i-th window. i To determine the number of cfDNA fragments aligned to the i-th window, R total This represents the total number of valid sequencing sequences in this sample. The standardized C0 i This indicates the number of cfDNA fragments that align to this window out of every million sequencing sequences.

[0044] 3. Screening for fetal chromosomal aneuploidy characteristics 3.1 Difference Analysis For samples from the true positive group, false positive group, and healthy control group, nonparametric tests were performed using the `wilcox.test` function in R software to compare differences in the 5' end base sequence frequency of cfDNA and the 15Mb sliding window normalized coverage among the groups. Multiple test corrections were performed using the Benjamini-Hochberg method, and a p-value <0.05 was considered statistically significant.

[0045] The screening criteria were: (1) there was a significant difference between the true positive group and the healthy control group (P<0.05 after correction); (2) there was no significant difference between the false positive group and the healthy control group (P≥0.05 after correction).

[0046] For the combined analysis of the three trisomy types (T13, T18, and T21), 16 cfDNA 5' end base sequences and 7 15Mb windows that met the criteria were obtained. For the individual analysis of trisomy 21 (with only T21 samples as positive), 21 cfDNA 5' end base sequences and 2 15Mb windows that met the criteria were obtained.

[0047] 3.2 Logistic Regression Screening Collect the ages (in years) of pregnant women in the sample and the calculated Z-scores. Perform binomial logistic regression analysis using the glm function in R software.

[0048] For the combined analysis of three types of trisomy, the frequencies of 16 differentially expressed cfDNA 5' end base sequences, maternal age, and Z-score were used as independent variables, and the sample category (true positive = 1, false positive + healthy control = 0) was used as the dependent variable, and logistic regression was performed. Features contributing significantly were selected from the p-values ​​(P < 0.05) of the regression coefficients, yielding the Z-score (P < 0.001), maternal age (P = 0.001), and three cfDNA 5' end base sequences: GAGC (P = 0.031), GAGG (P = 0.008), and TATT (P = 0.022).

[0049] For the trisomy 21 analysis, the frequencies of 5' end bases of 21 cfDNAs that met the differential criteria, the maternal age, and the Z-score were used as independent variables for logistic regression. The results were obtained as follows: Z-score (P<0.001), maternal age (P<0.001), and two 5' end bases of cfDNAs: TAGG (P=0.020) and TCCC (P=0.030).

[0050] 3.3 Diversity Analysis of 15Mb Window Coverage Based on Equation 1, the cfDNA coverage diversity value Tau(τ) for each sample aligned to different sliding windows across the entire genome (15 Mb) was calculated: (Formula 1) in,

[0051] Where Xi represents the cfDNA coverage in the i-th 15Mb window, and n is the total number of 15Mb windows in the genome. The Tau value is between 0 and 1, reflecting the uniformity of cfDNA coverage distribution across different windows.

[0052] For the combined analysis of the three types of trisomy, the standardized coverage, Tau value, maternal age, and Z-score of the seven 15Mb windows that met the differential criteria were used as independent variables in logistic regression. The features that contributed the most were: Z-score (P<0.001), maternal age (P=0.040), Tau value (P=0.002), and three 15Mb windows: chr18:30000000-45000000 (P<0.001), chr18:60000000-75000000 (P=0.003), and chr21:30000000-45000000 (P<0.001).

[0053] For the trisomy 21 analysis, the standardized coverage, Tau value, maternal age, and Z-score of two 15Mb windows that met the difference criteria were used as independent variables in logistic regression. The features that contributed the most were: Z-score (P<0.001), maternal age (P<0.001), Tau value (P<0.001), and the two 15Mb windows: chr21:30000000-45000000 (P<0.001) and chr21:45000000-46709983 (P<0.001).

[0054] 4. Construction and validation of the screening model 4.1 Dataset Partitioning The data in Table 1 were randomly divided into training and validation sets in a 1:1 ratio. To ensure the balance of the partitioning, stratified sampling was used to maintain a consistent proportion of true positives, false positives, and healthy controls in both the training and validation sets. The false positives and healthy controls were merged into a single negative sample. Specific statistics are shown in Table 2.

[0055] Table 2: Statistics on the Number of Model Datasets

[0056] 4.2 Model Construction and Validation Based on cfDNA 5' Terminal Base Motif Use the randomForest package in R to build a random forest classification model. Model parameter settings: number of trees ( n The number of features randomly selected when splitting each decision tree is 200 (mtry) and the random forest classification weight (classwt) is 1:1.

[0057] For the analysis of three types of trisomy, the input features are: Z-score, maternal age, GAGC frequency, GAGG frequency, and TATT frequency, for a total of five features. The model is trained using the training set data, and parameters are tuned using 5-fold cross-validation.

[0058] Model predictions were performed on the validation set. The model prediction results were compared with the results of amniocentesis karyotype analysis (gold standard). When using the Z-score alone as the criterion (in the NIPT field, |Z|>3 is usually used as the high-risk threshold), all 34 true positives in the validation set were correctly detected, and 10 false positives occurred among the 681 negatives (specificity 98.53%).

[0059] Using the model trained in this embodiment (input features include Z-score, maternal age, and frequencies of three cfDNA 5' end bases) for prediction, under the condition of zero false negatives (i.e., all 34 true positives were correctly detected), 3 out of 10 false positives in the validation set were correctly predicted as negative by the model (e.g., ...). Figure 1 ).

[0060] like Figure 1 As shown, the prediction results statistics table and confusion matrix on the left indicate that "confirmed negative" represents actually healthy pregnant women (681 cases in total), and "confirmed positive" represents actually sick pregnant women (34 cases in total). The prediction results using the Z-value alone in the left table are: 671 true negatives were correctly predicted as negative, 10 true negatives were incorrectly predicted as positive (10 false positives), 34 true positives were correctly detected, and 0 true positives were incorrectly misclassified as negative (0 false negatives). The table on the right shows the prediction results of the model of this invention: 674 true negatives were correctly predicted as negative, 7 true negatives were incorrectly predicted as positive (7 false positives), 34 true positives were correctly detected, and 0 true positives were incorrectly misclassified as negative (0 false negatives). Comparing the number of predicted positives in the two tables shows that the model of this invention reduces the number of false positives from 10 to 7.

[0061] For Trisomy 21 screening, the input features are: Z-score, maternal age, TAGG frequency, and TCCC frequency, a total of four features. The model is trained using the training set data. When using the Z-score alone, all 25 true positives in the validation set were correctly detected, and 7 false positives occurred out of 678 negatives (specificity 98.97%). Using the model trained in this example for prediction, under the condition of zero false negatives, 2 out of the 7 false positives were correctly predicted as negative (e.g., ...). Figure 2 ).

[0062] like Figure 2 As shown, the validation results of the above-mentioned Trisomy 21 screening are presented. Figure 2 The prediction results statistics table and confusion matrix on the left show that "confirmed negative" represents actually healthy pregnant women (678 cases in total), and "confirmed positive" represents actually sick pregnant women (25 cases in total). The prediction results using the Z-value alone in the left table are: 671 true negatives were correctly predicted as negative, and 7 true negatives were incorrectly predicted as positive (7 false positives). 25 true positives were correctly detected, and 0 true positives were incorrectly misclassified as negative (0 false negatives). The table on the right shows the prediction results of the model of this invention: 673 true negatives were correctly predicted as negative, and 5 true negatives were incorrectly predicted as positive (5 false positives). 25 true positives were correctly detected, and 0 true positives were incorrectly misclassified as negative (0 false negatives). Comparing the number of predicted positives in the two tables shows that the model of this invention reduces the number of false positives from 7 to 5.

[0063] The overall performance metrics of the model are shown in Table 3.

[0064] Table 3: Overall Performance Indicators of the Model

[0065] Among them, sensitivity = true positive / (true positive + false negative); specificity = true negative / (true negative + false positive); accuracy = (true positive + true negative) / total number of samples; positive predictive value = true positive / (true positive + false positive); negative predictive value = true negative / (true negative + false negative).

[0066] 4.3 Model Construction and Validation Based on 15Mb Window Coverage Use the same random forest parameter settings as in 4.2.

[0067] For the combined analysis of three types of trisomy, the input features are: Z-score, maternal age, Tau value, chr18 window coverage (30000000-45000000), chr18 window coverage (60000000-75000000), and chr21 window coverage (30000000-45000000), for a total of six features. The model is trained using the training set data.

[0068] When using the Z-score alone, 10 false positives were observed in the validation set. Using the model trained in this embodiment for prediction, assuming a false negative rate of 0, 6 out of the 10 false positives were correctly predicted as negative (e.g., ...). Figure 3 ).

[0069] like Figure 3 As shown, the prediction results statistics table and confusion matrix on the left indicate that "confirmed negative" represents actually healthy pregnant women (681 cases in total), and "confirmed positive" represents actually sick pregnant women (34 cases in total). The prediction results using Z-values ​​alone in the left table are: 671 true negatives were correctly predicted as negative, 10 true negatives were incorrectly predicted as positive (10 false positives), 34 true positives were correctly detected, and 0 true positives were incorrectly misclassified as negative (0 false negatives). The table on the right shows the prediction results of the model of this invention: 677 true negatives were correctly predicted as negative, 4 true negatives were incorrectly predicted as positive (4 false positives), 34 true positives were correctly detected, and 0 true positives were incorrectly misclassified as negative (0 false negatives). Comparing the predicted positive numbers in the left and right tables in the scatter plot of the sample Z-value distribution shows that the model of this invention reduces the number of false positives from 10 to 4.

[0070] For Trisomy 21 screening, the input features are: Z-score, maternal age, Tau-score, chr21:30000000-45000000 window coverage, and chr21:45000000-46709983 window coverage, for a total of 5 features. The model is trained using the training set data. When using the Z-score alone, 7 false positives appeared in the validation set. Using the model trained in this embodiment for prediction, under the condition of ensuring 0 false negatives, 5 out of the 7 false positives were correctly predicted as negative (e.g., ...). Figure 4 ).

[0071] like Figure 4 As shown, the prediction results statistics table and confusion matrix on the left indicate that "confirmed negative" represents actually healthy pregnant women (678 cases in total), and "confirmed positive" represents actually sick pregnant women (25 cases in total). The prediction results using Z-values ​​alone in the left table are: 671 true negatives were correctly predicted as negative, 7 true negatives were incorrectly predicted as positive (7 false positives), 25 true positives were correctly detected, and 0 true positives were incorrectly misclassified as negative (0 false negatives). The table on the right shows the prediction results of the model of this invention: 676 true negatives were correctly predicted as negative, 2 true negatives were incorrectly predicted as positive (2 false positives), 25 true positives were correctly detected, and 0 true positives were incorrectly misclassified as negative (0 false negatives). Comparing the predicted positive numbers in the left and right tables in the scatter plot of the sample Z-value distribution shows that the model of this invention reduces the number of false positives from 7 to 2.

[0072] The overall performance metrics of the model are shown in Table 4.

[0073] Table 4: Overall Performance Indicators of the Model

[0074] 5. Detection Method 5.1 Acquisition of Feature Data of the Sample to be Tested Taking a pregnant woman at 21 weeks + 1 day of gestation as an example, 10 mL of her peripheral blood was collected, and plasma and cfDNA were separated according to the method in Example 1.2. Low-depth whole genome sequencing was performed according to the method in Example 1.4.

[0075] After alignment and quality control, 3,791,086 valid sequencing sequences were obtained. Following the method in Example 2.2, the frequencies of 256 four-base motifs were statistically analyzed. The frequencies of five motifs—GAGC, GAGG, TATT, TAGG, and TCCC—were extracted. GAGC frequency = 5892 / 3791086 = 0.00155417 GAGG frequency = 9336 / 3791086 = 0.00246262 TATT frequency = 29464 / 3791086 = 0.00777192 TAGG frequency = 6132 / 3791086 = 0.00161748 TCCC frequency = 7834 / 3791086 = 0.00206643 Following the method in Example 2.3, the cfDNA coverage of the 15Mb sliding window was calculated. Standard coverage was extracted from four regions: The coverage of the chr18:30000000-45000000 area is calculated as 21407 / 3791086 × 1000000 = 5646.7 chr18:60000000-75000000 Area Coverage = 21180 / 3791086 × 1000000 = 5586.8 The coverage of the chr21:30000000-45000000 area is calculated as 20072 / 3791086 × 1000000 = 5294.5 The coverage of the region chr21:45000000-46709983 is 2242 / 3791086×1000000=591.4 The Z value was calculated according to the method in Example 1.4. The Z21 value for this sample was 4.76.

[0076] The pregnant woman is 24 years old.

[0077] 5.2 Model Prediction The above nine features (GAGC frequency, GAGG frequency, TATT frequency, TAGG frequency, TCCC frequency, chr18: 30000000-45000000 coverage, chr18: 60000000-75000000 coverage, chr21: 30000000-45000000 coverage, chr21: 45000000-46709983 coverage, Z21 value, and maternal age) are input into the constructed screening model (using a random forest model with the same parameters as in 4.2).

[0078] The model output is 0.

[0079] 5.3 Result Validation The pregnant woman subsequently underwent amniocentesis, and G-banding karyotype analysis of the amniotic fluid cells showed normal results. The model prediction results were consistent with the karyotype analysis results.

[0080] The above embodiments are merely illustrative of the technical concept and features of the present invention, intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly, and should not be construed as limiting the scope of protection of the present invention. It will be apparent to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects. The scope of the present invention is defined by the appended claims rather than the foregoing description, and thus all changes falling within the meaning and scope of the equivalents of the claims are intended to be included within the present invention.

Claims

1. A method for constructing a fetal chromosomal aneuploidy screening model, comprising the following steps: (1) Feature analysis steps: Obtain low-depth whole-genome sequencing data of cfDNA from pregnant women's plasma samples, compare it with the reference genome, calculate and measure the deviation of the proportion of sequencing data of the tested chromosome from the reference value of the normal population, i.e., the Z value, and analyze the frequency of cfDNA 5' end base motif and the coverage of cfDNA in the 15Mb sliding window on the genome of the same sample; The pregnant women's plasma samples include true positive group samples that are positive for fetal chromosomal aneuploidy by non-invasive prenatal testing and are still positive after amniocentesis, false positive group samples that are positive for fetal chromosomal aneuploidy by non-invasive prenatal testing but are negative after amniocentesis, and healthy control group samples that are negative for fetal chromosomal aneuploidy by non-invasive prenatal testing; (2) Feature screening steps: Based on the analysis results of step (1), statistical screening is performed on cfDNA 5' end base sequence and cfDNA coverage abnormal 15Mb sliding window that show significant differences between the true positive group and the healthy control group but no significant differences between the false positive group and the healthy control group to obtain candidate cfDNA fragment features. ( 3) Model construction steps: Using the pregnant woman's age, Z value and the candidate cfDNA fragment features obtained in step (2), a screening model is constructed through machine learning algorithm to obtain the optimal feature combination and the corresponding screening model.

2. The method for constructing a fetal chromosomal aneuploidy screening model according to claim 1, characterized in that, The candidate cfDNA fragment characteristics mentioned in step (2) include the 5' end base sequence frequencies of GAGC, GAGG, TATT, TAGG, and TCCC cfDNA, as well as the cfDNA coverage in the regions chr18:30000000-45000000, chr18:60000000-75000000, chr21:30000000-45000000, and chr21:45000000-46709983.

3. The method for constructing a fetal chromosomal aneuploidy screening model according to claim 1, characterized in that, The machine learning algorithm mentioned in step (3) is the random forest algorithm.

4. The method for constructing a fetal chromosomal aneuploidy screening model according to claim 1, characterized in that, It also includes step (4): based on the model results of step (3), classify and validate the true positive group, false positive group and healthy control group samples in the validation set.

5. A combination of features for screening fetal chromosomal aneuploidy, characterized in that, include: Frequency of 5' end base motifs of GAGC, GAGG, TATT, TAGG, and TCCC cfDNA; cfDNA coverage in the regions chr18:30000000-45000000, chr18:60000000-75000000, chr21:30000000-45000000, chr21:45000000-46709983; Z-score and maternal age.

6. The application of the feature combination according to claim 5 in the preparation of a detection product for fetal chromosomal aneuploidy screening.

7. A method for detecting fetal chromosomal aneuploidy, comprising: (1) Obtain the characteristic combination data of the pregnant woman to be tested, the characteristic combination including the 5' end base sequence frequency of GAGC, GAGG, TATT, TAGG, TCCCcfDNA, the cfDNA coverage in the regions chr18:30000000-45000000, chr18:60000000-75000000, chr21:30000000-45000000, chr21:45000000-46709983, as well as the Z value and the age of the pregnant woman; (2) Input the feature combination data obtained in step (1) into the screening model for processing to obtain the detection results; The screening model is obtained by the construction method described in any one of claims 1-4.

8. The method for detecting fetal chromosomal aneuploidy according to claim 7, characterized in that, When the screening model outputs a result of 0, it is determined that the fetus has a low risk of chromosomal aneuploidy; when the output result is 1, it is determined that the fetus has a high risk of chromosomal aneuploidy.

9. The method for detecting fetal chromosomal aneuploidy according to claim 7, characterized in that, The frequency of the 5' end base motif of the cfDNA is obtained by calculating the proportion of the 4-base motif at the 5' end of the effective sequencing sequence to the total number of effective sequencing sequences.

10. The method for detecting fetal chromosomal aneuploidy according to claim 7, characterized in that, The cfDNA coverage of the 15Mb sliding window is obtained by dividing the genome into 15Mb windows and calculating the number of cfDNA fragments that are aligned to each window for every 1,000,000 sequencing sequences.

Citation Information

Patent Citations

  • CN107133495B