Methods, devices and storage media for detecting fetal chromosomal aneuploidy

By combining fetal DNA concentration and chimerism using a machine learning model to calculate a new Z-value, the problems of low positive predictive value and high retest rate in existing fetal chromosomal aneuploidy abnormality detection are solved, achieving higher detection accuracy and stability.

CN115223654BActive Publication Date: 2025-12-02BGI GENOMICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210825534.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2025-12-02
Estimated Expiration
2042-07-13

AI Technical Summary

Technical Problem

Existing methods for detecting fetal chromosomal aneuploidy have problems such as low positive predictive value and high retest rate. In particular, the sensitivity and specificity of detection for Down syndrome, Edwards syndrome and Patau syndrome are insufficient, resulting in high false positive rate and high rate of repeated blood draws, which affects the health of pregnant women and the efficiency of testing.

Method used

A new Z-value (Znew) is calculated by combining fetal DNA concentration, Z-value, and chimerism using a machine learning model. A linear discriminant analysis model is then used to improve detection accuracy and reduce false positive and retest rates.

Benefits of technology

It improves the accuracy of fetal chromosomal aneuploidy detection, reduces false positive and retest rates, meets regulatory and clinical requirements, and enhances the stability and efficiency of test results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115223654B_ABST
    Figure CN115223654B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, and storage medium for detecting fetal chromosomal aneuploidy. The method for detecting fetal chromosomal aneuploidy includes calculating a new Z-value for the sample based on the fetal DNA concentration, Z-value, and chimerism of the cell-free DNA in the pregnant woman's blood; the new Z-value is used to determine whether fetal chromosomal aneuploidy has occurred. Chimerism is the ratio of abnormal fetal cells to all fetal cells. This application is the first to incorporate chimerism into the detection of fetal chromosomal aneuploidy, comprehensively considering three variables—fetal DNA concentration, chimerism, and Z-value—to calculate a new Z-value, which improves the accuracy of NIPT detection, provides excellent differentiation between true positive and false positive samples, and reduces false positives. The new Z-value conforms to a normal distribution, meeting current regulatory and clinical requirements, reducing data distribution volatility, thereby reducing the gray area rate, reducing the retest rate, and improving the stability of the test results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of fetal chromosomal aneuploidy detection technology, and in particular to a method, apparatus and storage medium for detecting fetal chromosomal aneuploidy. Background Technology

[0002] Fetal chromosomal aneuploidy means that the fetus's chromosomes are aneuploid. A normal fetus has 23 pairs (46) of chromosomes, which is euploid. If there are chromosome deletions or an increase in chromosomes, forming an aneuploidy, it indicates that there is an abnormality in the fetal chromosomes, namely fetal chromosomal aneuploidy.

[0003] Currently, the most common fetal chromosomal aneuploidies in clinical practice are Down syndrome, Edwards syndrome, and Patau syndrome.

[0004] Down syndrome (trisomy 21) is a genetic disorder caused by trisomy of chromosome 21. Common symptoms include developmental delay, distinctive facial features, and mild to moderate intellectual disability. Currently, there is no effective treatment for Down syndrome; only care and education can improve the quality of life for affected individuals. Besides Down syndrome, other common fetal chromosomal aneuploidies include Edwards syndrome (trisomy 18) and Patau syndrome (trisomy 13), both of which cause severe developmental abnormalities in children.

[0005] The molecular biological mechanism of Down syndrome is the nondisjunction of chromosome 21 during germ cell formation, resulting in three copies of chromosome 21 in the fertilized egg, which in turn leads to a series of abnormalities in molecular and developmental biological processes. Since there is currently no effective treatment for chromosomal aneuploidy syndromes, such as Down syndrome, and no specific behavioral or environmental factors have been identified as contributing to its development, the main approach is to prevent the birth of babies with Down syndrome or other serious genetic disorders through prenatal screening of pregnant women. This involves conducting relevant tests during pregnancy, and if the test results are positive or indicate a high risk, termination of pregnancy is performed to avoid the birth of a trisomy baby.

[0006] Traditional screening assesses trisomy risk using serological markers such as AFP, free β-hCG, uE3, and Inhibin-A. However, because serological markers are indirect indicators and cannot directly reflect fetal chromosomal aneuploidy, their sensitivity and specificity are poor. Around 2010, high-throughput sequencing technology emerged and became widespread. This technology allows for precise detection and quantification of cell-free DNA (cfDNA) in maternal plasma, enabling screening for chromosomal abnormalities, including trisomy 21, based on the relative abundance of target chromosomes (NIPT). In 2015, the *New England Journal of Medicine* published an article analyzing 15,841 samples in a prospective, multicenter clinical trial, demonstrating that NIPT significantly outperformed traditional screening, achieving a sensitivity and specificity exceeding 99.9%. In contrast, traditional serological screening methods had a sensitivity of only 78.9% and a specificity of only 94.6%, proving that NIPT greatly improves the effectiveness of screening for chromosomal aneuploidy syndromes, such as Down syndrome.

[0007] However, the performance of NIPT testing still needs improvement. According to a 2015 article by Zhang et al., an analysis of 112,669 NIPT test results with follow-up revealed two main problems with traditional NIPT testing performance: First, the positive predictive value (PPV) needs improvement. The article showed that the PPV for T21 was 92.2%, while for T18 it was 76.6%, and for T13 it was only 32.8%, indicating a high rate of false positives and a need for improvement in the PPV of traditional NIPT testing. Second, the retest rate is high. The article showed that 3,213 blood samples were drawn from 112,669 samples, a retest rate of 2.8%. A retest means that the initial NIPT test result was in the gray zone, thus failing to provide a negative or positive result, requiring a new blood sample for retesting. In this situation, the pregnant woman not only endures the pain of an extra blood draw, but more importantly, the time it takes for the NIPT test results to be available is prolonged, which may cause the pregnant woman to miss the optimal intervention period, posing a significant risk to her life and health.

[0008] Therefore, improving the positive predictive value of NIPT and reducing the retest rate are the key research areas and challenges in the detection of fetal chromosomal aneuploidy. Summary of the Invention

[0009] The purpose of this application is to provide an improved method, apparatus, and storage medium for detecting fetal chromosomal aneuploidy.

[0010] To achieve the above objectives, this application adopts the following technical solution:

[0011] The first aspect of this application discloses a method for detecting fetal chromosomal aneuploidy, comprising calculating a new Z-value for the test sample based on the concentration of fetal DNA, Z-value, and chimerism in the cell-free DNA of the pregnant woman's blood, and labeling it as Z. new The new Z-value is used to determine whether the fetal chromosomes of the sample to be tested have aneuploidy; the degree of mosaicism is the ratio of abnormal fetal cells to all fetal cells.

[0012] It should be noted that the key to the method for detecting fetal chromosomal aneuploidy in this application lies in calculating a new Z-value, commonly used and recognized in the field, from three variables: fetal DNA concentration, the unique indicator of this application (chimeric degree), and the traditional Z-value. new The "traditional Z-value" refers to the Z-value obtained using conventional methods; the "new Z-value" in this application refers to the Z-value calculated using three variables. The method in this application uses chimerism as the input variable for calculating the "new Z-value," which helps improve the accuracy of NIPT testing, provides excellent differentiation between true positive and false positive samples, and reduces false positives. The new Z-value conforms to a normal distribution, meeting current regulatory and clinical requirements; it also significantly reduces data distribution volatility, thereby lowering the gray zone rate, reducing the retest rate, and improving the stability of test results.

[0013] In one implementation of this application, a new Z-value of the test sample is calculated based on the fetal DNA concentration, Z-value, and chimerism in the cell-free DNA of the pregnant woman's blood. This includes inputting the fetal DNA concentration, Z-value, and chimerism into a fetal chromosomal aneuploidy abnormality detection model to obtain the model output value corresponding to the test sample, and mapping the model output value to obtain the new Z-value of the test sample. The fetal chromosomal aneuploidy abnormality detection model uses several samples with known fetal chromosomal conditions as training samples. The training samples include positive and negative samples of fetal chromosomal aneuploidy abnormalities. The model is trained using fetal DNA concentration, Z-value, and chimerism as inputs to obtain a model output value that comprehensively represents the fetal chromosomal condition using the three variables of fetal DNA concentration, Z-value, and chimerism.

[0014] It is understandable that obtaining a new Z-value through machine learning model training is only one implementation method of this application. It is not excluded that other calculation methods can be used to calculate the new Z-value of this application from fetal DNA concentration, Z-value, and chimerism.

[0015] In one implementation of this application, a new Z-value of the test sample is obtained by mapping the model output value, including calculating the new Z-value of the test sample based on the model output value of the test sample, the positive threshold, the negative threshold, and the median of the model output values ​​of all negative samples; wherein the positive threshold is the threshold of the model output value corresponding to the positive sample, and the negative threshold is the threshold of the model output value corresponding to the negative sample.

[0016] It should be noted that the model output value, or machine learning model-generated value, is a numerical value output by the fetal chromosomal aneuploidy abnormality detection model to assess fetal chromosomal aneuploidy abnormalities. This value cannot be thresholded based on statistical significance like the traditional Z-score; instead, thresholds are defined based on the characteristics of the training data. For example, a negative threshold is defined so that all true positive samples in the training data are not classified as negative, ensuring the model does not produce false negatives. A positive threshold is defined so that as many true positive samples as possible are classified as positive, while as few original false positive samples as possible are classified as positive, thereby reducing false positives and improving the performance of NIPT detection. The area between the positive and negative thresholds is a gray zone. To enable the model output value of the test sample to be directly used to determine the fetal chromosomal aneuploidy abnormality status, this application further projects the model output value as a new Z-score, namely Z0. new Furthermore, the experimental results show that the new Z value obtained by the printing of this application conforms to a normal distribution, and the center of the distribution is located at 0. Therefore, Z>3 can still be used as the criterion for a positive result, and Z<1.96 can still be used as the criterion for a negative result.

[0017] In one implementation of this application, the median of the model output value of negative samples is the median of the model output value of all negative samples obtained by re-inputting all negative training samples into the fetal chromosomal aneuploidy abnormality detection model.

[0018] In one implementation of this application, a new Z-value of the test sample is obtained by mapping the model output value, including the following mapping methods:

[0019] When the model output value of the sample to be tested is greater than the positive threshold, Z new =LD-cut p +3;

[0020] When the model output value of the sample to be tested is less than the positive threshold and greater than the negative threshold,

[0021]

[0022] When the model output value of the sample to be tested is less than the negative threshold,

[0023] In the above formula, Z newThe new Z value, LD is the model output value of the test sample, cut p The positive threshold is cut. n is the negative threshold, and Med is the median of the model output values ​​for all negative samples.

[0024] It should be noted that, in this application, the concept of obtaining a new Z value by imprinting the model output value is as follows:

[0025] 1) When the model output value is greater than the positive threshold, the final converted new Z value needs to be greater than 3, because clinically, the judgment of trisomy positivity is usually based on Z>3 as the threshold.

[0026] 2) When the model output value is in the gray area, the final converted new Z value needs to be between 1.96 and 3, because clinically, Z ~ [1.96, 3) is commonly used as the gray area range.

[0027] 3) When the model output value is less than the negative threshold, the final converted new Z value needs to be less than 1.96, because clinically, the judgment of trisomy negative is usually based on Z<1.96 as the threshold.

[0028] Furthermore, this application guarantees that the median of the new Z-value for negative samples is 0, because the median of a standard normal distribution should be equal to 0; therefore, the formula considers Med, i.e., the median of the model output values ​​for negative samples. According to the above imprinting formula, when the model output value equals the median of the model output values ​​for negative samples, the new Z-value is 0.

[0029] It should also be noted that the specific values ​​in the above formulas are data obtained from one implementation of this application; it is understood that if the training samples change, the data in the corresponding imprinting formula will also change; however, the basic principle of obtaining new Z values ​​through the imprinting formula remains unchanged.

[0030] In one implementation of this application, determining whether the fetal chromosome of the sample to be tested has aneuploidy based on the new Z value includes: a new Z value greater than 3 is considered positive, i.e., the fetal chromosome has aneuploidy; a new Z value less than 1.96 is considered negative, i.e., the fetal chromosome is normal.

[0031] In one implementation of this application, the machine learning model is a linear discriminant analysis (LDA) model.

[0032] In one implementation of this application, the abnormal fetal cell is a cell containing fetal chromosomal aneuploidy.

[0033] In one implementation of this application, the concentration of fetal DNA and the Z-value in cell-free DNA from the pregnant woman's blood are calculated using high-throughput sequencing data of cell-free DNA from the pregnant woman's blood.

[0034] In one implementation of this application, the degree of chimerism is calculated using Formula 1;

[0035] Formula 1

[0036] In Formula 1, Mosaic k Let fra represent the mosaicism of the k-th chromosome. k FF represents the relative fetal concentration of the k-th chromosome, and FF represents the fetal DNA concentration.

[0037] fra k The result was obtained using Formula 2.

[0038] Formula 2

[0039] In Formula 2, fra k The relative fetal concentration of chromosome k. This represents the average depth after correction for the k-th chromosome. This represents the average depth after correction for all autosomes;

[0040] In Formula 1 and Formula 2, the value of k ranges from 1 to 22;

[0041] Mosaic k A value of 0 indicates that the kth chromosome of the fetus is normal; Mosaic k A value of 1 indicates that the fetus's kth chromosome is completely trisomic; Mosaic k A value between 0 and 1 indicates that chromosome k in the fetus is mosaic. In this application, mosaicism of fetal chromosome k means that chromosome k in some fetal cells is in a trisomic state, while chromosome k in other fetal cells is not in a trisomic state. In principle, under a fixed fetal DNA concentration, if the fetus is completely trisomic, the trisomic signal in the maternal peripheral blood is stronger; if the fetus is mosaic trisomic, the trisomic signal in the maternal peripheral blood is weaker. Furthermore, a lower degree of mosaicism generally indicates a false positive due to data fluctuations, while a higher degree of mosaicism generally indicates a true positive.

[0042] In one implementation of this application, the average depth of each chromosome after correction and the average depth of all autosomes after correction are calculated using high-throughput sequencing data of cell-free DNA from the pregnant woman's blood.

[0043] In one implementation of this application, the method for detecting fetal chromosomal aneuploidy includes the following steps:

[0044] The data acquisition steps include acquiring high-throughput sequencing data of cell-free DNA from the blood of the pregnant woman to be tested;

[0045] The data processing steps include calculating the fetal DNA concentration and Z-value based on the high-throughput sequencing data of cell-free DNA from the pregnant woman's blood.

[0046] The steps for calculating chimerism include calculating the chimerism of each chromosome according to Formula 1;

[0047] The new Z-value calculation steps include calculating the new Z-value of the test sample based on the concentration of fetal DNA, Z-value, and chimerism in the cell-free DNA of the pregnant woman's blood.

[0048] The steps for determining fetal chromosomal aneuploidy include determining whether the chromosomes of the fetus being tested have aneuploidy based on the new Z-score.

[0049] It should be noted that the key to the method for detecting fetal chromosomal aneuploidy in this application lies in comprehensively considering three variables—fetal DNA concentration, mosaicism, and the traditional Z-score—through a fetal chromosomal aneuploidy detection model to obtain the model output value, and then converting this value into the Z-score, which is currently commonly used and recognized in the field, i.e., the new Z-score (Z0). new In this application, the "traditional Z-value" refers to the Z-value obtained by the "data processing step" using conventional methods. To better distinguish it from the "new Z-value," the Z-value obtained by the "data processing step" is called the "traditional Z-value," while the Z-value obtained by mapping the model output value is called the "new Z-value." Incorporating chimerism into the machine learning model helps improve the accuracy of NIPT detection, providing excellent differentiation between true positive and false positive samples and reducing false positives. The new Z-value conforms to a normal distribution, meeting current regulatory and clinical requirements. Furthermore, it significantly reduces data distribution volatility, thereby lowering the gray area rate, reducing the retest rate, and improving the stability of test results.

[0050] The second aspect of this application discloses a method for constructing a fetal chromosomal aneuploidy abnormality detection model, which includes using several samples with known fetal chromosomal conditions as training samples. The training samples include positive and negative samples of fetal chromosomal aneuploidy abnormalities. The model is trained using fetal DNA concentration, Z-value, and chimerism as inputs to obtain a model output value that comprehensively represents the fetal chromosomal condition using three variables: fetal DNA concentration, Z-value, and chimerism. The model obtained is the fetal chromosomal aneuploidy abnormality detection model.

[0051] It should be noted that the method for constructing the fetal chromosomal aneuploidy abnormality detection model in this application is actually the same method for constructing the fetal chromosomal aneuploidy abnormality detection model in the method for detecting fetal chromosomal aneuploidy abnormalities in this application. Therefore, the calculation methods for fetal DNA concentration, Z-value and chimerism can refer to the method for detecting fetal chromosomal aneuploidy abnormalities in this application, and will not be repeated here.

[0052] The third aspect of this application discloses an apparatus for detecting fetal chromosomal aneuploidy, comprising a new Z-value calculation module and a fetal chromosomal aneuploidy judgment module; the new Z-value calculation module includes calculating a new Z-value of the test sample based on the concentration of fetal DNA, Z-value, and chimerism in the cell-free DNA of the pregnant woman's blood; the chimerism is the ratio of abnormal fetal cells to all fetal cells; the fetal chromosomal aneuploidy judgment module includes judging whether the fetal chromosomes of the test sample have aneuploidy based on the new Z-value.

[0053] In one implementation of this application, the new Z-value calculation module further includes inputting fetal DNA concentration, Z-value, and chimerism into the fetal chromosomal aneuploidy abnormality detection model to obtain the model output value corresponding to the test sample, and obtaining a new Z-value of the test sample by mapping the model output value; wherein, the fetal chromosomal aneuploidy abnormality detection model uses several samples with known fetal chromosomal conditions as training samples, including positive and negative samples of fetal chromosomal aneuploidy abnormalities, and uses fetal DNA concentration, Z-value, and chimerism as input to perform machine learning model training, thereby obtaining the model; the model output value is used to comprehensively characterize the fetal chromosomal condition by combining the three variables of fetal DNA concentration, Z-value, and chimerism.

[0054] Therefore, in one implementation of this application, the apparatus further includes a model training module. This module uses several samples with known fetal chromosomal information as training samples. The training samples include positive and negative samples of fetal chromosomal aneuploidy. Using fetal DNA concentration, Z-score, and chimerism as inputs, a machine learning model is trained to obtain a model output value that comprehensively represents the fetal chromosomal information using three variables: fetal DNA concentration, Z-score, and chimerism. The resulting model is the fetal chromosomal aneuploidy detection model. Preferably, the machine learning model is a linear discriminant analysis model.

[0055] In one implementation of this application, the new Z-value calculation module includes a model output value analysis submodule and a Z-value imprinting submodule. The model output value analysis submodule includes a method for inputting the fetal DNA concentration, Z-value, and chimerism of the sample to be tested into a fetal chromosomal aneuploidy abnormality detection model to obtain the model output value corresponding to the sample to be tested. The Z-value imprinting submodule includes a method for calculating a new Z-value for the sample to be tested based on the model output value of the sample to be tested, as well as a positive threshold, a negative threshold, and the median of the model output values ​​of all negative samples. The positive threshold is the threshold value of the model output value corresponding to the positive sample, and the negative threshold is the threshold value of the model output value corresponding to the negative sample.

[0056] In one implementation of this application, the Z-value imprinting submodule obtains a new Z-value according to the following method.

[0057] When the model output value of the sample to be tested is greater than the positive threshold, Z new =LD-cut p +3;

[0058] When the model output value of the sample to be tested is less than the positive threshold and greater than the negative threshold,

[0059]

[0060] When the model output value of the sample to be tested is less than the negative threshold,

[0061] In the above formula, Z new The new Z value, LD is the model output value of the test sample, cut p The positive threshold is cut. n The negative threshold is denoted by Med, which is the median of the model output values ​​for all negative samples.

[0062] In one implementation of this application, the fetal chromosomal aneuploidy module determines whether the fetal chromosomes of the sample to be tested have aneuploidy based on a new Z-value, including: a new Z-value greater than 3 indicates a positive result, i.e., fetal chromosomal aneuploidy; a new Z-value less than 1.96 indicates a negative result, i.e., the fetal chromosomes are normal.

[0063] It should be noted that in the device of this application, the model training module can be used as needed. For example, if a fetal chromosomal aneuploidy abnormality detection model, positive threshold, negative threshold, and the median of the model output values ​​for all negative samples have already been obtained, other modules can directly call the model and data; therefore, it is not necessary to run the model training module for every test. Of course, if the training samples change, such as by adding training samples, it is recommended to run the model training module to further improve the model and various data.

[0064] It should also be noted that the device for detecting fetal chromosomal aneuploidy in this application is actually implemented through various modules to achieve the method for detecting fetal chromosomal aneuploidy in this application; therefore, the specific limitations of each module can be referred to the method for detecting fetal chromosomal aneuploidy in this application. For example, the calculation of fetal DNA concentration, Z-score, and mosaicism, specifically Z... new Calculation method, linear discriminant analysis model, and how to determine Z new For determining positive and negative results, the method for detecting fetal chromosomal aneuploidy in this application can be used as a reference.

[0065] The fourth aspect of this application discloses an apparatus for detecting fetal chromosomal aneuploidy, the apparatus comprising a memory and a processor; the memory includes a program for storing programs; the processor includes a method for detecting fetal chromosomal aneuploidy or a method for constructing a fetal chromosomal aneuploidy detection model of this application by executing the program stored in the memory.

[0066] It is understood that when the device of this application executes the program stored in the memory to implement the method for constructing the fetal chromosomal aneuploidy abnormality detection model of this application, the device of this application is actually a device for model construction. The model constructed by the device can be used to detect fetal chromosomal aneuploidy abnormalities according to the method of this application.

[0067] The fifth aspect of this application discloses a computer-readable storage medium storing a program that can be executed by a processor to implement the method for detecting fetal chromosomal aneuploidy abnormalities or the method for constructing a fetal chromosomal aneuploidy abnormality detection model of this application.

[0068] It is understood that when the program stored in the computer-readable storage medium of this application can be executed by a processor to implement the method for constructing the fetal chromosomal aneuploidy abnormality detection model of this application, the computer-readable storage medium of this application is actually a computer-readable storage medium for model construction. The computer-readable storage medium can be directly used to implement the construction of the fetal chromosomal aneuploidy abnormality detection model. The model constructed thereby can be used to detect fetal chromosomal aneuploidy abnormalities according to the method of this application.

[0069] Due to the adoption of the above technical solutions, the beneficial effects of this application are as follows:

[0070] This application discloses a method and apparatus for detecting fetal chromosomal aneuploidy, pioneering the inclusion of mosaicism in fetal chromosomal aneuploidy detection. It comprehensively considers three variables—fetal DNA concentration, mosaicism, and the traditional Z-score—to calculate a new Z-score. This method and apparatus improve the accuracy of NIPT testing, exhibiting excellent differentiation between true and false positive samples and reducing false positives. Furthermore, the new Z-score conforms to a normal distribution, meeting current regulatory and clinical requirements; it also reduces data distribution volatility, thereby lowering the gray zone rate, reducing retest rate, and improving the stability of test results. Attached Figure Description

[0071] Figure 1 This is a flowchart of a method for detecting fetal chromosomal aneuploidy in an embodiment of this application;

[0072] Figure 2 This is a structural block diagram of the device for detecting fetal chromosomal aneuploidy in the embodiments of this application;

[0073] Figure 3 This is a T13 chimerism analysis diagram of 10,240 samples in the embodiments of this application;

[0074] Figure 4 This is a QQ plot of the new Z-values ​​of chromosome 21 from 10,000 samples in this application embodiment;

[0075] Figure 5 This is a distribution map of the traditional Z-values ​​and the new Z-values ​​of chromosome 13 in 10,000 samples from embodiments of this application. Detailed Implementation

[0076] The present application will now be described in further detail with reference to specific embodiments and accompanying drawings. In the following embodiments, many details are described to facilitate a better understanding of the present application. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other devices, materials, or methods. In some cases, certain operations related to the present application are not shown or described in the specification to avoid obscuring the core parts of the application with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; a complete understanding of the related operations can be obtained from the description in the specification and general technical knowledge in the art.

[0077] This application innovatively uses chimerism as a variable in calculating a new Z-value, thereby improving the accuracy of NIPT detection. Therefore, this application proposes a method for detecting fetal chromosomal aneuploidy, comprising calculating a new Z-value for the test sample based on the concentration of fetal DNA in the cell-free DNA of the pregnant woman's blood, the Z-value, and the chimerism, and determining whether the fetal chromosomes in the test sample have aneuploidy based on the new Z-value; wherein, chimerism is the ratio of abnormal fetal cells to all fetal cells.

[0078] In one implementation of this application, a method for detecting fetal chromosomal aneuploidy is provided, such as... Figure 1 As shown, the specific steps include data acquisition step 11, data processing step 12, chimerism calculation step 13, new Z-value calculation step 14, and fetal chromosomal aneuploidy judgment step 15.

[0079] Step 11, which involves acquiring high-throughput sequencing data of cell-free DNA from the blood of the pregnant woman to be tested, includes this step. For example, in one implementation of this application, the sequencing data is a FastQ format file generated by the sequencer.

[0080] Data processing step 12 includes calculating the fetal DNA concentration, Z-value, average corrected depth of each chromosome, and average corrected depth of all autosomes based on the high-throughput sequencing data of the obtained cell-free DNA from the pregnant woman's blood. In one implementation of this application, this step includes common operations of a conventional NIPT procedure, specifically including the following:

[0081] A) Sequence alignment and filtering: The sequence information contained in the fastq format file generated by the sequencer is aligned to the human reference genome, such as GRCh37 / hg19, using publicly available software, such as BWA (0.7.7-r441). The sequence information is filtered out to remove poor alignment quality sequences, multiple alignment sequences, repetitive sequences, and imperfect alignment sequences, leaving only the unique alignment sequences. The coordinates and other information of each unique alignment sequence are stored in a bam format file.

[0082] B) Windowing and Data Correction: The human reference genome is divided into windows of approximately 60kb. The number of uniquely aligned sequences within each 60kb window is counted, serving as the original depth information for that window, i.e., the window depth. Further, GC correction and inter-sample correction are performed on the original depth of each window to obtain the corrected depth information (i.e., UR) for each window. The average of the corrected depths of all windows on a chromosome is then calculated to obtain the "average corrected depth of the k-th chromosome." This application calculates the average corrected depth for all autosomes.

[0083] C) Fetal DNA concentration calculation: This application uses different calculation methods for male and female fetuses, as detailed below:

[0084] The method for calculating male fetal concentration is as follows:

[0085] The male fetal concentration is determined by the proportion of the Y chromosome. The mean UR value of the Y chromosome window is divided by the mean UR value of the autosomes, and then multiplied by 2 to obtain the male fetal concentration FF.

[0086]

[0087] The method for calculating female fetal concentration is as follows:

[0088] Fetal concentration in female fetuses was estimated using a high-dimensional regression model based on the non-uniform distribution of fetal cell-free DNA across the genome. The underlying assumption was that, regardless of sex, the distribution characteristics of fetal cfDNA and maternal cfDNA across the genome differed. Therefore, fetal concentration estimated using the Y chromosome method for male fetuses was used as input to train the model. A neural network machine learning approach was then employed to construct the regression model, as detailed below:

[0089]

[0090] Where l is the layer number of the network, the first layer is the input layer, the last layer is the output layer (with only one neuron), and the middle layers are hidden layers. Let be the value of the j-th neuron in the l-th layer. This represents the value of the k-th neuron in the (l-1)-th layer. The connection weights are the connection weights from the k-th neuron in layer (l-1) to the j-th neuron in layer l. This represents the input bias of the j-th neuron in the l-th layer. The most common form of the function f is the rectified linear unit, i.e., f(x) = max(0,x). w and b are obtained during model training. When applying the model, the neuron values ​​are calculated layer by layer according to the above formula, and the neuron values ​​in the last layer are the predicted values ​​of the fetal concentration model.

[0091] D) In ​​the traditional calculation of Z-value, the depth of all intervals on a chromosome conforms to a normal distribution. Therefore, by using a chromosome as a reference, the Z-value of the chromosome to be tested can be calculated using the distribution of the interval depths of the chromosome to be tested. This Z-value can then be used as the basis for determining whether the chromosome is trisomic.

[0092] Specifically, the traditional method for calculating the Z-value in this application is as follows:

[0093] The corrected depth information (UR) of each window on an autosome follows a Poisson distribution. When the number of windows is large, it follows a normal distribution. For normal samples, there is no significant difference between the distribution of the UR of the tested chromosome and the distribution of the UR of the reference chromosome. For abnormal samples, there is a slight difference. The Z-test can be used to determine fetal chromosomal aneuploidy, as detailed below:

[0094]

[0095] in:

[0096] Mean value of UR on chromosome i;

[0097] The mean value of UR for chromosome j;

[0098] SD i : Represents the standard deviation of the UR of chromosome i;

[0099] SD j : Represents the standard deviation of the UR of chromosome j;

[0100] L i : Indicates the number of windows used to divide chromosome i;

[0101] L j : Indicates the number of windows used to divide chromosome j;

[0102] Z i : Indicates the significance of aneuploidy on chromosome i, reflecting the difference from euploidy.

[0103] The above formula compares the 22 autosomes within the same sample. This is based on the assumption that the vast majority of chromosomes in a sample should be normal diploid. Therefore, the target chromosome is compared 21 times with the remaining 21 chromosomes. If the target chromosome is normal diploid, the vast majority of the 21 Z-test values ​​should be close to 0, and averaging them yields a negative Z-value. Conversely, if the target chromosome is trisomic, the vast majority of the 21 Z-test values ​​will be much greater than 0, and averaging them yields a positive Z-value.

[0104] Step 13, which calculates the degree of chimerism, includes calculating the degree of chimerism for each chromosome based on the concentration of fetal DNA.

[0105] For example, the degree of mosaicism for each chromosome can be calculated using Formula 1:

[0106] Formula 1

[0107] In Formula 1, Mosaic kLet fra represent the mosaicism of the k-th chromosome. k FF represents the relative fetal concentration of chromosome k, where FF is the fetal DNA concentration; k The result was obtained using Formula 2:

[0108] Formula 2

[0109] In Formula 2, fra k The relative fetal concentration of chromosome k. This represents the average depth of the k-th chromosome after correction. This represents the average depth after correction for all autosomes; in Formula 1 and Formula 2, k ranges from 1 to 22.

[0110] Mosaic k A value of 0 indicates that chromosome k is normal; Mosaic k A value of 1 indicates that the fetus's chromosome k is completely trisomic; Mosaic k A value between 0 and 1 indicates that the fetal chromosome k is mosaic.

[0111] It should be noted that the calculation of chimerism and its inclusion in fetal chromosomal aneuploidy abnormalities are among the innovative improvements of this application. Studies show that fetal trisomy is not always a complete trisomy; that is, not every cell in the fetus is in a trisomic state. A situation where some fetal cells are in a trisomic state and others are not is called chimerism. Fetal chimerism can affect the detection results of NIPT. For example, with a fixed fetal DNA concentration, if the fetus has a complete trisomy, the trisomy signal in the maternal peripheral blood is stronger; if the fetus has a chimeric trisomy, the trisomy signal in the maternal peripheral blood is weaker. Since NIPT involves multiple steps, including plasma collection, preservation, transportation, cfDNA isolation, library construction, and sequencing, even slight fluctuations in any step can lead to fluctuations in the final test results. For negative samples, data fluctuations may result in a weak trisomy signal similar to low chimerism. Therefore, this application creatively proposes to quantitatively describe the degree of chimerism and further clarify the difference between the degree of chimerism of true positive samples and the degree of chimerism of weak trisomy signals caused by data fluctuations, so as to better distinguish between true positive and false positive samples.

[0112] The new Z-value calculation step 14 includes calculating the new Z-value of the test sample based on the concentration of fetal DNA, Z-value, and chimerism in the cell-free DNA of the pregnant woman's blood; wherein, chimerism is the ratio of abnormal fetal cells to all fetal cells.

[0113] For example, the new Z-value calculation step 14 is divided into a model output value analysis sub-step and a Z-value imprinting sub-step.

[0114] The model output value analysis sub-step includes inputting the fetal DNA concentration, traditional Z-score, and chimerism of the sample to be tested into the fetal chromosomal aneuploidy abnormality detection model to obtain the model output value corresponding to the sample. Specifically, the fetal chromosomal aneuploidy abnormality detection model uses several samples with known fetal chromosomal aneuploidy abnormalities as training samples, with fetal DNA concentration, traditional Z-score, and chimerism as inputs and the model output value as output, to perform machine learning model training to obtain the model.

[0115] It should be noted that machine learning model training is another innovative improvement in this application. Before training the model, this application found that there is a very good linear relationship between the three variables of fetal DNA concentration, chimerism, and traditional Z-value. Therefore, the three variables of fetal DNA concentration, chimerism, and traditional Z-value are put into the LDA (linear discriminant analysis) model for model training to obtain the trained model, namely the fetal chromosomal aneuploidy abnormality detection model.

[0116] The general form of the LDA model is as follows:

[0117] LD = W1a1 + W2a2 + ... + w k a k

[0118] Where w k Here, is the coefficient, i.e., the output value of the model obtained from model training, and 'a' is the coefficient. k Let be the variables, and be the sample information input to the model, which in this example are fetal concentration, traditional Z-score, and chimerism. Therefore, after model training, what we actually obtain are the coefficients of these three variables: fetal concentration, traditional Z-score, and chimerism. With these three coefficients, plus the sample's fetal concentration, traditional Z-score, and chimerism, we can obtain the result of the machine learning model, i.e., the model output value (LD value), through the above formula.

[0119] The Z-score imprinting step includes calculating a new Z-score for the test sample based on the model output value of the sample to be tested, the positive threshold, the negative threshold, and the median of the model output values ​​of all negative samples, and labeling it as Z. new .

[0120] It's important to note that the results generated by machine learning models no longer conform to a statistically significant distribution. Therefore, unlike traditional Z-scores, thresholds cannot be determined based on statistical significance. Instead, thresholds must be defined using the features of the training data. A negative threshold is set so that all true positive samples in the training data are not classified as negative, ensuring the model does not produce false negatives. A positive threshold is set so that as many true positive samples as possible are classified as positive, while as few original false positive samples as possible are classified as positive, thereby reducing false positives and improving NIPT detection performance. The area between the positive and negative thresholds is a gray zone.

[0121] The results generated by the machine learning model no longer conform to a statistically significant distribution. However, in actual clinical use, according to clinical usage habits and regulatory requirements, NIPT trisomy test results must be fed back in the form of Z-scores, with 3 as the positive threshold. How to transform the results of the non-statistically significant machine learning model into statistically significant Z-scores is the third innovative improvement of this application. In one implementation of this application, the machine learning model used is a linear model, which allows the final results generated by the machine learning model to maintain the distribution characteristics of traditional Z-scores. Therefore, this application creatively adopts an imprinting method to imprint the model output values ​​into new Z-scores. This not only improves the performance of NIPT detection but also makes the final results have distribution characteristics similar to Z-scores, that is, conforming to a normal distribution with a center of 0.

[0122] In one implementation of this application, the specific printing method is as follows:

[0123] When the model output value of the sample to be tested is greater than the positive threshold, Z new =LD-cut p +3;

[0124] When the model output value of the sample to be tested is less than the positive threshold and greater than the negative threshold,

[0125]

[0126] When the model output value of the sample to be tested is less than the negative threshold,

[0127] In the above formula, Z new The new Z value, LD is the model output value of the test sample, cut p The positive threshold is cut. n is the negative threshold, and Med is the median of the model output values ​​for negative samples.

[0128] Step 15 of the fetal chromosomal aneuploidy judgment step includes determining whether the chromosome of the fetus to be tested has aneuploidy based on the new Z value.

[0129] In one implementation of this application, the new Z value obtained by segmented mapping also conforms to a normal distribution, and the center of the distribution is located at 0; therefore, Z>3 can still be used as a positive judgment value and Z<1.96 as a negative judgment value.

[0130] Based on the method for detecting fetal chromosomal aneuploidy abnormalities in this application, this application proposes a method for constructing a fetal chromosomal aneuploidy abnormality detection model. This method includes using several samples with known fetal chromosomal characteristics as training samples, including positive and negative samples of fetal chromosomal aneuploidy abnormalities. Using fetal DNA concentration, Z-score, and chimerism as inputs, a machine learning model is trained to obtain a model output value that comprehensively represents the fetal chromosomal characteristics using three variables: fetal DNA concentration, Z-score, and chimerism. The resulting model is the fetal chromosomal aneuploidy abnormality detection model. The calculation methods for fetal DNA concentration, Z-score, and chimerism can all refer to the method for detecting fetal chromosomal aneuploidy abnormalities in this application, and will not be elaborated here.

[0131] Those skilled in the art will understand that all or part of the functions of the above methods can be implemented in hardware or by computer programs. When all or part of the functions of the above methods are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in storage media such as a server, another computer, disk, optical disk, flash drive, or portable hard drive, and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions of the above methods can be achieved.

[0132] Therefore, based on the method for detecting fetal chromosomal aneuploidy in this application, this application proposes an apparatus for detecting fetal chromosomal aneuploidy, including a new Z-value calculation module and a fetal chromosomal aneuploidy judgment module. The new Z-value calculation module includes calculating a new Z-value of the sample to be tested based on the concentration of fetal DNA, Z-value, and chimerism in the cell-free DNA of the pregnant woman's blood. Chimerism is the ratio of abnormal fetal cells to all fetal cells. The fetal chromosomal aneuploidy judgment module includes judging whether the fetal chromosomes of the sample to be tested have aneuploidy based on the new Z-value.

[0133] In one implementation of this application, the device for detecting fetal chromosomal aneuploidy, such as... Figure 2 As shown, it includes a data acquisition module 21, a data processing module 22, a chimerism calculation module 23, a model training module 24, a new Z-value calculation module 25, and a fetal chromosomal aneuploidy abnormality judgment module 26.

[0134] The data acquisition module 21 includes a method for acquiring high-throughput sequencing data of cell-free DNA from the blood of the pregnant woman to be tested. For example, it acquires FastQ format files generated by the sequencer.

[0135] The data processing module 22 includes functions for calculating fetal DNA concentration, conventional Z-value, average depth of each chromosome after correction, and average depth of all autosomes after correction, based on high-throughput sequencing data of cell-free DNA from the pregnant woman's blood. For example, the calculation of fetal DNA concentration, conventional Z-value, average depth of each chromosome after correction, and average depth of all autosomes can be performed with reference to existing conventional NIPT protocols.

[0136] The chimerism calculation module 23 includes calculating the chimerism of each chromosome based on the fetal DNA concentration.

[0137] For example, the degree of chimerism of each chromosome can be calculated using Formula 1;

[0138] Formula 1

[0139] In Formula 1, Mosaic k Let fra represent the mosaicism of the k-th chromosome. k FF represents the relative fetal concentration of the k-th chromosome, and FF represents the fetal DNA concentration.

[0140] fra k The result was obtained using Formula 2.

[0141] Formula 2

[0142] In Formula 2, fra k The relative fetal concentration of chromosome k. This represents the average depth after correction for the k-th chromosome. This represents the average depth after correction for all autosomes;

[0143] In Formula 1 and Formula 2, the value of k ranges from 1 to 22;

[0144] Mosaic k A value of 0 indicates that chromosome k is normal; Mosaic k A value of 1 indicates that the fetus's chromosome k is completely trisomic; Mosaick A value between 0 and 1 indicates that the fetal chromosome k is mosaic.

[0145] The model training module 24 includes using several samples with known fetal chromosomal information as training samples. The training samples include positive and negative samples of fetal chromosomal aneuploidy abnormalities. The machine learning model is trained with fetal DNA concentration, Z-score, and chimerism as inputs to obtain a model output value that comprehensively represents the fetal chromosomal information using three variables: fetal DNA concentration, Z-score, and chimerism. The model obtained is the fetal chromosomal aneuploidy abnormality detection model. After model training, the corresponding positive threshold is obtained using positive samples, the corresponding negative threshold is obtained using negative samples, and the median is obtained using the model output value of all negative samples.

[0146] The new Z-value calculation module 25 includes calculating a new Z-value for the sample based on the concentration of fetal DNA, Z-value, and chimerism in the cell-free DNA of the pregnant woman's blood; wherein, chimerism is the ratio of abnormal fetal cells to all fetal cells.

[0147] For example, the new Z-value calculation module 25 includes a model output value analysis submodule and a Z-value imprinting submodule; the model output value analysis submodule includes a method for inputting the fetal DNA concentration, Z-value, and chimerism of the sample to be tested into the fetal chromosomal aneuploidy abnormality detection model to obtain the model output value corresponding to the sample to be tested; the Z-value imprinting submodule includes a method for calculating a new Z-value of the sample to be tested based on the model output value of the sample to be tested, as well as the positive threshold, the negative threshold, and the median of the model output values ​​of all negative samples; the positive threshold is the threshold value of the model output value corresponding to the positive sample, and the negative threshold is the threshold value of the model output value corresponding to the negative sample.

[0148] The fetal chromosomal aneuploidy abnormality detection module 26 includes a method for determining whether the chromosomes of the fetus under test have aneuploidy abnormalities based on a new Z-value. For example, a new Z-value greater than 3 is considered positive, indicating fetal chromosomal aneuploidy abnormalities; a new Z-value less than 1.96 is considered negative, indicating normal fetal chromosomes.

[0149] Another implementation of this application provides an apparatus for detecting fetal chromosomal aneuploidy, the apparatus comprising a memory and a processor; the memory includes a program for storing a program; the processor includes a program for executing the program stored in the memory to implement the following method: calculating a new Z-value for the sample based on the concentration of fetal DNA, Z-value, and chimerism in the cell-free DNA of the pregnant woman's blood, and determining whether the fetal chromosomes of the sample are aneuploid based on the new Z-value; wherein, chimerism is the ratio of abnormal fetal cells to all fetal cells. Alternatively, the method may specifically implement the following steps: a data acquisition step, including acquiring high-throughput sequencing data of cell-free DNA from the pregnant woman's blood; a data processing step, including calculating fetal DNA concentration and conventional Z-value based on the acquired high-throughput sequencing data of cell-free DNA from the pregnant woman's blood; a chimerism calculation step, including calculating the chimerism of each chromosome based on the fetal DNA concentration; a model value analysis step, including inputting the fetal DNA concentration, conventional Z-value, and chimerism into a fetal chromosomal aneuploidy abnormality detection model to obtain the model output value corresponding to the sample; a Z-value imprinting step, including calculating a new Z-value based on the model output value of the sample, the positive threshold, the negative threshold, and the median of the model output value of the negative sample; and a fetal chromosomal aneuploidy abnormality judgment step, including determining whether the chromosomes of the fetus to be tested have aneuploidy abnormalities based on the new Z-value.

[0150] Alternatively, the device includes a memory and a processor; the memory includes a program for storing a program; the processor includes a program for executing the program stored in the memory to implement the following method: using several samples of known fetal chromosomal conditions as training samples, the training samples including positive and negative samples of fetal chromosomal aneuploidy abnormalities, using fetal DNA concentration, Z-score, and chimerism as inputs, training a machine learning model to obtain a model output value that comprehensively represents the fetal chromosomal condition using three variables: fetal DNA concentration, Z-score, and chimerism, thereby obtaining a model, namely, a fetal chromosomal aneuploidy abnormality detection model.

[0151] Another implementation of this application also provides a computer-readable storage medium, which includes a program that can be executed by a processor to implement the following method: calculating a new Z-value of the test sample based on the concentration of fetal DNA, Z-value, and chimerism in the cell-free DNA of the pregnant woman's blood, and determining whether the fetal chromosome of the test sample has aneuploidy based on the new Z-value; wherein, chimerism is the ratio of abnormal fetal cells to all fetal cells. Alternatively, the method may specifically implement the following steps: a data acquisition step, including acquiring high-throughput sequencing data of cell-free DNA from the pregnant woman's blood; a data processing step, including calculating fetal DNA concentration and conventional Z-value based on the acquired high-throughput sequencing data of cell-free DNA from the pregnant woman's blood; a chimerism calculation step, including calculating the chimerism of each chromosome based on the fetal DNA concentration; a model value analysis step, including inputting the fetal DNA concentration, conventional Z-value, and chimerism into a fetal chromosomal aneuploidy abnormality detection model to obtain the model output value corresponding to the sample; a Z-value imprinting step, including calculating a new Z-value based on the model output value of the sample, the positive threshold, the negative threshold, and the median of the model output value of the negative sample; and a fetal chromosomal aneuploidy abnormality judgment step, including determining whether the chromosomes of the fetus to be tested have aneuploidy abnormalities based on the new Z-value.

[0152] Alternatively, the storage medium includes a program that can be executed by a processor to implement the following method: using several samples with known fetal chromosomal information as training samples, including positive and negative samples of fetal chromosomal aneuploidy abnormalities, and using fetal DNA concentration, Z-score, and chimerism as inputs to train a machine learning model, and obtaining a model output value that comprehensively represents the fetal chromosomal information by three variables: fetal DNA concentration, Z-score, and chimerism, thereby obtaining a model, namely a fetal chromosomal aneuploidy abnormality detection model.

[0153] The method and apparatus of this application differ from the prior art in that:

[0154] (1) This application has created a unique detection index—chimerism. Research has found that chimerism has a good ability to distinguish between true positive and false positive samples reported by the current Z-value (i.e., traditional Z-value) method.

[0155] (2) This application integrates three variables: fetal concentration, the unique indicator of this application—chimerism, and the traditional Z-score. In one implementation method, linear discriminant analysis (LDA) is specifically selected as the machine learning model for model training and result determination. The study found that the determination results of this model can reduce false positives and improve detection effectiveness compared with the original results.

[0156] (3) All three variables used in this application have a linear relationship. The linear relationship is simple and clear, avoiding the complexity caused by too many variables and different dimensions and distribution characteristics among the variables. In one implementation method, a linear discriminant analysis (LDA) model is used for analysis. The model is simple and does not have the problem of overfitting.

[0157] (4) This application develops a new method for converting Z-values, which transforms statistically insignificant numerical values ​​obtained through machine learning into Z-values ​​that are commonly used in clinical practice and meet regulatory requirements, namely the new Z-value of this application. Furthermore, the new Z-value obtained through the Z-value conversion method of this application conforms to a normal distribution and can meet the requirements of current regulatory and clinical use.

[0158] (5) By comparing the new Z value of this application with the traditional Z value, it can be found that the new Z value greatly reduces the volatility of data distribution, reduces gray area and retesting, and improves the stability of detection results.

[0159] (6) The machine learning-based solution provided in this application, while considering multiple variables, utilizes real sample data accumulated by BGI Genomics for model training. This enables the model to learn and grasp the unique characteristics of BGI Genomics' own data due to factors such as experimental reagents and sequencing platforms, thus allowing it to be better applied to the data generated in BGI Genomics' current actual production. It can be understood that this application establishes a method for learning from personalized data, rather than being limited to BGI Genomics' own data.

[0160] This application first innovatively introduces a novel detection indicator—chimerism—and further integrates three variables: fetal DNA concentration, chimerism, and the traditional Z-score, overcoming the inaccuracy of traditional NIPT's reliance solely on the Z-score for trisomy assessment. Moreover, it uses a linear model to integrate these three variables, resulting in a simple model free from overfitting. Furthermore, this application develops a Z-score transformation method, namely Z-score imprinting, which converts meaningless numerical values ​​obtained from machine learning models into meaningful and clinically acceptable Z-scores. new This also reduces the gray area of ​​the traditional Z-value, reduces retesting, and improves the stability of the detection results.

[0161] It is understandable that, based on this application, more parameters could be used for model training and fetal chromosomal aneuploidy analysis, such as considering variables like gestational age and maternal age. Of course, with more variables, the corresponding machine learning model would also need to be changed, for example, by adopting a non-linear QDA model. Furthermore, the specific Z-value segmentation in this application can also be adjusted as needed.

[0162] Example 1

[0163] This example uses an established fetal chromosomal aneuploidy detection model to predict chromosomal abnormalities in samples with diagnostic / follow-up results. Specifically, a total of 108,293 samples were used for model training. These samples were divided into negative and positive categories upon entering the model training, but each category contained three karyotypes: negative samples included true negatives and false positives, and positive samples included true positives. Because the calculation methods for fetal concentration differ between male and female fetuses, resulting in differences in fetal concentration data characteristics, and since fetal concentration is one of the key variables in the model, two separate models were trained for males and females. The specific sample numbers are shown in Table 1.

[0164] Table 1 Samples used for model training

[0165] male female True positive 798 620 False positive 234 318 True negative 56864 49459 total 57896 50397

[0166] Table 2. Examples of sample data used for model training

[0167] fetal concentration Traditional Z-value Chimerism True positive 1 0.149 14.501 0.900 True positive 2 0.120 10.058 0.885 False positive 1 0.293 4.834 0.173 False positive 2 0.389 6.821 0.158 True negative 1 0.229 -0.810 -0.035 True Negative 2 0.120 -0.596 -0.049

[0168] In this example, the fetal DNA concentration, conventional Z-score, and chimerism of the training samples, as shown in Table 2, are input into the LDA model for training to obtain the model output values. The median of the machine learning values ​​obtained from the negative samples is taken as the "median of the model output values," which is the Med in the subsequent imprinting formula. The median calculated in this example is shown in Table 3.

[0169] Table 3 Median of model output values

[0170]

[0171]

[0172] Before printing, the distribution of true negative, false positive, and true positive samples is manually observed to determine the threshold for the LD value, ensuring that: 1. No true positive samples are classified as negative; 2. As many true positive samples as possible are classified as positive; 3. As few false positive samples as possible are classified as positive. The threshold for the LD value is determined based on these principles, which is the positive threshold (cut) in the printing formula. p ) and negative threshold (cut n The specific values ​​for this example are shown in Table 4.

[0173] Table 4 Thresholds for LD Values

[0174]

[0175] After imprinting, using clinically common values ​​of 1.96 and 3 as the new Z-value thresholds, the following imprinting method is obtained:

[0176] When the model output value is greater than the positive threshold, Znew =LD-cut p +3;

[0177] When the model output value is less than the positive threshold and greater than the negative threshold,

[0178]

[0179] When the model output value is less than the negative threshold

[0180] In the above formula, Z new That is, the new Z value, LD is the model output value, cut p The positive threshold is cut. n is the negative threshold, and Med is the median of the model output values ​​for negative samples.

[0181] A total of 10,240 samples were selected from those tested by BGI Genomics in actual clinical applications and underwent prenatal diagnosis / postnatal follow-up. These samples were tested using the traditional Z-score in actual clinical testing, and subsequent prenatal diagnosis / postnatal follow-up was conducted based on these results. Therefore, based on the test results and prenatal diagnosis / postnatal follow-up results of each sample, each sample can be classified into three categories: true positive, false positive, and true negative. Specific sample information is shown in Table 5.

[0182] The traditional method for calculating the Z-value is as follows:

[0183]

[0184] in:

[0185] Mean value of UR on chromosome i;

[0186] The mean value of UR for chromosome j;

[0187] SD i : Represents the standard deviation of the UR of chromosome i;

[0188] SD j : Represents the standard deviation of the UR of chromosome j;

[0189] L i : Indicates the number of windows used to divide chromosome i;

[0190] L j : Indicates the number of windows used to divide chromosome j;

[0191] Z i : Indicates the significance of aneuploidy on chromosome i, reflecting the difference from euploidy.

[0192] Table 5 shows the three-body detection results given by the traditional Z-score.

[0193]

[0194] Table 6. Sample data examples used for model testing

[0195] fetal concentration Traditional Z-value Chimerism True positive 1 0.067 5.824 0.885 True positive 2 0.146 11.714 0.808 False positive 1 0.188 3.602 0.205 False positive 2 0.187 4.186 0.252 True negative 1 0.115 1.137 0.090 True Negative 2 0.059 -1.125 -0.192

[0196] As can be seen, based on the traditional Z-score for detection, the positive predictive values ​​for T21, T18, and T13 are 0.86, 0.58, and 0.36, respectively, indicating a significant problem with false positives.

[0197] Using the fetal chromosomal aneuploidy detection model and method of this application, the mosaicism of the above 10240 samples was calculated. Taking T13 as an example, the results are as follows: Figure 3 As shown, chimerism can effectively distinguish between true positive, false positive, and true negative samples. Furthermore, the chimerism, fetal concentration, and traditional Z-score were input into the trained machine learning model, and a new Z-score was generated through Z-score mapping. Partial sample data used for model testing are shown in Table 6. The 10240 samples were re-evaluated using the new Z-score, with Z>3 considered positive and Z<1.96 considered negative, generating new test results, as shown in Table 7.

[0198] Table 7 shows the trisomy detection results from the improved method for detecting fetal chromosomal aneuploidy.

[0199]

[0200] The results in Table 7 show that using the new Z-value, all 14 false positives (T21), 33 false positives (T18), and 39 false positives (T13) were correctly identified as negative. At the same time, 87 true positives (T21), 45 true positives (T18), 22 true positives (T13), and 10,000 true negative samples were still correctly identified. Therefore, the positive predictive value of T21, T18, and T13 all reached 100%, with a sensitivity of 100% and a specificity of 100%. This significantly reduced false positives, improved PPV, and increased specificity while maintaining sensitivity.

[0201] Example 2

[0202] This example uses the established model to detect continuous samples from the production line.

[0203] Due to factors such as the collection of diagnostic / follow-up results, karyotyped samples are not continuous samples from a single center. Therefore, their data distribution characteristics cannot reflect the true distribution characteristics of the population, making it impossible to assess the true distribution characteristics of the new Z-value. Therefore, this study uses continuous samples received by a medical testing laboratory of BGI Genomics over a period of time to evaluate the distribution characteristics of the new Z-value and compare it with the traditional Z-value to demonstrate the true characteristics and patterns of the new Z-value in practical use.

[0204] A series of 10,000 consecutive samples from a single medical testing laboratory of BGI Genomics, who underwent clinical testing within a specific time period, were selected. Using the fetal chromosomal aneuploidy detection model and method described in this application, new Z-values ​​were calculated for these 10,000 samples. Taking the Z-value of chromosome 21 as an example, the distribution of the Z-value of chromosome 21 was examined to see if it conforms to a normal distribution. The results are as follows... Figure 4 As shown. Figure 4 The results showed that the new Z-values ​​of 10,000 samples from a single center over a continuous time period were basically located on the diagonal of the QQ plot. Among them, a few samples that deviated significantly from the diagonal of the QQ plot were positive samples with stronger signals. Figure 4 The new Z value shows that it has very good normality.

[0205] Further comparison of the distribution of the new Z-value with the traditional Z-value, taking the Z-value of chromosome 13 as an example, such as... Figure 5 As shown. Figure 5 The results show that, firstly, the center of the new Z-value distribution is closer to 0, indicating that the new Z-value conforms to a normal distribution centered at 0 more closely than the traditional Z-value. Secondly, the new Z-value distribution is more concentrated than the traditional Z-value, indicating that the new Z-value has lower volatility and better stability.

[0206] The new Z-value exhibits less fluctuation compared to the traditional Z-value, resulting in a reduction in the gray zone rate. This example further demonstrates this with a larger sample size. Specifically, 360,786 clinical samples from a single medical testing laboratory of BGI Genomics were analyzed throughout 2020, and these samples underwent a total of 383,306 tests. The new Z-value generated 785 T21 gray zones, 345 T18 gray zones, and 288 T13 gray zones in these 383,306 tests, with gray zone rates of 0.22%, 0.09%, and 0.08% for T21, T18, and T13, respectively. The overall gray zone rate for trisomy testing was 0.39%. In contrast, the traditional Z-value produced 3071 T21 gray areas, 4350 T18 gray areas, and 2335 T13 gray areas. The gray area rates of T21, T18, and T13 were 0.80%, 1.14%, and 0.61%, respectively. The overall gray area rate of the three-body detection was 2.55%, as shown in Table 8.

[0207] Table 8 Comparison of gray area sample number and gray area ratio between traditional Z-value and new Z-value.

[0208]

[0209] The results in Table 8 show that the new Z-value generated by the method of this application can reduce the gray area rate of three-body detection to about one-tenth of the previous level, significantly reducing retesting caused by gray areas and improving the detection performance of NIPT.

[0210] The above description, in conjunction with specific embodiments, provides a further detailed explanation of this application and should not be construed as limiting the specific implementation of this application to these descriptions. Those skilled in the art to which this application pertains can make several simple deductions or substitutions without departing from the concept of this application.

Claims

1. A method for detecting fetal chromosomal aneuploidy, characterized in that: This includes calculating a new Z-value for the test sample based on the concentration of fetal DNA, Z-value, and chimerism in the cell-free DNA of the pregnant woman's blood, and determining whether the fetal chromosome of the test sample has aneuploidy based on the new Z-value. The degree of chimerism is the ratio of abnormal fetal cells to all fetal cells. Based on the concentration, Z-value, and chimerism of fetal DNA in the cell-free DNA of the pregnant woman's blood, a new Z-value for the test sample is calculated. This includes inputting the fetal DNA concentration, Z-value, and chimerism into the fetal chromosomal aneuploidy abnormality detection model to obtain the model output value corresponding to the test sample, and then mapping the new Z-value of the test sample from the model output value. The fetal chromosomal aneuploidy abnormality detection model uses several samples with known fetal chromosomal information as training samples. The training samples include positive and negative samples of fetal chromosomal aneuploidy abnormalities. The model is trained using fetal DNA concentration, Z-score, and chimerism as inputs to obtain a model output value that comprehensively represents the fetal chromosomal information using three variables: fetal DNA concentration, Z-score, and chimerism. The fetal DNA concentration and Z-value were calculated based on high-throughput sequencing data of cell-free DNA from the pregnant woman's blood.

2. The method according to claim 1, characterized in that: The new Z-value of the test sample is obtained by mapping the model output value, including calculating the new Z-value of the test sample based on the model output value of the test sample, the positive threshold, the negative threshold, and the median of the model output values ​​of all negative samples. The positive threshold is the threshold value of the model output value corresponding to the positive sample, and the negative threshold is the threshold value of the model output value corresponding to the negative sample.

3. The method according to claim 2, characterized in that: The median of the model output values ​​for all negative samples is the median of the model output values ​​for all negative samples obtained by re-inputting all negative training samples into the fetal chromosomal aneuploidy abnormality detection model.

4. The method according to claim 2, characterized in that: New Z-values ​​of the test samples are obtained by imprinting the model output values, including the following imprinting methods: When the model output value of the sample to be tested is greater than the positive threshold, Z new = LD - cut p +3; When the model output value of the sample to be tested is less than the positive threshold and greater than the negative threshold, Z new = ; When the model output value of the sample to be tested is less than the negative threshold, Z new = ; In the above formula, Z new For the new Z value, LD The model output value for the sample to be tested. cut p The positive threshold, cut n The negative threshold, Med This is the median of the model output values ​​for all negative samples.

5. The method according to claim 1, characterized in that: The new Z-value is used to determine whether the fetal chromosomes of the sample being tested are aneuploid. A new Z-value greater than 3 is considered positive, indicating fetal chromosomal aneuploidy; a new Z-value less than 1.96 is considered negative, indicating normal fetal chromosomes.

6. The method according to claim 1, characterized in that: The machine learning model is a linear discriminant analysis model.

7. The method according to claim 1, characterized in that: The abnormal fetal cells are those containing fetal chromosomal aneuploidy.

8. The method according to any one of claims 1-7, characterized in that: The degree of fitting is calculated using Formula 1; Formula 1 In Formula 1, Mosaic k Let k be the degree of mosaicism of the k-th chromosome. fra k The relative fetal concentration of chromosome k. FF This refers to the concentration of fetal DNA. fra k The result was obtained using Formula 2. Formula 2 In Formula 2, fra k The relative fetal concentration of chromosome k. This represents the average depth of the k-th chromosome after correction. This represents the average depth after correction for all autosomes; In Formula 1 and Formula 2, the value of k ranges from 1 to 22; Mosaic k A value of 0 indicates that the fetus's kth chromosome is normal; Mosaic k A value of 1 indicates that the fetus's kth chromosome is completely trisomic; Mosaic k A value between 0 and 1 indicates that the fetus's kth chromosome is mosaic.

9. The method according to claim 8, characterized in that: The average depth of each corrected chromosome and the average depth of all corrected autosomes were calculated using high-throughput sequencing data of cell-free DNA from the pregnant woman's blood.

10. A method for constructing a fetal chromosomal aneuploidy abnormality detection model, characterized in that: The method involves using several samples with known fetal chromosomal information as training samples, including positive and negative samples of fetal chromosomal aneuploidy. The model is trained using fetal DNA concentration, Z-value, and chimerism as inputs to obtain a model output value that comprehensively represents the fetal chromosomal information using three variables: fetal DNA concentration, Z-value, and chimerism. The model obtained by training in this way is the fetal chromosomal aneuploidy detection model. The fetal DNA concentration and Z-value were calculated based on high-throughput sequencing data of cell-free DNA in the pregnant woman's blood; the chimerism was the ratio of abnormal fetal cells to all fetal cells.

11. The construction method according to claim 10, characterized in that: The abnormal fetal cells are those containing fetal chromosomal aneuploidy.

12. The construction method according to claim 10, characterized in that: The degree of fitting is calculated using Formula 1; Formula 1 In Formula 1, Mosaic k Let k be the degree of mosaicism of the k-th chromosome. fra k The relative fetal concentration of chromosome k. FF This refers to the concentration of fetal DNA. fra k The result was obtained using Formula 2. Formula 2 In Formula 2, fra k The relative fetal concentration of chromosome k. This represents the average depth after correction for the k-th chromosome. This represents the average depth after correction for all autosomes; In Formula 1 and Formula 2, the value of k ranges from 1 to 22; Mosaic k A value of 0 indicates that the fetus's kth chromosome is normal; Mosaic k A value of 1 indicates that the fetus's kth chromosome is completely trisomic; Mosaic k A value between 0 and 1 indicates that the fetus's kth chromosome is mosaic.

13. The construction method according to claim 12, characterized in that: The average depth of each corrected chromosome and the average depth of all autosomes were calculated based on high-throughput sequencing data of cell-free DNA from the pregnant woman's blood.

14. The construction method according to claim 10, characterized in that: The machine learning model is a linear discriminant analysis model.

15. A device for detecting fetal chromosomal aneuploidy, characterized in that: This includes a new Z-score calculation module and a module for determining fetal chromosomal aneuploidy. The new Z-value calculation module includes a tool for calculating a new Z-value for the sample based on the concentration of fetal DNA, Z-value, and chimerism in the cell-free DNA of the pregnant woman's blood; the chimerism is the ratio of abnormal fetal cells to all fetal cells. The fetal chromosomal aneuploidy module includes a tool for determining whether the fetal chromosomes of the sample to be tested have aneuploidy based on the new Z value. The new Z-value calculation module also includes a method for inputting fetal DNA concentration, Z-value and chimerism into the fetal chromosomal aneuploidy abnormality detection model to obtain the model output value corresponding to the sample to be tested, and obtaining the new Z-value of the sample to be tested by imprinting the model output value. The fetal chromosomal aneuploidy abnormality detection model uses several samples with known fetal chromosomal information as training samples. The training samples include positive and negative samples of fetal chromosomal aneuploidy abnormalities. The model is trained using fetal DNA concentration, Z-score, and chimerism as inputs. The model output value is used to comprehensively characterize the fetal chromosomal information by combining the three variables of fetal DNA concentration, Z-score, and chimerism.

16. The apparatus according to claim 15, characterized in that: It also includes a model training module, which uses several samples with known fetal chromosomal information as training samples. The training samples include positive and negative samples of fetal chromosomal aneuploidy. The machine learning model is trained with fetal DNA concentration, Z-value and chimerism as inputs to obtain a model output value that comprehensively represents the fetal chromosomal information by three variables: fetal DNA concentration, Z-value and chimerism. The model obtained is the fetal chromosomal aneuploidy abnormality detection model.

17. The apparatus according to claim 16, characterized in that: The machine learning model is a linear discriminant analysis model.

18. The apparatus according to claim 15, characterized in that: The new Z-value calculation module includes a model output value analysis submodule and a Z-value imprinting submodule. The model output value analysis submodule includes a function to input the fetal DNA concentration, Z-value, and chimerism of the sample to be tested into the fetal chromosomal aneuploidy abnormality detection model to obtain the model output value corresponding to the sample to be tested. The Z-value imprinting submodule includes a function to calculate a new Z-value for the sample to be tested based on the model output value of the sample to be tested, as well as a positive threshold, a negative threshold, and the median of the model output values ​​of all negative samples. The positive threshold is the threshold value of the model output value corresponding to the positive sample, and the negative threshold is the threshold value of the model output value corresponding to the negative sample.

19. The apparatus according to claim 18, characterized in that: The Z-value printing submodule obtains the new Z-value according to the following method. When the model output value of the sample to be tested is greater than the positive threshold, Z new = LD - cut p +3; When the model output value of the sample to be tested is less than the positive threshold and greater than the negative threshold, Z new = ; When the model output value of the sample to be tested is less than the negative threshold, Z new = ; In the above formula, Z new For the new Z value, LD The model output value for the sample to be tested. cut p The positive threshold, cut n The negative threshold, Med This is the median of the model output values ​​for all negative samples.

20. The apparatus according to claim 19, characterized in that: In the fetal chromosomal aneuploidy module, the new Z-value is used to determine whether the fetal chromosome of the sample to be tested has aneuploidy. This includes a new Z-value greater than 3 indicating a positive result, i.e., fetal chromosomal aneuploidy; and a new Z-value less than 1.96 indicating a negative result, i.e., fetal chromosomes are normal.

21. The apparatus according to claim 15, characterized in that: It also includes a data acquisition module for acquiring high-throughput sequencing data of cell-free DNA from the pregnant woman's blood in the sample to be tested.

22. The apparatus according to claim 15, characterized in that: It also includes a data processing module, which is used to calculate fetal DNA concentration and Z-value based on the high-throughput sequencing data of cell-free DNA from the pregnant woman's blood.

23. The apparatus according to claim 22, characterized in that: The data processing module also includes a function to calculate the average depth of each chromosome after correction and the average depth of all autosomes after correction, based on the high-throughput sequencing data of cell-free DNA from the blood of the pregnant woman to be tested.

24. The apparatus according to claim 15, characterized in that: It also includes a chimerism calculation module, which is used to calculate the chimerism of each chromosome according to Formula 1; Formula 1 In Formula 1, Mosaic k Let k be the degree of mosaicism of the k-th chromosome. fra k The relative fetal concentration of chromosome k. FF This refers to the concentration of fetal DNA. fra k The result was obtained using Formula 2. Formula 2 In Formula 2, fra k The relative fetal concentration of chromosome k. This represents the average depth of the k-th chromosome after correction. This represents the average depth after correction for all autosomes; In Formula 1 and Formula 2, the value of k ranges from 1 to 22; Mosaic k A value of 0 indicates that the fetus's kth chromosome is normal; Mosaic k A value of 1 indicates that the fetus's kth chromosome is completely trisomic; Mosaic k A value between 0 and 1 indicates that the fetus's kth chromosome is mosaic.

25. A device for detecting fetal chromosomal aneuploidy, characterized in that, The device includes: Memory, used to store programs; A processor is configured to implement the method for detecting fetal chromosomal aneuploidy as described in any one of claims 1-9 or the method for constructing a fetal chromosomal aneuploidy detection model as described in any one of claims 10-14 by executing a program stored in the memory.

26. A computer-readable storage medium, characterized in that: The method includes a program that can be executed by a processor to implement the method for detecting fetal chromosomal aneuploidy as described in any one of claims 1-9 or the method for constructing a fetal chromosomal aneuploidy detection model as described in any one of claims 10-14.

Citation Information

Patent Citations

  • Method for detecting fetal gene haplotype, and device thereof and storage medium

    CN113308548A

  • Method for reducing false positive and false negative in noninvasive prenatal detection based on semiconductor sequencing

    CN113593629A