Method, computing device, medium, and program product for detecting fetal chromosomal status
By filtering and Z-value calculation of fetal DNA fragments in traditional detection methods, the problems of low fetal DNA fragment concentration and maternal chromosome abnormalities are solved, which significantly improves the accuracy of the detection results.
Patent Information
- Application Number
- CN202510178534.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The traditional method used to detect fetal chromosome status has inaccurate results due to the low concentration of fetal DNA fragments and abnormal maternal chromosomes.
By filtering the sequencing data of the maternal peripheral blood sample to be tested during pregnancy, filtering data with the length of the DNA fragment within a predetermined length range is obtained, and Z value is calculated based on the initial data and filtered data to determine the fetal chromosome status.
The accuracy of the detection results for fetal chromosome status has been significantly improved, and the effects of low fetal DNA fragment concentration and maternal chromosome abnormalities have been reduced.
Smart Images

Figure CN119673286B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to bioinformatics processing, and specifically, to methods, computing devices, computer storage media, and computer program products for detecting fetal chromosomal status. Background Art
[0002] Currently, non-invasive prenatal testing (NIPT) for chromosomal aneuploidy is a non-invasive prenatal screening technique that uses next-generation sequencing (NGS) methods to detect cell-free fetal DNA (cffDNA) fragments in the peripheral blood of pregnant women during pregnancy. It is mainly used to detect fetal chromosomal aneuploidy and subchromosomal abnormalities.
[0003] In traditional methods for detecting fetal chromosomal status, sequencing is mainly performed on the peripheral blood of pregnant women, the DNA fragments mapped to different genomes are counted, and the fetal chromosomal status is calculated based on this count. Since there are both fetal DNA fragments and maternal DNA fragments in the peripheral blood of pregnant women, and the concentration (or proportion) of fetal DNA fragments is relatively low, the concentration (or proportion) of fetal DNA in the peripheral blood of pregnant women has a greater impact on the accuracy of the detection results. In addition, if there are abnormalities in the maternal chromosomes, it will also lead to inaccurate detection results for fetal chromosomes. Moreover, the sex chromosomes are relatively complex, so the proportion of false positives for fetal chromosomal abnormalities is relatively high.
[0004] In summary, the deficiencies of traditional solutions for detecting fetal chromosomal status are: inaccurate detection results for fetal chromosomal status due to the low concentration of fetal DNA fragments and the influence of maternal chromosomal abnormalities. Summary of the Invention
[0005] The present invention provides a method, a computing device, a computer storage medium, and a program product for detecting fetal chromosomal status, which can significantly improve the accuracy of the detection results for fetal chromosomal status.
[0006] According to a first aspect of the present invention, a method for detecting the chromosomal status of a fetus is provided. The method includes: comparing the obtained sequencing data of a peripheral blood sample of a pregnant woman to be tested with a human reference genome to obtain initial data that is paired-end compared to the human reference genome; filtering the DNA fragments in the initial data to obtain filtered data with DNA fragment lengths within a predetermined length range; calculating Z values for a target region based on the initial data and the filtered data respectively to obtain the Z value for the target region in the initial data and the Z value for the target region in the filtered data; and determining the chromosomal status of the fetus in the peripheral blood sample of the pregnant woman to be tested based on the obtained Z value for the target region in the initial data and the Z value for the target region in the filtered data.
[0007] In some embodiments, the method for detecting the chromosomal status of a fetus further includes: calculating abnormal content ratio characterization data for the initial data and the filtered data to obtain abnormal content ratio characterization data for the DNA fragments in the initial data and abnormal content ratio characterization data for the DNA fragments in the filtered data; and comparing the abnormal content ratio characterization data for the DNA fragments in the initial data and the abnormal content ratio characterization data for the DNA fragments in the filtered data to determine the source of chromosomal abnormalities based on the comparison result.
[0008] In some embodiments, filtering the DNA fragment lengths in the initial data to obtain filtered data with DNA fragment lengths within a predetermined length range includes: obtaining the start position and end position of the DNA fragments in the initial data that are paired-end compared to the human reference genome; calculating the length of each DNA fragment in the initial data based on the obtained start position and end position; determining whether the calculated length of each DNA fragment in the initial data is greater than a first predetermined size threshold and less than or equal to a second predetermined size threshold; in response to determining that the calculated length of the current DNA fragment in the initial data is greater than the first predetermined size threshold and less than or equal to the second predetermined size threshold, determining that the current DNA fragment in the initial data belongs to the filtered data within the predetermined size range; and in response to determining that the calculated length of the current DNA fragment in the initial data is less than or equal to the first predetermined size threshold or greater than the second predetermined size threshold, filtering out the current DNA fragment in the initial data.
[0009] In some embodiments, calculating the Z-value for a target region includes: for the initial data or the filtered data respectively, dividing the number of inserted fragments in the target region by the total number of inserted fragments in autosomes, so as to obtain the genomic characterization data of the genomic target region respectively; calculating the mean and standard deviation of the genomic characterization data of the genomic target region of the reference sample set; and calculating the Z-value of the target region in the initial data and the Z-value of the target region in the filtered data respectively based on the genomic characterization data of the genomic target region, the mean and the standard deviation of the genomic characterization data of the reference sample set.
[0010] In some embodiments, determining the fetal chromosome status of a maternal peripheral blood sample during the pregnancy to be tested includes: determining whether the Z-value of the target region in the filtered data belongs to a predetermined threshold range; and in response to determining that the Z-value of the target region in the filtered data belongs to the predetermined threshold range, determining that the fetal chromosome status of the maternal peripheral blood sample during the pregnancy to be tested is abnormal.
[0011] In some embodiments, the predetermined threshold range is greater than or equal to 3 or less than or equal to -3.
[0012] In some embodiments, determining the source of chromosomal abnormality based on the comparison result includes: comparing the abnormal content ratio characterization data of the DNA fragments in the filtered data with the abnormal content ratio characterization data of the DNA fragments in the initial data; in response to determining that the abnormal content ratio characterization data of the DNA fragments in the filtered data is greater than the abnormal content ratio characterization data of the DNA fragments in the initial data, determining that the chromosomal abnormality originates from the fetus; and in response to determining that the abnormal content ratio characterization data of the DNA fragments in the filtered data is equal to the abnormal content ratio characterization data of the DNA fragments in the initial data, determining that the chromosomal abnormality originates from both the mother and the fetus.
[0013] In some embodiments, determining the source of chromosomal abnormality based on the comparison result includes: in response to confirming that the Y chromosome ratio of the DNA fragments in the filtered data is greater than a predetermined fetal DNA fragment concentration threshold of the Y chromosome, determining that the fetal DNA of the maternal peripheral blood sample during the pregnancy to be tested originates from a male fetus; constructing a relationship graph between the abnormal content ratio characterization data of the X chromosome and the Y chromosome ratio of the maternal peripheral blood sample during the pregnancy to be tested; in response to determining that the relationship graph is within the range between the first threshold line and the second threshold line, determining that the maternal peripheral blood sample during the pregnancy to be tested is a sample with normal sex chromosomes; and in response to determining that the relationship graph is outside the range between the first threshold line and the second threshold line, determining that the maternal peripheral blood sample during the pregnancy to be tested is a sample with abnormal sex chromosomes.
[0014] In some embodiments, determining that a maternal peripheral blood sample during pregnancy to be tested is a sex chromosome abnormal sample includes: in response to determining that the relationship graph is outside the range between the first threshold line and the second threshold line and higher than the first threshold point, determining that the maternal peripheral blood sample during pregnancy to be tested has an abnormal karyotype of 47XXY or 47XYY; and in response to determining that the relationship graph is outside the range between the first threshold line and the second threshold line and lower than the second threshold point, determining that the maternal peripheral blood sample during pregnancy to be tested has an abnormal karyotype of 46XY / 45X or 46XY / 45Y. The second threshold point is lower than the first threshold point.
[0015] According to a second aspect of the present invention, there is also provided a computing device, which includes: a memory configured to store one or more computer programs; and a processor coupled to the memory and configured to execute one or more programs to cause the device to perform the method of the first aspect of the present invention.
[0016] According to a third aspect of the present invention, there is also provided a non-transitory computer-readable storage medium. Machine-executable instructions are stored on the non-transitory computer-readable storage medium, and when the machine-executable instructions are executed, the machine is caused to perform the method of the first aspect of the present invention.
[0017] According to a fourth aspect of the present invention, there is also provided a computer program product. The computer program product includes a computer program, and when the computer program is executed by a machine, the method of the first aspect of the present invention is performed.
[0018] The summary of the invention is provided to introduce a selection of concepts in a simplified form, which will be further described in the detailed implementation below. The summary of the invention is not intended to identify the key features or main features of the present invention, nor is it intended to limit the scope of the present invention. Brief Description of the Drawings
[0019] Figure 1 A schematic diagram of a system for detecting the chromosomal status of a fetus according to an embodiment of the present invention is shown.
[0020] Figure 2 A flowchart of a method for detecting the chromosomal status of a fetus according to an embodiment of the present invention is shown.
[0021] Figure 3 A frequency curve of fetal DNA fragments at different DNA fragment lengths is shown.
[0022] Figure 4 A flowchart of a method for determining the source of chromosomal abnormalities based on comparison results according to some embodiments of the present invention is shown.
[0023] Figure 5The flowchart of a method for determining the source of chromosomal abnormalities based on comparison results according to some other embodiments of the present invention is shown.
[0024] Figure 6 The flowchart of a method for determining the karyotype of deterministic chromosomal abnormalities according to an embodiment of the present invention is shown.
[0025] Figure 7 The flowchart of a method for determining the karyotype of deterministic chromosomal abnormalities according to an embodiment of the present invention is shown.
[0026] Figure 8 The relationship graph between the abnormal content ratio characterization data of the X chromosome and the ratio of the Y chromosome is schematically shown.
[0027] Figure 9 The relationship between the fetal ratio of chromosome Y of the DNA fragment in the initial data and the fetal ratio of chromosome Y of the DNA fragment in the initial data is shown.
[0028] Figure 10 The block diagram of an electronic device suitable for implementing the embodiments of the present invention is schematically shown.
[0029] In each of the drawings, the same or corresponding reference numerals represent the same or corresponding parts. Detailed embodiments
[0030] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0031] As used herein, the term "including" and its variants mean open inclusion, that is, "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.
[0032] As described above, the deficiencies of the traditional solution for detecting the chromosomal status of a fetus are: the detection results for the chromosomal status of the fetus are inaccurate due to the low concentration of the DNA fragments of the fetus and the influence of maternal chromosomal abnormalities.
[0033] It has been found through research that the sizes of DNA fragments of a fetus have certain characteristics. For example, the proportion of small fragments is relatively high. In the DNA fragments of the mother, the proportion of small fragments is relatively low. Therefore, the chromosomal status abnormality of the fetus can be analyzed based on the small fragments in the DNA fragments. Figure 1 FIG. 100 shows a schematic diagram of a system 100 for detecting the chromosomal status of a fetus according to an embodiment of the present invention.
[0034] Regarding the computing device 110, it is used, for example, to compare the obtained sequencing data of a maternal peripheral blood sample during the pregnancy to be tested with a human reference genome to obtain initial data that is paired-end compared to the human reference genome; filter the DNA fragment lengths in the initial data to obtain filtered data with DNA fragment lengths within a predetermined length range. The computing device 110 is also used, for example, to calculate the proportion of the insert fragment ratio for a target region based on the initial data and the filtered data respectively to obtain a Z value for the target region in the initial data and a Z value for the target region in the filtered data; and determine the chromosomal status of the fetus in the maternal peripheral blood sample during the pregnancy to be tested based on the obtained Z value for the target region in the initial data and the Z value for the target region in the filtered data.
[0035] In some embodiments, the computing device 110 may have one or more processing units, including dedicated processing units such as GPUs, FPGAs, and ASICs, as well as general-purpose processing units such as CPUs. Additionally, one or more virtual machines may be running on each computing device. The computing device 110 includes, for example: an initial data acquisition unit 112, a filtered data acquisition unit 114, a Z value acquisition unit 116 for the target region in the initial data and the filtered data, and a fetal chromosomal status determination unit 118. The initial data acquisition unit 112, the filtered data acquisition unit 114, the Z value acquisition unit 116 for the target region in the initial data and the filtered data, and the fetal chromosomal status determination unit 118 may be configured on one or more computing devices 110.
[0036] Regarding the initial data acquisition unit 112, it is used to compare the obtained sequencing data of a maternal peripheral blood sample during the pregnancy to be tested with a human reference genome to obtain initial data that is paired-end compared to the human reference genome.
[0037] Regarding the filtered data acquisition unit 114, it is used to filter the DNA fragment lengths in the initial data to obtain filtered data with DNA fragment lengths within a predetermined length range.
[0038] The Z-value acquisition unit 116 for the target region in the initial data and the filtered data is configured to calculate the Z-value for the target region based on the initial data and the filtered data respectively, so as to obtain the Z-value for the target region in the initial data and the Z-value for the target region in the filtered data.
[0039] Regarding the fetal chromosome status determination unit 118, it is configured to determine the fetal chromosome status of the maternal peripheral blood sample during the pregnancy to be tested based on the obtained Z-value for the target region in the initial data and the Z-value for the target region in the filtered data.
[0040] The following will be combined with Figure 2 、 Figure 3 and Figure 9 to describe the method for detecting the fetal chromosome status according to an embodiment of the present invention. Figure 2 FIG. shows a flowchart of a method 200 for detecting the fetal chromosome status according to an embodiment of the present invention. Figure 3 FIG. shows the frequency curves of the DNA fragments of the fetus at different DNA fragment lengths. Figure 9 FIG. shows the relationship between the proportion of fetuses with chromosome Y in the DNA fragments of the initial data and the proportion of fetuses with chromosome Y in the DNA fragments of the initial data. It should be understood that the method 200 can be executed, for example, at Figure 10 the electronic device 1000 described. It can also be executed at Figure 1 the computing device 110 described. It should be understood that the method 200 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this regard.
[0041] At step 202, the computing device 110 compares the obtained sequencing data of the maternal peripheral blood sample during the pregnancy to be tested with the human reference genome, so as to obtain the initial data that is paired-end compared to the human reference genome.
[0042] For example, the computing device 110 compares the sequencing data of the maternal peripheral blood sample during the pregnancy to be tested with the human reference genome via the BWA software, so as to obtain the initial data that is paired-end compared to the human reference genome.
[0043] The sequencing data of the maternal peripheral blood sample during the pregnancy to be tested (or simply referred to as the "sample to be tested") is, for example but not limited to, low-depth whole-genome sequencing data. For example, the sequencing depth <= 3X.
[0044] For example, the computing device 110 uses BWA or similar publicly available software to align the sequencing fastq data to a human reference genome sequence (e.g., hg19 or hg38) to generate alignment result data; and based on the alignment result data, retain the uniquely aligned inserted fragments; calculate the position of each inserted fragment in the reference genome (the position information includes, for example, the chromosome number and the genomic coordinates) and the number of times each position is sequenced.
[0045] At step 204, the computing device 110 filters the DNA fragment lengths in the initial data to obtain filtered data with DNA fragment lengths within a predetermined length range.
[0046] Regarding the method for obtaining the filtered data, it includes, for example: the computing device 110 obtains the start position and the end position of the DNA fragments in the initial data that are paired-end aligned to the human reference genome; based on the obtained start position and end position, calculates the length of each DNA fragment in the initial data; determines whether the calculated length of each DNA fragment in the initial data is greater than a first predetermined size threshold and less than or equal to a second predetermined size threshold; in response to determining that the calculated length of the current DNA fragment in the initial data is greater than the first predetermined size threshold and less than or equal to the second predetermined size threshold, determines that the current DNA fragment in the initial data belongs to the filtered data within the predetermined size range; and in response to determining that the calculated length of the current DNA fragment in the initial data is less than or equal to the first predetermined size threshold or greater than the second predetermined size threshold, filters out the current DNA fragment in the initial data. The following will be combined with Figure 4 Specifically illustrate the method for obtaining the filtered data with DNA fragment lengths within a predetermined length range, and will not be elaborated here.
[0047] Regarding the predetermined size range, it is, for example but not limited to, greater than or equal to a first predetermined size threshold and less than or equal to a second predetermined size threshold. Regarding the first predetermined size threshold, it is, for example, 170 bp, and in some embodiments, the first predetermined size threshold is, for example, 160 bp. Regarding the second predetermined size threshold, it is, for example, 30 bp.
[0048] As Figure 3 shown, the marker 310 indicates the first predetermined size threshold. The marker 312 indicates the second predetermined size threshold. The first predetermined size threshold is, for example, 30 bp. The second predetermined size threshold is, for example, 166 bp. The marker 320 indicates the distribution curve of the DNA fragment sizes of the high-concentration fetal DNA in the maternal peripheral blood sample during pregnancy. The marker 322 indicates the distribution curve of the DNA fragment sizes of the low-concentration fetal DNA in the maternal peripheral blood sample during pregnancy. As Figure 3As shown, in small-sized DNA fragments, for example, in DNA fragments with a size between a first predetermined size threshold and a second predetermined size threshold, there is more fetal DNA at a high concentration. In large-sized DNA fragments, for example, in DNA fragments with a size greater than the second predetermined size threshold, there is more fetal DNA at a low concentration. Therefore, it is necessary to screen out small-sized (or "small fragments") DNA fragments with a size between the first predetermined size threshold and the second predetermined size threshold for analysis based on the screened small-sized DNA fragments.
[0049] As Figure 9 shown, Figure 9 The abscissa is the fetal proportion of chromosome Y of the DNA fragment in the initial data, and the ordinate is the fetal proportion of chromosome Y of the DNA fragment in the initial data. Line 910 indicates the regression function (y = 1.56x + 5.41). The dashed line 912 indicates the function y = x.
[0050] At step 206, the computing device 110 calculates the Z value for the target region based on the initial data and the filtered data respectively, so as to obtain the Z value for the target region in the initial data and the Z value for the target region in the filtered data.
[0051] A method for calculating the Z value for a target region, for example, includes: for the initial data or the filtered data respectively, a computing device 110 divides the number of inserted fragments in the target region by the total number of inserted fragments in the autosomes to respectively obtain genomic characterization data for the genomic target region; calculates the mean and standard deviation of the genomic characterization data of the reference sample set target region in the reference sample set; and calculates the Z value of the target region in the initial data and the Z value of the target region in the filtered data respectively based on the genomic characterization data of the genomic target region, the mean of the genomic characterization data of the reference sample set genomic target region, and the standard deviation. Specifically, based on the DNA fragment data in the initial data, the computing device 110 divides the number of inserted fragments in the target region of the DNA fragments in the initial data by the total number of inserted fragments in the autosomes to obtain the genomic characterization data of the genomic target region; calculates the mean and standard deviation of the genomic characterization data of the reference sample set target region in the reference sample set; and calculates the Z value of the target region in the initial data based on the genomic characterization data of the genomic target region, the mean of the genomic characterization data of the reference sample set genomic target region, and the standard deviation. In addition, based on the DNA fragment data in the filtered data, the computing device 110 divides the number of inserted fragments in the target region of the DNA fragments in the filtered data by the total number of inserted fragments in the autosomes to obtain the genomic characterization data of the genomic target region; calculates the mean and standard deviation of the genomic characterization data of the reference sample set target region in the reference sample set; and calculates the Z value of the target region in the filtered data based on the genomic characterization data of the genomic target region, the mean of the genomic characterization data of the reference sample set genomic target region, and the standard deviation.
[0052] The following schematically shows an algorithm for calculating the Z value for a target region in conjunction with formula (1).
[0053]
[0054] In the above formula (1), GR represents the genomic characterization data of the genomic target region in the peripheral blood sample of the pregnant woman to be tested. This genomic characterization data is obtained by dividing the number of inserted fragments in the target region by the total number of inserted fragments in the autosomes. represents the mean of the genomic characterization data of the genomic target region in the reference sample, represents the standard deviation of the genomic characterization data of the genomic target region in the reference sample. Z score represents the Z value for the target region.
[0055] In some embodiments, the present invention can also calculate the inserted fragment Z value for the target region in conjunction with formula (2).
[0056]
[0057] In the above formula (2), ΔF represents the difference in the proportion of small-sized DNA fragments between the genomic target region and the reference chromosome (the reference chromosome refers to all autosomes except chromosomes 21, 18, and 13). represents the mean difference in the proportion of small-sized DNA fragments between the genomic target region and the reference chromosome. represents the standard deviation of the difference in the proportion of short DNA fragments between the genomic target region and the reference chromosome. Size Z score represents the Z value for the target region calculated based on some embodiments. The following schematically shows an algorithm for calculating the difference in the proportion of small-sized DNA fragments (i.e., ΔF) between the genomic target region and the reference chromosome in combination with formula (3).
[0058]
[0059] In the above formula (3), P(≤150) represents the proportion of sequencing fragments derived from the target region and having a size less than or equal to 150 bp. represents the proportion of sequencing fragments derived from the reference chromosome and having a size less than or equal to 150 bp. ΔF represents the difference in the proportion of small-sized DNA fragments between the genomic target region and the reference chromosome.
[0060] At step 208, the computing device 110 determines the fetal chromosome status of the maternal peripheral blood sample to be tested based on the obtained Z value for the target region in the initial data and the Z value for the target region in the filtered data.
[0061] Regarding the proportion of fetal DNA fragments, it is the proportion of cell-free fetal DNA (cffDNA) fragments in maternal plasma. It should be understood that in non-invasive fetal testing, the proportion of fetal cffDNA in maternal plasma is linearly correlated with the abnormal proportion of variations. The fetal free DNA concentration (fetal fraction, FF) is the percentage of cffDNA in the total amount of maternal cell-free DNA (cfDNA). It should be understood that FF is a major bioinformatics analysis quality control factor for non-invasive prenatal screening sequencing results. Low FF may lead to false negative results. Therefore, accurately estimating the proportion of fetal cffDNA in maternal plasma is very important. There are many methods to calculate the proportion of fetal cffDNA in maternal plasma.
[0062] The present invention can calculate the proportion of fetal DNA fragments (i.e., the proportion of fetal cffDNA in maternal plasma) based on the Y chromosome, DNA fragment size, and the SeqFF model, thereby more accurately calculating the proportion of fetal cffDNA in maternal plasma.
[0063] For example, for a female fetus, the present invention can use a linear regression model to combine the calculation results of a size-based method and the calculation results of a SeqFF model to calculate the combined concentration of fetal free DNA. For example, the following formula (4) illustrates an algorithm for calculating the combined concentration of fetal free DNA.
[0064]
[0065] In the above formula (4), BFF represents the combined concentration of free fetal DNA. SeqFF represents the concentration of free fetal DNA calculated based on the SeqFF model. Size represents the concentration of free fetal DNA calculated from the data filtered by DNA fragment size. Among them, SeqFF The model is used to directly calculate the proportion of fetal DNA from sequencing data of routine depth in non-invasive prenatal screening. Since SeqFF This high-dimensional model requires a large amount of sample data in the construction process. Therefore, when FF is lower than 5%, its performance is significantly reduced. Thus, the present invention can significantly increase FF through DNA fragment size filtering.
[0066] For a male fetus, its sex chromosome is inconsistent with that of the mother. Research shows that the correlation coefficient between the combined concentration of free fetal DNA BFF and the proportion of fetal DNA based on the Y chromosome is 0.94.
[0067] Regarding the method for determining the fetal chromosome status of a maternal peripheral blood sample during the pregnancy to be tested, there can be various methods. In some embodiments, the method for determining the fetal chromosome status includes: the computing device 110 determines whether the Z value of the target region in the filtered data belongs to a predetermined threshold range; in response to determining that the Z value of the target region in the filtered data belongs to the predetermined threshold range, it is determined that the fetal chromosome status of the maternal peripheral blood sample during the pregnancy to be tested is abnormal. Regarding the predetermined threshold range, for example, it is greater than or equal to 3, or less than or equal to -3.
[0068] In some embodiments, the method for determining the fetal chromosome status further includes: the computing device 110 compares the abnormal content ratio characterization data of the DNA fragments in the initial data and the abnormal content ratio characterization data of the DNA fragments in the filtered data, so as to determine the source of chromosomal abnormalities based on the comparison result.
[0069] For example, the present invention obtains 78 peripheral blood samples of pregnant women to be tested, and each sample carries a male fetus. For the sequencing data of each sample, the initialization data of the DNA fragments in the initial data that are paired-end aligned to the human reference genome is obtained, that is, the DNA fragment data in the initial data. The DNA fragment data in the initial data of the 78 peripheral blood samples of pregnant women to be tested constitutes the initial data (or simply referred to as "Tag"). Filtering is performed on the DNA fragment lengths in the initial data. For example, DNA fragments with lengths of 30 - 170 bp are selected to form a new sample data set, that is, the filtered data (or simply referred to as "Filter Tag"). Then, for each sample data, the fetal DNA concentration (or simply referred to as "Filter Tag FF") of the DNA fragments in the filtered data and the fetal DNA concentration (or simply referred to as "Tag FF") of the DNA fragments in the initial data are calculated according to the Y chromosome. Experimental data shows that the average of Filter Tag FF is 2 times that of Tag FF. For the case where the fetal DNA concentration in the peripheral blood samples of pregnant women to be tested is relatively low, Filter Tag FF even increases to 3 times that of Tag FF.
[0070] In the above solution, by obtaining the initial data that is paired-end aligned to the human reference genome; filtering the DNA fragment lengths in the initial data to obtain filtered data with DNA fragment lengths within a predetermined length range, the present invention can obtain data of small-sized DNA fragments with DNA fragment lengths within a predetermined length range, thereby effectively increasing the concentration of fetal free DNA in the peripheral blood samples of pregnant women to be tested. Additionally, by calculating the Z value for the target region based on the initial data and the filtered data respectively to obtain the Z value for the target region in the initial data and the Z value for the target region in the filtered data; and determining the fetal chromosome status of the peripheral blood samples of pregnant women to be tested based on the obtained Z value for the target region in the initial data and the Z value for the target region in the filtered data, the present invention can significantly reduce the inaccuracy of the detection results for the fetal chromosome status caused by the low concentration of fetal DNA fragments and the influence of maternal chromosomal abnormalities. Therefore, the present invention can significantly improve the accuracy of the detection results for the fetal chromosome status.
[0071] In some embodiments, method 200 further includes: the computing device 110 calculates the abnormal content ratio characterization data for the initial data and the filtered data.
[0072] The following will be combined with Figure 4 Describe method 400 for obtaining DNA fragment data in filtered data with DNA fragment lengths within a predetermined length range according to an embodiment of the present invention. Figure 4FIG. 400 is a flowchart of a method for obtaining DNA fragment data in filtered data where the length of the DNA fragment is within a predetermined length range according to an embodiment of the present invention. It should be understood that method 400 can be executed, for example, at Figure 10 the electronic device 1000 described. It can also be executed at Figure 1 the computing device 110 described. It should be understood that method 400 may further include additional actions not shown and / or actions shown may be omitted, and the scope of the present invention is not limited in this regard.
[0073] At step 402, the computing device 110 obtains the start position and end position of the DNA fragment in the initial data that is paired-end aligned to the human reference genome.
[0074] For example, the computing device 110 obtains the start position and end position of the DNA fragment in each initial data that is aligned to the DNA fragment of the human reference genome.
[0075] At step 404, the computing device 110 calculates the length of each DNA fragment in the initial data based on the obtained start position and end position.
[0076] The computing device 110 calculates the length of each DNA fragment in the initial data for each DNA fragment in the initial data based on the difference between the start position and the end position.
[0077] At step 406, the computing device 110 determines whether the length of each DNA fragment in the calculated initial data is greater than a first predetermined size threshold and less than or equal to a second predetermined size threshold.
[0078] Regarding the first predetermined size threshold, in some embodiments, it is, for example, but not limited to, 40bp.
[0079] Regarding the second predetermined size threshold, in some embodiments, it is, for example, but not limited to, 156bp.
[0080] At step 408, if the computing device 110 determines that the length of the current DNA fragment in the calculated initial data is greater than the first predetermined size threshold and less than or equal to the second predetermined size threshold, it determines that the current DNA fragment in the initial data belongs to the filtered data within the predetermined size range.
[0081] At step 410, if the computing device 110 determines that the length of the current DNA fragment in the calculated initial data is less than or equal to the first predetermined size threshold or greater than the second predetermined size threshold, it filters out the current DNA fragment in the initial data.
[0082] For example, the computing device 110 extracts DNA fragments in all initial data that have a length greater than or equal to a first predetermined size threshold (e.g., 30 bp) and less than or equal to a second predetermined size threshold (e.g., 170 bp) from the initial data, so as to generate filtered initial data, that is, filtered data. The computing device 110 filters out DNA fragments in the initial data that are less than the first predetermined size threshold (e.g., 30 bp) or greater than the second predetermined size threshold (e.g., 170 bp).
[0083] Experimental data show that the concentration of small-sized DNA fragments between the first predetermined size threshold and the second predetermined size threshold is more relevant to the concentration of cell-free fetal DNA. Therefore, the method 400 is beneficial to reducing the influence on the detection result of the fetal chromosome state caused by the low concentration of cell-free fetal DNA and the influence of maternal chromosomal abnormalities by selecting small-sized DNA fragments.
[0084] The following will be combined with Figure 5 to describe the method 500 for determining the source of chromosomal abnormalities based on the comparison result according to an embodiment of the present invention. Figure 5 The flowchart of the method 500 for determining the source of chromosomal abnormalities based on the comparison result according to some embodiments of the present invention is shown. It should be understood that the method 500 can be executed, for example, at Figure 10 the described electronic device 1000. It can also be executed at Figure 1 the described computing device 110. It should be understood that the method 500 may further include additional actions not shown and / or the shown actions may be omitted, and the scope of the present invention is not limited in this regard.
[0085] At step 502, the computing device 110 calculates aberration-containing fraction (AcF) characterization data for the initial data and the filtered data, so as to obtain aberration-containing fraction characterization data for DNA fragments in the initial data and aberration-containing fraction characterization data for DNA fragments in the filtered data.
[0086] Regarding chimeras, they are individuals used in genetics to refer to those with chimeric or mixed manifestations of different genetic traits, and also refer to one of the types of chromosomal abnormalities.
[0087] Regarding the aberration-containing fraction (AcF) characterization data, taking a positive sample of Down's syndrome (DS, also known as trisomy 21) as an example, in the case of non-chimeras, the AcF value is approximately equal to the fetal DNA concentration value.
[0088] Studies have shown that if only the fetus carries an abnormality, only the DNA molecules of fetal origin in the maternal peripheral blood during pregnancy will contain the abnormality, and the characterization data of the abnormal content ratio (i.e., AcF) will be approximately equal to the ratio of fetal DNA fragments in the maternal peripheral blood during pregnancy. Similarly, if only the mother carries an abnormality, only those DNA molecules derived from the mother in the maternal peripheral blood during pregnancy will contain the abnormality, and AcF will be equal to the ratio of maternal DNA fragments in the maternal peripheral blood during pregnancy. On the other hand, if both the mother and the fetus carry an abnormal content ratio, all DNA molecules in the maternal peripheral blood during pregnancy will be derived from cells containing the abnormality, and AcF will be 100%. In some embodiments, the characterization data of the abnormal content ratio can be calculated in combination with formula (5).
[0089]
[0090] In the above formula (5), AcF represents the characterization data of the abnormal content ratio. GR represents the genomic characterization data of the genomic target region in the maternal peripheral blood sample to be tested during pregnancy. This genomic characterization data is obtained by dividing the number of inserted fragments in the target region by the total number of inserted fragments in the autosomes. Represents the mean value of the genomic characterization data of the genomic target region in the reference sample.
[0091] At step 504, the computing device 110 compares the characterization data of the abnormal content ratio of the DNA fragments in the initial data and the characterization data of the abnormal content ratio of the DNA fragments in the filtered data, so as to determine the source of the chromosomal abnormality based on the comparison result.
[0092] It should be understood that by filtering the DNA fragment lengths to obtain the DNA fragment data in the filtered data with the DNA fragment lengths within a predetermined length range, theoretically, the percentage of fetal DNA will increase and the percentage of maternal DNA will decrease. Therefore, the concentration value or percentage of fetal DNA (for example, before filtration, the fetal DNA concentration is 10, and after filtration, the fetal DNA concentration is 15, increased to 1.5 times) will increase, and this increase ratio should be consistent with the increase ratio of the AcF value of the positive sample (for example, the AcF value of the positive sample in the case of non-mosaic Down syndrome increases from 10 to 15). It should be understood that if only the fetus carries an abnormality, the AcF value for the filtered data will increase. If only the mother carries an abnormality, the AcF value for the filtered data will decrease. If both the mother and the fetus carry an abnormality, the AcF value of the DNA fragment set after filtration will not change compared with the AcF value of the DNA fragment set before filtration. Thus, the present invention can infer the root cause of the abnormality by using the relationship between the characterization data of the abnormal content ratio of the DNA fragments in the initial data and the characterization data of the abnormal content ratio of the DNA fragments in the filtered data.
[0093] In some embodiments, regarding the method for determining the origin of chromosomal abnormalities based on comparison results, for example, it includes: the computing device 110 compares the characterization data of the abnormal content ratio of DNA fragments in the filtered data with the characterization data of the abnormal content ratio of DNA fragments in the initial data; in response to determining that the characterization data of the abnormal content ratio of DNA fragments in the filtered data is greater than the characterization data of the abnormal content ratio of DNA fragments in the initial data, it is determined that the chromosomal abnormality originates from the fetus; and in response to determining that the characterization data of the abnormal content ratio of DNA fragments in the filtered data is equal to the characterization data of the abnormal content ratio of DNA fragments in the initial data, it is determined that the chromosomal abnormality originates from both the mother and the fetus.
[0094] By adopting the above means, the present invention can accurately determine the origin of chromosomal abnormalities.
[0095] The following will be combined with Figure 6 Describe a method 600 for determining the origin of chromosomal abnormalities based on comparison results according to other embodiments of the present invention. Figure 6 The flowchart of a method 600 for determining the origin of chromosomal abnormalities based on comparison results according to other embodiments of the present invention is shown. It should be understood that the method 600 can be executed, for example, at Figure 10 the described electronic device 1000. It can also be executed at Figure 1 the described computing device 110. It should be understood that the method 600 may further include additional actions not shown and / or may omit the shown actions, and the scope of the present invention is not limited in this regard.
[0096] As mentioned above, by filtering for the DNA fragment length to obtain filtered data with DNA fragment lengths within a predetermined length range, theoretically, the percentage of fetal DNA will increase and the percentage of maternal DNA will decrease. The following will explain the reduction amount of the characterization data of the abnormal content ratio of DNA fragments in the filtered data when only the mother carries the abnormality in combination with formula (6).
[0097]
[0098] In the above formula (6), represents the proportion of fetal DNA fragments. represents the proportion of maternal DNA fragments. represents the proportion of filtered fetal DNA fragments. θ represents the increase ratio of the characterization data of the abnormal content ratio of DNA fragments in the filtered data relative to the characterization data of the abnormal content ratio of DNA fragments in the initial data when only the fetus carries the abnormality. represents the proportion of filtered maternal DNA fragments.F It represents the characterization data of the abnormal content ratio of the DNA fragments in the initial data when only the mother carries the abnormality.
[0099] It represents the characterization data of the abnormal content ratio of the DNA fragments in the filtered data of the abnormality (i.e., Filter AcF). It represents the reduction amount of the characterization data of the abnormal content ratio of the DNA fragments in the filtered data when only the mother carries the abnormality.
[0100] At step 602, if the computing device 110 determines that the characterization data of the abnormal content ratio of the DNA fragments in the filtered data is greater than the characterization data of the abnormal content ratio of the DNA fragments in the initial data, it is determined that the chromosomal abnormality originates from the fetus.
[0101] For example, if the following expression (7) holds, it is determined that the chromosomal abnormality originates from the fetus.
[0102]
[0103] In the above formula (7), It represents the characterization data of the abnormal content ratio of the DNA fragments in the filtered data. It represents the characterization data of the abnormal content ratio of the DNA fragments in the initial data. θ represents a predetermined multiple.
[0104] In some embodiments, the predetermined multiple is, for example, but not limited to, 1.8.
[0105] At step 604, if the computing device 110 confirms that the absolute value of the characterization data of the abnormal content ratio of the DNA fragments in the filtered data is less than the absolute value of the characterization data of the abnormal content ratio of the DNA fragments in the filtered data of the abnormality, it is determined that the chromosomal abnormality originates from the mother.
[0106] For example, if the following expression (8) holds, it is determined that the chromosomal abnormality originates from the mother.
[0107]
[0108] In the above formula (8), It represents the characterization data of the abnormal content ratio of the DNA fragments in the filtered data. It represents the characterization data of the abnormal content ratio of the DNA fragments in the filtered data of the abnormality.
[0109] At step 606, if the computing device 110 confirms that the absolute value of the abnormal content ratio characterization data of the DNA fragments in the filtered data is greater than or equal to the absolute value of the abnormal content ratio characterization data of the DNA fragments in the abnormal filtered data and less than or equal to the absolute value of the abnormal content ratio characterization data of the DNA fragments in the initial data, it is determined that the chromosomal abnormality originates from the fetus and the mother.
[0110] For example, if the following expression (9) holds, it is determined that the chromosomal abnormality originates from the fetus and the mother.
[0111]
[0112] In the above formula (9), represents the abnormal content ratio characterization data of the DNA fragments in the filtered data. represents the abnormal content ratio characterization data of the DNA fragments in the abnormal filtered data. represents the abnormal content ratio characterization data of the DNA fragments in the initial data. θ represents a predetermined multiple.
[0113] By adopting the above solution, the present invention can accurately determine the source of chromosomal abnormalities.
[0114] The present invention further includes: a method for determining the karyotype of sex chromosomal abnormalities.
[0115] The following will be combined with Figure 7 describe a method 700 for determining the karyotype of sex chromosomal abnormalities according to an embodiment of the present invention. Figure 7 FIG. shows a flowchart of a method 700 for determining the karyotype of sex chromosomal abnormalities according to an embodiment of the present invention. It should be understood that the method 700 can be executed, for example, at Figure 10 the described electronic device 1000. It can also be executed at Figure 1 the described computing device 110. It should be understood that the method 700 may further include additional actions not shown and / or may omit the shown actions, and the scope of the present invention is not limited in this regard.
[0116] It should be understood that sex chromosomal abnormalities (SCA) include various abnormal karyotypes, such as: 47XXY, 47XYY, 46XY / 45X, and 46XY / 45Y respectively. It should be understood that compared with the DNA data of female fetuses, the DNA data of male fetuses will contain Y chromosome data. Therefore, the present invention can use the concentration of fetal DNA fragments determined based on the Y chromosome (abbreviated as "ChrY%") to determine the gender.
[0117] At step 702, if computing device 110 confirms that the concentration of DNA fragments in the filtered data is greater than a predetermined Y-chromosome fetal DNA fragment concentration threshold, it is determined that the fetal DNA in the maternal peripheral blood sample to be tested is from a male fetus.
[0118] Regarding the predetermined Y-chromosome fetal DNA fragment concentration threshold, that is, the predetermined ChrY% threshold, which is, for example but not limited to, 2.
[0119] At step 704, computing device 110 constructs a relationship graph between the abnormal content ratio characterization data of the X chromosome and the Y chromosome ratio for the maternal peripheral blood sample during pregnancy to be tested.
[0120] Regarding the method of constructing the relationship graph, in some embodiments, computing device 110 constructs a relationship graph between the abnormal content ratio characterization data of the X chromosome and the Y chromosome ratio for the sample to be tested based on a linear regression function. The following formula (10) schematically shows the linear regression function for constructing the relationship graph between the abnormal content ratio characterization data of the X chromosome and the Y chromosome ratio for the DNA fragments in the initial data. The following formula (11) schematically shows the linear regression function for constructing the relationship graph between the abnormal content ratio characterization data of the X chromosome and the Y chromosome ratio for the DNA fragments in the filtered data.
[0121] y = -0.99x + 0.74 (10)
[0122] y = -1.03x + 0.90 (11)
[0123] In the above formulas (10) and (11), x represents the Y chromosome ratio of the sample to be tested. y represents the abnormal content ratio characterization data of the X chromosome of the sample to be tested. Figure 8A graphical relationship between the abnormal content ratio characterization data of the X chromosome and the Y chromosome ratio is schematically shown. Among them, the left graphical relationship 810 indicates the relationship between the abnormal content ratio characterization data of the X chromosome and the Y chromosome ratio for the DNA fragments in the initial data. Among them, the vertical dotted line 816 indicates the Y chromosome ratio when x = 0.75, that is, y = -0.99x + 0.74 = 0. The horizontal dotted line 818 indicates the quality control standard (cutoff). For example, the quality control standard for the graphical relationship 810 is that the abnormal content ratio characterization data of the X chromosome for the DNA fragments in the initial data is 4% (i.e., cutoff = 4%). The right graphical relationship 820 indicates the relationship between the abnormal content ratio characterization data of the X chromosome and the Y chromosome ratio for the DNA fragments in the filtered data. Among them, the vertical dotted line 826 indicates the Y chromosome ratio when x = 0.87, that is, y = -1.03x + 0.90 = 0. The horizontal dotted line 828 indicates the quality control standard (cutoff). For example, the quality control standard indicates that the abnormal content ratio characterization data of the X chromosome for the DNA fragments in the filtered data is 6% (i.e., cutoff = 6%)
[0124] It should be understood that the abnormal content ratio characterization data (i.e., AcF) of the X chromosome and the Y chromosome ratio (i.e., Y ChrY%) for the sample of the male fetus are linearly correlated because the AcF of the X chromosome is the proportion of fetal DNA fragments
[0125] In Figure 8 In the shown graphical relationship, the abscissa represents the Y chromosome ratio for the DNA fragments in the initial data or the filtered data, and the ordinate represents the abnormal content ratio characterization data of the X chromosome for the DNA fragments in the initial data or the filtered data
[0126] At step 706, if the computing device 110 determines that the graphical relationship is within the range between the first threshold line and the second threshold line, it determines that the maternal peripheral blood sample to be tested during pregnancy is a sample with normal sex chromosomes
[0127] As Figure 8 shown, in the graphical relationship 810 for the DNA fragments in the initial data, 812 indicates the first threshold line for the DNA fragments in the initial data, and 814 indicates the second threshold line for the DNA fragments in the initial data. If the relevant graph 810 is within the range between the first threshold line 812 and the second threshold line 814, it is determined that the maternal peripheral blood sample to be tested during pregnancy is a sample with normal sex chromosomes
[0128] The first threshold line 812 for the DNA fragments in the initial data is constructed, for example, based on the function illustrated by the following expression (12). The first threshold line 814 for the DNA fragments in the initial data is constructed, for example, based on the function illustrated by the following expression (13). The range between the first threshold line 812 and the second threshold line 814 is for normal samples. Its confidence interval is 99.5%.
[0129] y = -0.99x + 4.15 (12)
[0130] y = -0.99x - 2.67 (13)
[0131] In the above formulas (12) and (13), x represents the proportion of the Y chromosome in the sample to be tested. y represents the characterization data of the abnormal content proportion of the X chromosome in the sample to be tested.
[0132] In the relationship graph 820 for the DNA fragments in the filtered data, 822 indicates the first threshold line for the DNA fragments in the filtered data, and 824 indicates the second threshold line for the DNA fragments in the filtered data. If the relevant graph of the sample to be tested is within the range between the first threshold line 822 and the second threshold line 824 for the DNA fragments in the filtered data, it is determined that the peripheral blood sample of the pregnant woman to be tested is a sample with normal sex chromosomes. The first threshold line 822 for the DNA fragments in the filtered data is constructed, for example, based on the function illustrated by the following expression (14). The second threshold line 824 for the DNA fragments in the filtered data is constructed, for example, based on the function illustrated by the following expression (15).
[0133] y = -1.03x + 5.37 (14)
[0134] y = -1.03x - 3.57 (15)
[0135] In the above formulas (14) to (15), x represents the proportion of the Y chromosome in the sample to be tested. y represents the characterization data of the abnormal content proportion of the X chromosome in the sample to be tested.
[0136] At step 708, if the computing device 110 determines that the relationship graph is outside the range between the first threshold line and the second threshold line, it is determined that the peripheral blood sample of the pregnant woman to be tested is a sample with abnormal sex chromosomes.
[0137] At step 710, if the computing device 110 determines that the relationship graph is outside the range between the first threshold line and the second threshold line and is higher than the first threshold point, it is determined that the peripheral blood sample of the pregnant woman to be tested has an abnormal karyotype of 47XXY or 47XYY.
[0138] Regarding the first threshold point, for example, it is the first threshold point 815 for the DNA fragment in the initial data (i.e., point A in the relationship graph 810); or the first threshold point 825 for the DNA fragment in the filtered data (i.e., point A in the relationship graph 820).
[0139] In some embodiments, if ChrY% > 1.5 * BFF, the sample to be tested has an abnormal karyotype of 47XYY. If ChrY% < 1.5 * BFF, the sample to be tested has an abnormal karyotype of 47XXY. BFF can be calculated, for example, using the formula (4) mentioned above.
[0140] In some embodiments, the present invention can further distinguish the origin of the abnormal karyotypes 47XXY or 47XYY based on the like - mosaic value (or "Like - Mosaic value", LM value). The following formula (16) schematically shows the calculation method of the predetermined ratio (LM value).
[0141] LM = LenAB / LenBC (16)
[0142] In the above formula (16), LenAB represents the length of the line segment AB in the relationship graph (i.e., the distance between point A and point B), and LenBC represents the length of the line segment BC in the relationship graph (i.e., the distance between point B and point C). LM represents the predetermined ratio, which is equal to the ratio of the length of the line segment AB to the length of the line segment BC in the relationship graph. It should be understood that if the sample is a non - mosaic 47XXY abnormality, then theoretically at this concentration the sample should be at point C, but actually it is at point A. Therefore, the line segment AB represents the difference between the concentration calculated by ChrY and the concentration calculated by ChrX. The line segment BC represents the difference between the concentration calculated by ChrY and 0. AB / BC represents the mosaic ratio. If the fetus is normal, the LM value will be approximately equal to "0". If the fetus is 47XXY and is not essentially mosaic, the LM value will be approximately equal to "1". If the fetus has an abnormality, the LM value for the DNA fragment in the filtered data will not change relative to the LM value for the DNA fragment in the initial data. If the mother has an abnormality, the LM value for the DNA fragment in the filtered data will decrease relative to the LM value for the DNA fragment in the initial data. Therefore, the present invention can judge the origin of the abnormality based on the ratio of the LM value for the DNA fragment in the filtered data to the LM value for the DNA fragment in the initial data. The following formula (17) schematically shows the calculation method of the LM value ratio.
[0143]
[0144] In the above formula (17), RLM represents the ratio of the LM value for the DNA fragments in the filtered data to the LM value for the DNA fragments in the initial data. Filter Tag LM Represents the LM value for the DNA fragments in the filtered data. Tag LM Represents the LM value for the DNA fragments in the initial data.
[0145] In some embodiments, if it is determined that the RLM value is greater than 0.8, it is determined that the abnormality comes from the fetus.
[0146] At step 712, if the computing device 110 determines that the relationship graph is outside the range between the first threshold line and the second threshold line and below the second threshold point, it is determined that the maternal peripheral blood sample to be tested has an abnormal karyotype of 46XY / 45X or 46XY / 45Y. The second threshold point is lower than the first threshold point.
[0147] Regarding the second threshold point, for example, it is the second threshold point 817 for the DNA fragments in the initial data (i.e., point D in the relationship graph 810); or the second threshold point 827 for the DNA fragments in the filtered data (i.e., point D in the relationship graph 820).
[0148] In some embodiments, if ChrY% > 1.5 * BFF, the sample to be tested has an abnormal karyotype of 46XY / 45X. If ChrY% < 1.5 * BFF, the sample to be tested has an abnormal karyotype of 46XY / 45Y.
[0149] In some embodiments, the present invention can further distinguish the origin of the abnormal karyotypes 46XY / 45X or 46XY / 45Y based on a predetermined ratio. The following formula (18) schematically shows the calculation method of the predetermined ratio (LM value).
[0150] LM = LenDE / LenEF (18)
[0151] In the above formula (16), LenDE represents the length of the line segment DE in the relationship graph (i.e., the distance between point D and point E), and LenEF represents the length of the line segment EF in the relationship graph (i.e., the distance between point E and point F). LM represents the predetermined ratio.
[0152] By adopting the above means, the present invention can avoid a relatively high false positive rate of fetal chromosomal abnormalities caused by the complexity of sex chromosomes, and can effectively detect abnormal karyotypes of sex chromosomes and infer the origin of the abnormalities.
[0153] Figure 10 Schematically shows a block diagram of an electronic device 1000 suitable for implementing the embodiments of the present invention. The electronic device 1000 can be used to implement the executionFigure 2 , Figures 4 to 7 the methods 200, 400 to 700 shown. As Figure 10 shown, the electronic device 1000 includes a central processing unit (i.e., CPU 1001), which can perform various appropriate actions and processes according to computer program instructions stored in the read-only memory (i.e., ROM 1002) or computer program instructions loaded from the storage unit 1008 into the random access memory (i.e., RAM 1003). In the RAM 1003, various programs and data required for the operation of the electronic device 1000 can also be stored. The CPU 1001, ROM 1002, and RAM 1003 are connected to each other via the bus 1004. The input / output interface (i.e., I / O interface 1005) is also connected to the bus 1004.
[0154] Multiple components in the electronic device 1000 are connected to the I / O interface 1005, including: the input unit 1006, the output unit 1007, the storage unit 1008, and the CPU 1001 executes the various methods and processes described above, such as executing the methods 200, 400 to 700. For example, in some embodiments, the methods 200, 400 to 700 can be implemented as computer software programs, which are stored in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the CPU 1001, one or more operations of the methods 200, 400 to 700 described above can be performed. Alternatively, in other embodiments, the CPU 1001 can be configured to perform one or more actions of the methods 200, 400 to 700 by any other suitable means (e.g., by means of firmware).
[0155] It should be further noted that the present invention can be a method, a device, a system, and / or a computer program product. The computer program product can include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0156] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0157] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0158] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit may execute the computer-readable program instructions to implement various aspects of the present invention.
[0159] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0160] These computer-readable program instructions can be provided to a processor in a voice interaction device, a general purpose computer, a special purpose computer, or other programmable data processing device's processing unit, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram is produced. These computer-readable program instructions can also be stored in a computer-readable storage medium, which instructions cause a computer, a programmable data processing device, and / or other devices to operate in a particular manner, so that the computer-readable medium storing the instructions comprises a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0161] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0162] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each box of the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0163] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
[0164] The above are only alternative embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for detecting the chromosome status of a fetus, characterized in that: include: Comparing the obtained sequencing data of the maternal peripheral blood sample during the pregnancy to be tested to the human reference genome, so as to obtain the initial data of the paired-end comparison to the human reference genome; Filtering the DNA fragment length in the initial data to obtain filtered data with DNA fragment length within a predetermined length range; Calculating a Z value for the target area based on the initial data and the filtered data, respectively, so as to obtain a Z value for the target area in the initial data and a Z value for the target area in the filtered data; as well as Determine the fetal chromosome status of the maternal peripheral blood sample during the pregnancy to be tested based on the obtained Z value of the target region in the initial data and the Z value of the target region in the filtered data; The method further comprises: calculating abnormal content ratio characterization data AcF for the initial data and the filtered data, so as to obtain abnormal content ratio characterization data about the DNA fragments in the initial data and abnormal content ratio characterization data about the DNA fragments in the filtered data; and comparing the data characterizing the abnormal content ratio of DNA fragments in the initial data and the data characterizing the abnormal content ratio of DNA fragments in the filtered data, so as to determine the source of the chromosome abnormality based on the comparison result; Wherein, determining the source of the chromosomal abnormality based on the comparison result includes: in response to confirming that the Y chromosome ratio of the DNA fragments in the filtered data is greater than a predetermined Y chromosome fetal DNA fragment concentration threshold, determining that the fetal DNA of the maternal peripheral blood sample to be tested is derived from a male fetus; based on a linear regression function, constructing a relationship graph between the abnormal content ratio characterization data of the X chromosome and the Y chromosome ratio of the maternal peripheral blood sample to be tested; in response to determining that the relationship graph is within a range between a first threshold line and a second threshold line, determining that the maternal peripheral blood sample to be tested is a sample with normal sex chromosomes.
2. The method according to claim 1, characterized in that Filtering the DNA fragment length in the initial data to obtain filtered data with DNA fragment length within a predetermined length range includes: Obtain the starting and ending positions of the DNA fragments in the initial data of the paired-end comparison to the human reference genome; Based on the obtained starting position and ending position, the length of each DNA fragment in the initial data is calculated; Determining whether the calculated length of each DNA fragment in the initial data is greater than a first predetermined size threshold and less than or equal to a second predetermined size threshold; In response to determining that the calculated length of the current DNA fragment in the initial data is greater than a first predetermined size threshold and less than or equal to a second predetermined size threshold, determining that the current DNA fragment in the initial data is filtered data belonging to a predetermined size range; and In response to determining that the calculated length of the current DNA fragment in the initial data is less than or equal to a first predetermined size threshold, or greater than a second predetermined size threshold, the current DNA fragment in the initial data is filtered out.
3. The method according to claim 1, characterized in that Calculating the Z value about the target area includes: For the initial data or the filtered data, respectively, the number of inserted fragments in the target region is divided by the total number of inserted fragments in the autosome, so as to obtain the genomic characterization data of the genomic target region respectively; Calculate the mean and standard deviation of the reference set sample genome representation data for the target region of the reference set sample; and Based on the genome characterization data of the genome target region and the mean and standard deviation of the genome characterization data of the reference set samples, the Z value of the target region in the initial data and the Z value of the target region in the filtered data are calculated respectively.
4. The method according to claim 1, characterized in that Determination of the fetal chromosome status of the maternal peripheral blood sample during the pregnancy to be tested includes: Determining whether the Z value of the target area in the filtered data falls within a predetermined threshold range; and In response to determining that the Z value of the target region in the filtered data belongs to a predetermined threshold range, it is determined that the chromosome status of the fetus of the maternal peripheral blood sample during the pregnancy to be tested is abnormal.
5. The method according to claim 2, characterized in that: The predetermined threshold value range is greater than or equal to 3, or less than or equal to -3.
6. The method according to claim 1, characterized in that The sources of chromosomal abnormalities determined based on the comparison results include: comparing the data characterizing the abnormal content ratio of DNA fragments in the filtered data with the data characterizing the abnormal content ratio of DNA fragments in the initial data; In response to determining that the abnormal content ratio characterizing data about the DNA fragments in the filtered data is greater than the abnormal content ratio characterizing data about the DNA fragments in the initial data, determining that the chromosomal abnormality originates from the fetus; and In response to determining that the abnormal content ratio characterizing data regarding the DNA fragments in the filtered data is equal to the abnormal content ratio characterizing data regarding the DNA fragments in the initial data, it is determined that the chromosomal abnormality originates from the mother and the fetus.
7. The method according to claim 1, characterized in that The sources of chromosomal abnormalities determined based on the comparison results include: In response to determining that the relationship graph is outside the range between the first threshold line and the second threshold line, the maternal peripheral blood sample during the pregnancy to be tested is determined to be a sample with sex chromosome abnormality.
8. The method according to claim 7, characterized in that The maternal peripheral blood samples to be tested during pregnancy are confirmed to be samples with sex chromosome abnormalities, including: In response to determining that the relationship graph is outside the range of the first threshold line and the second threshold line and is higher than the first threshold point, determining that the maternal peripheral blood sample during the pregnancy to be tested is an abnormal karyotype of 47XXY or 47XYY; and In response to determining that the relationship graph is outside the range of the first threshold line and the second threshold line and is lower than the second threshold point, it is determined that the maternal peripheral blood sample to be tested is an abnormal karyotype of 46XY / 45X or 46XY / 45Y, and the second threshold point is lower than the first threshold point.
9. A computing device, characterized in that include: at least one processing unit; At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the device to perform the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: A computer program is stored on a computer-readable storage medium, and when the computer program is executed by a machine, the method according to any one of claims 1 to 8 is implemented.
11. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a machine, performs the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Detection method of chromosome copy number variation
CN105349678A
Optimization method for noninvasive prenatal detection of fetal chromosome copy number abnormality
CN117230165A