Non-invasive method for assessing the risk of uniparental disomy and related diseases in pregnant fetuses
By non-invasively screening cfDNA in the peripheral blood of pregnant women, combined with quality control and hidden Markov model analysis, the invasiveness and missed diagnosis problems of fetal uniparental disomy detection in existing technologies are solved, and high-accuracy and sensitive fetal uniparental disomy detection is achieved.
Patent Information
- Application Number
- CN202510934638.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Existing UPD detection technologies rely on invasive and complex procedures, making it difficult to accurately identify fetal uniparental disomy and its related disease risks, especially the risk of missed diagnosis in the detection of chromosomal UPD.
Through non-invasive screening, the sequencing data of cfDNA in the peripheral blood of pregnant women is utilized. In combination with the quality control module, analysis and calculation module, and output module, the hidden Markov model (HMM) is used to analyze the allele fractions of the fetus to identify fetal uniparental disomy and its related disease risks.
It achieves high-accuracy and high-sensitivity non-invasive detection of fetal uniparental disomy, simplifies the detection process, reduces the risk of invasive operations on the fetus, and improves the accuracy of detection.
Smart Images

Figure CN120431996B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical diagnosis, and in particular to a method for non-invasively assessing the risk of uniparental disomy and related diseases in a pregnant fetus. Background Art
[0002] Uniparental disomy (UPD) occurs in approximately 1 in 2,000 newborns. Although most chromosomal UPDs do not manifest clinically, UPDs on chromosomes 6, 7, 11, 14, 15, and 20 can lead to phenotypic abnormalities due to differences in parental gene expression or abnormal methylation of imprinted genes. These imprinting disorders have become important factors in fetal developmental abnormalities and birth defects. The pathogenicity of other chromosomal UPDs is primarily associated with homozygous mutations in genes associated with recessive genetic diseases.
[0003] UPD detection technologies are developing in a diverse range of ways: short tandem repeat (STR) analysis, as a classic method, can identify UPD type and parental origin with high sensitivity and specificity, but requires parental samples and has difficulty detecting segmental UPD; chromosomal microarrays (SNP arrays) containing SNP probes are highly effective initial screening tools, especially for identifying homodisomy by detecting large regions of homozygous heterozygosity (ROH) on single chromosomes, but cannot detect heterodisomy, and approximately one-third of cases may be missed due to the lack of ROH; SNP computational analysis based on family exome / genome sequencing (WESW / GS) can identify whole-chromosome UPD and segmental UPD greater than 10 Mb, but attention should be paid to interference from heterozygous deletions. Existing detection technologies all rely on invasive procedures such as chorionic villus sampling or amniocentesis to obtain fetal samples, and often require comparison of parental samples, which increases the complexity of the diagnostic process.
[0004] With the advancement of next-generation sequencing (NGS) technology, non-invasive prenatal screening (NIPS) based on cell-free fetal DNA (cfDNA) analysis has become an important means of detecting fetuses at high risk of chromosomal aneuploidy.
[0005] Therefore, developing a method for detecting uniparental disomy and the risk of related diseases in a pregnant fetus through non-invasive screening is of great significance to the field. Summary of the Invention
[0006] The present invention provides a method for detecting uniparental disomy and related disease risks in a pregnant fetus by non-invasive screening.
[0007] In a first aspect of the present invention, a detection device for determining fetal uniparental disomy (UPD) is provided, the device comprising:
[0008] (a) a data input module configured to input nucleic acid data of a sample to be analyzed, wherein the nucleic acid data is sequencing data of cfDNA in the peripheral blood of a pregnant woman, and the sequencing data includes sequencing data of maternal free deoxyribonucleic acid and sequencing data of fetal free deoxyribonucleic acid;
[0009] (b) a quality control module configured to perform quality control on the nucleic acid data to obtain nucleic acid data that meets quality control conditions; the quality control includes: sequencing depth, fetal fraction calculation, maternal CNV and polymorphism information site screening;
[0010] (c) an analysis and calculation module configured to analyze the nucleic acid data after quality control, the analysis including copy number analysis and allele score analysis; wherein, in the copy number analysis, if a paternal CNV or is detected, a fetal CNV is determined; if no CNV is detected, an allele score analysis is performed; in the allele score analysis, for the fetus, at a maternally homozygous SNP site (effective single nucleotide variant site), if the maternal specific allele is missing, it indicates that the fetus has a high risk of paternal UPD; if the paternal specific allele is missing, it indicates that the fetus has a high risk of maternal UPD;
[0011] (d) an output module, wherein the output module is configured to output the detection result of the fetal UPD.
[0012] In another preferred embodiment, the quality control conditions include: sequencing depth greater than threshold A (min_dep), fetal fraction greater than threshold B (min_ff), no maternal CNV detected, and the number of polymorphic information sites greater than threshold C.
[0013] In another preferred embodiment, in the analysis and calculation module, the allele fraction analysis includes: (c1) calculating the fetal fraction (FF), (c2) determining the fetal genotype, and (c3) detecting uniparental disomy of the fetus based on a hidden Markov model.
[0014] In another preferred embodiment, in (c1), it includes: analyzing the nucleic acid data of the sample to be analyzed to determine the fetal fraction.
[0015] In another preferred embodiment, in (c1), when the maternal genotype is homozygous (AA or BB) and the fetal genotype is heterozygous (AB) at the SNP site, the fetal fraction (FF) is calculated by the following steps:
[0016] (1) For a certain SNP site i, let N be the total number of sequencing reads for alleles A and B, NA i is the read length of allele A, NB iis the number of reads of allele B; then the fetal fraction (FF) of this site AAi or FF BBi ):
[0017] (1),
[0018] (2),
[0019] (3);
[0020] (2) Take the median of the FF values of all the SNP sites and calculate it as FF AA and FF BB , the calculation formula of the fetal fraction of the sample is as follows:
[0021] (4).
[0022] In another preferred embodiment, in the analysis module, the copy number variation analysis includes: probe design.
[0023] In another preferred embodiment, in (c2), determining the fetal genotype comprises:
[0024] (i) determining the B allele frequency afe of the candidate fetal genotype g based on the fetal score and the maternal genotype;
[0025] (ii) For each SNP site i in the fetus, the alternative allele (alt i ) is shown in formula (5):
[0026] (5);
[0027] Where n represents the total read count of site i, α and β are the parameters of the β-binomial distribution;
[0028] (iii) Using the B allele frequency afe of the candidate fetal genotype g determined in (i), calculate the parameters of the β-binomial distribution, as shown in equations (6) and (7):
[0029] (6),
[0030] (7);
[0031] And, assign a weight W to each candidate fetal genotype g g ;
[0032] (iv) Calculate the raw score of the candidate fetal genotype: For each SNP i of the fetus, the raw score S(i) is calculated as shown in formula (8):
[0033] (8)
[0034] Where B() is the β function; is the binomial coefficient, i.e. x in (ii); W g is the weight corresponding to the candidate fetal genotype g; alt i is the read count of the alternative allele; n is the total read count of site i; the values of α and β are calculated by equations (6) and (7) in (iii);
[0035] (v) Determining the fetal genotype: selecting the candidate fetal genotype with the highest raw score as the fetal genotype of the current SNP site.
[0036] In another preferred embodiment, the W g The values are selected from the following table:
[0037] .
[0038] In another preferred embodiment, in (c3), it includes:
[0039] (c31) Obtaining the observation sequence V: For each SNP in the query region (referring to the interval where UPD needs to be detected, see Table 1 for details), if the maternal genotype is homozygous, generate the observation sequence V consisting of the observed values according to the maternal genotype and the fetal genotype obtained in (c2) and compare them with the following table;
[0040]
[0041] (c32) deriving a hidden state sequence X using a hidden Markov model (HMM) based on the observation sequence V, thereby obtaining a UPD detection result of the fetus; the hidden state is selected from the group consisting of diploid (D), maternal uniparental disomy (UPDM), paternal uniparental disomy stage I (UPDPI), and paternal uniparental disomy stage II (UPDPII);
[0042] The HMM is defined by a parameter set λ=(A; B; π), where A is the state transition probability, B is the observation probability matrix, and π is the initial state distribution.
[0043] In another preferred embodiment, in (c32), the state transition probability matrix used by the HMM is shown in the following table:
[0044] .
[0045] In another preferred embodiment, in (c32), the observation probability matrix used by the HMM is shown in the following table:
[0046] .
[0047] In another preferred embodiment, in (c32), the initial probability distribution used by the HMM is shown in the following table:
[0048] .
[0049] In another preferred embodiment, in (c32), the method further includes: calculating the proportion of each state in the final hidden state sequence X, thereby obtaining the UPD detection result of the fetus;
[0050] Among them, when the proportion of D is greater than the preset threshold, the detection result of the fetal UPD is: D;
[0051] When the proportion of D is less than the preset threshold, the following judgment is made: if the sum of the proportions of UPDPI and UPDPII is greater than the proportion of UPDM, UPDM is corrected to UPDPI or UPDPII, and the detection result of the fetal UPD is: paternal UPD; if the sum of the proportions of UPDPI and UPDPII is less than the proportion of UPDM, UPDPI or UPDPII is corrected to UPDM, and the detection result of the fetal UPD is: maternal UPD.
[0052] In another preferred embodiment, in (c32), if the UPD test result of the fetus is paternal UPD, the UPD subtype can be further determined by the following method:
[0053] The proportions of UPDPI and UPDPII were compared, and the state with the higher proportion was designated as the final UPD type for that region.
[0054] In a second aspect of the present invention, a method for analyzing fetal UPD data is provided, the method comprising the following steps:
[0055] (a) providing data, wherein the data includes nucleic acid data of a sample to be analyzed, wherein the nucleic acid data is sequencing data of cfDNA in the peripheral blood of a pregnant woman, and the sequencing data includes sequencing data of maternal free deoxyribonucleic acid and sequencing data of fetal free deoxyribonucleic acid;
[0056] (b) Data quality control: performing quality control on the nucleic acid data to obtain nucleic acid data that meets quality control requirements; the quality control includes sequencing depth, fetal fraction calculation, maternal CNV, and polymorphic information site screening;
[0057] (c) Analysis and calculation: The nucleic acid data after quality control are analyzed, including copy number analysis and allele fraction analysis. In the copy number analysis, if a fetal CNV is detected, the fetal CNV is determined; if no CNV is detected, the allele fraction analysis is performed. In the allele fraction analysis, if the maternal homozygous informative SNP site (effective single nucleotide variant site) is missing, it indicates that the fetus has a high risk of paternal UPD; if the paternal specific allele is missing, it indicates that the fetus has a high risk of maternal UPD.
[0058] (d) Output results: Output the fetal UPD assessment results obtained after analysis and calculation.
[0059] It should be understood that within the scope of the present invention, the above-mentioned technical features of the present invention and the technical features described in detail below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be listed here one by one. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 Examples of allele fractions for fetal uniparental disomy (UPD) are shown: (A) Maternal uniparental heterodimy: caused by maternal meiosis I (MI) nondisjunction and trisomy rescue; (B) Maternal uniparental homodimy: caused by maternal meiosis II (MII) nondisjunction and trisomy rescue; (C) Paternal uniparental heterodimy: caused by paternal meiosis I (MI) nondisjunction and trisomy rescue; (D) Paternal uniparental homodimy: caused by paternal meiosis II (MII) nondisjunction and trisomy rescue; (E) Mixed maternal uniparental disomy: caused by maternal meiosis I (MI) recombination, MI / meiosis II (MII) nondisjunction, and trisomy rescue; (F) Segmental maternal uniparental heterodimy; (G) Segmental maternal uniparental homodimy: homozygosity of specific chromosomal regions due to mitosis recombination; (H) Paternally derived mixed uniparental disomy: caused by paternally derived meiosis I (MI) recombination, MI / meiosis II (MII) nondisjunction, and trisomy rescue; (I) segmented paternally derived uniparental heterodimy; (J) segmented paternally derived uniparental homodimy: caused by mitotic recombination resulting in homozygosity of specific chromosomal regions.
[0061] Figure 2A flow chart for UPD detection according to the present invention is shown. Fetal UPD detection requires a combination of read depth (RD) and SNP allele fraction analysis. Quality control (QC) includes monitoring for adequate read depth, the proportion of fetal free DNA, maternal copy number variation (CNV), and calculating the number of informative loci where the fetus is heterozygous and the mother is homozygous. After passing QC, read depth analysis and allele fraction analysis are performed to detect fetal UPD.
[0062] Figure 3 The study demonstrated the clinical validation of NIPS-UPD (non-invasive prenatal uniparental disomy) using maternal plasma samples. Four maternal blood samples were tested, and all four showed positive results for uniparental disomy on whole or segmental chromosomes in the fetus.
[0063] Figure 4 Figure 3 shows the detection of fetal uniparental disomy (UPD) by combining sequencing depth and SNP allele fraction: (A) Case P1: The fetus had maternal UPD of the whole chromosome 7 (chr7); (B) Case P2: The fetus had trisomy 7; (C) The fetus had paternal UPD of the long arm of chromosome 14, region 32 (chr14q32); (D) Case P3: The fetus had maternal UPD of the long arm of chromosome 15, regions 11 to 13 (chr15q11q13).
[0064] Figure 5 The probe distribution of the present invention is shown.
[0065] Figure 6 The UPD probe distribution of the present invention is shown. DETAILED DESCRIPTION
[0066] After extensive and in-depth research, the present inventors have developed a highly accurate and sensitive method for detecting fetal uniparental disomy (UPD). Based on their research, the present invention integrates chromosome dose and SNP allele fraction analysis, and employs a hidden Markov model to detect UPD in fetal cfDNA. This method is highly accurate and can serve as an important tool for screening UPD and related diseases. This work has led to the completion of the present invention.
[0067] the term
[0068] In order to more easily understand the present disclosure, some terms are first defined. As used in this application, unless otherwise expressly provided herein, each of the following terms should have the meaning given below. Other definitions are set forth throughout the application.
[0069] The term "about" can refer to a value or composition that is within an acceptable error range for the particular value or composition as determined by one of ordinary skill in the art, which will depend in part on how the value or composition is measured or determined.
[0070] As used herein, the terms "comprising" or "including" may be open, semi-closed, or closed. In other words, the terms also include "consisting essentially of" or "consisting of."
[0071] As used herein, unless otherwise indicated, any concentration range, percentage range, ratio range, or integer range should be understood to include the value of any integer within the range and, where appropriate, fractional values thereof (e.g., tenths and hundredths of an integer).
[0072] As used herein, the term "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0073] As used herein, "monosomy self-rescue", "monosomy rescue" or "Monosomy Rescue" have the same meaning and are used interchangeably, and all refer to the formation of a disomy in a monosomic embryo by duplicating a single chromosome.
[0074] As used herein, “trisomy self-rescue”, “trisomy rescue” or “Trisomy Rescue” have the same meaning and can be used interchangeably, referring to the restoration of trisomic embryos to normal disomy by the random loss of a chromosome.
[0075] The monosomy self-rescue and trisomy self-rescue are both biological mechanisms by which embryos autonomously correct abnormal chromosome numbers.
[0076] As used herein, the term "read depth" refers to the target number of sequencing reads obtained per sample.
[0077] As used herein, the terms "fetal-specific allele fraction (BAF)" and "B-allele fraction" are used interchangeably to refer to the overall BAF of a fetal-specific SNP.
[0078] Uniparental disomy
[0079] Uniparental disomy (UPD) occurs when a pair of homologous chromosomes or chromosome segments are derived entirely from a single parent, with no genetic material from the other parent. Based on the parental origin, UPD can be divided into maternal UPD (matUPD, UPDM) and paternal UPD (patUPD, UPDP).
[0080] Paternal uniparental disomy includes uniparental disomy at paternal stage I (UPDPI) and uniparental disomy at paternal stage II (UPDPII). UPDPI refers to a fetal chromosome pair or part of a chromosome derived from two different homologous chromosomes from the father (meiosis I error); UPDPII refers to a chromosome pair or part of a chromosome derived from the same paternal chromosome copy (meiosis II error).
[0081] Its types are further divided into two types: homologous disomy (iUPD) refers to the two homologous chromosomes being exact copies of the same parental chromosome; heterologous disomy (hUPD) refers to the two homologous chromosomes being derived from the same parent but representing different parental homologous chromosomes.
[0082] Complete hUPD typically results from nondisjunction during meiosis I (MI), while complete iUPD results from nondisjunction during meiosis II (MII) or mitotic errors. Mixed UPD refers to the presence of both iUPD and hUPD segments on the same chromosome, often due to meiotic recombination. Segmental UPD involves only a subset of chromosomes and is often triggered by mechanisms such as mitotic recombination or trisomy rescue.
[0083] The main mechanisms of UPD formation include trisomy / monosomy self-rescue, gamete complementation and post-fertilization errors.
[0084] Hidden Markov Model
[0085] The Hidden Markov Model (HMM) is a probability-based statistical model used to infer hidden states in sequential data. Its core concept is to assume that a system has a set of hidden states that are not directly observable. These states transition according to the Markov property (the current state depends only on the previous state), and each state generates an observable output. It has applications in fields such as speech recognition, natural language processing, and bioinformatics.
[0086] Detection method of the present invention
[0087] The present invention provides a method for analyzing fetal UPD data, the method comprising the following steps:
[0088] (a) providing data, wherein the data includes nucleic acid data of a sample to be analyzed, wherein the nucleic acid data is sequencing data of cfDNA in the peripheral blood of a pregnant woman, and the sequencing data includes sequencing data of maternal free deoxyribonucleic acid and sequencing data of fetal free deoxyribonucleic acid;
[0089] (b) Data quality control: performing quality control on the nucleic acid data to obtain nucleic acid data that meets quality control requirements; the quality control includes sequencing depth, fetal fraction calculation, maternal CNV, and polymorphic information site screening;
[0090] (c) Analysis and calculation: The nucleic acid data after quality control are analyzed, including copy number analysis and allele fraction analysis. In the copy number analysis, if a fetal CNV is detected, the fetal CNV is determined; if no CNV is detected, the allele fraction analysis is performed. In the allele fraction analysis, if the maternal homozygous informative SNP site (effective single nucleotide variant site) is missing, it indicates that the fetus has a high risk of paternal UPD; if the paternal specific allele is missing, it indicates that the fetus has a high risk of maternal UPD.
[0091] (d) Output results: Output the fetal UPD assessment results obtained after analysis and calculation.
[0092] The main advantages of the present invention include:
[0093] (a) Compared with traditional biochemical methods and ultrasound screening, NIPS is more accurate in detecting fetal UPD by analyzing cfDNA in maternal blood.
[0094] The present invention will be further described below in conjunction with specific examples. It should be understood that these examples are intended to illustrate the present invention only and are not intended to limit the scope of the invention. The experimental methods in the following examples, for which specific conditions are not specified, are generally performed under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or according to the conditions recommended by the manufacturer. Unless otherwise indicated, percentages and parts are by weight.
[0095] General Methods
[0096] Extraction and sequencing of cell-free DNA
[0097] Cell-free DNA extraction
[0098] Maternal plasma was processed using a standardized two-step centrifugation procedure. First, at least 1.8 mL of maternal peripheral blood was collected and centrifuged at 1,600 g for 15 minutes at 4°C to initially separate plasma. Subsequently, the plasma was transferred to a fresh tube and centrifuged a second time (16,000 g for 10 minutes at 4°C) to completely remove residual cellular debris. Finally, circulating free DNA (cfDNA) was extracted using a magnetic bead-based serum / plasma circulating DNA extraction kit (Cat. No. YDP710; Tiangen Biochemical Technology, Beijing, China) in strict accordance with the manufacturer's instructions.
[0099] Next-generation sequencing (NGS) library construction and data analysis process
[0100] During next-generation sequencing (NGS) library preparation, extracted circulating free DNA (cfDNA) was end-repaired and adapter-ligated using an adapter ligation reagent with unique molecular identifiers (UMIs) (Cat. No. 1002212; Nona Biotech, Nanjing, China). Sample-specific barcodes were introduced by PCR, and PCR products were quantified using the Qubit™ 1X Double-Stranded DNA High Sensitivity Detection Kit (Cat. No. Q33231; Thermo Fisher Scientific, USA). At least 400 ng of PCR product was used for subsequent target enrichment. After pooling 12 to 36 samples, hybridization capture was performed using probes according to the manufacturer's protocol (Heristar LLC, USA). Hybridized DNA was purified using Dynamag-270 magnetic beads (Cat. No. 65306; Thermo Fisher Scientific).
[0101] To construct sequencing libraries, secondary PCR amplification (Cat. No. 1002212; Nona Biopharmaceuticals) was performed and quantified using the Qubit™ Single-Stranded DNA Detection Kit (Cat. No. Q10212; Thermo Fisher Scientific). Subsequently, a single-stranded circular DNA library was prepared using the MGI-Easy Circularization Kit (Cat. No. 1000005260; MGI, Shenzhen, China). DNA nanospheres were generated by rolling circle amplification (RCA) according to the manufacturer's instructions (Cat. No. 1000012554; MGI), and quantified using the Qubit™ Single-Stranded DNA Detection Kit (Cat. No. Q10212; Thermo Fisher Scientific). Sequencing was performed on the MGISEQ-2000 platform (Cat. No. 1000012554; MGI) in paired-end 100 bp (PE100) mode. Detailed experimental procedures are described in the literature.
[0102] Raw FASTQ data were filtered and UMI preprocessed using Fastp (v0.21.0, https: / / github.com / OpenGene / fastp). Quality-controlled FASTQ files were aligned to the hg38 human reference genome using BWA (v0.7.17-r1188, https: / / github.com / lh3 / bwa) and sorted using Samtools (v1.19, https: / / github.com / samtools / samtools / releases / ). Finally, consensus BAM files were generated using Gencore (v0.15.0, https: / / github.com / OpenGene / gencore).
[0103] SNP chip analysis
[0104] Amniocytes were collected, and genomic DNA was extracted using the QIAamp DNA Mini Kit (Qiagen, Valencia, California, USA). To identify potential genomic imbalances, SNP microarray analysis was performed using a commercial 750K microarray (CytoScan® 750K Array; Affymetrix, Santa Clara, California, USA) according to the manufacturer's instructions. DNA was fragmented and hybridized to the array, which was then washed with buffer and scanned using a laser scanner. The microarray is equipped with probes for over 750,000 CNV markers and 200,000 genotypeable SNPs, enabling high-resolution copy number variation detection, precise breakpoint estimation, and identification of regions of homozygosity (ROH). Data were subsequently analyzed using Chromosome Analysis Suite (ChAS) V3.2 software (Affymetrix, USA).
[0105] Methylation analysis
[0106] To confirm uniparental disomy (UPD), methylation-specific multiplex ligation-dependent probe amplification (MS-MLPA) analysis was performed according to the manufacturer's instructions (SALSA® MLPA® Probemix ME034 Multi-locus Imprinting, MRC-Holland, https: / / www.mrcholland.com / product / ME034). This assay detects abnormal methylation patterns in 14 differentially methylated regions (DMRs) across eight chromosomes.
[0107] The probe mix contains a total of 40 probes, 27 of which are methylation-specific. Each methylation-specific probe is designed with an HhaI recognition site, enabling the assessment of the methylation status of sequences known to be methylated on either the paternal or maternal allele.
[0108] Target regions include three probes each for H19 and PEG3, two probes each for KCNQ1OT1, MEST, MEG3, MEG8, SNRPN, PLAGL1, GRB10, and ZNF597, and five probes for the GNAS complex locus. In addition to methylation analysis, all probes provide information on potential CNVs. Eleven reference probes are included in the probe mix to detect regions outside the target chromosomal loci, ensuring robust normalization and quality control. Additionally, two digestion control probes are included to verify the integrity of the HhaI restriction enzyme digestion in the MS-MLPA reaction. For the chr15q11 region, copy number variation and methylation status are detected using the SALSA® MLPA® Probemix ME028 kit.
[0109] Example 1
[0110] 1.1 Research subjects
[0111] A total of 15 samples from pregnant women with singleton pregnancies at 12 weeks of gestation or longer were included in this study. These samples were leftover from women who had undergone amniocentesis or chorionic villus sampling for standard prenatal diagnosis. Plasma was obtained from the mothers before the invasive prenatal diagnostic procedure. For all cases with available pregnancy outcome data, results from microarray-based comparative genomic hybridization, sequencing data, and / or methylation analysis were collected for validation studies.
[0112] 1.2 Probe design
[0113] Probes were specifically designed to target single nucleotide polymorphisms (SNPs) throughout chromosomes and in chromosomal regions frequently associated with aneuploidy and microdeletions, including chromosomes 1, 2, 4, 5, 8, 9, 11, 13, 15, 16, 17, 18, 21, 22, X, and Y (Table 1).
[0114] Table 1 Screening for chromosomal abnormalities
[0115]
[0116] In addition, additional enrichment was observed for chromosomes associated with pathogenic imprinting disorders, including transient neonatal diabetes mellitus (chr6:142,440,300-145,508,262), Silver-Russell syndrome (chr7:128,986,171-132,168,748; chr11:495,166-4,385,775), Beckwith-Wiedemann syndrome (chr11:495, 166-4,385,775), Temple syndrome / Kagami-Ogata syndrome (chr14:99,226,892-103,063,452), Prader-Willi syndrome / Angelman syndrome (chr15:23,065,674-26,365,088), pseudohypoparathyroidism type 1B (chr20:57,318,919-60,411,192) (Table 2).
[0117] Table 2 Probe design area
[0118] UPD: uniparental disomy.
[0119] The probe design information of the present invention is as follows Figure 5-6 shown.
[0120] Among them, the information of 60 probes designed by the present invention for detecting UPD is shown in Table 3.
[0121] Table 3
[0122]
[0123]
[0124] Chromosomal coordinates are based on the GRCh38 reference genome. To ensure optimal performance, probes must be located within the target region, where the GC content ranges from 30% to 70% and common copy number variations (CNVs) reported in the Database of Genomic Variants (http: / / dgv.tcag.ca / dgv / app / home, last accessed March 31, 2023) are absent (population frequency >1%). The minor allele frequencies of the aforementioned SNPs range from 30% to 70% in East Asian, African, and non-Finnish European populations to ensure sufficient polymorphism across diverse populations. To reduce probe hybridization bias between the reference and alternative alleles, one of the four possible nucleotides (A, C, G, and T) was selected at the locus corresponding to the target SNP to minimize the difference in melting temperature between the reference and alternative alleles. Furthermore, quality control measures require a read depth of ≥200× for >99.0% of the target loci to ensure sufficient coverage for reliable detection in this assay.
[0125] 1.3 Calculation of fetal fraction (FF)
[0126] At any biallelic locus, the maternal and fetal genotypes are homozygous, and two possible combinations (AA-AB and BB-AB; A: variant allele, B: reference allele) can be used to calculate the fetal fraction. For a certain SNP site i, let N be the total number of sequencing reads for alleles A and B, NA i is the read length of allele A, NB i is the read length of allele B.
[0127] At a locus where the maternal genotype is homozygous (AA or BB) and the fetal genotype is heterozygous (AB), the fetal fraction (FF) of the locus can be calculated using the following formula: AAi or FF BBi ):
[0128] (1)
[0129] (2)
[0130] (3)
[0131] The median FF value of all informative SNP sites is taken as FF AA and FF BB , then the fetal fraction calculation formula of the sample is:
[0132] (4)
[0133] 1.4 Determination of fetal genotype
[0134] Based on the B allele frequency (BAF) value, the maternal genotype can be determined using the following rules:
[0135] When BAF < 0.25, the maternal genotype is BB;
[0136] When 0.25 ≤ BAF ≤ 0.75, the maternal genotype is BA (heterozygous);
[0137] When BAF>0.75, the maternal genotype is AA.
[0138] For fetal genotype:
[0139] When BAF<0.1 or BAF>0.99, the fetal genotype was determined to be homozygous, corresponding to BB or AA, respectively;
[0140] For BAF values in the intermediate range (0.1 ≤ BAF ≤ 0.99), the algorithm previously established by the present inventors was used for more accurate fetal genotype inference (Xu C et al. 2022).
[0141] To determine the probability score for each SNP, the following steps are performed:
[0142] (i) Calculate the fetal fraction (FF) using the method described above and then substitute the FF value into Table 4 to determine the expected allele frequencies for subsequent calculations.
[0143] Table 4 Expected allele frequencies according to genotype
[0144]
[0145] Note: Genotype: maternal-fetal genotype; BAF: B allele
[0146] FF: fetal fraction.
[0147] (ii) Define SNP parameters and their distribution: For each SNP site i (where alt i represents the read count of the alternative allele, and n is the total read count), and the β-binomial distribution is used to model its likelihood function:
[0148] (5)
[0149] Among them, x represents alt i ;
[0150] (iii) Assign parameters based on genotype: Using the maternal genotype and considering the three possible fetal genotypes (denoted by g), the expected allele frequencies (afe) are obtained from Table 3. This value is used to calculate the parameters of the beta-binomial distribution:
[0151] (6)
[0152] (7)
[0153] In addition, each genotype g is assigned a weight (W) from Table 5 based on the maternal and fetal genotypes.
[0154] Table 5 Weight of each genotype
[0155]
[0156] Note: Genotype: maternal-fetal genotype; D: diploid; UPD: uniparental disomy;
[0157] M: maternal; P: paternal; I: meiosis I; II: meiosis II.
[0158] (iv) Calculate the raw score for each state: For each SNP 𝑖, its raw score S(i) is calculated as follows:
[0159]
[0160] Where B() is the β function; is the binomial coefficient (i.e. x in step (iii)); is the weight corresponding to the maternal and fetal genotype g. i represents the read count of the alternative allele, and n is the total read count;
[0161] (v) Determine the fetal genotype: Based on the genotype g with the highest score among all genetic states, this genotype is determined as the most likely fetal genotype at the current SNP site.
[0162] In summary, this method integrates prior knowledge of maternal and fetal genotypes through a probabilistic model to calculate the likelihood score for each SNP locus. By integrating the parameters of the β-binomial distribution, expected allele frequencies, and genotype-specific weights, it ensures the accuracy and robustness of fetal genotype determination.
[0163] 1.5 Detection of uniparental disomy (UPD)
[0164] To identify uniparental disomy (UPD), this study developed a hidden Markov model (HMM)-based analytical strategy. This approach relies on the availability of fetal genotype information, specifically focusing on loci where the maternal genotype is homozygous. By incorporating this information, the HMM framework can accurately detect UPD events, effectively distinguishing normal inheritance patterns from abnormal genomic configurations.
[0165] The HMM is defined by the following core components:
[0166] T: length of the observation sequence;
[0167] N: the number of hidden states in the model;
[0168] M: the number of observation symbols;
[0169] Q: set of hidden states, Q = {q0, q1, …, q_(N-1)} (e.g., q0 = normal two-body, q1 = UPDM, q2 = UPDPI, q3 = UPDPII);
[0170] V: set of possible observation values, V = {0, 1, …, M-1} (e.g., genotype coding such as 0=BB-BB, 1=BB-BA, 2=BB-AA);
[0171] A: state transition probability matrix (Table 6), A[i][j] represents the probability of transitioning from state q_i to q_j;
[0172] B: Observation probability matrix (Table 7), B[i][k] represents the probability of generating observation value k in state q_i;
[0173] π: initial state distribution (Table 8), π[i] represents the probability that the model is in state q_i at the initial moment;
[0174] O: Observation sequence, O = {O0, O1, …, O_(T-1)} (actually observed genotype data sequence).
[0175] Table 6 State transition probability
[0176]
[0177] Note: D: diploid; UPD: uniparental disomy; M: maternal;
[0178] P: paternal origin; I: meiosis I; II: meiosis II.
[0179] Table 7 Observation probability matrix
[0180] Note: D: diploid; UPD: uniparental disomy; M: maternal;
[0181] P: paternal origin; I: meiosis I; II: meiosis II.
[0182] Table 8 Initial probability distribution
[0183]
[0184] Note: D: diploid; UPD: uniparental disomy; M: maternal;
[0185] P: paternal origin; I: meiosis I; II: meiosis II.
[0186] For this analysis: V = {0, 1, 2, 3, 4, 5} (possible observations), Q = {“D”, “UPDM”, “UPDPI”, “UPDPII”} (different states), the state transition probability matrix A is defined as:
[0187]
[0188] For a SNP in the query region (i.e., the interval where UPD needs to be detected, see Table 1 for details), if the maternal genotype is homozygous, it is converted into the corresponding state label through the predefined mapping table (Table 9) and the label is added to the observation sequence O.
[0189] Table 9 Conversion relationship between state labels and genotypes
[0190]
[0191] Hidden Markov Model (HMM) Analysis and State Inference: A hidden Markov model is defined by a parameter set λ = (A; B; π), where A is the transition probability, B is the observation probability (emission probability), and π is the initial state distribution. Given an observation sequence O, the HMM can infer the corresponding hidden state sequence X.
[0192] If the most common state in X is D (i.e., the proportion of D is greater than a threshold), the fetal UPD test result is output as normal diploidy, meaning the fetus does not have UPD. If the proportion of D is less than the threshold, and the most common state in X is UPDPI or UPDPII (i.e., if the proportion of UPDPI and UPDPII in X is greater than the proportion of UPMD), the state UPDM is updated (corrected to UPDPI or UPDPII) to match the most common state, and the fetal test result is determined to be paternal UPD. If the proportion of D is less than the threshold, and the most common state in X is UPMD (i.e., if the proportion of UPMD in X is greater than the proportion of UPDPI and UPDPII), UPDPI or UPDPII is corrected to UPDM, resulting in a fetal UPD test result of maternal UPD.
[0193] If the most common states are UPDPI and UPDPII (that is, the number of states of UPDPI and UPDPII exceeds half of the total number of states in the state sequence X), perform the following steps:
[0194] (1) Comparison of the total probability scores of type I (UPDPI) and type II (UPDPII) SNPs;
[0195] (2) The state with the higher total probability score (type I or type II) is designated as the final uniparental disomy (UPD) type for the region.
[0196] This process, through iterative optimization and combining dominant UPD patterns with probability scores, gradually improves the accuracy of analysis results. Its core significance lies in accurately identifying uniparental disomy (UPD) events, ensuring reliable results and providing a basis for interpreting genomic variation (such as the source of chromosomal abnormalities).
[0197] Uniparental disomy (UPD) refers to a genetic abnormality in which both homologous chromosomes in an offspring are derived from a single parent. UPDPI, UPDPII, and UPDM may represent different subtypes or states of UPD (specific definitions should be considered in context). This method, through probabilistic modeling and dynamic adjustment of hidden Markov models, effectively improves the detection accuracy of complex genetic events such as UPD, providing an important tool for genomic research and clinical diagnosis.
[0198] 1.6 Experimental Results
[0199] Fetal uniparental disomy detection based on chromosome dosage and SNP allele fraction
[0200] Cell-free DNA (cfDNA) extracted from maternal plasma is a mixture of maternal and fetal genomic DNA. Fetal heterozygosity can be effectively used to detect chromosomal imbalance abnormalities at biallelic loci where the mother is homozygous. Based on Mendelian inheritance, the allele fraction distribution patterns of different UPD types have been systematically analyzed ( Figure 1 AJ in ).
[0201] Based on targeted, high-depth sequencing data, this study developed a novel algorithm that integrates sequencing depth with allele fraction data to simultaneously analyze chromosome dosage and SNP information to accurately identify UPD events. This method also reduces analytical error by detecting maternal and fetal copy number variations (CNVs) as a quality control step.
[0202] Main process of fetal UPD detection ( Figure 2 )as follows:
[0203] Quality control (QC): Evaluate sequencing data quality, including sequencing depth, fetal fraction (FF) calculation, maternal CNV analysis, and screening of informative sites (SNP sites where the maternal genotype is homozygous and the fetus is heterozygous).
[0204] Copy number analysis: If fetal CNV is detected, the fetal CNV is inferred; if no CNV is detected, proceed to the next step.
[0205] Allele fraction analysis: At informative sites (SNP sites), if the maternal allele is missing, it indicates the fetus's paternal UPD, and if the paternal allele is missing, it indicates the fetus's maternal UPD.
[0206] Example 2 Clinical validation of the NIPS-UPD method
[0207] This study developed a NIPS-UPD method based on chromosome dosage and SNP allele scoring and evaluated its clinical validity using maternal plasma samples validated by SNP microarray analysis and fetal genome sequencing.
[0208] In this example, a total of 4 singleton pregnancies were included ( Figure 3 ), the method of the present invention (NIPS-UPD) was used for auxiliary diagnosis, and 4 cases were detected to be UPD positive on whole chromosomes or segmental chromosomes ( Figure 4 AD in the , including:
[0209] chr7 (1 case), chr7q32 (1 case), chr14q32 (1 case), chr15q11q13 (1 case).
[0210] Finally, 4 NIPS-UPD-positive cases were confirmed as true positive by fetal genome sequencing ( Table 10 ).
[0211] Table 10 Summary of clinical information and test results of the 4 cases included in this study
[0212]
[0213] Example 3 Identification of UPD-related diseases based on MS-MLPA
[0214] UPD in imprinted regions of the genome can disrupt the balance of gene expression, leading to a variety of imprinting disorders. Currently, nine UPD-related imprinting disorders are known (Table 11).
[0215] Table 11 Overview of human diseases associated with UPD
[0216]
[0217] In this example, four cases identified known pathogenic UPD regions using NIPS-UPD (Table 9), suggesting an increased risk of UPD-related diseases. To confirm their pathogenicity, they were further analyzed using methylation-specific multiplex ligation-dependent probe amplification (MS-MLPA), the gold standard for diagnosing UPD-related imprinting diseases.
[0218] MS-MLPA results showed that all four fetuses tested had pathogenic UPD, involving the following diseases:
[0219] Silver-Russell syndrome (SRS) (subjects P1, P1);
[0220] Kagami-Ogata syndrome (KOS) (subject P3);
[0221] Prader-Willi syndrome (PWS) (subject P4);
[0222] These results fully verified the reliability and clinical practicality of the NIPS-UPD method.
[0223] discuss
[0224] The clinical manifestations of uniparental disomy (UPD) are highly heterogeneous, with effects ranging from asymptomatic to classic autosomal recessive (AR) or syndromic imprinting disorders, depending on the parental origin and the chromosome or segment affected. While most chromosomal UPDs do not cause clinical abnormalities, UPDs involving chromosomes 6, 7, 11, 14, 15, and 20 (known to be associated with imprinting disorders) can result in severe adverse outcomes in offspring.
[0225] Therefore, early detection and molecular genetic analysis are crucial for the accurate diagnosis and clinical management of UPD-related diseases. This study developed and validated a novel noninvasive prenatal screening method (NIPS-UPD) that integrates chromosome dose assessment with SNP allele fraction analysis to detect UPD in cell-free fetal DNA (cfDNA). The results confirmed the high sensitivity of this method in detecting UPD in fetuses, providing a potential screening tool for UPD-related diseases.
[0226] The core advantage of NIPS-UPD lies in its simultaneous analysis of chromosome dosage and allele fraction, significantly improving the accuracy of UPD testing compared to existing non-invasive methods. Traditional techniques (such as STR markers and SNP chip analysis) typically require parental samples for verification. However, NIPS-UPD uses high-depth sequencing data and a hidden Markov model (HMM) framework to infer UPD regions directly from cfDNA without parental samples. Furthermore, the introduction of quality control steps such as maternal copy number variation (CNV) testing further enhances the reliability of the results. This method allows doctors and pregnant women to quickly obtain fetal genetic information, avoiding the anxiety, infection, and miscarriage risks associated with invasive testing. With a testing cycle of only two weeks, it facilitates the timely development of prenatal and postpartum management plans and optimizes clinical outcomes.
[0227] Clinical validation based on retrospective maternal plasma samples further supports the potential application of NIPS-UPD in prenatal diagnosis. Of the four cases analyzed, four were accurately identified as whole-chromosomal or segmental UPD positive, and all were confirmed as true positive by fetal genome sequencing.
[0228] In addition, the results of NIPS-UPD were validated by methylation-specific MLPA (MS-MLPA) analysis: 4 cases with known pathogenic UPD regions were tested by MS-MLPA ( Figure 5 ), diagnosed with imprinting disorders such as Silver-Russell syndrome (SRS), Kagami-Ogata syndrome (KOS) and Prader-Willi syndrome (PWS), highlighting the important value of incorporating NIPS-UPD into routine prenatal screening, especially for pregnancies at high risk of imprinting disorders.
[0229] This invention is funded by the following projects: Discovery of Pathogenic Genes for Neurodevelopmental Abnormalities in Children (2020YFA0804001); National Key R&D Program (2023YFC2705600); Capital Clinical Characteristic Diagnosis and Treatment Technology Research and Translational Application Project (Z221100007422012); Beijing Municipal Hospital Management Center "Yanfan" Project (ZLRK202329).
[0230] All documents mentioned in this application are incorporated herein by reference, just as if each document were incorporated herein by reference individually. It should also be understood that after reading the above teachings of the present invention, those skilled in the art may make various changes or modifications to the present invention, and that such equivalents also fall within the scope of the claims appended hereto.
Claims
1. A detection device for determining fetal uniparental disomy (UPD), characterized in that: The device comprises: (a) a data input module configured to input nucleic acid data of a sample to be analyzed, wherein the nucleic acid data is sequencing data of cfDNA in the peripheral blood of a pregnant woman, and the sequencing data includes sequencing data of maternal free deoxyribonucleic acid and sequencing data of fetal free deoxyribonucleic acid; (b) a quality control module configured to perform quality control on the nucleic acid data to obtain nucleic acid data that meets quality control conditions; the quality control includes: sequencing depth, fetal fraction calculation, maternal CNV and polymorphism information site screening; (c) an analysis and calculation module configured to analyze the nucleic acid data after quality control, the analysis including fetal copy number analysis and allele score analysis; wherein, in the copy number analysis, if a fetal CNV is detected, the fetal CNV is determined; if no fetal CNV is detected, the allele score analysis is performed; in the allele score analysis, for the fetus, if the maternal specific allele is missing at a SNP site that is homozygous for the mother, it indicates that the fetus has a high risk of paternal UPD; if the paternal specific allele is missing, it indicates that the fetus has a high risk of maternal UPD; (d) an output module configured to output the detection result of the fetal UPD; Wherein, in the analysis and calculation module, the allele fraction analysis includes: (c1) calculating the fetal fraction FF, (c2) determining the fetal genotype, and (c3) detecting uniparental disomy of the fetus based on a hidden Markov model; In (c1), when the maternal genotype is homozygous AA or BB and the fetal genotype is heterozygous AB at the SNP site, the fetal fraction FF is calculated by the following steps: (1) For a certain SNP site i, let N be the total number of sequencing reads for alleles A and B, NA i is the read length of allele A, NB i is the number of reads of allele B; then the fetal fraction FF of this site AAi or FF BBi : (1), (2), (3); (2) Take the median of the FF values of all the SNP sites and calculate it as FF AA and FF BB , the calculation formula of the fetal fraction of the sample is as follows: (4); In (c2), include: (i) determining the B allele frequency afe of the candidate fetal genotype g based on the fetal score and the maternal genotype; (ii) For each SNP site i in the fetus, construct the alternative allele alt of the site using the β-binomial distribution i The likelihood function of is shown in formula (5): (5); Where n represents the total read count of site i, α and β are the parameters of the β-binomial distribution; (iii) Using the B allele frequency afe of the candidate fetal genotype g determined in (i), calculate the parameters of the β-binomial distribution, as shown in equations (6) and (7): (6), (7); And, assign a weight W to each candidate fetal genotype g g ; (iv) Calculate the raw score of the candidate fetal genotype: For each SNP i of the fetus, the raw score S(i) is calculated as shown in formula (8): (8) Where B() is the β function; is the binomial coefficient, i.e. x in (ii); W g is the weight corresponding to the candidate fetal genotype g; alt i is the read count of the alternative allele; n is the total read count of site i; the values of α and β are calculated by equations (6) and (7) in (iii); (v) determining the fetal genotype: selecting the candidate fetal genotype with the highest raw score as the fetal genotype at the current SNP site; In (c3), include: (c31) Obtaining an observation sequence V: For each SNP in the query region or the interval where UPD needs to be detected, if the maternal genotype is homozygous, generate an observation sequence V consisting of the observed values based on the maternal genotype and the fetal genotype obtained in (c2) and comparing them with the following table; (c32) Deriving a hidden state sequence X using a hidden Markov model HMM based on the observation sequence V to obtain a UPD detection result of the fetus; the hidden state is selected from the following group: diploid D, maternal uniparental disomy UPDM, paternal stage I uniparental disomy UPDPI, and paternal stage II uniparental disomy UPDPII.
2. The detection device according to claim 1, wherein The quality control conditions include: sequencing depth greater than a preset threshold A, fetal fraction greater than a preset threshold B, no maternal CNV detected, and the number of polymorphic information sites greater than a preset threshold C.
3. The detection device according to claim 1, wherein The copy number analysis includes: probe design.
4. The detection device according to claim 1, wherein The W g The values are selected from the following table: ; Among them, the left side of "-" is the maternal genotype, and the right side is the fetal genotype. D is normal diploid, UPDM is maternal uniparental disomy, UPDPI is paternal stage I uniparental disomy, and UPDPII is paternal stage II uniparental disomy.
5. The detection device according to claim 1, wherein The HMM is defined by a parameter set λ=(A; B; π), where A is the state transition probability matrix, B is the observation probability matrix, and π is the initial probability distribution; In addition, the state transition probability matrix is shown in the following table: ; The observation probability matrix is shown in the following table: ; The initial probability distribution is shown in the following table: ; Among them, D is normal diploid, UPDM is maternal uniparental disomy, UPDPI is paternal stage I uniparental disomy, and UPDPII is paternal stage II uniparental disomy.
6. The detection device according to claim 1, wherein In (c32), it also includes: calculating the proportion of each state in the final hidden state sequence X, thereby obtaining the UPD detection result of the fetus; Among them, when the proportion of D is greater than the preset threshold, the detection result of the fetal UPD is: D; When the proportion of D is less than the preset threshold, the following judgment is made: if the sum of the proportions of UPDPI and UPDPII is greater than the proportion of UPDM, UPDM is corrected to UPDPI or UPDPII, and the detection result of the fetal UPD is: paternal UPD; if the sum of the proportions of UPDPI and UPDPII is less than the proportion of UPDM, UPDPI or UPDPII is corrected to UPDM, and the detection result of the fetal UPD is: maternal UPD.
7. The detection device according to claim 6, characterized in that In (c32), if the fetal UPD test result is paternal UPD, the UPD subtype is further determined by the following method: The proportions of UPDPI and UPDPII were compared, and the state with the higher proportion was designated as the final UPD type for that region.
Citation Information
Patent Citations
Family-based low-depth sequencing method for detecting chromosome monoparental disomes
CN113593644A
Method for non-invasively evaluating zygotic property and fetal fraction of double-fetal pregnant fetuses
CN119230101A