A method for detecting whole genome ROH based on ultra-low depth sequencing

By detecting whole-genome ROH using ultra-low depth sequencing and calculating the heterozygosity index HetRR and Z value using PGT-A raw data, the problems of high detection complexity and cost in existing technologies are solved, achieving efficient and economical whole-genome ROH screening and improving the accuracy and safety of embryo screening.

CN122135778APending Publication Date: 2026-06-02GUANGZHOU PHIL MEDICAL TESTING CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU PHIL MEDICAL TESTING CO LTD
Filing Date
2026-04-28
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies require additional experimental steps or sequencing costs when detecting whole-genome ROH before embryo implantation, leading to operational complexity and economic burden. Furthermore, the incidence of whole-genome ROH is low, making routine screening neither economical nor efficient.

Method used

Using ultra-low depth sequencing, high-frequency SNP sites in the population are extracted from the raw PGT-A sequencing data, a sample SNP site set is constructed, the heterozygosity index HetRR and its standard score Z value are calculated, and a threshold is set to judge the whole genome ROH detection results without the need for additional experimental operations.

Benefits of technology

The process of simultaneously performing whole-genome ROH detection in the routine PGT-A procedure reduces costs, improves screening accuracy and safety, avoids mistransfer of high-risk blastocysts, and has strong scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135778A_ABST
    Figure CN122135778A_ABST
Patent Text Reader

Abstract

This invention provides a method for detecting genome-wide ROH based on ultra-low depth sequencing. The method first acquires the raw PGT-A sequencing data of the sample, extracts high-frequency SNP loci from the population, and screens out loci on autosomes with sequencing depths of 2-4X to construct a locus set. Then, the heterozygosity index (HetRR) and its standard score (Z) are calculated. Based on thresholds, the results are judged as follows: HetRR < 0.1 and Z ≤ -3.0 indicates genome-wide ROH; 0.1 ≤ HetRR < 0.5 and Z ≤ -3.0 indicates ROH chimerism; HetRR ≥ 1.05 and Z ≥ 3.0 suggests triploidy or DNA contamination; the rest are normal. This method can utilize sequencing data as low as 0.01X and can be simultaneously detected in the routine PGT-A process without additional operations or costs, improving the accuracy of embryo screening and possessing high clinical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of preimplantation genetic testing technology in assisted reproduction, specifically a method for detecting whole-genome ROH based on ultra-low depth sequencing. Background Technology

[0002] Chromosomal abnormalities are a significant cause of implantation failure and pregnancy loss. To improve pregnancy success rates, preimplantation genetic testing for aneuploidy (PGT-A) is widely used in in vitro fertilization (IVF) to screen for and avoid transferring aneuploid embryos, aiming to increase live birth rates. Currently, the mainstream clinical PGT-A methods are mainly based on next-generation sequencing (NGS) platforms, relying on whole-genome amplification and chromosome copy number comparison analysis to accurately identify whole-chromosome or fragmentary aneuploidy. However, genome-wide chromosomal abnormalities, such as triploidy and genome-wide homozygous regions (ROH), can lead to hydatidiform mole, embryonic lethality, and developmental disorders, but are not included in routine PGT-A screening.

[0003] A homozygous region (ROH), also known as a loss of heterozygosity, refers to a region on a chromosome where a contiguous sequence of alleles is missing a heterozygous allele. When the entire chromosome set is involved, it is called a genome-wide homozygous region (genome-wide ROH). When a detected ROH is confirmed to be inherited from only one parent, it is called uniparental disomy, which is often caused by the nondisjunction of sister chromatids during meiosis.

[0004] Several existing technologies exist for detecting ROH. For example, patent CN120099172A proposes a method including 178 pairs of STR primers and a matching kit for identifying uniparental diploids at the chromosome and whole-genome levels; patent CN113337600A introduces a low-depth sequencing technology based on enzyme digestion to detect ROH; and patent CN115798580A demonstrates an integrated genome analysis strategy combining genotype imputation and low-depth sequencing. However, the former two require additional experimental steps beyond the standard PGT-A procedure, significantly increasing operational complexity and detection costs; the latter, while requiring no additional experiments, demands approximately 10 times more sequencing data, also increasing the economic burden. It is noteworthy that the incidence of whole-genome ROH in human blastocysts is extremely low (approximately 0.06%), therefore, dedicated whole-genome ROH screening before embryo implantation is generally considered unnecessary. However, considering that whole-genome ROH blastocysts often present as diploid, there is a potential risk of implantation complications.

[0005] Therefore, there is an urgent need in this field for a technical solution that can simultaneously, economically, and efficiently perform whole-genome ROH detection during routine PGT-A procedures, without requiring additional experimental steps or significant detection costs. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a method for detecting whole-genome ROH based on ultra-low depth sequencing, which solves the problems of existing ROH detection requiring additional experiments or increased sequencing volume and high cost.

[0007] To address the aforementioned technical problems, embodiments of the present invention provide the following technical solution: a method for detecting whole-genome ROH based on ultra-low depth sequencing, comprising the following steps: S1. Obtain the raw PGT-A sequencing data of the sample to be tested; S2. Extract high-frequency single nucleotide polymorphism (SNP) sites from the original sequencing data, and screen out SNP sites located on autosomes with a sequencing depth in the range of 2-4X to construct a sample SNP site set, where X is the sequencing depth unit. S3. Based on the SNP locus set, calculate the heterozygosity index HetRR and its standard score Z value of the sample to be tested, and determine the whole genome ROH detection result of the sample to be tested by setting the threshold of HetRR and Z value.

[0008] Furthermore, in step S2, the method for obtaining the sample SNP site set includes: S21. Raw Sequencing Data Quality Control and Alignment: The raw sequencing data is subjected to quality control to remove adapter sequences and low-quality bases, obtaining purified sequence data; the purified sequence data is aligned to a specified version of the human reference genome to generate a sequence alignment file; the sequence alignment file is processed for format conversion, sorting, and removal of PCR repetitive sequences to generate a standardized alignment file for subsequent analysis; simultaneously, quality control indicators, including alignment rate, number of unique aligned reads, and average sequencing depth, are statistically analyzed based on the processing. S22. Extracting high-frequency SNP loci from the population: Obtain human SNP loci and their allele frequency information from the publicly available dbSNP database; screen for high-frequency SNP loci in the population with a minor allele frequency (MAF) between 0.2 and 0.8 and a genomic position interval of at least 50 bp; the genomic coordinate information of the selected high-frequency SNP loci is converted into a standard BED format file for subsequent SNP genotype identification and analysis. S23. Constructing a sample SNP locus set: Based on the standardized alignment file described in step S21, within the genomic coordinate range of the SNP locus list file described in step S22, retrieve the locus genotype information; from the retrieved genotype information, screen out SNP loci located on autosomes with a sequencing depth of 2-4X; perform data cleaning on the screened loci, remove invalid reads containing insertions, deletions, or unknown bases, and retain valid reads containing only definite base information to construct a sample SNP locus set for subsequent analysis.

[0009] Furthermore, in step S3 above, the method for calculating the HetRR heterozygosity index of the sample to be tested based on the SNP locus set includes: S31. Based on the sample SNP site set, read the sequencing depth, reference base number and variant base number of each site, determine whether the site is heterozygous, and count the number x of heterozygous SNP sites. S32. Based on the statistical results, calculate the heterozygosity index of the sample to be tested using Equations I to III: Formula I; Formula II; Formula III; Where HetRR is the heterozygosity index of the sample to be tested, HetRRe is the expected value of the heterozygosity index, HetRRo is the observed heterozygosity index of the sample to be tested, n is the total number of high-frequency SNP sites in the population of the SNP site set, x is the number of heterozygous sites in the SNP site set, pi and qi are the population frequencies corresponding to the reference base and the variant base of the high-frequency SNP site of the i-th population, respectively, pi+qi=1, 1≤i≤n.

[0010] Furthermore, in step S3 above, the method for calculating the standard score Z value of the heterozygosity index HetRR of the sample to be tested is as follows: S33. Based on a reference dataset consisting of negative samples confirmed to be free of whole-genome ROH and exogenous DNA contamination, calculate the mean m and standard deviation SD of the heterozygosity index of the dataset. S34. Based on the mean m and standard deviation SD, calculate the standardized score Z_score of the heterozygosity index of the test sample relative to the reference dataset: Formula IV; Furthermore, in step S3 above, the criteria for judging the whole-genome ROH detection results of the sample to be tested by setting thresholds for HetRR and Z values ​​include: when HetRR < 0.1 and Z ≤ -3.0, it is judged as whole-genome ROH; when 0.1 ≤ HetRR < 0.5 and Z ≤ -3.0, it is judged as whole-genome ROH chimerism; when HetRR ≥ 1.05 and Z ≥ 3.0, it suggests the possible presence of triploidy or DNA contamination; all other cases are judged as normal.

[0011] Furthermore, the average depth of the raw PGT-A sequencing data of the sample to be tested was as low as 0.01X.

[0012] The beneficial effects of the above-described technical solution of the present invention are as follows: This invention enables simultaneous, economical, and efficient GW-ROH detection within the conventional PGT-A process, without additional experimental or sequencing costs. It utilizes existing ultra-low depth NGS data (≥0.01X) to simultaneously screen for aneuploidy and GW-ROH, improving the accuracy and safety of embryo screening and preventing the misimplantation of seemingly chromosomally normal but actually high-risk blastocysts. Furthermore, its seamless integration with existing clinical workflows makes it highly scalable, providing a truly economical, efficient, and reliable solution for the co-detection of such rare but serious genetic abnormalities in routine embryo screening. Attached Figure Description

[0013] Figure 1 This is a flowchart of the method for detecting whole-genome ROH based on ultra-low depth sequencing in Example 1; Figure 2 This is an overview of the ROH results of the test sample using the present invention in a specific implementation. Detailed Implementation

[0014] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0015] Example 1: A method for detecting whole-genome ROH based on ultra-low depth sequencing This embodiment details how to utilize ultra-low depth whole-genome sequencing data generated by the conventional PGT-A workflow to simultaneously detect homozygous regions (GW-ROH) in embryonic samples without adding any extra experimental procedures. A flowchart of the method is shown below. Figure 1 Specifically, it includes the following steps: S1. Obtain the raw PGT-A sequencing data of the sample to be tested.

[0016] Embryonic trophoblast cell samples that have undergone biopsy and routine PGT-A testing were selected, and their corresponding raw single-end or paired-end sequencing data were obtained (typically with a sequencing depth of 0.01X to 0.1X). ​​The raw sequencing data was stored in FASTQ format and contained the base sequence of each read and its corresponding quality score.

[0017] S2. Extract high-frequency single nucleotide polymorphism (SNP) sites from the original sequencing data, and screen out SNP sites located on autosomes with a sequencing depth in the range of 2-4X to construct a sample SNP site set.

[0018] S21. Raw Sequencing Data Quality Control and Alignment. The raw sequencing data undergoes quality control. Trimmomatic software is used to remove adapter sequences and low-quality bases to obtain purified sequence data. The purified sequence data is then aligned to a specified version of the human reference genome (e.g., hg19 or hg38) using BWA MEM software. The alignment results are output in SAM format. Samtools software is used to convert the SAM file to BAM format. During this process, key quality control indicators are simultaneously calculated, including alignment rate (MapRate), total reads aligned to the reference genome (TotalReads), unique reads aligned to the reference genome (UniqueReads), duplicated reads, and mean sequencing depth (MeanDepth), to ensure data reliability and provide a quality basis for subsequent GW-ROH analysis.

[0019] S22. Extracting high-frequency SNP loci from the population. To avoid misidentification of genotypes due to rare variants or low-frequency loci, this invention specifically limits the use of "high-frequency SNPs from the population" as the genetic markers for ROH detection. Specifically, human SNP loci and their allele frequency information are obtained from the publicly available dbSNP database (https: / / ftp.ncbi.nih.gov / snp / organisms / human_9606 / VCF / All_20180418.vcf.gz); high-frequency SNP loci from the population with a minor allele frequency (MAF) between 0.2 and 0.8 and a genomic location interval of at least 50 bp are screened; the genomic coordinate information of the screened high-frequency SNP loci is converted into a standard BED format file for subsequent SNP genotype identification (Call SNP) analysis.

[0020] S23. Constructing a sample SNP locus set. Based on the normalized alignment file (BAM) generated in step S21, within the genomic coordinates of the BED format high-frequency SNP locus list obtained in S22, use a genotype identification tool (such as Samtoolsmpileup) to accurately retrieve the genotype information of each SNP locus. Strict quality filtering parameters are set during the retrieval process, for example, a minimum base quality score of 30 and a minimum alignment quality score of 25. From the retrieved genotype information, further screen out SNP loci located on autosomes with sequencing depths between 2-4X. Perform data cleaning on the screened loci, removing invalid reads containing insertions, deletions, or unknown bases, and retaining valid reads containing only clearly defined base information to construct a sample SNP locus set for subsequent analysis.

[0021] S3. Based on the SNP locus set, calculate the heterozygosity index (HetRR) and its standard score (Z value) of the sample to be tested, and determine the GW-ROH detection result of the sample to be tested by setting thresholds for HetRR and Z value.

[0022] The methods for obtaining the heterozygosity index (HetRR) of the test sample include: S31. Based on the sample SNP locus set constructed in step S23, sequentially read the total sequencing depth, the number of reads supporting reference genome bases, and the number of reads supporting variant bases for each SNP locus. Determine the heterozygous status of each SNP locus according to a preset genotype determination threshold (e.g., requiring the variant base ratio to be between 20% and 80%). Count the total number of heterozygous SNP loci in the test sample, denoted as x, and let the total number of all high-frequency SNP loci in the locus set be n. If the total number n of all high-frequency SNP loci in the test sample locus set is less than 800, the reliability of subsequent GW-ROH analysis results cannot be guaranteed, and the analysis results of such samples will be marked as "unevaluable" or "NA".

[0023] S32. Based on the statistical results, calculate the heterozygosity index of the sample to be tested using the following three formulas: ; ; ; HetRR is the heterozygosity index of the sample to be tested. e HetRR is the expected value of the heterozygosity index. o Let n be the heterozygosity index observed in the sample to be tested, n be the total number of high-frequency SNP loci in the population of the SNP locus cluster, x be the number of heterozygous loci in the SNP locus cluster, and p be the heterozygosity index observed in the sample to be tested. i and q ip represents the population frequencies of the reference base and the variant base at the high-frequency SNP site of the i-th population, respectively. i +q i =1, 1≤i≤n; The methods for obtaining the standard score (Z-score) of the heterozygosity index (HetRR) of the sample to be tested include: S33. Construct a negative reference dataset consisting of at least 300 PGT-A negative samples confirmed to be free of GW-ROH and exogenous DNA contamination. Following the methods described in S1-S3 above, calculate the heterozygosity index (HetRR) for each negative sample in the reference dataset, and calculate the mean m and standard deviation SD of the heterozygosity index (HetRR) for the reference dataset. S34. Based on the mean and standard deviation of the reference dataset, calculate the standardized score (Z_score) of the heterozygosity index of the test sample relative to the reference dataset: ; HetRR is the heterozygosity index of the test sample, m is the mean of the heterozygosity index of the reference dataset, and SD is the standard deviation of the heterozygosity index of the reference dataset.

[0024] Preferably, in steps S32 and S34 above, the criteria for judging the GW-ROH detection results of the sample to be tested by setting thresholds for HetRR and Z values ​​include: When HetRR < 0.1 and Z ≤ -3.0, it is determined to be GW-ROH; when 0.1 ≤ HetRR < 0.5 and Z ≤ -3.0, it is determined to be GW-ROH chimerism; when HetRR ≥ 1.05 and Z ≥ 3.0, it suggests possible triploidy or DNA contamination; all other cases are considered normal.

[0025] Conclusion: In this embodiment, the raw PGT-A sequencing data is not only used for routine chromosome copy number variation (CNV) analysis, but also serves as the input data source for subsequent ROH testing, achieving the technical effect of "sequencing once, dual use", significantly reducing testing costs and improving clinical efficiency.

[0026] Example 2: This example is based on the method for detecting whole-genome ROH using ultra-low depth sequencing described in Example 1. A set of PGT-A data from an assisted reproductive center was analyzed for whole-genome ROH to verify the feasibility, accuracy and clinical applicability of the method in practical applications.

[0027] A dataset of 577 samples that had completed routine PGT-A testing was selected. The corresponding sequencing data were generated by the Illumina sequencing platform and consisted of paired-end ultra-low-depth whole-genome sequencing data with an average sequencing depth of approximately 0.021 ± 0.007.

[0028] Following steps S1 to S3 as described in Example 1, the raw FASTQ data of the above 577 samples were analyzed, and the results are shown in Table 1 and... Figure 2 : Table 1. ROH analysis results of 577 samples

[0029] “NA” indicates that the test result of this sample does not meet the quality control standards.

[0030] To verify this, we performed single nucleotide polymorphism (SNP) microarray analysis on the remaining biopsy cells from this GW-ROH sample and the associated parental peripheral blood samples. The SNP array results confirmed that the sample was homozygous across the entire genome (i.e., GW-ROH), and genotypic analysis confirmed that it was a uniparental isodisomy.

[0031] Conclusion: This embodiment demonstrates that, based on the method described in Example 1, even under ultra-low depth sequencing conditions (0.021±0.007), whole-genome ROH in PGT-A samples can be identified efficiently and accurately, combining cost advantages with clinical applicability, and is suitable for whole-genome ROH screening in large-scale assisted reproductive scenarios.

[0032] The test results include: identification data of the product structure and effect data corresponding to the advantages in Part 2. Please provide attached figures as graphic evidence.

[0033] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for detecting whole-genome ROH based on ultra-low depth sequencing, characterized in that, S1. Obtain the raw PGT-A sequencing data of the sample to be tested; S2. Extract high-frequency single nucleotide polymorphism (SNP) sites from the original sequencing data, and screen out SNP sites located on autosomes with a sequencing depth in the range of 2-4X to construct a sample SNP site set, where X is the sequencing depth unit. S3. Based on the SNP locus set, calculate the heterozygosity index HetRR and its standard score Z value of the sample to be tested, and determine the whole genome ROH detection result of the sample to be tested by setting the threshold of HetRR and Z value.

2. The method for detecting whole-genome ROH based on ultra-low depth sequencing according to claim 1, characterized in that, In step S2, the method for obtaining the sample SNP site set includes: S21. Raw Sequencing Data Quality Control and Alignment: The raw sequencing data is subjected to quality control to remove adapter sequences and low-quality bases, obtaining purified sequence data; the purified sequence data is aligned to a specified version of the human reference genome to generate a sequence alignment file; the sequence alignment file is processed for format conversion, sorting, and removal of PCR repetitive sequences to generate a standardized alignment file for subsequent analysis; simultaneously, quality control indicators, including alignment rate, number of unique aligned reads, and average sequencing depth, are statistically analyzed based on the processing. S22. Extracting high-frequency SNP loci from the population: Obtain human SNP loci and their allele frequency information from the publicly available dbSNP database; screen for high-frequency SNP loci in the population with a minor allele frequency (MAF) between 0.2 and 0.8 and a genomic position interval of at least 50 bp; the genomic coordinate information of the selected high-frequency SNP loci is converted into a standard BED format file for subsequent SNP genotype identification and analysis. S23. Constructing a sample SNP locus set: Based on the standardized alignment file described in step S21, within the genomic coordinate range of the SNP locus list file described in step S22, retrieve the locus genotype information; from the retrieved genotype information, screen out SNP loci located on autosomes with a sequencing depth of 2-4X; perform data cleaning on the screened loci, remove invalid reads containing insertions, deletions, or unknown bases, retain valid reads containing only definite base information, and construct a sample SNP locus set for subsequent analysis.

3. The method for detecting whole-genome ROH based on ultra-low depth sequencing according to claim 1, characterized in that, In step S3 above, the method for calculating the HetRR heterozygosity index of the sample to be tested based on the SNP locus set includes: S31. Based on the sample SNP site set, read the sequencing depth, reference base number and variant base number of each site, determine whether the site is heterozygous, and count the number x of heterozygous SNP sites. S32. Based on the statistical results, calculate the heterozygosity index of the sample to be tested using Equations I to III: Formula I; Formula II; Formula III; Where HetRR is the heterozygosity index of the sample to be tested, HetRRe is the expected value of the heterozygosity index, HetRRo is the observed heterozygosity index of the sample to be tested, n is the total number of high-frequency SNP sites in the population of the SNP site set, x is the number of heterozygous sites in the SNP site set, pi and qi are the population frequencies corresponding to the reference base and the variant base of the high-frequency SNP site of the i-th population, respectively, pi+qi=1, 1≤i≤n.

4. The method for detecting whole-genome ROH based on ultra-low depth sequencing according to claim 1, characterized in that, In step S3 above, the method for calculating the standard score Z value of the heterozygosity index HetRR of the sample to be tested is as follows: S33. Based on a reference dataset consisting of negative samples confirmed to be free of whole-genome ROH and exogenous DNA contamination, calculate the mean m and standard deviation SD of the heterozygosity index of the dataset. S34. Based on the mean m and standard deviation SD, calculate the standardized score Z_score of the heterozygosity index of the test sample relative to the reference dataset: Formula IV.

5. The method for detecting whole-genome ROH based on ultra-low depth sequencing according to claim 1, characterized in that, In step S3 above, the criteria for judging the whole-genome ROH detection results of the sample to be tested by setting the thresholds for HetRR and Z values ​​include: when HetRR < 0.1 and Z ≤ -3.0, it is judged as whole-genome ROH; when 0.1 ≤ HetRR < 0.5 and Z ≤ -3.0, it is judged as whole-genome ROH chimerism; when HetRR ≥ 1.05 and Z ≥ 3.0, it suggests that triploidy or DNA contamination may exist; all other cases are judged as normal.

6. The method for detecting whole-genome ROH based on ultra-low depth sequencing according to claim 1, characterized in that, The average depth of the raw PGT-A sequencing data of the test samples was as low as 0.01X.