A method, system, product, and apparatus for preimplantation genetic diagnosis
By using single-molecule long-read sequencing technology on a single platform to obtain sequencing data of embryos and parents, the genetic origin of haplotypes in embryos can be determined, solving the problems of high cost and insufficient information utilization in existing technologies, and realizing efficient, accurate and integrated preimplantation genetic testing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2026-01-13
- Publication Date
- 2026-04-17
AI Technical Summary
Existing single-molecule long-read sequencing technologies are costly in preimplantation genetic testing and cannot fully utilize genomic information or effectively utilize phase information between embryonic SNP sites, resulting in inaccurate test results.
Preimplantation genetic testing (PGD) is performed on a single platform using single-molecule long-read sequencing technology. By acquiring sequencing data of the embryos to be tested and their parents, parental haplotype information is determined, and similarity is calculated to determine the haplotype genetic origin of the embryo samples. This is combined with chromosomal structural rearrangement and single-gene genetic disease detection.
Improving detection accuracy and reducing costs under low-depth sequencing conditions enables integrated preimplantation genetic testing, allowing for rapid, simple, and efficient determination of whether an embryo has genetic defects.
Smart Images

Figure CN121506262B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics detection technology, specifically relating to a method, system, product, and equipment for preimplantation genetic testing of embryos. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Preimplantation genetic testing (PGT) is an important measure in assisted reproductive technology to avoid the occurrence of hereditary diseases.
[0004] In recent years, with the development of single-molecule long-read sequencing technology, clinical practice has gradually begun to use single-molecule long-read sequencing for preimplantation genetic testing (PGT for monogenic defect, PGT-M) and preimplantation genetic testing (PGT for structural rearrangement, PGT-SR). However, given the high cost of single-molecule long-read sequencing at present, this method usually only performs single-molecule long-read sequencing on both parents or one affected parent, and then performs pathogenic mutation verification and single nucleotide polymorphism (SNP) microarray / next-generation sequencing on the embryo, performing linkage analysis through upstream and downstream SNPs of the variant site. However, this approach still requires high SNP microarray / next-generation sequencing data depth in the embryo sample, and different sequencing strategies need to be selected for parental and embryo samples, placing high demands on clinicians. In addition, this approach also has the problem of not being able to fully utilize genomic information: the number of available single nucleotide polymorphism sites is relatively small. This method can only construct haplotypes of embryos based on SNP sites where one parent is homozygous and the other is heterozygous. It cannot fully utilize single nucleotide polymorphism (SNP) sites on the genome, and it fails when the number of available upstream and downstream SNP sites is insufficient. Furthermore, it utilizes limited site information. This method uses SNP microarray / next-generation sequencing technology to detect embryo samples, naturally losing the phase information between embryonic SNP sites that uniquely characterizes the genetic origin of haplotypes; in other words, it cannot utilize the phase information between embryonic SNP sites. Summary of the Invention
[0005] To address the aforementioned problems, this invention proposes a method, system, product, and device for preimplantation genetic testing. This invention enables comprehensive, integrated preimplantation genetic testing on a single platform, even with limited family samples, thereby further enhancing the applicability and clinical value of preimplantation genetic testing technology.
[0006] According to some embodiments, the present invention adopts the following technical solution:
[0007] A method for preimplantation genetic testing includes the following steps:
[0008] Obtain sequencing sequence data of the embryo sample to be tested and its paternal and maternal samples, wherein the sequencing sequence data includes single-molecule long-read sequencing sequences of the corresponding samples obtained using single-molecule long-read sequencing technology;
[0009] Based on the sequencing sequence data of the paternal and maternal samples, haplotype information of each parent was determined;
[0010] The sequencing data of each parental haplotype and the embryo samples to be tested are filtered to determine the common single nucleotide polymorphism sites and corresponding base types between the sequencing sequences of the embryo samples to be tested and the corresponding parental haplotypes. The ratio of the number of common single nucleotide polymorphism sites with the same base type to the total number of common single nucleotide polymorphism sites is calculated to obtain the similarity between the sequencing sequences of the embryo samples to be tested and the corresponding parental haplotypes. Based on the similarity, the parental haplotype source of each sequencing sequence of the embryo samples to be tested in the target region is determined.
[0011] Based on the identified parental haplotype origin, the haplotype genetic origin of the embryo sample to be tested within the target region is determined.
[0012] As an alternative implementation method, the process of determining haplotype information of each parent based on the sequencing sequence data of the paternal sample and the sequencing sequence data of the maternal sample includes: determining the single nucleotide polymorphism site information in the genomes of the paternal and maternal samples based on the alignment results of the sequencing sequences of the paternal and maternal samples with the reference genome sequence.
[0013] Based on the single nucleotide polymorphism site information, the haplotypes of each parent were constructed, and the haplotype information of each parent was determined;
[0014] The parent haplotype information includes the single nucleotide polymorphism sites covered by the corresponding parent haplotype and the base types of the corresponding parent haplotypes corresponding to the single nucleotide polymorphism sites.
[0015] As an alternative implementation method, the target area is determined based on at least one of the following information:
[0016] Coordinates of gene mutation sites obtained in advance or based on sequencing sequence data from paternal or maternal samples;
[0017] The coordinates of structural variation breakpoints obtained in advance or based on sequencing sequence data from paternal or maternal samples; and the pre-specified target gene regions.
[0018] As an optional implementation, the process of filtering the haplotype information of each parent and the sequencing sequence data of the embryo samples to be tested includes: filtering the sequencing sequences of the embryo samples to be tested that fall into the target region according to preset filtering conditions, wherein the preset filtering conditions include:
[0019] Remove sequencing sequences whose alignment quality scores do not meet the preset alignment quality score threshold;
[0020] Remove bases from the sequencing sequence whose base mass fraction does not meet the preset base mass fraction threshold;
[0021] Remove sequencing sequences that do not meet the preset coverage threshold for the number of base sites that satisfy the preset base mass fraction threshold in the single nucleotide polymorphism sites covered by the target region.
[0022] As an alternative implementation method, the process of calculating the ratio of the number of common single nucleotide polymorphism sites with the same base type to the number of common single nucleotide polymorphism sites, and obtaining the similarity between the sequencing sequence of the embryo sample to be tested and the corresponding parental haplotype, includes:
[0023] Based on the alignment results of the sequencing sequence of the embryo sample to be tested and the reference genome sequence, the sequencing sequence of the embryo sample to be tested that falls into the target region is extracted.
[0024] For each sequencing sequence of the embryo sample to be tested within the target region, and for each parental haplotype, based on the corresponding sequencing sequence of the embryo sample to be tested and the corresponding parental haplotype information, the common single nucleotide polymorphism sites between the sequencing sequence of the embryo sample to be tested and the corresponding parental haplotype, and the base types at the common single nucleotide polymorphism sites between the sequencing sequence of the embryo sample to be tested and the corresponding parental haplotype are determined.
[0025] When the number of common single nucleotide polymorphism sites is not zero, for each of the common single nucleotide polymorphism sites, determine whether the sequencing sequence of the embryo sample to be tested and the corresponding parent haplotype have the same base type at the common single nucleotide polymorphism site, and count the number of common single nucleotide polymorphism sites with the same base type.
[0026] The ratio of the number of common single nucleotide polymorphism sites with the same base type to the number of common single nucleotide polymorphism sites is calculated to obtain the similarity between the sequencing sequence of the embryo sample to be tested and the haplotype of the parent.
[0027] When the number of shared single nucleotide polymorphism sites is zero, the similarity between the corresponding sequencing sequence of the embryo sample to be tested and the corresponding parent haplotype is zero.
[0028] As an alternative implementation, the process of determining the parental haplotype origin of each sequencing sequence of the embryo sample to be tested within the target region based on the similarity includes:
[0029] For each sequencing sequence of the embryo sample to be tested in the target area, based on the similarity between the sequencing sequence of the embryo sample to be tested and each parental haplotype, the parental haplotype with the highest similarity to the sequencing sequence of the embryo sample to be tested and its corresponding first similarity, and the parental haplotype with the second highest similarity to the sequencing sequence of the embryo sample to be tested and its corresponding second similarity are determined.
[0030] If the number of shared single nucleotide polymorphism (SNP) sites between the corresponding sequencing sequence of the embryo sample to be tested and the parental haplotype with the highest similarity in the target region meets a preset threshold for the number of shared SNP sites, the first similarity meets a preset threshold for similarity, and the difference between the first similarity and the second similarity meets a preset threshold for similarity difference, then the parental haplotype with the highest similarity is determined to be the source of the parental haplotype of the corresponding sequencing sequence.
[0031] As a further defined implementation, the process of determining the parental haplotype origin of each sequencing sequence of the embryo sample to be tested within the target region based on the similarity further includes:
[0032] For each sequencing sequence of the embryo sample to be tested within the target region, if the number of shared single nucleotide polymorphism sites between the corresponding sequencing sequence of the embryo sample and the parental haplotype with the highest similarity in the target region does not meet the preset threshold for the number of shared single nucleotide polymorphism sites, or if the first similarity does not meet the preset similarity threshold or the difference between the first similarity and the second similarity does not meet the preset similarity difference threshold, then the parental haplotype source will not be assigned to the corresponding sequencing sequence.
[0033] As an alternative implementation, the process of determining the haplotype genetic origin of the embryo sample to be tested in the target region based on the determined parental haplotype origin includes:
[0034] For each parental haplotype, count the number of sequencing sequences of the embryo samples to be tested that are assigned to that parental haplotype within the target region;
[0035] In response to the number of sequencing sequences of the embryo sample to be tested assigned to either the father or the mother that meets a preset support threshold, the parental haplotype that meets the preset support threshold is determined to be the haplotype genetic source of the embryo sample to be tested in the target region.
[0036] In response to the fact that the number of sequencing sequences of the embryo samples to be tested for each haplotype assigned to the father or mother does not meet the preset support threshold, it is determined that no haplotype originating from the parent is detected in the target region of the embryo sample to be tested.
[0037] As a further defined implementation, in response to the fact that the number of sequencing sequences of the embryo samples to be tested for each haplotype assigned to the father or mother does not meet the preset support threshold, and the corresponding parent carries a deletion variant in the target region, the haplotype of the parent carrying the deletion variant in the target region is determined to be the haplotype genetic source of the embryo sample to be tested in the target region.
[0038] As an alternative implementation, the method further includes the following steps: determining the single-gene genetic disease and / or chromosomal structural rearrangement carrier status of the embryo sample to be tested based on the parental haplotype information and the haplotype genetic origin of the embryo sample to be tested.
[0039] As a further defined implementation, the process of determining the carrier status of a single-gene genetic disease and / or chromosomal structural rearrangement in the embryo sample to be tested, based on parental haplotype information and the haplotype genetic origin of the embryo sample to be tested, includes:
[0040] Based on the parental haplotype information, determine the parental pathogenic chain haplotype carrying pathogenic variants and / or chromosomal structural rearrangements;
[0041] In response to the fact that the haplotype genetic origin of the embryo sample to be tested includes the parental pathogenic chain haplotype, it is determined that the embryo sample to be tested carries a single-gene genetic disease and / or chromosomal structural rearrangement;
[0042] In response to the fact that the haplotype of the embryo sample to be tested does not contain the parental pathogenic chain haplotype, it is determined that the embryo sample to be tested does not carry a single-gene genetic disease and / or chromosomal structural rearrangement.
[0043] As an alternative implementation, the method further includes the following step: determining the chromosomal euploidy of the embryo sample to be tested based on the sequencing sequence data of the embryo sample to be tested.
[0044] As a further defined implementation, the process of determining the chromosomal euploidy of the embryo sample to be tested based on the sequencing sequence data of the embryo sample to be tested includes:
[0045] Based on the preset window length, the reference genome is divided into multiple consecutive equal-length intervals;
[0046] Based on the sequencing sequence data of the embryo samples to be tested, the number of aligned sequencing sequences within each reference genomic region is counted;
[0047] Based on the number of aligned sequencing sequences within each reference genome region, the copy number prediction result for each reference genome region is determined;
[0048] Calculate the degree of difference in copy number prediction results between adjacent reference genome intervals, and merge adjacent intervals with a degree of difference below a preset merging threshold;
[0049] The chromosome euploidy test results of the embryo sample to be tested are determined based on the merged intervals.
[0050] A preimplantation genetic testing system includes:
[0051] The sequencing sequence data acquisition unit is used to acquire sequencing sequence data of the embryo sample to be tested and its paternal and maternal samples. The sequencing sequence data includes single-molecule long-read sequencing sequences of the corresponding samples obtained using single-molecule long-read sequencing technology.
[0052] The parental haplotype phasing unit is used to determine the haplotype information of each parent based on the sequencing sequence data of the paternal sample and the sequencing sequence data of the maternal sample.
[0053] The parental haplotype origin determination unit is used to filter the information of each parental haplotype and the sequencing sequence data of the embryo sample to be tested, determine the common single nucleotide polymorphism sites and corresponding base types between the sequencing sequence of the embryo sample to be tested and the corresponding parental haplotype, calculate the ratio of the number of common single nucleotide polymorphism sites with the same base type to the total number of common single nucleotide polymorphism sites, obtain the similarity between the sequencing sequence of the embryo sample to be tested and the corresponding parental haplotype, and determine the parental haplotype origin of each sequencing sequence of the embryo sample to be tested in the target region based on the similarity.
[0054] The embryo haplotype genetic origin determination unit is used to determine the haplotype genetic origin of the embryo sample to be tested within the target region based on the determined parental haplotype origin.
[0055] As an alternative implementation, it also includes:
[0056] A preimplantation single-gene genetic disease / chromosomal structural rearrangement detection unit is used to determine the parental pathogenic chain haplotype carrying pathogenic variants and / or chromosomal structural rearrangements based on the parental haplotype information.
[0057] In response to the fact that the haplotype genetic origin of the embryo sample to be tested includes the parental pathogenic chain haplotype, it is determined that the embryo sample to be tested carries a single-gene genetic disease and / or chromosomal structural rearrangement;
[0058] In response to the fact that the haplotype of the embryo sample to be tested does not contain the parental pathogenic chain haplotype, it is determined that the embryo sample to be tested does not carry a single-gene genetic disease and / or chromosomal structural rearrangement.
[0059] As an alternative implementation, it also includes:
[0060] The preimplantation aneuploidy detection unit is used to divide the reference genome into multiple consecutive equal-length intervals according to a preset window length;
[0061] Based on the sequencing sequence data of the embryo samples to be tested, the number of aligned sequencing sequences within each reference genomic region is counted;
[0062] Based on the number of aligned sequencing sequences within each reference genome region, the copy number prediction result for each reference genome region is determined;
[0063] Calculate the degree of difference in copy number prediction results between adjacent reference genome intervals, and merge adjacent intervals with a degree of difference below a preset merging threshold;
[0064] The chromosome euploidy test results of the embryo sample to be tested are determined based on the merged intervals.
[0065] A computer program product includes a computer program that, when executed by a processor, performs the steps in the above-described method.
[0066] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the steps in the method described above.
[0067] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0068] This invention fully leverages the advantages of long read length and high accuracy of single-molecule long-read sequencing technology to determine the haplotype genetic origin and carrier status of single-gene genetic diseases and / or chromosomal structural rearrangements in embryos with limited family samples and low embryo sequencing depth (e.g., 1×). Since a single sequence from single-molecule long-read sequencing can contain a sufficient number of SNP loci, and the sequence itself contains the phase relationships between SNP loci, compared to traditional preimplantation genetic testing methods based on microarrays / next-generation sequencing, single-molecule long-read sequencing not only increases the number of available loci but also provides another dimension of information, including the phase relationships between loci.
[0069] This invention utilizes the characteristics of single-molecule long-read sequencing, enabling the assignment of parental haplotype sources to each sequencing sequence using only low-depth sequencing data based on the SNP site information contained in each sequencing sequence, thereby significantly reducing detection costs and improving detection accuracy.
[0070] The method of this invention can perform integrated detection of preimplantation genetic disorders (PGT-A), single-gene genetic diseases (PGT-M), and chromosomal structural rearrangements (PGT-SR) on a single platform, providing a one-stop, rapid, simple, efficient, and accurate determination of whether there are genetic defects in the embryo to be implanted.
[0071] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0072] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0073] Figure 1 This is a schematic diagram of a system for implementing an integrated preimplantation genetic testing method based on single-molecule long-read sequencing, according to an embodiment of the present invention.
[0074] Figure 2 This is a flowchart of an integrated preimplantation genetic testing method based on single-molecule long-read sequencing according to an embodiment of the present invention.
[0075] Figure 3 This is a flowchart of a method for determining the chromosomal euploidy of an embryo sample according to an embodiment of the present invention;
[0076] Figure 4 This is a flowchart of a method for determining the haplotype genetic origin of an embryo sample to be tested, according to an embodiment of the present invention;
[0077] Figure 5 This is a flowchart of a method for determining the parental haplotype origin of the sequencing sequence of an embryo sample according to an embodiment of the present invention;
[0078] Figure 6 This is a block diagram of an electronic device according to an embodiment of the present invention;
[0079] Among them, 110 is the computing device, 111 is the sequencing sequence data acquisition unit, 112 is the preimplantation aneuploidy detection unit, 113 is the parental haplotype phasing unit, 114 is the embryo haplotype genetic origin determination unit, and 115 is the preimplantation single-gene genetic disease / chromosomal structural rearrangement detection unit.
[0080] 120. Network; 130. Sequencing equipment; 140. Server;
[0081] 1001, CPU; 1002, ROM; 1003, RAM; 1004, Bus; 1005, I / O Interface; 1006, Input Unit; 1007, Output Unit; 1008, Storage Unit; 1009, Communication Unit. Detailed Implementation
[0082] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0083] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0084] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0085] Where there is no conflict, the embodiments and features described in this application may be combined with each other.
[0086] This embodiment provides an integrated preimplantation genetic testing method based on single-molecule long-read sequencing. By utilizing the high accuracy and long read advantage of single-molecule long-read sequencing, integrated preimplantation genetic testing can be achieved.
[0087] Specifically, the method includes: acquiring sequencing sequence data of the embryo sample to be tested and its paternal and maternal samples, wherein the sequencing sequence data includes single-molecule long-read sequencing sequences of the sample obtained by single-molecule long-read sequencing technology; determining haplotype information of each parent based on the sequencing sequence data of the paternal sample and the sequencing sequence data of the maternal sample; calculating the similarity between each sequencing sequence of the embryo sample to be tested and each parental haplotype within a target region based on the parental haplotype information and the sequencing sequence data of the embryo sample to be tested, thereby determining the parental haplotype origin of each sequencing sequence of the embryo sample to be tested within the target region; and determining the haplotype genetic origin of the embryo sample to be tested within the target region based on the parental haplotype origin of the sequencing sequence of the embryo sample to be tested within the target region. This invention enables integrated detection of preimplantation genetic aneuploidy (PGT-A), single-gene genetic diseases (PGT-M), and chromosomal structural rearrangements (PGT-SR) on a single platform, providing a one-stop, rapid, simple, efficient, and accurate determination of whether an embryo has genetic defects.
[0088] Figure 1 A schematic diagram of an exemplary system for implementing an integrated preimplantation genetic testing method based on single-molecule long-read sequencing, according to an embodiment of the present invention, is shown. Figure 1 As shown, the system includes a computing device 110. In some embodiments, the system also includes a sequencing device 130, a server 140, and a network 120. In some embodiments, the computing device 110, the server 140, and the sequencing device 130 can interact with each other via wired or wireless means (e.g., network 120).
[0089] Regarding sequencing device 130, it is used, for example, to sequence a sample to be tested to generate a sequencing sequence of the sample; and to send the generated sequencing sequence data of the sample to computing device 110. In some embodiments, the sequencing sequence data of the sample to be tested may also be sent to computing device 110 by server 140. In some embodiments, sequencing device 130 may be a sequencing device based on single-molecule long-read sequencing technology. Sequencing device 130 is, for example, but not limited to, sequencers based on single-molecule real-time (SMRT) sequencing technology developed by Pacific Biosciences (PacBio) (such as RSI, RS II, and Sequel series sequencers), and sequencers based on protein nanopore sequencing technology developed by Oxford Nanopore Technologies (ONT) (such as MinION, GridION, and PromethlON series sequencers).
[0090] Regarding server 140, it is used, for example, to store or provide sequence information of a reference genome obtained from a public database. For example, it provides human reference genome sequence information, which is, for example, but not limited to, human reference genome GRCh37 or human reference genome GRCh38.
[0091] Regarding the computing device 110, it is used, for example, for integrated preimplantation genetic testing based on single-molecule long-read sequencing. Specifically, the computing device 110 is used to acquire single-molecule long-read sequencing data of the embryo sample to be tested, and to determine the chromosome euploidy of the embryo sample based on the acquired sequencing data. The computing device 110 is also used to acquire single-molecule long-read sequencing data of the paternal and maternal samples of the embryo sample to be tested, and to determine the haplotype information of each parent based on the acquired sequencing data. The computing device 110 is also used to calculate the similarity between each sequencing sequence of the embryo sample to be tested and each parental haplotype within a target region based on the determined parental haplotype information and the acquired sequencing data of the embryo sample to be tested, thereby determining the parental haplotype origin of each sequencing sequence of the embryo sample to be tested within the target region; and based on the parental haplotype origin of the sequencing sequence of the embryo sample to be tested within the target region, to determine the haplotype genetic origin of the embryo sample to be tested within the target region. The computing device 110 is also used to determine the single-gene genetic disease and / or chromosomal structural rearrangement carrier status of the embryo sample to be tested based on the parental haplotype information and the haplotype genetic origin of the embryo sample to be tested.
[0092] In some implementations, computing device 110 may have one or more processing units, including dedicated processing units such as GPUs, FPGAs, and ASICs, and general-purpose processing units such as CPUs. Computing device 110 may be an integration of multiple physical servers, an integration of multiple processing units, etc. Additionally, one or more virtual machines may run on each computing device 110. A specific structure of computing device 110 may be, for example, combined as follows: Figure 6 As shown. In some implementations, computing device 110 and server 140 can be integrated together or set up separately.
[0093] In some implementations, the computing device 110 includes, for example, a sequencing data acquisition unit 111, a preimplantation aneuploidy detection unit 112, a parental haplotype phasing unit 113, an embryo haplotype genetic origin determination unit 114, and a preimplantation single-gene genetic disease / chromosomal structural rearrangement detection unit 115. The sequencing data acquisition unit 111, the preimplantation aneuploidy detection unit 112, the parental haplotype phasing unit 113, the embryo haplotype genetic origin determination unit 114, and the preimplantation single-gene genetic disease / chromosomal structural rearrangement detection unit 115 can be configured on one or more computing devices 110.
[0094] Regarding the sequencing sequence data acquisition unit 111, it is used to acquire sequencing sequence data of the embryo sample to be tested and its paternal and maternal samples. The sequencing sequence data includes single-molecule long-read sequencing sequences of the sample obtained by single-molecule long-read sequencing technology.
[0095] Regarding the preimplantation aneuploidy detection unit 112, it is used to determine the chromosomal neuploidy of the embryo sample to be tested based on the sequencing sequence data of the embryo sample to be tested.
[0096] Regarding the parent haplotype phasing unit 113, it is used to determine the haplotype information of each parent based on the sequencing sequence data of the paternal sample and the sequencing sequence data of the maternal sample.
[0097] Regarding the embryo haplotype genetic origin determination unit 114, it is used to determine the haplotype genetic origin of the embryo sample to be tested in the target region based on the parental haplotype information and the sequencing sequence data of the embryo sample to be tested.
[0098] Regarding the preimplantation single-gene genetic disease / chromosomal structural rearrangement detection unit 115, it is used to determine the single-gene genetic disease and / or chromosomal structural rearrangement carrier status of the embryo sample to be tested based on the parental haplotype information and the haplotype genetic origin of the embryo sample to be tested.
[0099] The following will combine Figure 2 This invention describes an integrated preimplantation genetic testing method for embryos based on single-molecule long-read sequencing, according to embodiments of the present invention. Figure 2 A flowchart illustrating an integrated preimplantation genetic testing method for embryos based on single-molecule long-read sequencing, according to an embodiment of the present invention, is shown. It should be understood that the integrated preimplantation genetic testing method for embryos based on single-molecule long-read sequencing can, for example, be used in… Figure 6 It can be executed at the described device, or it can be performed at... Figure 1 The described computing device 110 performs the operation. It should be understood that the integrated preimplantation genetic testing method for embryos based on single-molecule long-read sequencing may also include additional actions not shown and / or omit the actions shown, and the scope of the invention is not limited in this respect.
[0100] In step 201, computing device 110 acquires sequencing sequence data of the embryo sample to be tested and its paternal and maternal samples.
[0101] Regarding the embryo sample to be tested, it may be an embryo biopsy sample obtained based on assisted reproductive technology, such as a blastomere biopsy sample or a trophoblast cell biopsy sample from the blastocyst stage.
[0102] The sequencing data of the embryo sample to be tested and its paternal and maternal samples, including, for example, single-molecule long-read sequencing sequences of the samples obtained by single-molecule long-read sequencing technology.
[0103] Single-molecule long-read sequencing technologies include, but are not limited to, single-molecule real-time (SMRT) sequencing technology developed by Pacific Biosciences (PacBio) and protein nanopore-based sequencing technology developed by Oxford Nanopore Technologies (ONT). The core of SMRT sequencing technology lies in its carrier, the SMRT Cell. Each SMRT Cell contains millions of zero-mode waveguide pores (ZMWs) on its surface, with a DNA polymerase and a template molecule to be tested fixed at the bottom of each ZMW. During the sequencing reaction, when a fluorescently labeled nucleotide substrate is incorporated into the nascent strand by the polymerase, the excitation light at the bottom of the pore excites the fluorescent label on the nucleotide substrate. The fluorescence signal is then recorded by a monitoring system, thus obtaining base information. Nanopore sequencing platforms, on the other hand, directly read sequence information by measuring the characteristic current changes caused by different bases in the DNA molecule passing through protein nanopores. This technology supports real-time analysis of ultra-long reads (e.g., >100 kb).
[0104] Compared to short-read sequencing (e.g., NGS), the single-molecule long-read sequencing technology used in this invention allows a single read to span a longer genomic distance (e.g., 10 kb-100 kb), thus directly covering multiple single nucleotide polymorphism (SNP) sites. Utilizing the linkage information carried on long-read sequences, highly continuous and accurate haplotypes can be directly constructed from low-depth (e.g., 1×) sequencing data without relying on complex statistical inferences. For PGT-SR detection, long reads can also traverse genomic regions rich in repetitive sequences or structurally complex regions, enabling more precise identification of breakpoints in structural variations.
[0105] In some implementations, the sequencing sequences of the embryo sample to be tested, as well as its paternal and maternal samples, may undergo preliminary quality control and filtering, including but not limited to processes such as low-quality read removal, sequencing error correction, and redundant sequence removal. Those skilled in the art should understand that the aforementioned quality control and filtering can be implemented using any suitable quality control and filtering tools, and their specific parameters can be flexibly set according to the sequencing platform, sample characteristics, and experimental objectives.
[0106] In some embodiments, the sequencing data of the embryo sample to be tested, its paternal sample, and its maternal sample may include alignment results obtained after sequence alignment with a reference genome sequence, respectively. The alignment results include single-molecule long-read sequencing sequences of the sample obtained via single-molecule long-read sequencing technology, and the positional information of each sequencing sequence aligned to the reference genome, which can be used for one or more downstream analyses. In some embodiments, the alignment results can be obtained using sequence alignment algorithms suitable for long-read sequencing data, such as, but not limited to, tools like BWA-MEM, Minimap2, or BLASR. Specifically, computing device 110 can align the sequencing sequences (e.g., in FASTQ format) with a reference genome sequence (e.g., a GRCh37 or GRCh38 version of the human reference genome) to obtain alignment results data (e.g., a BAM or SAM file). Through this alignment operation, the raw sequencing data (e.g., in FASTQ format) is mapped onto genomic coordinates, thereby generating alignment results data (e.g., a BAM or SAM format file) containing positional information to determine the positional information of each sequencing sequence aligned to the reference genome. Those skilled in the art should understand that the above comparison operation can be implemented using any suitable comparison tool and algorithm, and its specific parameters can be adjusted according to actual needs.
[0107] In step 202, the computing device 110 determines the chromosomal euploidy of the embryo sample to be tested based on the sequencing sequence data of the embryo sample to be tested.
[0108] A method for determining the chromosomal neuploidy of a test embryo sample includes, for example: a computing device 110 divides a reference genome into multiple consecutive equal-length intervals based on a preset window length; based on the sequencing sequence data of the test embryo sample, counts the number of aligned sequencing sequences within each reference genome interval; based on the number of aligned sequencing sequences within each reference genome interval, determines the copy number prediction result for each reference genome interval; calculates the degree of difference in the copy number prediction results of adjacent reference genome intervals; based on the degree of difference in the copy number prediction results of adjacent reference genome intervals, merges adjacent intervals with a difference degree lower than a preset merging threshold, and determines the chromosomal neuploidy detection result of the test embryo sample based on the merged intervals. The following will combine... Figure 3 The methods used to determine the chromosomal euploidy of the embryo sample to be tested are detailed in detail here and will not be repeated.
[0109] In step 203, the computing device 110 determines the haplotype information of each parent based on the sequencing sequence data of the paternal sample and the sequencing sequence data of the maternal sample. The construction of parental haplotypes provides a reference for tracing the genetic origin of haplotypes in subsequent embryo samples to be tested.
[0110] Specifically, methods for determining parental haplotype information include, for example, variation detection and haplotype phasing.
[0111] Based on the alignment results of the paternal and maternal sequencing sequences with a reference genome sequence, the computing device 110 performs variant detection to determine single nucleotide polymorphism (SNP) sites in the paternal and maternal genomes. During this process, the computing device 110 can use variant detection tools such as, but not limited to, GATK HaplotypeCaller, Samtools mpileup, and DeepVariant to analyze the alignment results and identify SNP sites and their corresponding genotypes in the parental genome. Those skilled in the art will understand that homozygous SNP sites typically do not provide useful information for haplotype phasing. Therefore, the computing device 110 can further identify heterozygous SNP sites. Considering sequencing errors, the computing device 110 can identify heterozygous SNP sites based on whether the maximum base frequency at the SNP site falls within a preset base frequency range. The maximum base frequency refers to the proportion of the number of sequencing sequences supporting the most frequent base at that site to the total effective sequencing depth of that site. This can usually be calculated by dividing the number of sequencing sequences that support the most frequently occurring bases by the total effective sequencing depth of that site.
[0112] The preset base frequency range is, for example, but not limited to, greater than or equal to a first preset base frequency threshold and less than or equal to a second preset base frequency threshold. The first preset base frequency threshold is, for example, 10%, and in some embodiments, it is, for example, 20%. The second preset base frequency threshold is, for example, 90%. In some embodiments, it is, for example, 80%. It should be understood that the preset base frequency range can be adjusted according to the experimental needs of those skilled in the art. Single nucleotide polymorphism sites within this preset base frequency range can be identified as heterozygous single nucleotide polymorphism sites for subsequent monomer phasing.
[0113] Most mammalian genomes are diploid, consisting of two sets of chromosome haplotypes inherited from the father and mother, respectively. Haplotype reconstruction aims to resolve the unique nucleotide sequence combinations of these two haplotypes. Since different haplotypes have different distinguishing significance in genetic linkage analysis and genetic origin determination, accurate parental haplotype reconstruction plays a crucial role in determining the genetic origin of embryonic haplotypes. Compared to short-read sequencing, single-molecule long-read sequencing can provide richer phase information from a single sequencing sequence, thereby improving the accuracy and continuity of haplotype phasing. Therefore, after obtaining the heterozygous single nucleotide polymorphism (SNP) site information, the computing device 110 uses the phase information carried in the long-read sequencing sequence, combined with the heterozygous SNP site information in the parental genome, to phase the parental haplotypes using haplotype phasing software such as, but not limited to, Whatshap, Margin, and hiphase, to construct each parental haplotype. The haplotype phasing algorithm utilizes information from heterozygous single nucleotide polymorphism (SNP) sites and the characteristic that long-read sequencing sequences simultaneously cover multiple SNP sites to infer the linkage relationships of these SNP sites on the same chromosome. For example, if a long-read sequencing sequence simultaneously covers sites A and B, and sites A and B are both heterozygous SNP sites, and a read has no mutation at site A (marked as 0) and a mutation at site B (marked as 1), then it is inferred that A=0 and B=1 are located on the same haplotype. It should be understood that sites with a maximum base frequency outside a preset base frequency range are considered "homozygous" sites, and these "homozygous" sites generally do not provide useful information during haplotype genotyping.
[0114] The final determined parental haplotype information can be represented as a structured dataset, specifically including, for example, a series of single nucleotide polymorphism (SNP) sites covered by the parental haplotype, and the specific base type (e.g., A, T, C, or G) corresponding to each of the said sites. In some embodiments, to subsequently determine the haplotype genetic origin of embryo samples, the computing device 110 can further screen the SNP sites included in the parental haplotype information. Specifically, to establish a uniform evaluation benchmark to ensure comparability when subsequently calculating the similarity between the sequencing sequences of embryo samples and each parental haplotype, only SNP sites commonly covered by each parental haplotype are retained in the parental haplotype information. Furthermore, in some embodiments, when all four haplotypes from both parents have the same base type at a certain SNP site, that site typically does not provide useful information for subsequently distinguishing the parental haplotype origin of the embryo sequencing sequences. Therefore, optionally, single nucleotide polymorphisms (SNPs) whose base types are identical in all four haplotypes of both parents are not included in the parental haplotype information. It should be understood that for a given SNP, even if it is homozygous in both parents, but the base types at that SNP differ between the parents (e.g., for a given SNP, the paternal genotype is CC and the maternal genotype is TT), that site can still be used to distinguish the parental haplotype origin of the embryo sequencing sequence and should be retained in the parental haplotype information.
[0115] Those skilled in the art should understand that the above-mentioned variant detection and phasing steps can be implemented using any combination of tools and algorithms that conform to bioinformatics standards, and the specific parameters can be flexibly adjusted and optimized according to the error rate characteristics of the sequencing data, sequencing depth, and actual analysis needs.
[0116] In step 204, the computing device 110 determines the genetic origin of the haplotype in the target region of the embryo sample under test based on the parental haplotype information determined in step 203 and the sequencing data of the embryo sample under test obtained in step 201. This step aims to accurately track whether the embryo has inherited a specific risk haplotype or a normal haplotype from the parent.
[0117] In some implementations, computing device 110 first identifies a target region of the genome. The target region may be identified based on at least one of the following:
[0118] 1) Coordinates of gene mutation sites: For the detection of single-gene genetic diseases (PGT-M), the target region is usually based on the genomic coordinates of the pathogenic mutation (such as point mutation, small insertion or deletion), extending to a certain range upstream and downstream (e.g., 2 Mb upstream and downstream or the entire gene region).
[0119] 2) Structural Variation Breakpoint Coordinates: For chromosomal structural rearrangement (PGT-SR) detection, such as balanced translocations or Roche translocations, the target region is typically the breakpoint of the structural variation and its flanking regions (e.g., extending 2 Mb upstream and downstream from the breakpoint or extending to the entire gene region). Leveraging the advantage of single-molecule long-read sequencing across breakpoints, the breakpoint coordinates can be precisely located, thereby defining the region containing breakpoint connectivity information as the analytical target area.
[0120] The coordinate information mentioned above can be obtained in advance through clinical diagnosis, or it can be newly discovered in this process based on long-read sequencing data of the parent or mother sample through a variant detection algorithm;
[0121] 3) Pre-specified target gene region: For certain routine screening items or specific gene loci (such as HLA typing, common deletion regions in thalassemia, etc.), the computing device 110 can directly load a preset genomic interval as the target region.
[0122] After determining the target region, the computing device 110 executes a method for determining the haplotype genetic origin of the embryo sample to be tested within the target region. This method includes, for example,: calculating the similarity between each sequencing sequence of the embryo sample to be tested and each parental haplotype within the target region, based on the parental haplotype information and the sequencing sequence data of the embryo sample to be tested, thereby determining the parental haplotype origin of each sequencing sequence of the embryo sample to be tested within the target region; and determining the haplotype genetic origin of the embryo sample to be tested within the target region based on the parental haplotype origin of the sequencing sequence of the embryo sample to be tested within the target region. The following will combine... Figure 4 The detailed methods for determining the haplotype genetic origin of the embryo sample to be tested within the target region will not be elaborated here.
[0123] In step 205, the computing device 110 determines the single-gene genetic disease and / or chromosomal structural rearrangement carrier status of the embryo sample to be tested based on the parental haplotype information determined in step 203 and the haplotype genetic origin of the embryo sample to be tested determined in step 204.
[0124] In some implementations, a method for determining the carrier status of a single-gene genetic disease and / or chromosomal structural rearrangement in a test embryo sample includes, for example, two sub-steps: first, computing device 110 identifies a parental haplotype carrying the pathogenic variant and / or chromosomal structural rearrangement; then, it is determined whether the test embryo sample has inherited the parental haplotype.
[0125] Specifically, firstly, the computing device 110 determines the parental pathogenic chain haplotype based on the parental haplotype information, indicating that the parent carries pathogenic variants and / or chromosomal structural rearrangements. For example, for pre-inputted breakpoints of pathogenic variant sites and / or structural variations, or for newly discovered breakpoints of pathogenic variant sites and / or structural variations based on long-read sequencing data from the father or mother via a variant detection algorithm, it determines whether the breakpoint of the pathogenic variant site and / or structural variation is located on a certain parental haplotype, and marks the parental haplotype containing the breakpoint of the pathogenic variant site and / or structural variation as the parental pathogenic chain haplotype.
[0126] Subsequently, the computing device 110 matches the haplotype genetic origin of the embryo sample to be tested with the parental pathogenic chain haplotype: in response to the haplotype genetic origin of the embryo sample to be tested containing the parental pathogenic chain haplotype, the computing device 110 determines that the embryo sample to be tested carries a single-gene genetic disease and / or chromosomal structural rearrangement; in response to the haplotype genetic origin of the embryo sample to be tested not containing the parental pathogenic chain haplotype, the computing device 110 determines that the embryo sample to be tested does not carry a single-gene genetic disease and / or chromosomal structural rearrangement, and determines that the embryo is genetically normal or does not carry the target pathogenic variant.
[0127] It should be understood that although the above description illustrates the process for determining chromosomal euploidy in the embryo sample to be tested in a specific order (i.e., the PGT-A testing procedure, such as...), Figure 3 (As shown) and the process used to determine the haplotype genetic origin and single-gene genetic disease and / or chromosomal structural rearrangement carrier status of the embryo sample to be tested (i.e., the PGT-M / SR testing procedure, such as...). Figure 4 and Figure 5 (As shown), but this is for illustrative purposes only and not a limitation of the invention. Since the above analysis procedures are all based on the sequencing data of the same set of embryo samples and their parental samples obtained in step 201, the computing device 110 can execute the PGT-A and PGT-M / SR testing procedures in parallel or sequentially as needed. Regardless of whether parallel or sequential execution is used, the analysis results of the two testing procedures do not interfere with each other and together constitute a comprehensive genetic assessment result of the embryo sample to be tested.
[0128] This invention utilizes single-molecule long-read sequencing data from paternal, maternal, and embryo samples from a family pedigree to construct parental haplotypes and analyze sequence similarity. By tracing the genetic origin of the haplotypes in the embryo samples to be implanted, it determines whether the embryo carries a single-gene genetic disease and / or chromosomal structural rearrangement. Compared to existing technologies, this invention enables integrated detection of preimplantation genetic defects (Aneuploidy, single-gene genetic diseases, and chromosomal structural rearrangements) on a single platform, providing a rapid, simple, efficient, and accurate one-stop solution for determining whether an embryo to be implanted has genetic defects.
[0129] The following will combine Figure 3 This invention describes a method for determining the chromosomal euploidy of an embryo sample to be tested, according to an embodiment of the present invention. Figure 3 A flowchart illustrating a method for determining chromosomal euploidy in an embryo sample according to an embodiment of the present invention is shown. It should be understood that the above method can, for example, be used in... Figure 6 It can be executed at the described device, or it can be performed at... Figure 1 The described computing device 110 performs the operation. It should be understood that the above method may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the invention is not limited in this respect.
[0130] In step 301, the computing device 110 divides the reference genome into multiple consecutive equal-length intervals based on a preset window length. The preset window length is selected to balance the resolution and statistical stability of copy number detection. Those skilled in the art can dynamically adjust the preset window length according to actual needs, and a preferred preset window length is, for example, 100 Kb.
[0131] In step 302, the computing device 110, based on the sequencing sequence data of the embryo sample to be tested obtained in the preceding steps, uses bioinformatics statistical tools (e.g., but not limited to software such as bedtools and samtools) to count the number of aligned sequencing sequences within each of the reference genomic regions. It should be understood that the above statistics can be implemented using any suitable tools and algorithms, and sequencing sequences with poor alignment quality can be pre-filtered during the statistics process to improve counting accuracy.
[0132] Subsequently, at step 303, the computing device 110 determines the copy number prediction result for each reference genomic region based on the number of aligned sequencing sequences in each reference genomic region of the embryo sample to be tested. Depending on whether an external control sample is available, the determination of the copy number prediction result for each reference genomic region can be performed using a reference system or without a reference system.
[0133] Regarding the reference-free approach, normalization is primarily achieved using the sequencing depth distribution of the sample itself. Specifically, the computing device 110 divides chromosomes into autosomes and sex chromosomes for separate processing. For each region on an autosome, the computing device 110 first calculates the median or mean number of aligned sequencing sequences across all autosomal regions of the embryo sample to be tested as a baseline. The copy number of each region is calculated as the ratio of the number of aligned sequencing sequences in that region to the autosomal baseline. For sex chromosomes, the copy number is calculated as the ratio of the number of aligned sequencing sequences in the region containing that sex chromosome to the median or mean number of aligned sequencing sequences across all regions of the sex chromosome. It should be understood that, based on biological principles, for male embryo samples, the median or mean copy number of the X and Y chromosomes is theoretically approximately half the copy number of autosomes; for female embryo samples, the median or mean copy number of the X chromosome is theoretically consistent with the copy number of autosomes, while the theoretical copy number of the Y chromosome is zero.
[0134] In other embodiments, the computing device 110 uses a reference system to determine the copy number prediction for each reference genomic region. This method includes pre-constructing a reference system and performing copy number detection based on the reference system. Constructing the reference system includes: acquiring sequencing sequence data from multiple normal samples (e.g., healthy individuals with normal karyotypes) used to construct the reference system, performing the same alignment and region division as described above, and counting the number of reference sequences within each region. Subsequently, the computing device 110 calculates a baseline value for the number of reference sequencing sequences for each region, such as the median or mean of the number of sequencing sequences from multiple reference samples in that region. It should be understood that for autosomes, there is theoretically no difference between males and females when constructing the reference system, and male and female data can be combined to construct a unified reference system. For sex chromosomes, since male and female copy numbers are naturally different, male and female reference systems should be constructed separately. Specifically, the construction method involves calculating the number of reference sequencing sequences for each region of the sex chromosome at the median or mean level as the baseline value for that region. When performing copy number detection with a reference system, the computing device 110 uses the ratio of the number of sequencing sequences in each interval of the embryo sample to be tested to the number of reference sequencing sequences in that interval in the reference system as the copy number of that interval.
[0135] After determining the copy number prediction results for each interval, to improve the quality of the intervals supporting the copy number prediction algorithm, the computing device 110 can also filter the copy number data of the intervals according to preset filtering rules. For example, intervals with abnormal GC content or abnormal sequencing depth outliers can be removed to improve the accuracy and reliability of subsequent copy number prediction algorithms.
[0136] Since the copy number values of a single interval may fluctuate randomly, statistical algorithms are needed to identify continuous copy number states. Therefore, in step 304, the computing device 110 calculates the degree of difference in the copy number prediction results of adjacent reference genome intervals and merges adjacent intervals with a difference degree lower than a preset merging threshold, thereby smoothing the discrete interval data into continuous chromosomal segment events, and thus identifying large fragment amplifications or deletions. This process can be implemented using algorithms such as circular binary segmentation (CBS) or hidden Markov models (HMM).
[0137] Finally, at step 305, the computing device 110 determines the chromosomal neuploidy detection result of the embryo sample to be tested based on the merged intervals. In some embodiments, the computing device 110 may filter out small fragments with lengths lower than a preset length threshold to eliminate interference from random noise. The preset length threshold may be, for example, 4 Mb. It should be understood that the preset length threshold can be adjusted according to the experimental needs of those skilled in the art. Subsequently, based on the final copy number value of each chromosome fragment (e.g., close to 2 indicates diploidy, close to 3 indicates trisomy, and close to 1 indicates monosomy), the computing device 110 determines whether the embryo sample to be tested has aneuploidy abnormalities.
[0138] The following will combine Figure 4 This invention describes a method for determining the haplotype genetic origin of an embryo sample within a target region according to an embodiment of the present invention. Figure 4 A flowchart illustrating a method for determining the haplotype genetic origin of a test embryo sample within a target region, according to an embodiment of the present invention, is shown. It should be understood that the above, for example, can be applied to… Figure 6 It can be executed at the described device, or it can be performed at... Figure 1 The described computing device 110 performs the operation. It should be understood that the above method may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the invention is not limited in this respect.
[0139] In step 401, the computing device 110 extracts the sequencing sequence of the embryo sample to be tested that falls into the target region based on the alignment result data of the sequence alignment between the sequencing sequence of the embryo sample to be tested and the reference genome sequence.
[0140] Based on the target region determined in the aforementioned steps, the computing device 110 retrieves and extracts sequencing sequences that fall within the target range from the alignment result file.
[0141] It should be understood that the extracted sequencing sequences are not limited to sequences completely contained within the target region, but also include sequences that cross the boundary of the target region or only partially overlap with the target region. As long as a sequencing sequence shows a coordinate range that overlaps with the target region in the alignment results, and the overlapping region contains potentially valid genetic markers (such as SNP sites), the computing device 110 can extract it as a candidate sequence.
[0142] In some implementations, to ensure the reliability of the input data and reduce the interference of sequencing noise on subsequent analysis, after extracting the sequencing sequences of the embryo samples falling within the target region, the computing device 110 typically also performs a step of filtering the sequencing sequences of the embryo samples within the target region according to preset filtering conditions. The preset filtering conditions include at least one of the following:
[0143] 1) Remove sequencing sequences whose alignment quality score (MapQ) does not meet the preset alignment quality score threshold. The alignment quality score reflects the reliability of a sequencing sequence aligned to a specific location in the reference genome. If the alignment quality of a sequencing sequence is too low (e.g., MapQ < 20), it means that the sequence may come from a repetitive region in the genome or a highly homologous non-target region. Retaining these sequences may lead to the introduction of erroneous information into the analysis. Those skilled in the art can dynamically set the alignment quality score threshold according to the actual situation. The preset alignment quality score threshold is, for example, 20.
[0144] 2) Remove bases in the sequencing sequence whose base quality fraction (BaseQ) does not meet a preset base quality fraction threshold. If the base quality at a certain site is too low (e.g., BaseQ < 10), the base may be a reading error by the sequencing instrument rather than a real biological variation. In this step, the computing device does not necessarily remove the entire sequence, but rather "ignores" specific base sites in the sequence whose base quality fraction does not meet the preset base quality fraction threshold, thus preventing them from participating in subsequent similarity comparisons. Those skilled in the art can dynamically set this base quality fraction threshold according to actual conditions; the preset base quality fraction threshold is, for example, 10.
[0145] 3) Remove sequencing sequences whose number of base sites satisfying the preset base quality fraction threshold in the single nucleotide polymorphism sites covered by the sequencing sequence in the target region does not meet the preset coverage site number threshold. If a sequencing sequence covers only a very small number (e.g., 0 or 1) base sites in the target region that satisfy the preset base quality fraction threshold, it indicates that the sequencing sequence carries insufficient effective information, which may affect the subsequent determination of the parental haplotype origin. Those skilled in the art can dynamically set the coverage site number threshold according to the actual situation, and the preset coverage site number threshold is, for example, 2.
[0146] In step 402, the computing device 110 calculates the similarity between each sequencing sequence of the embryo sample to be tested and each parent haplotype within the target region, based on the parent haplotype information determined in step 203 and the sequencing sequence data of the embryo sample to be tested extracted in step 401.
[0147] The calculation of the similarity between each sequencing sequence of the embryo sample and each parent haplotype involves, for example, the computing device 110 first determining, for each sequencing sequence of the embryo sample to be tested within the target region, and for each of the parent haplotypes, based on the sequencing sequence of the embryo sample to be tested and the parent haplotype information, the common single nucleotide polymorphism sites between the sequencing sequence of the embryo sample to be tested and the parent haplotype, and the base types at the common single nucleotide polymorphism sites between the sequencing sequence of the embryo sample to be tested and the parent haplotype.
[0148] Subsequently, in response to the non-zero number of shared single nucleotide polymorphism (SNP) sites, the computing device 110 determines, for each of the shared SNP sites, whether the base types of the sequencing sequence of the embryo sample to be tested and the parent haplotype are consistent at the shared SNP site, and counts the number of shared SNP sites with consistent base types; calculates the ratio of the number of shared SNP sites with consistent base types to the total number of shared SNP sites, and obtains the similarity between the sequencing sequence of the embryo sample to be tested and the parent haplotype.
[0149] When the number of shared single nucleotide polymorphism sites is zero, due to the lack of a basis for comparison, the computing device 110 defines the similarity between the sequencing sequence of the embryo sample to be tested and the parental haplotype as zero. This means that the sequence cannot provide evidence to support the haplotype.
[0150] Specifically, let a long read sequencing sequence of the embryo sample to be tested be... It can be formally represented as a triple:
[0151] ;
[0152] in:
[0153] Unique identifier for long read sequencing sequences
[0154] The set of single nucleotide polymorphism sites covered by long-read sequencing sequences. .in This represents the coordinates of a single nucleotide polymorphism site on the reference genome.
[0155] Mapping of base types at various single nucleotide polymorphism sites in long-read sequencing sequences: ;
[0156] For any site , Indicator sequencing sequence at single nucleotide polymorphism sites The type of base at that location. If This indicates that the sequencing sequence did not cover single nucleotide polymorphism sites. or single nucleotide polymorphism sites Sequencing quality below a preset threshold is marked as unqualified.
[0157] Regarding the single-type collection of parents Defined as Each monomer type is represented as a binary tuple:
[0158] ;
[0159] in:
[0160] Monotype The set of single nucleotide polymorphism sites included
[0161] Monotype Mapping of base types at each single nucleotide polymorphism site. .
[0162] It should be understood that, since the sites covered by the haplotypes of the parents intersect, any haplotype The included single nucleotide polymorphism (SNP) sites are identical. It should be understood that for any given site, if the base type is identical across all haplotypes, that site provides no useful information for subsequently distinguishing the origin of the embryonic sequencing sequence. Therefore, alternatively, these sites may be excluded from the included set of SNP sites.
[0163] For embryo sequencing sequences And parental monotype The set of common sites is defined as follows:
[0164] ;
[0165] That is, embryo sequencing sequences were excluded. Single nucleotide polymorphisms (SNPs) that are not covered in the sequencing data or whose sequencing quality is below a preset threshold are marked as unqualified SNPs.
[0166] For embryo sequencing sequences and monomeric The set of matching sites, which consists of sites where the bases corresponding to alleles in the common site set are identical, is defined as follows:
[0167] ;
[0168] The following explanation of sequencing sequences is based on formula (1). With single type Methods for calculating similarity.
[0169]
[0170] In the above formula (1), Representative sequencing sequence and haplotype The similarity, with values ranging from 1 to 10. . Indicates sequencing sequence and haplotype The number of sites with identical bases at corresponding sites. Representative sequencing sequence and haplotype The number of sites that are commonly included.
[0171] In particular, when When, define Researchers in this field should understand that similarity The higher the value, the better the sequencing sequence. Derived from single type The greater the likelihood, the better. Theoretically, if the sequencing sequence does indeed originate from a certain haplotype, the similarity should be close to 1 (the actual value may be slightly lower considering sequencing errors).
[0172] At step 403, the computing device 110 determines the parental haplotype origin of each sequencing sequence of the embryo sample to be tested within the target region based on the similarity between each sequencing sequence of the embryo sample to be tested within the target region and each parental haplotype.
[0173] A method for determining the parental haplotype source of each sequencing sequence of a test embryo sample within a target region includes, for example,: for each sequencing sequence of the test embryo sample within the target region, based on the similarity between the sequencing sequence of the test embryo sample and each parental haplotype, determining the parental haplotype with the highest similarity to the sequencing sequence of the test embryo sample and its corresponding first similarity, and the parental haplotype with the second highest similarity to the sequencing sequence of the test embryo sample and its corresponding second similarity; in response to 1) the number of shared single nucleotide polymorphism (SNP) sites between the sequencing sequence of the test embryo sample and the parental haplotype with the highest similarity in the target region meets a preset threshold for the number of shared SNP sites, 2) the first similarity meets a preset similarity threshold, and 3) the difference between the first similarity and the second similarity meets a preset similarity difference threshold, determining the parental haplotype with the highest similarity as the parental haplotype source of the sequencing sequence. The following will combine... Figure 5 The detailed methods for determining the haplotype genetic origin of the embryo sample to be tested within the target region will not be elaborated here.
[0174] At step 404, computing device 110 determines the haplotype genetic origin of the embryo sample to be tested within the target region based on the parental haplotype origin of the sequencing sequence of the embryo sample to be tested within the target region determined in step 403.
[0175] Specifically, the computing device 110 counts the number of sequencing sequences (i.e., the number of supporting sequences) of the embryo samples to be tested allocated to that parental haplotype within the target region for each parental haplotype (i.e., paternal haplotype 1, paternal haplotype 2, maternal haplotype 1, and maternal haplotype 2). These statistics intuitively reflect whether there is genetic material in the embryo genome that is highly consistent with a specific parental haplotype.
[0176] Based on the above statistical results, the computing device 110 determines the haplotype genetic origin of the embryo sample to be tested based on a preset support threshold. The preset support threshold is used to distinguish between real genetic signals and possible sequencing noise or background interference. Those skilled in the art can dynamically set this support threshold according to the sequencing depth. For example, for 10× sequencing data, the support threshold can be set to 10; for 1× sequencing data, the support threshold can be set to 1.
[0177] In normal diploid embryonic development, theoretically, the two alleles (or haplotypes) of the embryo should each originate from a haplotype from the father and a haplotype from the mother. Therefore, in the determination process:
[0178] In response to the number of sequencing sequences of the embryo sample to be tested assigned to either the paternal or maternal haplotype satisfying the preset support threshold, the computing device 110 determines that the parental haplotype satisfying the preset support threshold is the haplotype genetic source of the embryo sample to be tested in the target region. In some embodiments, if the number of supporting sequences for two haplotypes of the same parent both satisfy the preset support threshold (e.g., an unbalanced translocation occurs), the computing device 110 may label both haplotypes as the haplotype genetic source of the embryo sample to be tested in the target region. In other embodiments, such as in routine single-gene disease testing, the haplotype with a significantly dominant number of supporting sequences (e.g., the largest number that satisfies the support threshold) is typically identified as the sole genetic source.
[0179] Furthermore, in response to the fact that the number of sequencing sequences of the embryo samples to be tested for each haplotype of the father or mother does not meet the preset support threshold, the computing device 110 first determines that no haplotype originating from the parent is detected in the target region of the embryo sample to be tested. This means that a chromosome segment deletion has occurred. In this case, in response to the fact that the number of sequencing sequences of the embryo samples to be tested for each haplotype of the father or mother does not meet the preset support threshold, and the parent carries a deletion variant in the target region, the computing device 110 determines that the haplotype of the parent carrying the deletion variant in the target region is the haplotype genetic source of the embryo sample to be tested in the target region.
[0180] Through the above scheme, the present invention can accurately trace the haplotype genetic origin of the embryo sample to be implanted, which is beneficial to accurately determine the single-gene genetic disease and / or chromosomal structural rearrangement carrier status of the embryo to be implanted.
[0181] The following will combine Figure 5 A method for determining the parental haplotype origin of the sequencing sequence of an embryo sample to be tested, according to embodiments of the present invention, is described. Figure 5 A flowchart illustrating a method for determining the parental haplotype origin of a sequencing sequence of an embryo according to an embodiment of the present invention is shown. It should be understood that the above method can, for example, be used in... Figure 6 It can be executed at the described device, or it can be performed at... Figure 1 The described computing device 110 performs the operation. It should be understood that the above method may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the invention is not limited in this respect.
[0182] To ensure the accuracy of parental haplotype source allocation, this invention introduces a set of thresholds for evaluating matching confidence. These thresholds include, but are not limited to, similarity thresholds. There is a threshold for the number of single nucleotide polymorphism sites. and similarity difference threshold Among them, a threshold for the number of shared single nucleotide polymorphism sites is set. This is to prevent sequencing sequences from being too short or having too few sites matching the parental haplotype; a similarity threshold is set. This is to perform quality filtering on the sequencing sequences, eliminating false positive matches caused by excessive sequencing errors; and to set a similarity difference threshold. This is to avoid incorrect sequencing sequence allocation due to insufficient differentiation when two parent haplotypes are highly similar.
[0183] In step 501, the computing device 110, for each sequencing sequence of the embryo sample to be tested within the target region, sorts the similarity between the sequencing sequence of the embryo sample to be tested and each parental haplotype, based on the similarity calculated in the preceding steps, in descending order. Based on the sorting results, the computing device 110 determines the parental haplotype with the highest similarity to the sequencing sequence of the embryo sample to be tested and its corresponding first similarity (denoted as ). ), and the parental haplotype with the second highest similarity to the sequencing sequence of the embryo sample to be tested, and its corresponding second similarity (denoted as ), and so on. ).
[0184] Subsequently, the computing device 110 enters the multi-condition verification stage (steps 502 to 504) to determine whether to assign the sequencing sequence to the parent haplotype with the highest similarity.
[0185] At step 502, the computing device 110 determines whether the number of shared SNP sites between the sequencing sequence of the embryo sample to be tested and the parent haplotype with the highest similarity in the target region meets a preset threshold for the number of shared SNP sites. In some embodiments, the determination logic is to determine whether the number of shared SNP sites is greater than or equal to... .
[0186] At step 503, computing device 110 determines the first similarity. Does it meet the preset similarity threshold? In some embodiments, the determination logic is to determine a first similarity. Is it greater than or equal to? .
[0187] At step 504, computing device 110 determines the first similarity. Similarity to the second Does the difference meet the preset similarity difference threshold? In some embodiments, the determination logic is to determine a first similarity. Similarity to the second The difference ( Is it greater than or equal to? .
[0188] At step 505, if the computing device 110 determines that all the above conditions are met, that is, that the number of shared SNP sites is greater than or equal to a preset threshold for the number of shared SNP sites. Determine the first similarity Greater than or equal to the preset similarity threshold And determine the first similarity Similarity to the second The difference is greater than or equal to the preset similarity difference threshold. Then the computing device 110 determines that the parent haplotype with the highest similarity is the parent haplotype source of the sequencing sequence.
[0189] In some embodiments, regarding the preset threshold for the number of common single nucleotide polymorphism sites For example, setting it to 2, regarding the preset similarity threshold. For example, setting it to 0.6 relates to the preset similarity difference threshold. For example, it can be set to 0.1. Those skilled in the art can dynamically set one or more of the above thresholds according to the actual situation.
[0190] Conversely, at step 506, if the computing device 110 determines that any of the above conditions are not met, for example, if it determines that the number of shared SNP sites is less than a preset threshold for the number of shared SNP sites. Or determine the first similarity. Less than the preset similarity threshold Or determine the first similarity. Similarity to the second The difference is less than the preset similarity difference threshold. If the sequencing sequence is not assigned a parental haplotype source, the computing device 110 will not assign a parental haplotype source. In some implementations, these sequences are labeled as having an "obscure" source.
[0191] It should be understood that, despite Figure 5 The above description shows the thresholds for the number of common single nucleotide polymorphism sites in a specific order. Preset similarity threshold and preset similarity difference threshold The verification process is described, but this is only for illustrative purposes and is not intended to limit the scope of the invention.
[0192] In the actual algorithm implementation or execution logic of the computing device 110, the execution order of steps 502, 503, and 504 is not necessarily fixed. For example, the computing device 110 may first determine whether the first similarity meets the preset similarity threshold (step 503), and then determine whether the number of shared single nucleotide polymorphism sites meets the threshold (step 502); or, the computing device 110 may first determine whether the similarity difference meets the condition (step 504). In addition, in some parallel computing architectures, the computing device 110 may also judge these three conditions simultaneously (in parallel).
[0193] As long as the sequencing sequence of the embryo sample to be tested simultaneously meets the above three preset conditions (i.e., the number of total SNP sites ≥ 100%) , ,and Regardless of the order in which these three conditions are judged, the computing device 110 will determine that the sequence originates from the parent haplotype with the highest similarity. Conversely, if any one of the conditions is not met, regardless of at which step the unmet condition is detected, the computing device 110 will not assign a specific parent haplotype source to the sequencing sequence. Those skilled in the art should understand that any adjustment to the order of these judgment logics falls within the scope of protection of this invention.
[0194] It should be understood that there may be a high degree of similarity between haplotypes of different parents. By introducing the above-mentioned triple preset thresholds, this invention can effectively filter out low-quality, insufficient information or ambiguous sequencing sequences, significantly reduce misassignment caused by sequencing errors or high similarity of haplotypes, and thus help improve the accuracy of determining the genetic origin of embryo haplotypes.
[0195] The following three specific examples illustrate the practical application and detection effect of the integrated preimplantation genetic testing method of the present invention.
[0196] Example 1: NKX2-1 c.1040del family
[0197] 1. Parental variant carrier status
[0198] Mother: Negative
[0199] Father: NKX2-1 c.1040del interlocking
[0200] 2. Analysis Strategy
[0201] Single-molecule long-read sequencing was performed on both parents and three embryos to complete PGT integrated detection, detect aneuploidy and pathogenic variants in the embryos, and linkage analysis was performed on the haplotypes of the parents and the embryo samples to exclude embryos carrying pathogenic variants.
[0202] 3. Mutation detection
[0203] Sequencing results analysis showed that the male partner carried the virus. NKX2-1 c.1040del chimeric variant: In this embodiment, it can be determined by IGV genome breakpoint visualization software that the variant carrying status of the parents and embryos is that embryo 1 carries the pathogenic variant, while embryos 2 and 3 do not carry the pathogenic variant.
[0204] 4. Parental haplotype phase fixation
[0205] Haplotype typing was performed on both the male and female partners, and the pathogenic haplotype H2 in the male partner was identified. Subsequently, based on... NKX2-1 c.1040del classifies the male's pathogenic haplotype H2 into two subtypes, H2a and H2b.
[0206] 5. Embryo haplotype genetic origin testing
[0207] Subsequently, linkage analysis was performed using this detection method to confirm the haplotype inheritance of the three embryos, and the results are shown in Table 1. Embryos 1 and 2 carried the paternal pathogenic haplotype H2, with embryo 1 inheriting the paternal H2b haplotype. NKX2-1 c.1040del, embryo 2 inherited the father's H2a haplotype and is not a carrier. NKX2-1 c.1040del. Embryo 3 did not carry the pathogenic haplotype. The PGT integration results for this family are shown in Table 2.
[0208] Table 1: NKX2-1 c.1040del Family Linkage Analysis Results
[0209]
[0210] Table 2: NKX2-1 c.1040del family PGT results
[0211]
[0212] Example 2: t(X; 19) (p22.31; p13.1) family
[0213] 1. Parental variant carrier status
[0214] Mother: 46,XX, t(X; 19)(p22.31; p13.1)
[0215] Father: 46, XY
[0216] 2. Analysis Strategy
[0217] Single-molecule long-read sequencing was performed on both parents and three embryos to complete PGT integrated detection, detect aneuploidy and balanced translocation in the embryos, and perform linkage analysis based on the haplotypes formed by the parents and the embryo samples to exclude embryos carrying balanced translocations.
[0218] 3. Mutation detection
[0219] Structural variation analysis showed translocations between the female's chromosome 19 and X chromosome. In this embodiment, the translocation carrier status of the parents and the embryo can be determined by locating the breakpoints using IGV genome breakpoint visualization software.
[0220] 4. Parental haplotype phase fixation
[0221] The monotypic types of the male and female individuals were classified, and the pathogenic monotypic type of the female individual was identified.
[0222] 5. Embryo haplotype genetic origin testing
[0223] Subsequently, linkage analysis was performed using this detection method to confirm the haplotype inheritance status of the three embryos, and the results are shown in Table 3. Embryos 1, 2, and 3 carried the pathogenic haplotype, with embryo 1 carrying two haplotypes from the mother's X chromosome, indicating that embryo 1 carried an unbalanced translocation. The PGT integration results for this family are shown in Table 4.
[0224] Table 3: Results of t(X; 19) (p22.31; p13.1) pedigree linkage analysis
[0225]
[0226] Table 4: t(x; 19) (p22.31; p13.1) family PGT results
[0227]
[0228] Example 3: FBN1 E1-2del family
[0229] 1. Parental variant carrier status
[0230] Mother: FBN1 E1-2del
[0231] Father: Negative
[0232] 2. Analysis Strategy
[0233] Single-molecule long-read sequencing was performed on both parents and three embryos to complete PGT integrated detection, detect aneuploidy and copy number deletion in the embryos, and perform linkage analysis based on the haplotypes formed by the parents and the embryo samples to exclude embryos carrying copy number deletion.
[0234] 3. Mutation detection
[0235] Structural variation analysis showed that the mother carried the deletion: chr15: 48625733-48708356del (82624bp). In this embodiment, the copy number deletion carrier status of the parents and the embryo can be determined by using IGV genome breakpoint visualization software.
[0236] 4. Parental haplotype phase fixation
[0237] The monotypic types of the male and female individuals were classified, and the pathogenic monotypic type of the female individual was identified.
[0238] 5. Embryo haplotype genetic origin testing
[0239] Subsequently, linkage analysis was performed using this detection method to confirm the haplotype inheritance status of the three embryos. The results are shown in Table 5. Embryo 3 did not carry the undeleted haplotype of the female parent, therefore embryo 3 carries the copy number deletion variant. The PGT integration results for this family are shown in Table 6.
[0240] Table 5: Results of Family Linkage Analysis for FBN1 E1-2del
[0241]
[0242] Table 6. PGT Results of FBN1 E1-2del Families
[0243]
[0244] Therefore, the method provided by this invention can also be used as a verification means to verify the discovered CNVs through chain analysis.
[0245] By adopting the above scheme, the present invention can complete the integrated detection of PGT-A, PGT-M and PGT-SR, and determine whether an embryo has genetic defects in a one-stop, fast, simple, efficient and accurate manner.
[0246] Figure 6 The schematic diagram illustrates an example device suitable for implementing embodiments of the present invention. For example, the device may be as follows: Figure 1 The computing device 110 shown. For example, the device may be used to implement execution Figures 2 to 5 The apparatus for the method shown.
[0247] like Figure 6As shown, the device includes a central processing unit (i.e., CPU 1001), which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (i.e., ROM 1002) or loaded from storage unit 1008 into random access memory (i.e., RAM 1003). The RAM 1003 may also store various programs and data required for the operation of the electronic device. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. An input / output interface (i.e., I / O interface 1005) is also connected to bus 1004.
[0248] Multiple components in the device are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, mouse, microphone, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless transceiver, etc. The communication unit 1009 allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0249] The various processes and procedures described above can be executed by CPU 1001. For example, in some embodiments, the various processes described above can be implemented as computer software programs stored in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the device via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by CPU 1001, one or more actions of the methods described above can be performed. Alternatively, in other embodiments, CPU 1001 can be configured to perform one or more actions of the methods by any other suitable means (e.g., by means of firmware).
[0250] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made by those skilled in the art without creative effort within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for preimplantation genetic testing, characterized in that, Includes the following steps: Obtain sequencing sequence data of the embryo sample to be tested and its paternal and maternal samples, wherein the sequencing sequence data includes single-molecule long-read sequencing sequences of the corresponding samples obtained using single-molecule long-read sequencing technology; Based on the sequencing sequence data of the paternal and maternal samples, haplotype information of each parent was determined; The sequencing data of each parental haplotype and the embryo samples to be tested are filtered to determine the common single nucleotide polymorphism sites and corresponding base types between the sequencing sequences of the embryo samples to be tested and the corresponding parental haplotypes. The ratio of the number of common single nucleotide polymorphism sites with the same base type to the total number of common single nucleotide polymorphism sites is calculated to obtain the similarity between the sequencing sequences of the embryo samples to be tested and the corresponding parental haplotypes. Based on the similarity, the parental haplotype source of each sequencing sequence of the embryo samples to be tested in the target region is determined. Based on the identified parental haplotype origin, determine the haplotype genetic origin of the embryo sample to be tested within the target region; The process of calculating the ratio of the number of shared single nucleotide polymorphism (SNP) sites with the same base type to the total number of shared SNP sites, and obtaining the similarity between the sequencing sequence of the embryo sample to be tested and the corresponding parental haplotype, includes: Based on the alignment results of the sequencing sequence of the embryo sample to be tested and the reference genome sequence, the sequencing sequence of the embryo sample to be tested that falls into the target region is extracted. For each sequencing sequence of the embryo sample to be tested within the target region, and for each parental haplotype, based on the corresponding sequencing sequence of the embryo sample to be tested and the corresponding parental haplotype information, the common single nucleotide polymorphism sites between the sequencing sequence of the embryo sample to be tested and the corresponding parental haplotype, and the base types at the common single nucleotide polymorphism sites between the sequencing sequence of the embryo sample to be tested and the corresponding parental haplotype are determined. When the number of common single nucleotide polymorphism sites is not zero, for each of the common single nucleotide polymorphism sites, determine whether the sequencing sequence of the embryo sample to be tested and the corresponding parent haplotype have the same base type at the common single nucleotide polymorphism site, and count the number of common single nucleotide polymorphism sites with the same base type. The ratio of the number of common single nucleotide polymorphism sites with the same base type to the number of common single nucleotide polymorphism sites is calculated to obtain the similarity between the sequencing sequence of the embryo sample to be tested and the haplotype of the parent. When the number of shared single nucleotide polymorphism sites is zero, the similarity between the corresponding sequencing sequence of the embryo sample to be tested and the corresponding parent haplotype is zero.
2. The preimplantation genetic testing method for embryos as described in claim 1, characterized in that, The process of determining haplotype information of each parent based on the sequencing sequence data of the paternal and maternal samples includes: determining the single nucleotide polymorphism site information in the genomes of the paternal and maternal samples based on the alignment results of the sequencing sequences of the paternal and maternal samples with the reference genome sequence. Based on the single nucleotide polymorphism site information, the haplotypes of each parent were constructed, and the haplotype information of each parent was determined; The parent haplotype information includes the single nucleotide polymorphism sites covered by the corresponding parent haplotype and the base types of the corresponding parent haplotypes corresponding to the single nucleotide polymorphism sites.
3. The preimplantation genetic testing method for embryos as described in claim 1, characterized in that, The target area is determined based on at least one of the following: Coordinates of gene mutation sites obtained in advance or based on sequencing sequence data from paternal or maternal samples; The coordinates of structural variation breakpoints obtained in advance or based on sequencing sequence data from paternal or maternal samples; and the pre-specified target gene regions.
4. The preimplantation genetic testing method as described in claim 1, characterized in that, The process of filtering the haplotype information of each parent and the sequencing sequence data of the embryo samples to be tested includes: filtering the sequencing sequences of the embryo samples to be tested that fall into the target region according to preset filtering conditions, wherein the preset filtering conditions include: Remove sequencing sequences whose alignment quality scores do not meet the preset alignment quality score threshold; Remove bases from the sequencing sequence whose base mass fraction does not meet the preset base mass fraction threshold; Remove sequencing sequences that do not meet the preset coverage threshold for the number of base sites that satisfy the preset base mass fraction threshold in the single nucleotide polymorphism sites covered by the target region.
5. The method of claim 1, wherein the method is a preimplantation genetic diagnosis method. The process of determining the parental haplotype origin of each sequencing sequence of the embryo sample to be tested within the target region based on the similarity includes: For each sequencing sequence of the embryo sample to be tested in the target area, based on the similarity between the sequencing sequence of the embryo sample to be tested and each parental haplotype, the parental haplotype with the highest similarity to the sequencing sequence of the embryo sample to be tested and its corresponding first similarity, and the parental haplotype with the second highest similarity to the sequencing sequence of the embryo sample to be tested and its corresponding second similarity are determined. If the number of shared single nucleotide polymorphism (SNP) sites between the corresponding sequencing sequence of the embryo sample to be tested and the parental haplotype with the highest similarity in the target region meets a preset threshold for the number of shared SNP sites, the first similarity meets a preset threshold for similarity, and the difference between the first similarity and the second similarity meets a preset threshold for similarity difference, then the parental haplotype with the highest similarity is determined to be the source of the parental haplotype of the corresponding sequencing sequence.
6. The preimplantation genetic testing method for embryos as described in claim 5, characterized in that, The process of determining the parental haplotype origin of each sequencing sequence of the embryo sample to be tested within the target region based on the similarity also includes: For each sequencing sequence of the embryo sample to be tested within the target region, if the number of shared single nucleotide polymorphism sites between the corresponding sequencing sequence of the embryo sample and the parental haplotype with the highest similarity in the target region does not meet the preset threshold for the number of shared single nucleotide polymorphism sites, or if the first similarity does not meet the preset similarity threshold or the difference between the first similarity and the second similarity does not meet the preset similarity difference threshold, then the parental haplotype source will not be assigned to the corresponding sequencing sequence.
7. The preimplantation genetic testing method for embryos as described in claim 1, characterized in that, The process of determining the haplotype genetic origin of the embryo sample to be tested within the target region, based on the identified parental haplotype origin, includes: For each parental haplotype, count the number of sequencing sequences of the embryo samples to be tested that are assigned to that parental haplotype within the target region; In response to the number of sequencing sequences of the embryo sample to be tested assigned to either the father or the mother that meets a preset support threshold, the parental haplotype that meets the preset support threshold is determined to be the haplotype genetic source of the embryo sample to be tested in the target region. In response to the fact that the number of sequencing sequences of the embryo samples to be tested for each haplotype assigned to the father or mother does not meet the preset support threshold, it is determined that no haplotype originating from the parent is detected in the target region of the embryo sample to be tested.
8. The preimplantation genetic testing method for embryos as described in claim 7, characterized in that, In response to the fact that the number of sequencing sequences of the embryo samples to be tested for each haplotype assigned to the father or mother does not meet the preset support threshold, and the corresponding parent carries a deletion variant in the target region, the haplotype of the parent carrying the deletion variant in the target region is determined to be the haplotype genetic source of the embryo sample to be tested in the target region.
9. The method of claim 1, wherein the sample is obtained from a chorionic villus biopsy. It also includes the following steps: Based on the haplotype information of the parents and the haplotype genetic origin of the embryo sample to be tested, the single-gene genetic disease and / or chromosomal structural rearrangement carrier status of the embryo sample to be tested is determined.
10. The preimplantation genetic testing method for embryos as described in claim 9, characterized in that, The process of determining the carrier status of single-gene genetic diseases and / or chromosomal structural rearrangements in embryo samples based on parental haplotype information and the haplotype genetic origin of the embryo sample to be tested includes: Based on the parental haplotype information, determine the parental pathogenic chain haplotype carrying pathogenic variants and / or chromosomal structural rearrangements; In response to the fact that the haplotype genetic origin of the embryo sample to be tested includes the parental pathogenic chain haplotype, it is determined that the embryo sample to be tested carries a single-gene genetic disease and / or chromosomal structural rearrangement; In response to the fact that the haplotype of the embryo sample to be tested does not contain the parental pathogenic chain haplotype, it is determined that the embryo sample to be tested does not carry a single-gene genetic disease and / or chromosomal structural rearrangement.
11. The method of claim 1, wherein the sample is obtained from a chorionic villus biopsy. It also includes the following steps: Based on the sequencing data of the embryo sample to be tested, the chromosomal euploidy of the embryo sample to be tested is determined.
12. The method of claim 11, wherein the method is a preimplantation genetic diagnosis method. The process of determining the chromosomal euploidy of the embryo sample based on its sequencing data includes: Based on the preset window length, the reference genome is divided into multiple consecutive equal-length intervals; Based on the sequencing sequence data of the embryo samples to be tested, the number of aligned sequencing sequences within each reference genomic region is counted; Based on the number of aligned sequencing sequences within each reference genome region, the copy number prediction result for each reference genome region is determined; Calculate the degree of difference in copy number prediction results between adjacent reference genome intervals, and merge adjacent intervals with a degree of difference below a preset merging threshold; The chromosome euploidy test results of the embryo sample to be tested are determined based on the merged intervals.
13. A preimplantation genetic testing system, employing the method of claim 1, characterized in that, include: The sequencing sequence data acquisition unit is used to acquire sequencing sequence data of the embryo sample to be tested and its paternal and maternal samples. The sequencing sequence data includes single-molecule long-read sequencing sequences of the corresponding samples obtained using single-molecule long-read sequencing technology. The parental haplotype phasing unit is used to determine the haplotype information of each parent based on the sequencing sequence data of the paternal sample and the sequencing sequence data of the maternal sample. The parental haplotype origin determination unit is used to filter the information of each parental haplotype and the sequencing sequence data of the embryo sample to be tested, determine the common single nucleotide polymorphism sites and corresponding base types between the sequencing sequence of the embryo sample to be tested and the corresponding parental haplotype, calculate the ratio of the number of common single nucleotide polymorphism sites with the same base type to the total number of common single nucleotide polymorphism sites, obtain the similarity between the sequencing sequence of the embryo sample to be tested and the corresponding parental haplotype, and determine the parental haplotype origin of each sequencing sequence of the embryo sample to be tested in the target region based on the similarity. The embryo haplotype genetic origin determination unit is used to determine the haplotype genetic origin of the embryo sample to be tested within the target region based on the determined parental haplotype origin.
14. A system for preimplantation genetic diagnosis of an embryo as claimed in claim 13, further comprising include: A preimplantation single-gene genetic disease / chromosomal structural rearrangement detection unit is used to determine the parental pathogenic chain haplotype carrying pathogenic variants and / or chromosomal structural rearrangements based on the parental haplotype information. In response to the fact that the haplotype genetic origin of the embryo sample to be tested includes the parental pathogenic chain haplotype, it is determined that the embryo sample to be tested carries a single-gene genetic disease and / or chromosomal structural rearrangement; In response to the fact that the haplotype of the embryo sample to be tested does not contain the parental pathogenic chain haplotype, it is determined that the embryo sample to be tested does not carry a single-gene genetic disease and / or chromosomal structural rearrangement.
15. A system for preimplantation genetic diagnosis of an embryo as claimed in claim 13, further comprising include: The preimplantation aneuploidy detection unit is used to divide the reference genome into multiple consecutive equal-length intervals according to a preset window length; Based on the sequencing sequence data of the embryo samples to be tested, the number of aligned sequencing sequences within each reference genomic region is counted; Based on the number of aligned sequencing sequences within each reference genome region, the copy number prediction result for each reference genome region is determined; Calculate the degree of difference in copy number prediction results between adjacent reference genome intervals, and merge adjacent intervals with a degree of difference below a preset merging threshold; The chromosome euploidy test results of the embryo sample to be tested are determined based on the merged intervals.
16. A computer program product, characterised in that, It includes a computer program, which, when executed by a processor, performs the steps of the method according to any one of claims 1-12.
17. An electronic device, characterized by It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the steps of the method according to any one of claims 1-12.
Citation Information
Patent Citations
Genetic detection method before embryo implantation based on three-generation sequencing
CN117230175A
Method for variation detection before embryo implantation
CN117925820A