Method for determining base types at predetermined sites in embryonic cell chromosomes and its application
The whole genome sequencing of the genomic DNA samples of both spouses was performed using single-tube long fragment reading technology (stLFR), and haploid linkage analysis was performed in combination with the embryonic cell sequencing results. This solved the problems of long time consumption and high complexity in existing technologies, and achieved efficient detection of single gene diseases and chromosomal abnormalities across the entire genome of the embryo.
Patent Information
- Application Number
- CN202080095705.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2040-05-22
AI Technical Summary
Existing preimplantation genetic testing methods are time-consuming and complex to operate, and are unable to effectively detect chromosomal abnormalities and small copy number variations. In addition, the diagnosis of serious lethal single-gene genetic diseases is difficult to achieve by relying on family samples.
The single-tube long fragment reading technology (stLFR) is used to perform whole-genome sequencing on the genomic DNA samples of both spouses to determine the haplotype information. The haplotype linkage analysis is then performed in combination with the embryonic cell sequencing results to directly detect whether the embryo carries pathogenic sites and chromosomal abnormalities, simplifying the operation process.
It has achieved the detection of single-gene diseases across the entire genome of the embryo, reduced the complexity of detection and the difficulty of sample collection, improved detection efficiency, and can accurately diagnose single-gene genetic diseases and chromosomal abnormalities.
Smart Images

Figure GPA0000325093160000111 
Figure GPA0000325093160000121 
Figure HPA0000325093180000011
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular to a method, a device and an application thereof for determining the base type of a predetermined site in an embryonic cell chromosome. Background Art
[0002] Preimplantation Genetic Testing (PGT) refers to the testing of embryos cultured in vitro using in vitro fertilization-embryo transfer (IVF) technology for single gene diseases and chromosomal abnormalities. Because genetic testing of embryos is performed before implantation, normal embryos can be selected for transplantation.
[0003] Monogenic diseases are genetic disorders caused by mutations in a pair of alleles (genes that control opposing traits at the same location on a pair of homologous chromosomes), and their transmission follows Mendel's laws of inheritance. According to the World Health Organization (WHO), there are currently over 10,000 known monogenic diseases, and the global prevalence of all monogenic diseases at birth is approximately 10 / 1,000. Furthermore, existing technologies lack effective treatments for most monogenic diseases, making effective testing essential to prevent the pregnancy and birth of children with monogenic diseases. Preimplantation testing for monogenic diseases allows embryo testing to select embryos free of the disease for transplantation, thereby preventing the inheritance of monogenic diseases.
[0004] Some studies have shown that approximately 50% of embryos formed by in vitro fertilization have chromosomal abnormalities, which may lead to early embryo loss, spontaneous abortion and stillbirth, and is one of the important reasons limiting the success of IVF. Therefore, it is necessary to test the chromosome number and structural abnormalities of in vitro cultured embryos, so as to select embryos with normal chromosomes and implant them into the uterus, in order to improve the patient's implantation success rate.
[0005] However, the current methods for preimplantation genetic testing still need to be improved, and detection methods with simple experimental operations, short processes and wide detection range are still to be developed. Summary of the Invention
[0006] This application is based on the inventor's discovery and understanding of the following facts and problems:
[0007] Single nucleotide polymorphism (SNP)-based haplotype analysis is currently the most commonly used method for determining whether an embryo has inherited a single-gene disorder. Techniques used in SNP-based haplotype linkage analysis include targeted region capture sequencing, which uses SNPs near the causative gene for haplotype linkage analysis, and karyomapping. STR-based linkage analysis also exists, but these require preliminary screening for relevant markers, resulting in a time-consuming and complex process. These techniques are no longer widely used as mainstream methods. Targeted region capture sequencing and haplotype linkage analysis use targeted region capture sequencing to identify SNPs within the causative gene and in upstream and downstream sequences for haplotype linkage analysis. Karyomapping utilizes microarrays to identify genome-wide SNPs, which are then used to construct haplotypes. These haplotype linkage analysis is then performed to determine whether the embryo has inherited the causative locus. Targeted region capture sequencing and haplotype linkage analysis can only detect specific single-gene disorders and cannot detect chromosomal abnormalities in embryos. Karyomapping technology can only rely on haplotypes and cannot directly detect pathogenic sites. It can only detect non-whole embryos and cannot detect small copy number variations.
[0008] In addition, it is worth noting that the above technologies all require linkage analysis of relatively complete family samples (such as Figure 1 ), where family samples include but are not limited to 1) a couple and their children, where the children are probands (the first individual to be diagnosed with the disease) or do not carry the disease-causing gene at all; 2) a couple and their father or mother, where the parents carry the disease-causing gene or are probands; 3) a couple and their siblings, where the siblings carry the disease-causing gene or are probands. However, collecting family samples is difficult, especially for certain serious, lethal single-gene genetic diseases. It is difficult to obtain a proband sample, making it difficult to construct a haplotype and perform linkage analysis, making it impossible to diagnose the single-gene genetic disease in the embryo.
[0009] Therefore, the present invention has developed a method for determining the base type of a predetermined site in the chromosome of an embryonic cell.
[0010] The single-tube long fragment reading technology (stLFR) used in the present invention can co-label long-fragment DNA, and after sequencing, the short-read information is restored to the corresponding long-fragment DNA information using the label. Through the long-fragment information read, the single sample can be directly haplotyped. In view of the current dependence on proband samples in the diagnosis of single-gene diseases before embryo implantation, the customization of detection methods, and the low throughput and high price of existing technologies, the present invention aims to develop an efficient and practical universal method for diagnosing single-gene genetic diseases in pre-implantation embryos. The present invention creatively uses the single-tube long fragment reading (stLFR) technology to perform whole-genome sequencing on the genomic DNA samples of couples undergoing pre-implantation diagnosis of embryos, obtains the haplotype information of the parents, and then analyzes the haplotypes of the pathogenic sites carried by the couple, so that the linkage relationship between the haplotype and the pathogenic site can be directly determined without the need for whole-genome sequencing information of other family samples or proband samples. The embryo biopsy cell sample is then subjected to ordinary whole-genome sequencing, and haplotype linkage analysis is used to determine whether the embryo has inherited the pathogenic site (the analysis process is as follows). Figure 2 This method can accurately diagnose designated single-gene diseases and all other known single-gene diseases, and can also accurately detect embryonic chromosomal abnormalities and gene copy number variations (CNVs).
[0011] To this end, based on the above findings, in a first aspect of the present invention, a method for determining the base type of a predetermined site in an embryonic cell chromosome is proposed. According to an embodiment of the present invention, the method comprises: (1) determining the linked haploid type block of the predetermined site based on the sequencing results of the embryonic parent, the linked haploid type block including the predetermined site and the base type of the predetermined site-linked site in the parent; (2) determining the sequence information of at least a portion of the embryonic genome based on the sequencing results of the embryonic cell, the at least a portion of the embryonic genome including the predetermined site; and (3) correcting the base type of the predetermined site in the embryonic cell based on the linked haploid type block of the predetermined site to obtain the base type of the predetermined site. According to the method of the embodiment of the present invention, based on the sequencing results of the embryo's parents, the parental haplotype information can be directly determined. Without the need for family and proband samples, the mutation information inspection and data collection of all single-gene disease gene sites in the entire genome of the embryo can be completed, and the base information of all single-gene genetic diseases in the entire genome of the embryo can be detected. There is no need to develop new methods for specific diseases, thereby reducing the difficulty of operation, reducing the complexity of the process, reducing the difficulty of sample collection, and improving the detection efficiency, thus preparing for scientific research.
[0012] According to an embodiment of the present invention, the above method may further include at least one of the following additional technical features:
[0013] According to an embodiment of the present invention, the embryonic cells are at the blastocyst stage, and the embryonic parents include at least one of the mother and father of the embryo. According to an embodiment of the present invention, only the genome sequencing data of the embryonic parents is required to complete the mutation information examination and data collection of all single-gene disease gene loci within the entire genome of the embryo, without requiring family or proband samples, thus reducing the difficulty of sampling. This allows for the diagnosis of certain serious lethal single-gene genetic diseases or the collection of certain serious lethal single-gene information. At the same time, the embryo can also be tested for chromosomal abnormalities based on the embryo's genomic information.
[0014] According to an embodiment of the present invention, the sequencing results of the embryonic cells in step (2) are derived from sequencing of 1 to 10 cells.
[0015] According to an embodiment of the present invention, the sequencing results of the embryonic parents are obtained by long-fragment sequencing.
[0016] According to an embodiment of the present invention, the sequencing read length of the long fragment sequencing is PE100 and / or PE150.
[0017] According to an embodiment of the present invention, the sequencing fragment length of the sequencing result of the embryonic cell is PE100 and / or PE150.
[0018] According to embodiments of the present invention, the method can be used to detect one to two blastomeres or three to eight trophoblast cells at the target cell or blastocyst stage of embryos cultured in vitro for three or five days. This method requires only a small number of embryonic cells for detection, and conventional whole-genome sequencing of embryonic cells is performed without the need for whole-genome amplification, reducing the probability of allele decoupling and, consequently, the occurrence of data errors.
[0019] According to an embodiment of the present invention, step (1) further includes: (1-1) performing genome single-tube long fragment reading technology (stLFR) sequencing on the blood sample of the embryonic parent; (1-2) based on the comparison of the sequencing result with the reference genome sequence, using GATK software, determining mutation information, the mutation information including at least one of SNP and insertion / deletion markers (Indel); (1-3) using Hapcut2 software, based on the mutation information obtained in step (1-2), assembling the haplotype of the embryonic parent; and (1-4) based on the predetermined site, selecting the linked haplotype block on the haplotype, optionally, the linked haplotype block corresponding to a length of 10,000 to 90 megabytes on the reference genome. According to an embodiment of the present invention, the stLFR technology is used to perform whole-genome sequencing on the genomic DNA samples of the parents to obtain the haplotype information of the father and the mother, and then the pathogenic loci carried by both parents are analyzed, so that the linkage relationship between the haplotype and the pathogenic locus can be directly determined, and the subsequent embryo sequencing results can be corrected to determine whether the embryo inherits the haplotype linked to the pathogenic locus. There is no need for whole-genome sequencing information of other samples or proband samples, which simplifies the operation process and improves the detection efficiency.
[0020] According to an embodiment of the present invention, the predetermined site is located in the COL1A1 gene.
[0021] In the second aspect of the present invention, the present invention proposes a method for determining CNV variation regions based on the embryonic genome, according to an embodiment of the present invention: (1) the embryonic genome reference sequence is divided into multiple windows, and the number of sequencing reads falling into each window is counted; (2) for each window, the starting point or the end point of the window is used as a dividing point, and based on the difference of two numerical sets consisting of the numerical values of the sequencing reads of the window on both sides of the dividing point, multiple initial breakpoints are determined; (3) based on the multiple initial breakpoints, multiple secondary windows are determined in the embryonic genome reference sequence, and the number of sequencing reads in each of the multiple secondary windows is determined; (4) based on the difference of two numerical sets consisting of the number of sequencing reads of the secondary windows on both sides of the initial breakpoint, the final breakpoint position is determined to determine the CNV variation region. According to the method of the embodiment of the present invention, the merging of windows is not limited to two rounds of statistics, that is, after determining the secondary window, the difference of the sequencing read values is still counted according to the two windows on the left and right of the endpoint, and whether it is a true breakpoint is determined based on the significance of the statistical difference, and whether it is a deletion or a duplication is determined based on the numerical value size of the reading segment in the window interval, and the detection accuracy is determined based on the window size. According to the method of the embodiment of the present invention, the method can not only detect the base information of all single-gene genetic diseases in the entire genome of the embryo, but also does not require the development of new methods specifically for specific diseases, thereby reducing the difficulty of operation.
[0022] In a third aspect, the present invention proposes an apparatus for determining the base type of a predetermined site in an embryonic cell chromosome. According to an embodiment of the present invention, the apparatus comprises: a linked haploid type block determination module, for determining the linked haploid type block of the predetermined site based on the sequencing results of the embryonic parents, the linked haploid type block including the base types of the sites linked to the positioning point and the predetermined site in the parents; an embryonic cell sequence information determination module, the embryonic cell sequence information determination module being connected to the linked haploid type block determination module, for determining the sequence information of at least a portion of the embryonic genome based on the sequencing results of the embryonic cells, the at least a portion of the embryonic genome including the predetermined site; and a predetermined site correction module, the predetermined site correction module being connected to the embryonic cell sequence information determination module, for correcting the base type of the predetermined site in the embryonic cell based on the linked haploid type block of the predetermined site, so as to obtain the base type of the predetermined site. The device for determining the base type at a predetermined site in the chromosome of an embryonic cell according to an embodiment of the present invention is suitable for executing the method for determining the base type at a predetermined site in the chromosome of an embryonic cell as described above, thereby being able to effectively determine whether the embryo carries the genetic information characteristics of a single-gene genetic disease and whether it carries chromosomal abnormality characteristics, simplifying the operating procedures and steps, improving detection efficiency, and preparing for scientific research.
[0023] According to an embodiment of the present invention, the above device may also have the following additional technical features:
[0024] According to an embodiment of the present invention, the linked haplotype block determination module further includes: a long fragment sequencing unit for performing genome stLFR sequencing on the blood sample of the embryonic parent; a comparison unit for determining mutation information based on the comparison of the sequencing result with the reference genome sequence using GATK software, wherein the mutation information includes at least one of SNP and Indel; a haplotype construction unit for assembling the haplotype of the embryonic parent based on the mutation information using Hapcut2 software; and a linked haplotype block determination unit for selecting the linked haplotype block on the haplotype based on the predetermined site. Optionally, the linked haplotype block corresponds to a length of 10,000 to 90 megabytes on the reference genome.
[0025] In a fourth aspect of the present invention, the present invention proposes a device for determining CNV variation regions based on the embryonic genome. According to an embodiment of the present invention, the device includes: a window division module, which is used to divide the genome reference sequence into multiple windows and count the number of sequencing reads falling into each window; an initial breakpoint determination module, which is connected to the window division module, and is used to determine multiple initial breakpoints for each window, using the starting point or end point of the window as a dividing point, based on the difference between two value sets consisting of the numerical values of the number of sequencing reads of the window on both sides of the dividing point; a secondary window division module, which is connected to the initial breakpoint determination module, and is used to determine multiple secondary windows in the genome reference sequence based on the multiple initial breakpoints, and determine the number of sequencing reads in each of the multiple secondary windows; a CNV variation region determination module, which is connected to the secondary window division module, and is used to determine the final breakpoint position based on the difference between the two value sets consisting of the number of sequencing reads of the secondary windows on both sides of the initial breakpoint, so as to determine the CNV variation region.
[0026] In its fifth aspect, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of a method for determining the base type of a predetermined site in an embryonic cell chromosome. This method effectively implements the aforementioned method for determining the base type of a predetermined site in an embryonic cell chromosome, thereby effectively determining whether an embryo carries genetic information characteristic of a single-gene genetic disease and whether it carries chromosomal abnormalities.
[0027] In the sixth aspect of the present invention, the present invention proposes a computer device comprising a processor and a memory; wherein the processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to implement the method of determining the base type of a predetermined site in the chromosome of an embryonic cell.
[0028] In a seventh aspect of the present invention, a computer program product is provided. When the instructions in the computer program product are executed by a processor, the method of determining the base type of a predetermined site in an embryonic cell chromosome is performed.
[0029] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:
[0031] Figure 1 Inferring the linked haplotype of the pathogenic locus by relying on family samples according to an embodiment of the present invention;
[0032] Figure 2 A preimplantation diagnostic test process according to an embodiment of the present invention that does not require a family sample or a proband sample;
[0033] Figure 3 A schematic diagram of a process for determining the base type of a predetermined site in an embryonic cell chromosome according to an embodiment of the present invention;
[0034] Figure 4 Schematic diagram of a flow chart of a method for determining a linked haplotype block of a predetermined site according to an embodiment of the present invention;
[0035] Figure 5 Schematic diagram of a process for determining CNV variation regions based on the embryo genome according to an embodiment of the present invention;
[0036] Figure 6 is a block diagram of a device for determining base types at predetermined sites in embryonic cell chromosomes according to an embodiment of the present invention;
[0037] Figure 7 is a block diagram of a linked haplotype block determination module according to an embodiment of the present invention;
[0038] Figure 8 This is a block diagram of an apparatus for determining CNV variation regions based on the embryo genome according to an embodiment of the present invention.
[0039] Detailed Description of the Invention
[0040] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to be used to explain the present invention, but should not be understood as limiting the present invention.
[0041] Explanation of terms
[0042] Unless otherwise specified, the terms "first", "second", "third" and the like used in this article are for the purpose of distinction for the convenience of description and are not intended to imply or explicitly indicate any difference in order or importance between them. It also does not mean that the content defined by the terms "first", "second", "third" and the like consists of only one component.
[0043] In the present invention, unless otherwise specified or limited, the terms "installed," "connected," "connect," "fixed," etc. should be understood in a broad sense. For example, they can refer to fixed connection, detachable connection, or integration; mechanical connection, electrical connection; direct connection, or indirect connection through an intermediate medium; internal communication between two components, or interaction between two components, unless otherwise specified. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0044] It should be noted that the single nucleotide polymorphism (SNP) used in this article refers to a DNA sequence polymorphism caused by a variation of a single nucleotide at the genomic level.
[0045] It should be noted that the insertion / deletion marker (Indel) used in this article refers to the difference between two parents in the whole genome, where a certain number of nucleotides are inserted or deleted in the genome of one parent relative to the other parent.
[0046] It should be noted that the CNV analysis used in this article includes chromosomal microduplication / microdeletion, aneuploidy, polyploidy, uniparental disomy, mosaicism, and family linkage analysis, and is particularly focused on the detection and interpretation of complex chromosomal diseases / genomic diseases from a molecular genetics perspective, and the discovery of the relationship between copy number abnormalities and phenotypes.
[0047] On the one hand, the present invention provides a method for determining the base type of a predetermined site in an embryonic cell chromosome. According to an embodiment of the present invention, referring to Figure 3 , the method comprising:
[0048] S100, determining a linked haploid type block of the predetermined site based on sequencing results of the embryo's parents, wherein the linked haploid type block includes the predetermined site and base types of sites linked to the predetermined site in the parents;
[0049] S200, determining sequence information of at least a portion of the embryonic genome based on the sequencing results of the embryonic cells, wherein at least a portion of the embryonic genome includes the predetermined site; and
[0050] S300: Correcting the base type of the predetermined site in the embryonic cell based on the linked haplotype block of the predetermined site to obtain the base type of the predetermined site.
[0051] According to the method of the embodiment of the present invention, based on the sequencing results of the father and mother of the embryo, the sequencing results are filtered and the data are compared to obtain SNP / Indel information, and the haploid assembly is performed using the SNP / Indel information to determine the parental haplotype information; based on the sequencing results of the embryo, the sequencing results are filtered, the chromosomes are split and the single chromosome is corrected, and the variation detection of SNP and Indel is performed based on the comparison results of the sequencing results with the reference genome sequence, the comparison results are annotated, the target region and the SNP information upstream and downstream are extracted, and the haplotype result is generated according to the family information. Then, a statistical analysis is performed based on the SNP / Indel in the target region of the embryo data and the upstream and downstream 1M region thereof, and based on the information of the haplotype of the pathogenic site linked to the single-tube long fragment data analysis of the couple before, it is judged whether the embryo inherits the haplotype linked to the pathogenic site, thereby judging whether the embryo inherits the pathogenic site.
[0052] According to an embodiment of the present invention, the embryonic cells are at the blastocyst stage, and the embryonic parent includes at least one of the mother and father of the embryo. According to an embodiment of the present invention, the embryonic cells can be blastomere-stage cells or blastocyst-stage trophoblast cells; the embryonic parent samples are obtained from blood samples of the embryonic father and mother, and genomic long-segment DNA is extracted.
[0053] According to an embodiment of the present invention, the sequencing results of the embryonic cells in step (2) are derived from 1 to 10 cell sequencing, and the sequencing results of the embryonic parents are obtained by long-fragment sequencing, the sequencing read length of the long-fragment sequencing is PE100 and / or PE150, and the sequencing fragment length of the sequencing results of the embryonic cells is PE100 and / or PE150.
[0054] According to an embodiment of the present invention, referring to Figure 4 , the step S100 further includes:
[0055] S110, performing genome single-tube long fragment read technology (stLFR) sequencing on the blood sample of the embryo's parents;
[0056] S120, based on the alignment of the sequencing results with the human reference genome sequence h37d5, using GATK software to determine mutation information, wherein the mutation information includes at least one of SNP and Indel;
[0057] S130, using Hapcut2 software, assembling the haplotype of the embryo parent based on the mutation information obtained in step S120; and
[0058] S140, based on the predetermined site, selecting the linked haploid type block on the haplotype, optionally, the linked haploid type block corresponds to a length of 10,000 to 90 megabytes on the reference genome.
[0059] According to an embodiment of the present invention, stLFR technology is used to construct a library for the extracted genomic long fragment DNA. The main steps are to fragment the DNA using transposase, add a joint, and then combine with the magnetic beads with labels, connect the labels to the DNA, and then perform PCR amplification after adding another joint to obtain a library read by a special single-tube long fragment, and sequence the built library. After obtaining the data, the molecular tags are split, and SOAPnuke software is used to filter the raw data with sequencing joints, too high a proportion of N bases, too high a proportion of A bases, and sequences with low sequencing quality, and statistical basic data. The filtered data is then compared using BWA software, and the comparison index is obtained using SAMtools software, and the bam file after sorting and de-duplication is obtained. The detection of SNPs and Indels is performed using GATK software based on the comparison result of the sequencing result with the reference genome sequence to obtain SNP / Indel information. Based on the SNP / Indel information, haploid assembly is performed using Hapcut2 software to obtain haplotype information. The haplotype of the pathogenic gene carried by the couple is then determined based on the pathogenic locus. Because pathogenic loci are generally heterozygous, the mutant and wild-type types correspond to different haplotypes. Therefore, the haplotype block linked to the pathogenic locus is determined based on the haplotype block where the pathogenic mutation is located.
[0060] According to an embodiment of the present invention, the predetermined site is located in the COL1A1 gene.
[0061] In the second aspect of the present invention, the present invention proposes a method for determining CNV variation regions based on the embryo genome. According to an embodiment of the present invention, referring to Figure 5 , the method comprising:
[0062] S1000, dividing the embryonic genome reference sequence into multiple windows, and counting the number of sequencing reads falling into each window;
[0063] S2000, for each window, using the start point or the end point of the window as a demarcation point, and determining a plurality of initial breakpoints based on differences between two sets of values consisting of the numbers of sequencing reads in the window on both sides of the demarcation point;
[0064] S3000, based on the multiple initial breakpoints, determining multiple secondary windows in the embryonic genome reference sequence, and determining the number of sequencing reads in each of the multiple secondary windows;
[0065] S4000: Determine a final breakpoint position based on the difference between two sets of values formed by the number of sequencing reads of the secondary window on both sides of the initial breakpoint, so as to determine the CNV variation region.
[0066] According to a specific embodiment of the present invention, chromosome abnormality analysis is carried out to the whole genome sequencing data of embryo, first SOAPnuke software is utilized to filter sequencing data, then BWA software is utilized to compare, and from comparison result, the sequence of unique comparison is picked out, after removing repetitive sequence, for subsequent analysis. Reference genome is randomly interrupted according to the length of sequence measured, simulated machine sample data is generated afterwards and it is re-aligned on reference genome, ensure that each window contains 100K read, overlapping region between adjacent windows contains 20K read, finally whole genome is divided into 131290 windows, other window lengths can also be selected according to the difference that falls into window read number. With single window as unit, the fluctuation value of the degree of depth in all windows of horizontal comparison is eliminated the larger window of fluctuation, afterwards GC content is pressed 0.01 interval, the depth value fluctuation in statistical interval, data are corrected, then batch correction is carried out according to the data after correction, eliminate batch difference. According to the filtered window information and the corresponding depth value, the detection is performed. First, the breakpoint coordinates on the genome are found to obtain the detection P value corresponding to each window. All P values are sorted and non-significant window positions are removed to obtain the initial breakpoint set B = {b1, b2, b3...}. For the breakpoints obtained in the above steps, the depth values in the intervals on the left and right sides of the adjacent breakpoints are statistically analyzed for two rounds to obtain a new P value corresponding to each breakpoint. Based on the above breakpoint P value, for a certain breakpoint, statistical tests are performed on the left and right breakpoint intervals respectively, and non-significant breakpoints are deleted in the loop, and the mean of the P value and depth value of each breakpoint interval is obtained. Finally, the significance of the breakpoint P value is used to determine whether it is a real breakpoint, the depth value is used to determine whether it is a deletion or duplication, and the detection accuracy is determined according to the size of the breakpoint interval. In this embodiment, the P value is 1e-10, the deletion threshold is 0.7, the duplication threshold is 1.3, and an interval greater than 16M is selected as the final copy number variation interval.
[0067] In the third aspect of the present invention, the present invention provides a device for determining the base type of a predetermined site in an embryonic cell chromosome, which is used to implement a method for determining the base type of a predetermined site in an embryonic cell chromosome, with reference to Figure 6The device includes: a linked haplotype block determination module 100, for determining the linked haplotype block of the predetermined site based on the sequencing results of the embryonic parents, the linked haplotype block including the base types of the sites linked to the positioning point and the predetermined site in the parents; an embryonic cell sequence information determination module 200, the linked haplotype block determination module 100 is connected to the embryonic cell sequence information determination module 200, for determining the sequence information of at least a portion of the embryonic genome based on the sequencing results of the embryonic cells, the at least a portion of the embryonic genome including the predetermined site; and a predetermined site correction module 300, the embryonic cell sequence information determination module 200 is connected to the predetermined site correction module 300, for correcting the base type of the predetermined site in the embryonic cell based on the linked haplotype block of the predetermined site, so as to obtain the base type of the predetermined site.
[0068] According to an embodiment of the present invention, referring to Figure 7 The linked haplotype block determination module 100 further includes: a long fragment sequencing unit 110, for performing genome stLFR sequencing on the blood sample of the embryo parent; an alignment unit 120, the long fragment sequencing unit 110 is connected to the alignment unit 120, for determining mutation information based on the alignment of the sequencing result with the reference genome sequence using GATK software, wherein the mutation information includes at least one of SNP and Indel; a haplotype construction unit 130, the alignment unit 120 is connected to the haplotype construction unit 130, for assembling the haplotype of the embryo parent based on the mutation information using Hapcut2 software; and a linked haplotype block determination unit 140, the haplotype construction unit 130 is connected to the linked haplotype block determination unit 140, for selecting the linked haplotype block on the haplotype based on the predetermined site. According to an embodiment of the present invention, the length of the linked haplotype block on the reference genome is 10,000 to 90 megabytes.
[0069] In the fourth aspect of the present invention, the present invention proposes a device for determining CNV variation regions based on the embryo genome. According to an embodiment of the present invention, reference is made to Figure 8The apparatus includes: a window division module 1000, configured to divide a genomic reference sequence into multiple windows and count the number of sequencing reads falling into each window; an initial breakpoint determination module 2000, connected to the window division module 1000, configured to determine, for each window, a plurality of initial breakpoints based on the difference between two sets of values consisting of the number of sequencing reads in the window on both sides of the dividing point, using the starting point or the end point of the window as a dividing point; a secondary window division module 3000, connected to the initial breakpoint determination module 2000, configured to determine, based on the multiple initial breakpoints, multiple secondary windows in the genomic reference sequence and determine the number of sequencing reads in each of the multiple secondary windows; and a CNV variation region determination module 4000, connected to the secondary window division module 3000, configured to determine a final breakpoint position based on the difference between two sets of values consisting of the number of sequencing reads in the secondary window on both sides of the initial breakpoint, so as to determine the CNV variation region.
[0070] In a fifth aspect, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of a method for determining base typing at predetermined sites in embryonic chromosomes and a method for determining CNV variant regions based on the embryonic genome. This method effectively implements the aforementioned method for determining base typing at predetermined sites in embryonic chromosomes, thereby effectively determining whether an embryo carries genetic information characteristic of a single-gene genetic disease and whether it carries chromosomal abnormalities.
[0071] In the sixth aspect of the present invention, the present invention proposes a computer device comprising a processor and a memory; wherein the processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to implement the method for determining the base type of a predetermined site in the chromosome of an embryonic cell and the method for determining the CNV variation region based on the embryonic genome.
[0072] The present invention provides a computer program product. When the instructions in the computer program product are executed by a processor, the method for determining the base type of a predetermined site in an embryonic cell chromosome is performed.
[0073] Those skilled in the art will understand that the features and advantages described above for the method of determining the base type of a predetermined site in the chromosome of an embryonic cell and the method of determining the CNV variation region based on the embryonic genome are all applicable to the computer-readable storage medium, computer device and computer program product, and will not be repeated here.
[0074] The scheme of the present invention will be explained below with reference to the examples. Those skilled in the art will understand that the following examples are only used to illustrate the present invention and should not be regarded as limiting the scope of the present invention. Where specific techniques or conditions are not specified in the examples, the techniques or conditions described in the literature in this field (for example, reference to "Molecular Cloning Experiment Guide" by J. Sambrook et al., translated by Huang Peitang et al., 3rd edition, Science Press) or in accordance with the product instructions are used. Reagents or instruments used without indicating the manufacturer are all conventional products that can be purchased commercially, for example, they can be purchased from MGI.
[0075] Example 1 Preimplantation Diagnosis of Osteogenesis Imperfecta
[0076] Experimental Overview: A couple, the wife suffering from osteogenesis imperfecta type I, caused by the c.769G>A mutation in the COL1A1 gene, and the husband appearing normal, sought to conceive a healthy child through preimplantation embryo testing. Through in vitro fertilization, the couple achieved a total of five embryos.
[0077] 5 ml of whole blood was drawn from the husband and wife, and genomic DNA greater than 40 kb in length was extracted using the QIAGEN MagAttract HMW kit. Library preparation and sequencing were performed using the stLFR technology, and the data were analyzed to obtain haplotype information linked to the pathogenic locus. After embryonic cell biopsy, whole-genome amplification products were amplified using the QIAGEN REPLI-g Single Cell Kit. The whole-genome amplification products of the embryonic cells were constructed and sequenced using conventional whole-genome library construction methods to obtain the whole-genome information of the embryo. By analyzing the data of the target gene COL1A1 and its upstream and downstream, it was determined whether the embryo inherited the haplotype linked to the pathogenic locus, thereby determining whether the embryo inherited the pathogenic locus. The whole-genome information was also used to analyze whether the embryo contained other single-gene pathogenic loci, chromosomal abnormalities, and CNVs.
[0078] Experimental samples: Whole blood of husband and wife, 5 embryos
[0079] Experimental steps:
[0080] 1) Family samples included 5 mL of whole blood from the husband and wife, and long-fragment genomic DNA was extracted using the QIAGEN MagAttract HMW kit according to the kit instructions;
[0081] 2) Use the MGIEasy stLFR library preparation kit to construct libraries of the husband and wife's long-fragment genomic DNA according to the kit instructions.
[0082] 3) After the stLFR library was constructed, sequencing was performed on a BGISEQ-500RS sequencer using the BGISEQ-500RS High-Throughput Sequencing Reagent Set (stLFR). The sequencing type was PE100+42.
[0083] 4) After the data is downloaded, analyze the data. First, use molecular barcode splitting software to split the molecular tags. Then filter low-quality reads (based on two criteria: the average quality value of the bases in the sequence and the proportion of the number of N bases contained. Reads that meet either or both of the following criteria are filtered out). The filtered sequences are aligned using BWA. After the alignment is complete, SNP / Indel detection is performed using GATK. Based on the detection results, haplotype construction is performed using Hapcut2 software.
[0084] 5) Based on the information of the pathogenic site c.769G>A on the wife's COL1A1 gene, it was determined that the mutant A base existed in the wife's haplotype M2, and thus the haplotype linked to the pathogenic site was determined to be M2 (see Table 1).
[0085] Table 1: COLIA1 genotype analysis results
[0086]
[0087] 6) Five embryos were biopsied at the blastomere stage. Whole-genome amplification (WGA) was performed on the biopsied cells using the QIAGEN REPLI-g Single Cell Kit. Libraries were constructed from the WGA products using the MGIEasy DNA Library Preparation Kit, following the kit instructions. After library construction, sequencing was performed using the BGISEQ-500RS, using PE100+10 sequencing.
[0088] 7) After the data was downloaded, it was filtered to remove low-quality reads (based on two criteria: the average quality value of the bases in the sequence and the proportion of the number of N bases contained. Reads that met either or both of the criteria, "base quality value less than or equal to 20" and "number of N bases greater than or equal to 5," were filtered out). The filtered sequences were aligned using BWA. After alignment, SNP / Indel detection was performed using GATK. The SNP and Indel information was used for haplotype linkage analysis.
[0089] 8) SAMtools was used to count the base types of SNPs and classify the haploid sources in the embryos based on the haplotypes of the couple that had been analyzed. The wife's M2 haplotype was known to be linked to the pathogenic locus. After analysis, it was found that embryos 1, 4, and 5 inherited the M1 haplotype that did not carry the pathogenic locus, while embryos 2 and 3 inherited the M2 haplotype that carried the pathogenic locus (see Table 1). Therefore, it was determined that embryos 1, 4, and 5 were normal embryos and did not carry the pathogenic gene, and embryos 2 and 3 were diseased embryos.
[0090] 9) Perform chromosome abnormality analysis on the embryo sample, select the unique matched sequence from the comparison result of the whole genome data of the embryo sample, remove the repeated sequence, divide the window, and correct the data depth in the window according to the GC content; then use the two-step method to determine the breakpoint position. The first step is to traverse the windows in the sample one by one, select the same number of windows on the left and right ends of the window for run test, and obtain the detection P value corresponding to each window. Sort all P values and remove non-significant window positions to obtain the initial breakpoint set B = {b1, b2, b3...}. Based on the breakpoints obtained in the above steps, perform two rounds of statistics on the depth values in the intervals on the left and right ends of the adjacent breakpoints to obtain a new P value corresponding to each breakpoint. Based on the above breakpoint P value, a certain breakpoint is statistically tested in the left and right breakpoint intervals, and insignificant breakpoints are deleted in the cycle. The mean P value and depth value of each breakpoint interval were obtained. Based on the detection P value of 1e-10, the deletion threshold of less than 0.7, and the duplication threshold of greater than 1.3, the interval greater than 16M was selected as the final copy number variation interval. The output results are shown in Table 2. It was found that only embryos 1 and 5 did not have chromosomal abnormalities.
[0091] Table 2 CNV analysis results
[0092]
[0093] 10) Ultimately, based on the results of the OI and chromosomal abnormality tests, it was determined that only embryos 1 and 5 did not have the OI pathogenic mutation and chromosomal abnormality, and one of these two embryos could be selected for transplantation.
[0094] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0095] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A method for determining the base type of a predetermined site in an embryonic cell chromosome, characterized in that: The method is for non-diagnostic purposes and comprises: (1) determining the linked haploid type block of the predetermined site based on the sequencing results of the embryo's parents, wherein the linked haploid type block includes the predetermined site and the base types of the sites linked to the predetermined site in the parents; (2) determining sequence information of at least a portion of the embryonic genome based on the sequencing results of the embryonic cells, wherein at least a portion of the embryonic genome includes the predetermined site; and (3) correcting the base type of the predetermined site in the embryonic cell based on the linked haplotype block of the predetermined site to obtain the base type of the predetermined site; The embryonic cells are in the blastomere or blastocyst stage, the embryonic parents are the mother and / or father of the embryo, and the sequencing results of the embryonic parents are obtained by stLFR long fragment sequencing.
2. The method according to claim 1, characterized in that The sequencing results of the embryonic cells in step (2) are derived from sequencing 1 to 10 cells.
3. The method according to claim 1, characterized in that The sequencing read length of the stLFR long fragment sequencing is PE100 and / or PE150.
4. The method according to claim 1, wherein The sequencing read length of the sequencing result of the embryonic cell is PE100 and / or PE150.
5. The method according to claim 1, wherein Step (1) further comprises: (1-1) performing genome stLFR sequencing on a blood sample of the embryo's parents; (1-2) Based on the alignment of the sequencing result with the reference genome sequence, using GATK software, determining mutation information of the sequencing result, wherein the mutation information includes at least one of SNP and Indel; (1-3) using Hapcut2 software to assemble the haplotype of the embryonic parent based on the mutation information obtained in step (1-2); and (1-4) Selecting the linked haplotype block on the haplotype based on the predetermined site.
6. The method according to claim 1, characterized in that The linked haplotype block corresponds to a length of 10,000 to 90 megabytes on the reference genome.
7. The method according to claim 1, characterized in that The predetermined site is located in the COL1A1 gene.
8. A device for determining the base type of a predetermined site in an embryonic cell chromosome, characterized in that: include: a linked haplotype block determination module, configured to determine the linked haplotype block of the predetermined site based on the sequencing results of the embryo's parents, wherein the linked haplotype block includes the predetermined site and the base types of the sites linked to the predetermined site in the parents; an embryonic cell sequence information determination module, the embryonic cell sequence information determination module being connected to the linked haplotype block determination module and configured to determine sequence information of at least a portion of the embryonic genome based on sequencing results of the embryonic cells, the at least a portion of the embryonic genome including the predetermined site; as well as A predetermined site correction module is connected to the embryonic cell sequence information determination module and is used to correct the base type of the predetermined site in the embryonic cell based on the linked haploid type block of the predetermined site, so as to obtain the base type of the predetermined site.
9. The device according to claim 8, characterized in that The linked haplotype block determination module further comprises: A long fragment sequencing unit, used for performing genome stLFR sequencing on the blood sample of the embryo's parents; An alignment unit, wherein the long fragment sequencing unit is connected to the alignment unit and is used to determine mutation information based on the alignment of the sequencing result with the reference genome sequence using GATK software, wherein the mutation information includes at least one of SNP and Indel; A haplotype construction unit, wherein the comparison unit is connected to the haplotype construction unit and is used to assemble the haplotype of the embryo parent based on the mutation information using Hapcut2 software; and A linked haplotype block determining unit, wherein the haplotype constructing unit is connected to the linked haplotype block determining unit and is configured to select the linked haplotype block on the haplotype based on the predetermined site.
10. The device according to claim 8, characterized in that The linked haplotype block corresponds to a length of 10,000 to 90 megabytes on the reference genome.
11. A device for determining CNV variation regions based on embryonic genome, characterized in that: include: A window division module is used to divide the embryonic genome reference sequence into multiple windows and count the number of sequencing reads falling into each window; an initial breakpoint determination module, the window division module being connected to the initial breakpoint determination module, configured to determine, for each window, a plurality of initial breakpoints based on the difference between two sets of values consisting of the numbers of sequencing reads in the window on either side of the demarcation point, using the start point or the end point of the window as a demarcation point; a secondary window division module, connected to the initial breakpoint determination module, for determining a plurality of secondary windows in the embryonic genome reference sequence based on the plurality of initial breakpoints, and determining the number of sequencing reads in each of the plurality of secondary windows; The CNV variation region determination module, the secondary window division module is connected to the CNV variation region confirmation module, and is used to determine the final breakpoint position based on the difference between the two numerical value sets composed of the number of sequencing reads of the secondary window on both sides of the initial breakpoint, so as to determine the CNV variation region.
12. A non-transitory storable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
13. A computer device, characterized in that: The method comprises a processor and a memory; wherein the processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to implement the method for determining the base type of a predetermined site in an embryonic cell chromosome as described in any one of claims 1 to 7.
14. A computer program product, when the instructions in the computer program product are executed by a processor, performs the method for determining the base type of a predetermined site in an embryonic cell chromosome according to any one of claims 1 to 7.
Citation Information
Patent Citations
Methods, systems, and computer-readable media for determining base information of predetermined regions in an embryonic genome
CN105051208B
Base linkage intensity analyzing method and gene typing method and system
CN110021351A
Genetic abnormality screening method for embryos
CN110628891A