A family-independent haplotype identification method

Through Hi-C technology and SNP detection, the structural variant chain haplotype of embryos is directly identified from a single sample, solving the detection problem of relying on family information in the prior art, and achieving efficient and low-cost embryo variant identification and blocking.

CN119339789BActive Publication Date: 2025-08-19YIKON GENOMICS SHANGHAI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411507561.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-25
Publication Date
2025-08-19
Estimated Expiration
2044-10-25

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently identify structural variant chain haplotypes in embryos without relying on family information, resulting in high application thresholds and high cost for eugenic and eugenic testing, and cannot meet the needs of a wide range of subjects.

Method used

Hi-C technology is used for low-deep sequencing, and local haplotypes and variant-linked haplotypes are constructed through chromosomal conformation capture technology. Combined with SNP detection, structural variation breakpoints and linkage SNPs are directly identified from a single sample, local chain SNP blocks are constructed, and the variant status of the embryo is inferred.

Benefits of technology

It realizes efficient and low-cost identification of embryonic structural variant chain haplotypes without relying on family information, expands the scope of subjects for variant blockade, and improves the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339789B_ABST
    Figure CN119339789B_ABST
Patent Text Reader

Abstract

The present invention provides a method for identifying linked haplotypes of structural variations, particularly in embryos. This method, independent of pedigree information, is particularly suitable for identifying embryos lacking pedigree information during pre-implantation screening and / or prenatal screening to determine whether they carry chromosomal structural variations, thereby implementing variant blocking. The present invention also provides a test product for implementing the aforementioned method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of genetic variation detection, and more particularly to a method for identifying structural variation-linked haplotypes, especially embryonic structural variation-linked haplotypes, that is family-independent. Background Art

[0002] In the biomedical field, chromosomal structural variation is a common chromosomal genetic defect. Variation types primarily include deletions, duplications, insertions, inversions, and translocations, with a frequency of occurrence potentially reaching 10% or higher. Structural variation influences gene function by altering gene structure, gene expression, and the three-dimensional structure of chromosomes, and is associated with many serious genetic diseases. Although structural variations such as balanced translocations, Robertsonian translocations, inversions, insertions, and deletions are more common in carriers and are typically asymptomatic, carriers are more likely to produce unbalanced gametes, making them more susceptible to infertility, recurrent miscarriage, and congenital malformations, autism, intellectual disability, and other disorders in their offspring (or their offspring's offspring). Therefore, the identification of structural variations and the prevention of their transmission to future generations (e.g., selecting and identifying embryos without the variation for embryo transfer) are of paramount importance in eugenics and assisted reproductive technology.

[0003] In current practice, embryos with copy number abnormalities are readily detected through whole-genome sequencing and microarrays. However, embryos with normal copy number but carrying variants such as inversions, translocations, and insertions are significantly more difficult to detect through direct testing. An effective identification strategy involves first identifying the haplotypes of the variant carrier in both parents, then testing and comparing the haplotypes of the embryos to determine whether they carry the same structural variant. Haplotype identification typically involves performing linkage analysis on the carrier's family information, including samples of the carrier's parents, offspring, or embryos with copy number imbalance. This allows for more accurate identification of SNPs linked to the variant, distinguishing between the variant haplotype and the normal haplotype, and selecting embryos with the normal haplotype that do not carry the variant for transfer. However, this approach has practical limitations, particularly in terms of genetic pedigree, such as the need to screen a certain number of embryos with copy number imbalance and / or sample family members. In addition to the long testing time and high cost, many subjects simply lack access to available family information, making it impossible to determine the variant-associated haplotype and thus achieving accurate and effective variant blocking during fertility.

[0004] Chromosome conformation capture technologies, such as Hi-C technology, use high-throughput sequencing and bioinformatics analysis methods to obtain the spatial relationship between chromosomes across the entire genome and have been widely used in genome assembly and structural variation detection. With the help of Hi-C technology and haplotype assembly technology, full-length haplotypes at the chromosome level can be assembled. In theory, variant carriers can be detected by constructing full-length haplotypes of the genome and constructing linked haplotypes, but this is very expensive and difficult to apply to clinical testing.

[0005] There are still gaps in the existing technology in the hope of better meeting the testing needs for the purpose of eugenics and good parenting. Summary of the Invention

[0006] The present invention uses Hi-C technology to construct local haplotypes and variant-linked haplotypes based on low-depth sequencing and a single sample, without the need for other embryos or carrier family information, thereby expanding the range of subjects for mutation blocking in embryo transplantation.

[0007] In a first aspect, the present invention provides a method for identifying linked haplotypes of structural variations, the method comprising the steps of:

[0008] a1. Obtain a biological sample from a structural variant carrier and extract genomic DNA;

[0009] a2. The extracted genomic DNA was constructed and sequenced by chromosome conformation capture technology;

[0010] a3. Align the genome sequence obtained in step a2 to the human reference genome;

[0011] a4. Determine one or more structural variation breakpoints based on the sequence aligned to the reference genome in step a3 and the characteristic data signals obtained by sequencing using chromosome conformation capture technology;

[0012] a5. For the sequence aligned to the reference genome in step a3, based on the one or more breakpoint information detected in step a4, detect SNPs upstream and downstream of the one or more breakpoints, and retain heterozygous SNP results;

[0013] a6. For the sequences aligned to the reference genome in step a3, extract their paired-end sequencing information and filter them according to their genomic location. The screening criteria are to retain paired-end sequences whose ends can be respectively aligned to the upstream and downstream breakpoint positions corresponding to the normal chromosome without mutation as the normal chain characteristic sequence, and to retain paired-end sequences that can be respectively aligned to the upstream and downstream breakpoint positions of the newly generated chromosome after mutation as the variant chain characteristic sequence. Sequences that do not meet the above criteria are removed.

[0014] a7. Detect SNPs in the sequences retained in step a6, and retain the homozygous SNP results;

[0015] a8. Take the intersection of the SNP sites retained in step a5 and the SNP sites retained in step a7, and compare the effective depths of each site in step a5 and step a7. Retain the site with an effective depth in step a5 greater than that in step a7 as the haplotype anchor point, where the homozygous base obtained in step a7 at this site is the characteristic base representing the structural variation haplotype, and the other of the heterozygous bases obtained in step a5 is the characteristic base representing the normal chromosome;

[0016] a9. Assemble and construct locally linked SNP blocks for the sequences aligned to the reference genome in step a3;

[0017] a10. In the locally linked SNP block obtained in step a9, locate the haplotype anchor point obtained in step a8. For the block containing the anchor SNP, use the base detected in step a8 as a reference and use the SNP allele linked to that base as the linked SNP allele representing the structural variant haplotype, and the other as the linked SNP allele representing the normal haplotype.

[0018] In some embodiments, the chromosome conformation capture technology in step a2 is Hi-C technology.

[0019] In some embodiments, the reference genome in step a3 is hg19, hg38, T2T, etc.

[0020] In some embodiments, the alignment is performed in step a3 using genome alignment tool software. In some preferred embodiments, the software is BWA, samtools, Bowtie2, HiC-Pro and / or hisat2.

[0021] In some embodiments, the breakpoints are detected in step a4 by Hi-C data analysis software. In some preferred embodiments, the software is juicer.

[0022] In some embodiments, the one or more structural variation breakpoints detected in step a4 include normal chain breaks and / or variant chain breaks. In some embodiments, the variant chain breaks include variant position breakpoints and / or variant sequence breakpoints.

[0023] In some embodiments, the software used to detect SNPs in step a5 is GATK or freebayes.

[0024] In some embodiments, the range upstream and downstream of the breakpoint in step a5 is preferably 10 Mbp upstream and downstream of a balanced translocation breakpoint, 10 Mbp downstream of a Robertson ectopic breakpoint, 5 Mbp upstream and downstream of an inversion breakpoint, 10 Mbp upstream and downstream of an insertion variant breakpoint, and 10 Mbp upstream and downstream of a deletion variant breakpoint.

[0025] In some embodiments, the range upstream and downstream of the breakpoint in step a6 is 10 Mbp upstream and downstream of the balanced translocation breakpoint, 10 Mbp downstream of the Robertson ectopic breakpoint, 5 Mbp upstream and downstream of the inversion breakpoint, 10 Mbp upstream and downstream of the insertion variant breakpoint, and 10 Mbp upstream and downstream of the deletion variant breakpoint.

[0026] In some embodiments, in step a9, haploid assembly analysis software is used to assemble and construct locally linked SNP blocks. In some preferred embodiments, the software is HapCUT or HapCUT2, more preferably HapCUT2.

[0027] In some embodiments, the structural variation is selected from: structural rearrangement and copy number abnormality (CNV), preferably structural rearrangement. In some specific embodiments, the structural rearrangement structural variation is selected from translocation (balanced translocation, unbalanced translocation, reciprocal translocation, non-reciprocal translocation, Robertsonian translocation and complex translocation), inversion (intra-arm inversion and inter-arm inversion), and insertion (forward insertion and inversion insertion). In some embodiments, the copy number abnormality structural variation is selected from deletion and duplication. In some embodiments, the structural variation can be a chromosome structure change in which sequence breakpoints exist in isochromosomes, ring chromosomes, etc.

[0028] In some embodiments, the sequencing is low-depth paired-end sequencing. In some embodiments, the sequencing depth is not less than 1X, not less than 2X, not less than 3X, not less than 4X, or not less than 5X. In some embodiments, the sequencing depth is not higher than 15X, not higher than 14X, not higher than 13X, not higher than 12X, not higher than 11X, not higher than 10X, not higher than 9X, not higher than 8X, not higher than 7X, not higher than 6X, or not higher than 5X. In some embodiments, the sequencing depth is not less than 3X, about 30M reads, wherein preferably, paired-ended sequencing is used. In some specific embodiments, the amount of sequencing data is about 40M reads. In some specific embodiments, the amount of sequencing data is about 30M reads to about 40M reads.

[0029] In a second aspect, the present invention further provides a method for identifying the linked haplotype status of a structural variation in an embryo, the method comprising the steps of:

[0030] b1. Sampling: obtain embryo biopsy cells and peripheral blood samples from both parents;

[0031] b2. Using the method for identifying structural variation-linked haplotypes according to the first aspect of the present invention, identifying structural variation-linked haplotypes of the structural variation carrier in the parent, and obtaining one or more breakpoint information associated with the variation and two allele information: a linked SNP allele representing the structural variation haplotype and a linked SNP allele representing the normal haplotype;

[0032] b3. Extract DNA from embryo biopsy cells and peripheral blood samples from the non-carrier parent, amplify it, create a library, and then sequence it;

[0033] b4. Align the genome sequences of the two samples measured in step b3 to the same human reference genome;

[0034] b5. For the sequence aligned to the reference genome in step b4, SNPs are detected upstream and downstream of the breakpoints according to the carrier's mutation-related breakpoint information;

[0035] b6. Intersecting the variant haplotype-linked SNP result obtained in step a10 of the method of the first aspect of the present invention with the SNP result obtained in step b5 according to their genomic positions to infer which allele of the carrier the embryo has inherited at that position. Thus, for each SNP site in the embryo, independently determining whether its genotype is consistent with the variant haplotype-linked SNP allele or the normal haplotype SNP allele;

[0036] b7. Count the number of SNP sites determined in step b6 upstream and downstream of the structural variation-associated breakpoint. If the number of SNP sites consistent with the variant haplotype-linked SNP allele is greater than the number of sites consistent with the normal haplotype, the embryo is considered to carry the structural variation haplotype obtained in step a10 of the first aspect, and is presumed to carry the variation. Otherwise, it is presumed not to carry the variation.

[0037] In some embodiments, the biopsied cells of the embryo in step b1 are single or multiple cells from the trophoblast of the blastocyst. In some preferred embodiments, the blastocyst is a healthy, high-quality blastocyst cultured to day 5-8 after oocyte retrieval.

[0038] In some embodiments, the sequencing in step b3 is performed by NGS sequencing after single-cell amplification and library construction using the MALBAC method.

[0039] In some embodiments, the range upstream and downstream of the breakpoint in step b5 is preferably 10 Mbp upstream and downstream of a balanced translocation breakpoint, 10 Mbp downstream of a Robertson ectopic breakpoint, 5 Mbp upstream and downstream of an inversion breakpoint, 10 Mbp upstream and downstream of an insertion variation breakpoint, and 10 Mbp upstream and downstream of a deletion variation breakpoint.

[0040] In some embodiments, the range upstream and downstream of the breakpoint in step b6 is preferably 10 Mbp upstream and downstream of a balanced translocation breakpoint, 10 Mbp downstream of a Robertson ectopic breakpoint, 5 Mbp upstream and downstream of an inversion breakpoint, 10 Mbp upstream and downstream of an insertion variation breakpoint, and 10 Mbp upstream and downstream of a deletion variation breakpoint.

[0041] In some embodiments, the inference method in step b6 is: when the non-carrier (i.e., the carrier spouse) carries a homozygous SNP, no matter whether the embryo is homozygous or heterozygous at this site, there must be an allele that is a homozygous base from the non-carrier, then the other allele in the embryo can be identified as an allele from the carrier; and when the embryo carries a homozygous SNP, even if the spouse's base is uncertain, it can be considered that one of the alleles in the homozygous SNP is inherited from the carrier.

[0042] In a third aspect, the present invention provides a detection product for implementing the aforementioned method of the present invention, which comprises one or more of the following modules:

[0043] (1) gDNA extraction module: used to extract genomic DNA from samples (e.g., cells);

[0044] (2) Chromosome conformation capture module: used to fix the chromosome conformation in the sample, that is, to fix the interactions between spatially contacted or adjacent DNA fragments and / or with proteins, for example, by covalent cross-linking, so that they can be detected in subsequent steps;

[0045] (3) Sequencing library construction module: used to generate whole genome DNA sequencing library;

[0046] (4) Sequencing module: used for low-depth paired-end sequencing of the sequencing library;

[0047] (5) Alignment module: used to align the whole genome sequencing results generated by the sequencing module to the reference genome;

[0048] (6) Structural variation assessment module: used to assess the chromosomal structural variation of sample genomic DNA from sequencing data;

[0049] (7) SNP detection module: used to detect SNP sites and their genotypes from sequencing data;

[0050] (8) Typing module: used to compare and screen related SNP sites and identify haplotype-related linked SNP alleles.

[0051] In some preferred embodiments, the chromosome conformation capture module and / or sequencing library construction module is based on Hi-C technology.

[0052] In some preferred embodiments, the detection product is a detection system. In some preferred embodiments, the detection system includes a computer and / or storable software. In some preferred embodiments, the computer is connected to a sequencer.

[0053] In a fourth aspect, the present invention provides the use of the identification method / detection method according to the present invention, or the detection product according to the present invention for pre-embryo implantation screening and / or prenatal screening, or for the preparation of a product for pre-embryo implantation screening and / or prenatal screening. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 Shown is a spatial contact matrix obtained after genomic HiC sequencing of a sample from a subject carrying a balanced chromosomal translocation.

[0055] Figure 2( Figure 2a and Figure 2b ) shows the results of genotyping of haplotype-related linked SNP sites for both parents (one of whom carries a balanced translocation), as well as the results of genotyping of embryonic cells of their candidate offspring, where Figure 2a and Figure 2b A portion of representative SNP sites are shown.

[0056] Figure 3 Figure 3a and Figure 3b ) shows the results of genotyping of haplotype-related linked SNP sites for both parents (one of whom carries Robertson translocation), as well as the results of genotyping of embryonic cells of their candidate offspring, where Figure 3a and Figure 3b A selection of representative SNPs are shown. Since the short arm of the chromosome is lost after translocation, the SNPs are indicated as the distance downstream of the breakpoint.

[0057] Figure 4 Shown is a spatial contact matrix obtained after genomic HiC sequencing of a sample from a subject carrying a chromosomal inversion.

[0058] Figure 5 The results of genotyping of haplotype-associated linked SNP loci for both parents (one of whom carries the inversion) and the results of genotyping of embryonic cells of their candidate offspring are displayed.

[0059] Figure 6 Shown is a spatial contact matrix obtained after genomic HiC sequencing of a sample from a subject carrying a chromosomal insertion variant.

[0060] Figure 7Displays the results of haplotype-linked SNP genotyping for both parents (one carrying the insertion variant) and the genotypes of their candidate offspring's embryos. SNPs are indicated by the distance upstream and downstream of the insertion site.

[0061] Figure 8 Schematic diagram showing copy number balanced insertion variants and their corresponding deletion variants, as well as the identification of the normal chain and variant chain in the two types of variants.

[0062] Figure 9 Shown is a spatial contact matrix obtained after genomic HiC sequencing of a sample from a subject carrying a chromosomal insertion variant.

[0063] Figure 10 Displays the results of haplotype-linked SNP genotyping for both parents (one carrying the insertion variant) and the genotypes of their candidate offspring's embryos. SNPs are indicated by the distance upstream and downstream of the insertion site. Detailed Description of the Invention

[0064] definition

[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. See, for example, Lackie, DICTIONARY OF CELL AND MOLECULARBIOLOGY, Elsevier (4th ed. 2007); Sambrook et al., MOLECULAR CLONING, A LABORATORY MANUAL, Cold Spring Harbor Laboratory Press (Cold Spring Harbor, New York 1989).

[0066] Any methods, devices and materials similar or equivalent to those described herein can be used in the practice of the present invention. The definitions provided herein are intended to help understand certain terms frequently used herein and do not limit the scope of the present invention.

[0067] As used herein, the term "comprise" and its various variations, such as "include" and "comprising," when preceding a step or element, are used to indicate that the addition of other steps or elements is optional and non-exclusive. As used herein, it is understood that the term "comprise" or its various variations also encompasses solutions consisting of the recited steps or elements.

[0068] As used herein, the term "about" when referring to a specific numerical value refers to the typical error range for that numerical value known to those skilled in the art. It should be understood that this expression encompasses a reference to the specific numerical value itself. Thus, for example, reference to "about X" also encompasses a specific reference to the specific numerical value X itself.

[0069] In this article, the terms "chromosome" and "chromatin" are used interchangeably to refer to a complex of chromosomes comprising all or part of a cell's genome. The genome of a cell is generally characterized by its karyotype, which is a collection of all chromosomes comprising the cell's genome. The genome of a cell may comprise one or more chromosomes. In humans, each chromosome has a short arm (called "p") and a long arm (called "q"). Each chromosome arm is divided into regions or bands, which can be seen in conventional karyotype analysis using a microscope. Chromosome bands are labeled p1, p2, p3, etc., and are counted from the centromere toward the telomere. The higher resolution sub-bands within the band are sometimes also used to identify regions in the chromosome. Sub-bands are also numbered from the centromere toward the telomere. Information about chromosome bands and chromosome nomenclature can be found in Strachan, T. and Read, AP 1999. Human Molecular Genetics, 2nd edition New York: John Wiley & Sons, pp. 37-39.

[0070] In this article, term " allele " or " allele " usually means the diverse candidate DNA sequence dna (alternative form) on the identical entity locus in the dna segment (for example, on homologous chromosomes).In diploid cells or organisms, two alleles (identical (homozygous) or different (heterozygous)) of a particular gene occupy corresponding locus on a pair of homologous chromosomes." difference " between different alleles can correspond to the single nucleotide difference (for example, SNP) of the sequence on the entity locus, also can correspond to many groups of differences, for example, many groups of differences of linkage, for example haploid (type). Conceptually, allele can mean the dna sequence dna that differs between a plurality of identical entity loci found on the homologous chromosomes in single cell or individual organism, also can mean or differ on the dna sequence dna (allelic variant) on the identical entity locus in a plurality of cells or organisms, if not otherwise specified, mainly relate to the former in this article.

[0071] As used herein, the term "chromosome structural variation", also referred to as "chromatin structural variation", refers to the structural differences in the chromosomes of an individual organism relative to the chromosomes of other individuals within the same species. The structural differences can be structural rearrangements and copy number variations of chromosomes; and encompass chromosome structural changes of various sizes, for example, 1-5M, 5-10M, more than 10M, or even a large portion of an individual chromosome, such as half, one-third, or three-quarters of the structure. Non-limiting examples of types of chromosome structural variation include translocations, reciprocal translocations, non-reciprocal translocations, balanced translocations, unbalanced translocations, complex translocations, inversions, insertions, deletions, duplications, repeat amplifications, and combinations of two or more of the same or different types of variation.

[0072] As used herein, the term "chromosomal structural rearrangement" or "structural rearrangement" refers to a change in the order and position of DNA sequences on a chromosome. A structural rearrangement can be a change in the position of a chromosomal DNA sequence, such as a translocation; a change in orientation, such as an inversion; or a change in type, such as the insertion of a nonhomologous chromosome segment. A structural rearrangement can be an interchromosomal rearrangement or an intrachromosomal rearrangement.

[0073] In this article, term " translocation " refers to chromosome breakage and its whole or part is reconnected on different chromosome segments, for example, between non-homologous chromosomes and / or between two or more positions on the same chromosome, the chromosome segment exchange. Translocation usually causes the abnormal approach of two chromosome regions that are not adjacent to each other. Depend on the fracture and the reconnection position of translocation, translocation can not have influence on gene expression, or can affect the expression of single gene, or can affect the expression of multiple genes. The type of translocation includes, but is not limited to, reciprocal translocation, non-reciprocal translocation, Robertsonian translocation, balanced translocation, unbalanced translocation.

[0074] As used herein, the term "balanced translocation" refers to a translocation in which no increase or decrease in genetic material occurs during the exchange of genetic material.

[0075] As used herein, the term "unbalanced translocation" refers to a form of translocation in which a loss of genetic material occurs during the exchange of genetic material.

[0076] As used herein, the term "reciprocal translocation" refers to the exchange of segments between two chromosomes. Carriers of balanced reciprocal translocations are often phenotypically normal, but may develop unbalanced chromosomal translocations in their gametes, leading to infertility, miscarriage, or phenotypic abnormalities in offspring.

[0077] As used herein, the term "nonreciprocal translocation" refers to the unidirectional transfer of genetic material from one chromosome to another nonhomologous chromosome.

[0078] As used herein, the term "Robertsonian translocation" or "Robertsonian translocation" refers to a translocation caused by the breakage of two chromosomes with acrocentric centromeres at or near the centromere. The long arms of the two chromosomes rejoin to form the translocated chromosome, while the two short arms form a minichromosome that is often lost. Robertsonian translocations typically occur between chromosomes 13, 14, 15, 21, and 22 in humans, with an incidence of approximately 1.23 per 1,000 individuals and accounting for approximately 2-3% of infertile individuals. Carriers of Robertsonian translocations have only 45 chromosomes but often have a normal phenotype.

[0079] As used herein, the term "complex translocation" refers to a translocation involving two or more chromosomes and two or more breakpoints. Complex translocations can be balanced or unbalanced.

[0080] As used herein, the term "inversion" refers to a structural rearrangement within a chromosome that changes the direction of the DNA sequence within a chromosome. In this process, the broken fragments produced by two breaks of the same chromosome reconnect after being reversed 180 degrees. Inversions can be intra-arm inversions and inter-arm inversions. Intra-arm inversions refer to inversions that occur on the same arm (e.g., long arm or short arm) of a chromosome. Inter-arm inversions refer to inversions of chromosome segments that include the centromere region.

[0081] In this article, the term "insertion" refers to the process by which a DNA segment on a chromosome moves from its original location to another location on the same or different chromosome. Insertion mutations typically do not result in an increase or decrease in genetic material; therefore, the genetic material of insertion mutations is usually balanced.

[0082] As used herein, the term "breakpoint" refers to a site or region of a chromosome where a chromosome break occurs during a structural change such as a translocation or inversion.

[0083] As used herein, the term "copy number variation" (CNV) refers to a change in the number of copies of a chromosomal DNA sequence. A CNV can be a copy number variation of an entire chromosome or a chromosome segment, such as an increase or decrease in the copy number.

[0084] As used herein, the term "read" or "read(s)" refers to a short nucleotide sequence generated by any sequencing process described herein or known in the art. In paired-end sequencing, a pair of reads (also called a read pair) is generated from the two ends of the sequenced nucleic acid. Pair-end reads obtained by paired-ended sequencing can be abbreviated as PEreads.

[0085] In this article, the term "read length" refers to the nucleotide length of the read. The length of a sequence read is often associated with a specific sequencing technology. For example, next-generation sequencing (NGS) can provide sequence reads ranging from tens to hundreds of base pairs (bp) in length.

[0086] As used herein, the term "haplotype" or "haplotype" refers to a combination of two or more genetic variant sites that coexist on a single chromosome. Typically, the combination is inherited entirely from a set of alleles from one parent of both parents. In a haplotype, the probability of the simultaneous occurrence of specific allele types at the two or more variant sites is higher than the frequency of random occurrence. This relationship is called linkage disequilibrium, which is based on the non-random association between alleles, i.e., alleles at different loci are not inherited independently, but exhibit a certain degree of linkage. Linkage disequilibrium on the same chromosome is a haplotype.

[0087] In this article, "sequencing data volume" is expressed in units of reads. For paired-end sequencing, a pair of PE reads is counted as one read for the purpose of data calculation. For example, if 20 million reads are obtained through pair-ended sequencing, the sequencing data volume can be expressed as 20M.

[0088] In this article, the term "sequencing depth" refers to the ratio of the total number of bases sequenced to the size of the genome being sequenced. Sequencing depth can be calculated using the following formula:

[0089] Sequencing depth = read length × total number of reads / measured genome sequence length.

[0090] In this article, unless otherwise specified in the context, the terms "detecting SNPs", "SNP detection" and "SNP calling" have the same meaning known in the art and can be used interchangeably, that is, identifying single nucleotide polymorphisms (SNPs) at various positions on the reference genome in the sequencing result data based on the sequencing information, and the output results include the genomic coordinates of the SNP, the homozygous / heterozygous genotype and the corresponding base information.

[0091] As used herein, the term "joined DNA fragment" or "joined fragment" in connection with chromosome conformation capture refers to a joined DNA fragment containing an internal junction point, formed by end-joining adjacent digested DNA fragments in a covalently cross-linked DNA / protein complex following enzymatic digestion during the chromosome conformation capture process. Accordingly, as used herein, the term "unjoined DNA fragment" or "unjoined fragment" refers to an enzymatically digested DNA fragment lacking an internal junction point, produced when the enzymatically digested DNA fragments do not undergo end-joining. In conventional chromosome conformation capture techniques, it is generally believed that distinguishing and separating joined DNA fragments from unjoined fragments is necessary, and this distinction and separation is typically performed by biotin-labeling the ends of the digested DNA and removing the terminal biotin labels of the unjoined fragments after ligation. However, the present inventors surprisingly discovered that, as shown in the Examples herein, according to some embodiments of the present method, when constructing a sequencing library using the enzymatic digestion products, it is not necessary to distinguish and remove such unjoined fragments from the joined fragments, and good sequencing library quality can still be obtained. It should be understood herein that “connected fragments” also encompass derivative fragments generated from the fragments during library construction (e.g., small fragments generated by mechanical or enzymatic action), as long as the derivative fragments retain the internal connection sites characteristic of the connected fragments from which they are derived. Similarly, “unconnected fragments” also encompass derivative fragments generated from the fragments during library construction (e.g., small fragments generated by mechanical or enzymatic action during library construction), as long as the derivative fragments retain at least one enzyme digestion end characteristic of the unconnected fragments from which they are derived, or enzyme digestion ends that are correspondingly modified, labeled, and / or repaired in the case of end modification, labeling, and / or end repair. Therefore, as will be understood by those skilled in the art, the term “unconnected fragments” encompasses the terminal small DNA fragments and their derivatives that retain the enzyme digestion ends generated by interrupting the fragment, but does not encompass internal small DNA fragments or their derivatives generated by interrupting the fragment.

[0092] In this article, chromosome conformation capture technology, such as Hi-C technology, refers to a technology based on chromosome conformation capture (chromatin conformation capture) and sequencing (such as NGS) to capture the interaction information between chromosomal DNA fragments in cells. High-throughput chromosome conformation capture technology (such as Hi-C technology) digests and reconnects chromosome fragments that are close in space and performs high-throughput sequencing on them to determine the spatial interactions between different sites on the chromosome. This technology typically includes two stages, i.e., the construction of a sequencing library in the first stage and the sequencing and bioinformatics analysis in the second stage. In the first stage, sequencing library construction generally involves: by a chromosome conformation capture method, chromosomal DNA / protein is covalently cross-linked, nuclease digested and reconnected to capture chromosomal DNA fragments that interact in space, and generate a sequencing library. In the second stage, the library is sequenced to obtain chromosome conformation capture sequencing data, which is also referred to as sequencing data in this article.

[0093] In this article, spatial contact matrix, sometimes also referred to as genome interaction matrix or Hi-C contact matrix in this article, refers to a two-dimensional matrix for representing the interaction between DNA sequences in a genome based on chromosome conformation capture sequencing data, which can generally be used to study the spatial organization of chromosome three-dimensional structure and genome. The rows and columns of the matrix represent the chromosome position intervals on chromosome coordinates; the matrix elements are the number of paired end sequence fragments falling into the corresponding row and column interaction intervals (in this article, also referred to as "contact positions"), also referred to as contact intensity / density or interaction frequency (IF). The spatial contact matrix can be a numerical matrix or can be presented in the form of a visual spatial contact matrix diagram (contact map). In this spatial contact matrix diagram, the change in contact intensity can be reflected by the gradual change of color, thus presenting color blocks of different color intensities on matrix positions with different contact intensities. In some embodiments, the spatial contact matrix is implemented or generated by computer software. In some embodiments, the computer software is juicer (see Durand NC et al. Juicer Provides a One-Click System for Analyzing Loop-Resolution Hi-C Experiments. Cell Syst. 2016 Jul; 3(1): 95-8.).

[0094] As used herein, the term "module" refers to a reagent, a set of reagents, a component, a set of components, and / or a device or system that can be used to implement the function of the module. It should be understood that the composition of the module is not subject to the constraints of a specific physical form, as long as it can achieve the desired function. Depending on the intended function, a module can be a combination of reagents and / or devices that implement the function; or it can be a software object or routine (for example, as an independent thread) that is executed on a single computing system (for example, a computer program, a tablet computer (PAD), one or more processors). For example, the chromosome conformation capture module according to the present invention can be composed of reagents or a set of reagents for implementing covalent cross-linking, enzymatic digestion and reconnection of chromosome DNA / protein, and optionally a device that can be used to ensure that this process is implemented under specific conditions. For another example, the SNP detection module of the present invention can be a program stored on a computer-readable medium, which contains a computer program logic or code portion for implementing the SNP detection.

[0095] Method of the present invention

[0096] The present invention improves the existing method mainly from three aspects:

[0097] 1. This method simultaneously detects precise structural variant breakpoints during the haplotype construction step for variant carriers, eliminating the need to rely on karyotype or CNV inference to obtain breakpoint location information. This simplifies the analysis process, enabling variant detection and genetic disruption to be completed in a single test. Furthermore, the highly accurate breakpoints obtained by this approach reduce the possibility of inaccurate results due to recombination near the breakpoints.

[0098] 2. This method does not require family information to infer haplotypes, which expands the scope of application and reduces the probability of detection failure.

[0099] 3. In the conventional variant blocking process, when there is no family information, third-generation sequencing and other technologies are usually used to complete haplotype construction. In comparison, this method greatly reduces the detection cost and can obtain more effective SNP sites than third-generation sequencing.

[0100] In one embodiment of the present invention, a method for identifying structural variation-linked haplotypes is provided, the method comprising one or more of the following steps:

[0101] a1. Obtain a biological sample from a structural variant carrier and extract genomic DNA;

[0102] a2. The extracted genomic DNA was constructed and sequenced by chromosome conformation capture technology;

[0103] a3. Align the genome sequence obtained in step a2 to the human reference genome;

[0104] a4. Determine one or more structural variation breakpoints based on the sequence aligned to the reference genome in step a3 and the characteristic data signals obtained by sequencing using chromosome conformation capture technology;

[0105] a5. For the sequence aligned to the reference genome in step a3, based on the one or more breakpoint information detected in step a4, detect SNPs upstream and downstream of the one or more breakpoints, and retain heterozygous SNP results;

[0106] a6. For the sequences aligned to the reference genome in step a3, extract their paired-end sequencing information and filter them according to their genomic location. The screening criteria are to retain paired-end sequences whose ends can be respectively aligned to the upstream and downstream breakpoint positions corresponding to the normal chromosome without mutation as the normal chain characteristic sequence, and to retain paired-end sequences that can be respectively aligned to the upstream and downstream breakpoint positions of the newly generated chromosome after mutation as the variant chain characteristic sequence. Sequences that do not meet the above criteria are removed.

[0107] a7. Detect SNPs in the sequences retained in step a6, and retain the homozygous SNP results;

[0108] a8. Take the intersection of the SNP sites retained in step a5 and the SNP sites retained in step a7, and compare the effective depth of each site in step a5 and step a7. The site with an effective depth greater than that in step a7 in step a5 is retained as the haplotype anchor point, wherein at this site, the homozygous base obtained in step a7 is the characteristic base representing the structural variant haplotype, and the other of the heterozygous bases obtained in step a5 is the characteristic base representing the normal chromosome;

[0109] a9. Assemble and construct locally linked SNP blocks for the sequences aligned to the reference genome in step a3;

[0110] a10. In the locally linked SNP block obtained in step a9, locate the haplotype anchor point obtained in step a8. For the block containing the anchor SNP, use the base detected in step a8 as a reference and use the SNP allele linked to that base as the linked SNP allele representing the structural variant haplotype, and the other as the linked SNP allele representing the normal haplotype.

[0111] In some embodiments, the chromosome conformation capture technology in step a2 is Hi-C technology.

[0112] In some embodiments, the reference genome in step a3 is hg19, hg38, T2T, etc.

[0113] In some embodiments, the alignment is performed in step a3 using genome alignment tool software. In some preferred embodiments, the software is BWA, samtools, Bowtie2, HiC-Pro and / or hisat2.

[0114] In some embodiments, the breakpoints are detected in step a4 by Hi-C data analysis software. In some preferred embodiments, the software is juicer.

[0115] In some embodiments, the one or more structural variation breakpoints detected in step a4 include normal chain breaks and / or variant chain breaks. In some embodiments, the variant chain breaks include variant position breakpoints and / or variant sequence breakpoints.

[0116] In some embodiments, the software used to detect SNPs in step a5 is GATK or freebayes.

[0117] In some embodiments, the range upstream and downstream of the breakpoint in step a5 is preferably 10 Mbp upstream and downstream of a balanced translocation breakpoint, 10 Mbp downstream of a Robertson ectopic breakpoint, 5 Mbp upstream and downstream of an inversion breakpoint, 10 Mbp upstream and downstream of an insertion variant breakpoint, and 10 Mbp upstream and downstream of a deletion variant breakpoint.

[0118] In some embodiments, the range upstream and downstream of the breakpoint in step a6 is 10 Mbp upstream and downstream of the balanced translocation breakpoint, 10 Mbp downstream of the Robertson ectopic breakpoint, 5 Mbp upstream and downstream of the inversion breakpoint, 10 Mbp upstream and downstream of the insertion variant breakpoint, and 10 Mbp upstream and downstream of the deletion variant breakpoint.

[0119] In some embodiments, in step a9, haploid assembly analysis software is used to assemble and construct locally linked SNP blocks. In some preferred embodiments, the software is HapCUT or HapCUT2, more preferably HapCUT2.

[0120] In some embodiments, the structural variation is selected from: structural rearrangement and copy number abnormality (CNV), preferably structural rearrangement. In some specific embodiments, the structural rearrangement structural variation is selected from translocation (balanced translocation, unbalanced translocation, reciprocal translocation, non-reciprocal translocation, Robertsonian translocation and complex translocation), inversion (intra-arm inversion and inter-arm inversion), and insertion. In some embodiments, the copy number abnormality structural variation is selected from deletion and duplication. In some embodiments, the structural variation can be a chromosome structure change with sequence breakpoints such as isochromosomes and ring chromosomes.

[0121] In some embodiments, the sequencing is a low-depth double-end sequencing. In some embodiments, the sequencing depth is not less than 1X, not less than 2X, not less than 3X, not less than 4X, or not less than 5X. In some embodiments, the sequencing depth is not higher than 15X, not higher than 14X, not higher than 13X, not higher than 12X, not higher than 11X, not higher than 10X, or not higher than 9X. In some embodiments, the sequencing depth is not less than 3X, about 30M reads, wherein preferably, pair-ended sequencing is used. In some specific embodiments, the sequencing data amount is about 40M reads.

[0122] Sequencing

[0123] sample

[0124] The methods according to the present invention can be applied to any sample containing nucleated cells. Such samples can be obtained from suitable subjects, such as eukaryotic individuals and mammalian individuals. In some embodiments, the subject is a mammal, preferably a human. In some embodiments, the subject is a non-diseased or diseased individual. In some embodiments, the subject is a human individual to be subjected to prenatal screening.

[0125] The sample can be a sample directly isolated or obtained from a single or multiple subjects or their parts. Examples of samples include, but are not limited to, body fluids or tissues from mammalian individuals, including but not limited to blood or blood products (e.g., serum, plasma, platelets, buffy coat, etc.), umbilical cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, spinal fluid, lavage fluid (e.g., lung, stomach, peritoneum, catheter, ear, arthroscopic), biopsy samples, isolated cells (lymphocytes, placental cells, stem cells, bone marrow-derived cells, embryonic cells such as blastocyst trophoblast cells), or a combination thereof.

[0126] In some embodiments or steps, the sample is a mammalian nucleated cell, body fluid or tissue sample. In some preferred embodiments or steps, the sample is an isolated peripheral blood cell sample. In some preferred embodiments or steps, the sample is an isolated peripheral blood mononuclear cell (PBMC) sample. In some preferred embodiments or steps, the sample is a blastocyst trophoblast cell sample. In some embodiments or steps, the sample comprises 0.1x10 6 Up to 10x10 6 nucleated cells, for example, 0.5x10 6 1x10 6 2x10 6 3x10 6 4x10 6 5x10 66x10 6 7x10 6 8x10 6 9x10 6 10x10 6 In some embodiments or steps, the sample comprises 1x10 6 Up to 5x10 6 nucleated cells, for example, 1x10 6 Up to 3x10 6 nucleated cells, e.g., no more than 2 x 10 6 In some embodiments or steps, the sample contains a very small number of cells, for example, no more than 100 cells, no more than 90 cells, no more than 80 cells, no more than 70 cells, no more than 60 cells, no more than 50 cells, no more than 40 cells, no more than 30 cells, no more than 20 cells, no more than 10 cells, no more than 9 cells, no more than 8 cells, no more than 7 cells, no more than 6 cells, no more than 5 cells, no more than 4 cells, no more than 3 cells, no more than 2 cells, no more than 1 cell. Therefore, in some embodiments or steps, the present invention uses single-cell sequencing technology / platform.

[0127] In some embodiments, genomic DNA extracted from a sample is purified, for example, using magnetic beads. In some specific embodiments, the purification is performed before sequencing.

[0128] Library construction

[0129] Sequencing libraries can be established using methods known in the art.

[0130] In order to generate the fragment size suitable for sequencing library construction, a variety of mechanical, chemical and / or enzymatic methods can be used to fragment or shear the DNA in the enzyme digestion product to the desired length. Fragmentation can be performed by ultrasonic treatment, or by brief exposure to DNA enzymes or a mixture of one or more restriction enzymes, or transposases or nicking enzymes. Many enzyme shearing kits are commercially available, for example, from Takara. After fragmentation, the size of the fragments can be screened using methods known in the art; and sequencing adapters can be added.

[0131] In some embodiments, therefore, the method of the present invention comprises: fragmenting the DNA in the enzymatic digestion product according to the present invention and screening for fragments of 150bp-600bp (preferably 300-500bp) for library construction. Preferably, the fragment screening uses a two-step magnetic bead addition to recover fragments of appropriate size. The fragment screening can be performed by adjusting the magnetic bead ratio. In one embodiment, the DNA fragment screening is performed by adding magnetic beads to the fragmented DNA product in two steps at a 0.6x magnetic bead ratio and a 0.4x magnetic bead ratio.

[0132] Sequencing

[0133] Double-end sequencing methods known in the art can be used for sequencing. In the case of double-end sequencing, typically, the two ends of the nucleic acid fragment are sequenced with suitable read lengths that are long enough for each read (e.g., the reads at the two ends of the fragment template) to be mapped to the reference genome. For example, the read length of the double-end read can be from about 10 continuous nucleotides to about 500 continuous nucleotides, from about 10 continuous nucleotides to about 400 continuous nucleotides, from about 10 continuous nucleotides to about 300 continuous nucleotides, from about 50 continuous nucleotides to about 200 continuous nucleotides, from about 100 continuous nucleotides to about 200 continuous nucleotides, or from about 100 continuous nucleotides to about 150 continuous nucleotides.

[0134] In some embodiments, a paired-end sequencing method is used. In some embodiments, the sequencing depth is less than 6x, for example, the sequencing depth is about 0.1x, 0.3x, 0.5x, 0.8x, 1x, 1.5x, 2x, 2.5x, 3x, 3.5x, 4x, 4.5x, and 5x.

[0135] In some preferred embodiments, the sequencing method is MALBAC sequencing, see Chenghang Zong et al., Genome-Wide Detection of Single-Nucleotide and Copy-Number Variations of aSingle Human Cell. Science 338, 1622-1626 (2012).

[0136] Detection of chromosome structural variations

[0137] In this article, chromosome structural variations (structural variations) refer to structural changes in the genome. These changes can include large-scale rearrangements, inversions, insertions, translocations, duplications or deletions, as well as small-scale structural variation events, such as copy number abnormalities (CNVs, such as duplications or deletions) of 1M-50M. Since the primary structure is the basis of spatial structure, abnormalities in the primary structure of DNA often lead to changes in the 3D structure of chromosomes. Therefore, chromosome capture technology can be used to identify these chromosome structural variations.

[0138] In chromosome conformation capture sequencing data (e.g., Hi-C sequencing data), the interaction frequency between chromosomal regions varies with genomic distance in a power law distribution. When structural variation causes changes in genomic structure, two chromosomal regions that were originally far apart actually move closer to each other, causing their interaction frequency to increase, reaching the level of interaction between close-range chromosomal regions, which significantly deviates from the situation in normal chromosomes. For example, the deletion of a genomic fragment will cause the two ends of the deleted fragment to move closer in spatial position, resulting in an increase in the interaction frequency of the chromosomal regions flanking the deletion breakpoint. When the corresponding Hi-C reads are mapped to the reference genome, this increase in interaction frequency will be manifested in the spatial contact matrix heat map as a significantly darker color block in the interaction interval corresponding to the chromosomal regions on both sides of the breakpoint. Similarly, any other form of chromosomal structural change (including but not limited to duplication, insertion, translocation, inversion) will lead to an increase or decrease in the interaction frequency of the corresponding chromosomal regions, and the quantitative assessment of this change is a key indicator for detecting whether there is structural variation in chromatin.

[0139] After establishing a sequencing library and obtaining sequencing data according to the method of the present invention, a genomic interaction matrix (e.g., a Hi-C contact map, for example, using Juicer software) that records spatial contact information can be generated from the sequencing data according to methods known in the art to determine the type and location of structural variations and breakpoints on the sample to be tested.

[0140] In some cases, the processing process of chromosome conformation capture sequencing data includes: sequence mapping, data filtering, and data correction. The reference genome used for sequence mapping can be any one of human reference genomes GRCh37, GRCh38, T2T-CHM13, etc. For example, the PEreads sequence obtained by pair-ended sequencing can be compared with the reference genome to determine that only the paired reads that are uniquely matched to the reference sequence at both ends are obtained. After data filtering, the final valid data (i.e., valid contacts data) for contact analysis is obtained. The ratio of this valid contact data to all sequencing data is an important indicator for evaluating the quality of the chromosome conformation capture sequencing library. In some embodiments, the chromosome conformation capture library sequencing data obtained according to the inventive method has at least 50%, preferably at least 60% or 65% valid contacts.

[0141] In some embodiments, structural variation detection includes two main steps: (a) constructing a control reference system using healthy samples, and (b) performing variation detection on the test samples. In some embodiments, structural variation detection is performed as shown in Figure 2.

[0142] In some embodiments, detecting a variation in a sample to be tested comprises the following steps:

[0143] (1) Convert the sequencing data into a numerical matrix that records spatial contact information, such as using Juicer

[0144] (https: / / doi.org / 10.1016 / j.cels.2016.07.002) Sequencing data were preprocessed and normalized to obtain the spatial contact value matrix, which contains the normalized contact positions and contact intensities;

[0145] (2) Correcting the spatial numerical matrix obtained in step (1). For example, in some embodiments, this step includes:

[0146] (i) processing the spatial contact value matrix using quality control information. Preferably, in some embodiments, for each sample to be tested, the spatial contact matrix is divided into bins of equal size of 1 Mb by rows and columns, and the interaction strength value between each pair of bins is recorded. The interaction strength of each bin is divided by the long range ratio of the sample to obtain the spatial contact strength value of the bin;

[0147] (ii) correcting the spatial contact value matrix using a matrix reference system constructed from the control reference system, preferably, in some embodiments, by subtracting the contact intensity baseline value at each position of the matrix reference system from the corresponding position of the value matrix; and

[0148] (iii) Convert the corrected spatial contact matrix into a data format, for example, by using HiSV ( https: / / doi.org / 10.1371 / journal.pcbi.1010760 )Convert spatial contact information into a numerical matrix for subsequent data analysis;

[0149] (3) Inspecting chromosome structural variation on the standardized and corrected spatial contact matrix. In some embodiments, the inspection includes determining the position of the variation and identifying the direction of the variation. In other embodiments, the chromosome structural variation is a chromosome structural rearrangement and / or copy number variation, for example, selected from the group consisting of: translocation, balanced and unbalanced translocation, Robertsonian translocation, intra-arm and inter-arm inversion, insertion, and CNV;

[0150] (4) Optionally, checking copy number variation (CNV) based on sequencing data and spatial contact value matrix;

[0151] (5) Optionally, visualize the detected structural variation results and output variation information.

[0152] In some embodiments, testing for a balanced chromosomal translocation comprises the following steps:

[0153] (i) On a spatial contact matrix (e.g., a spatial contact matrix before correction using a matrix reference frame), for example, by hic_breakfinder( https: / / github.com / dixonlab / hic_breakfinder and doi:10.1038 / s41588-018-0195-8), preliminary calculation of the balanced translocation breakpoint positions;

[0154] (ii) On the corrected spatial contact matrix, based on the initially obtained balanced translocation breakpoint positions, a hypothetical breakpoint position is set, and the difference index score is calculated based on the difference in contact density on both sides of the hypothetical breakpoint;

[0155] (iii) The position of the hypothetical breakpoint is stepped upstream and downstream (for example, 5 Mb in steps of 100 kb), and the difference index score is recalculated for the new hypothetical breakpoint position obtained after each step. The position of the highest difference index score that meets the following conditions at the end of the step is determined as the actual breakpoint position: if the highest score is higher than a preset score threshold (for example, the middle value of the highest difference index score obtained for a sample that actually carries the translocation and the highest difference index score obtained for a sample that does not carry the translocation), the breakpoint is judged to be positive and reported as the test result; otherwise, the breakpoint is judged to be a false positive and is not reported as the test result.

[0156] In some embodiments, the testing for balanced chromosomal translocation further comprises the steps of:

[0157] (iv) Based on the standardized and corrected spatial contact matrix, for each combination of two chromosomes, infer whether the centromere translocation hotspot region (for example, starting from the short arm side of the centromere and extending 5 Mb towards the short arm telomere, and extending 5 Mb from the long arm side of the centromere towards the long arm telomere) meets the characteristics of a balanced translocation (for example, the contact intensity between the short arm of chromosome A and the short arm of chromosome B, and the contact intensity between the long arm of chromosome A and the long arm of chromosome B, are significantly greater than those between the long arm of chromosome A and chromosome B, respectively).

[0158] For chromosome regions meeting this characteristic, the difference index score is inferred according to steps (i) to (iii) above, and the actual breakpoint position is calculated.

[0159] In some embodiments, the Robertson translocation test includes the following steps: on a standardized and corrected spatial contact matrix, for all combinations of chromosomes 13, 14, 15, 21, and 22, the spatial contact intensity in the matrix interaction interval corresponding to each chromosome combination and the difference in spatial contact intensity with other combinations are determined, and a difference index score is calculated; if the score is higher than a preset score threshold (for example, the middle value of the highest difference index score obtained for a sample that actually carries Robertson translocation and the highest difference index score obtained for a sample that does not carry Robertson translocation), it is determined that the corresponding chromosome combination has Robertson translocation; otherwise, it is determined to be a false positive and is not reported as a test result.

[0160] In some embodiments, the chromosome structural abnormality inspection includes an inspection for inversion. In some embodiments, the inspection includes: searching for two matrix positions with high contact intensity within the same chromosome on a standardized and corrected spatial contact matrix, judging whether the position meets the inversion breakpoint characteristics (for example, the contact intensity of breakpoint 1 upstream and breakpoint 2 upstream, breakpoint 1 downstream and breakpoint 2 downstream, are significantly greater than the contact intensity of breakpoint 1 upstream and breakpoint 2 downstream, breakpoint 1 downstream and breakpoint 2 upstream, then it is judged to be an inversion); the position that meets the inversion breakpoint characteristics is inferred from the breakpoint, and the difference index score is calculated. If the score is greater than the threshold, the breakpoint is determined to be an inversion breakpoint.

[0161] In some embodiments, the examination of chromosome structural abnormalities includes examination of insertions. The examination involves searching for locations in the standardized and corrected spatial contact matrix where the contact strength between different chromosomes is high, determining whether they meet the characteristics of an insertion (e.g., a fragment within chromosome A has significantly increased contact strength with chromosome B). Similarly, the difference index score is inferred according to the above steps to serve as the criterion for determining insertion variation.

[0162] In some embodiments, the present invention further provides a method for identifying the linked haplotype status of a structural variation in an embryo, the method comprising the steps of:

[0163] b1. Sampling: obtain embryo biopsy cells and peripheral blood samples from both parents;

[0164] b2. Using the method for identifying structural variation-linked haplotypes according to the first aspect of the present invention, identifying structural variation-linked haplotypes of the structural variation carrier in the parent, and obtaining one or more breakpoint information associated with the variation and two allele information: a linked SNP allele representing the structural variation haplotype and a linked SNP allele representing the normal haplotype;

[0165] b3. Extract DNA from embryo biopsy cells and peripheral blood samples from the non-carrier parent, amplify it, create a library, and then sequence it;

[0166] b4. Align the genome sequences of the two samples measured in step b3 to the same human reference genome;

[0167] b5. For the sequence aligned to the reference genome in step b4, SNPs are detected upstream and downstream of the breakpoints according to the carrier's mutation-related breakpoint information;

[0168] b6. Intersecting the variant haplotype-linked SNP result obtained in step a10 of the method for identifying structural variation-linked haplotypes according to the first aspect of the present invention with the SNP result obtained in step b5 according to their genomic positions to infer which allele of the carrier is inherited by the embryo at that position. Thus, for each SNP site in the embryo, independently determining whether its genotype is consistent with the variant haplotype-linked SNP allele or the normal haplotype SNP allele;

[0169] b7. Count the number of SNP sites determined in step b6 upstream and downstream of the structural variation-associated breakpoint. If the number of SNP sites consistent with the variant haplotype-linked SNP allele is greater than the number of sites consistent with the normal haplotype, the embryo is considered to carry the structural variation haplotype obtained in step a10 of the first aspect, and is presumed to carry the variation. Otherwise, it is presumed not to carry the variation.

[0170] In some embodiments, the biopsied cells of the embryo in step b1 are single or multiple cells from the trophoblast of the blastocyst. In some preferred embodiments, the blastocyst is a healthy, high-quality blastocyst cultured to day 5-8 after oocyte retrieval.

[0171] The following examples are described to assist understanding of the present invention. The examples are not intended to, and should not be interpreted in any way as, limiting the scope of protection of the present invention. Example

[0172] Example 1: Inference of balanced translocation

[0173] 1. Subject information: The female carries a balanced translocation 46,XX,t(7;20)(q31.32;p12.3), the male carries a 46,XY translocation, and three embryos were obtained by in vitro fertilization to be tested.

[0174] 2. Peripheral blood was obtained from both the woman and the man, and genomic DNA (gDNA) was extracted. The man's DNA was digested and fragmented using the Aowei Biotech TL674-03 pre-loaded magnetic bead-based universal DNA extraction kit, followed by whole-genome amplification using the MALBAC two-step method (Suzhou Xukang Medical Technology Co., Ltd.'s universal gene sequencing sample processing kit). The woman, who was a carrier, was then subjected to Hi-C sequencing and high-throughput sequencing.

[0175] 3. For the female partner's Hi-C data, the sequences were aligned to the reference coordinate system hg19 using BWA (0.7.17-r1188), and the contact matrix was generated using Juicer (2.0) (see Figure 1). Subsequently, according to the matrix analysis of the variant breakpoints, the breakpoints were obtained as chr7:121900000 and chr20:5300000, including two normal chain breakpoints (chr7:121900000; chr20:5300000) and two translocation variant chain breakpoints (chr7:121900000-chr20:5300000; chr20:5300000-chr7:121900000).

[0176] 4. For the female partner's Hi-C data, detect SNPs across the entire genome and retain heterozygous SNP results.

[0177] 5. The female partner's Hi-C data were screened according to the following criteria: paired sequences spanning the normal chain breakpoint, with one side position being chr7:111900000-121900000 and the other side position being chr7:121900000-131900000; or paired sequences with one side position being chr20:0-5300000 and the other side position being chr20:5300000-15300000. SNPs were detected for the screened sequences within the above position range, and the homozygous SNP results were retained as the characteristic sites of the normal chain haplotype.

[0178] 6. For the female partner's Hi-C data, screen according to the following criteria: paired sequences spanning the translocation variant breakpoint, one side position is chr7:111900000-121900000, the other side position is chr20:0-5300000; or one side position is chr7:121900000-131900000, the other side position is chr20:5300000-15300000. Detect SNPs in the selected sequences within the above position range, and retain the homozygous SNP results as the variant chain haplotype characteristic site.

[0179] 7. Intersect the homozygous and heterozygous SNPs in the female carrier, retain the shared sites, and determine the pathogenic strand. For example, at position chr7:121219538 (rs191194352), if the pathogenic strand result is T / T and the heterozygous result is T / G, this means that the two haplotypes corresponding to the SNPs at this site are T / G, respectively. The translocation haplotype corresponds to T, and the normal haplotype corresponds to C. Similarly, for SNPs corresponding to the normal strand, the pathogenic strand can also be determined based on the base pair opposite the heterozygous base pair.

[0180] Table 1 below shows some of the common sites (not all of them can be listed due to space limitations).

[0181] Table 1: Some effective SNP sites in Example 1

[0182]

[0183]

[0184] 8. For the female partner's Hi-C data, local haplotype blocks were constructed to obtain other SNPs linked to the identified normal / translocation SNPs (Table 1).

[0185] 9. Extract embryonic blastocyst trophoblast cells and perform single-cell whole genome amplification using the MALBAC two-step method (gene sequencing universal sample processing kit from Suzhou Xukang Medical Technology Co., Ltd.).

[0186] 10. Perform genotyping tests on the male partner's gDNA and embryo amplification products to obtain the genotype of specific SNP sites.

[0187] 11. Extract SNP information within 10M upstream and downstream of the variant breakpoint and select sites where the variant carrier is heterozygous and the normal partner is homozygous, or where the variant carrier is heterozygous and the embryo is homozygous, for further analysis. Sites that meet these criteria are considered valid haplotype sites, and carriers are classified into a series of variant and normal linked haplotypes according to step 6 (Table 1).

[0188] 12. After haplotype construction is completed, embryos can be divided into those carrying structural variant haplotypes and those carrying normal haplotypes based on the haplotypes carried by the embryos and the phenotypic information of their parents (see Figure 2a and Figure 2b Embryos 1 and 3 carry the maternal mutation (inherited dark brown haplotype), while embryo 2 does not carry the maternal mutation and carries the normal haplotype (inherited light brown haplotype).

[0189] 13. The embryo amplification products were also screened for aneuploidy using CNV-seq, and no abnormal chromosome copy number was found (Table 2).

[0190] 14. Further verification by family linkage analysis showed that embryos 1 and 3 carried balanced maternal translocations, while embryo 2 did not, which was consistent with the results of the non-family analysis (Table 2).

[0191] Table 2: Test results statistics in Example 1

[0192] Sample name Aneuploidy test results Structural variation detection results Embryo 1 46,XN Carrying translocation Embryo 2 46,XN Do not carry Embryo 3 46,XN Carrying translocation

[0193] Example 2: Inference of Robertsonian translocation

[0194] 1. Subject information: The female carries Robertsonian translocation 45,XX,der(13;14)(q10;q10), the male carries 46,XY, and four embryo samples are to be tested.

[0195] 2. Sampling, library construction, sequencing, and typing were performed on the pedigree following the same procedures as in Example 1. When screening for homozygous SNPs in carriers, the paired sequences with one side located at chr13:19400000-29400000 and the other side located at chr14:20500000-30500000 were used as variant strands for SNP calling. There are no available normal strand SNPs for the Roche variant strand.

[0196] 3. According to the haplotype carried by the embryo and the phenotypic information of its parents, the embryo is divided into those carrying structural variant haplotypes and those carrying normal haplotypes ( Figure 3a and 3b Each shows a portion of the loci.) Embryos 2, 3, and 4 carry the maternal mutation (inherited dark brown haplotype), while embryo 1 does not carry the maternal mutation and carries the normal haplotype (inherited light brown haplotype).

[0197] 4. The embryo amplification products were also screened for aneuploidy using CNV-seq, and no abnormal chromosome copy number was found (Table 3).

[0198] 5. Family linkage analysis confirmed that embryos 2, 3, and 4 carried maternal balanced translocations, while embryo 1 did not, which was consistent with the results of non-family linkage analysis (Table 3).

[0199] Table 3: Test results statistics in Example 2

[0200] Sample name Aneuploidy test results Structural variation detection results Embryo 1 46,XN Do not carry Embryo 2 46,XN Carrying translocation Embryo 3 46,XN Carrying translocation Embryo 4 46,XN Carrying translocation

[0201] Example 3: Inference of inversion

[0202] 1. Subject information: The female carries the inversion 46,XX,inv(5)(p13.1q34), the male carries the inversion 46,XY, and there are five embryo samples to be tested.

[0203] 2. Sampling, library construction, sequencing, and typing were performed on the family according to the process of Example 1. For the female's Hi-C data, a contact matrix was generated using Juicer ( Figure 4 ), and according to the matrix analysis of the mutation breakpoints, the breakpoints chr5:40900000 and chr5:165200000 (the reference coordinate system is hg19) were obtained, including two normal chain breakpoints (chr5:40900000; chr5:165200000) and two inversion mutation chain breakpoints (chr5:40900000-chr5:165200000; chr5:165200000-chr5:40900000).

[0204] 3. When screening for homozygous SNPs on the carrier side, take the paired sequences with one side at chr5:35900000-40900000 and the other side at chr5:40900000-45900000, and the paired sequences with one side at chr5:160200000-165200000 and the other side at chr5:165200000-170200000 as the normal chain for SNP calling.

[0205] 4. When screening for homozygous SNPs on the carrier side, take the paired sequences with one side at chr5:35900000-40900000 and the other side at chr5:160200000-165200000, and the paired sequences with one side at chr5:40900000-45900000 and the other side at chr5:165200000-170200000 as the variant chains for SNP calling.

[0206] 5. According to the haplotype carried by the embryo and the phenotypic information of its parents, the embryo is divided into those carrying structural variant haplotypes and those carrying normal haplotypes ( Figure 5 Embryos 1, 2, 3, and 4 are carriers of the maternal mutation (inherited dark brown haplotype), while embryo 5 does not carry the haplotype and carries the normal haplotype (inherited light brown haplotype).

[0207] 6. The embryo amplification products were also screened for aneuploidy using CNV-seq, and no abnormal chromosome copy number was found (Table 4).

[0208] 7. Family linkage analysis confirmed that embryos 1, 2, 3, and 4 carried maternal inversions, while embryo 5 did not, which was consistent with the results of non-family linkage analysis (Table 4).

[0209] Table 4: Test results statistics in Example 3

[0210] Sample name Aneuploidy test results Structural variation detection results Embryo 1 46,XN Carrying inverted position Embryo 2 46,XN Carrying inverted position Embryo 3 46,XN Carrying inverted position Embryo 4 46,XN Carrying inverted position Embryo 5 46,XN Do not carry

[0211] Example 4: Inference of insertion mutations

[0212] 1. Subject information: The female carries the insertion variant 46,XX,ins(17;20)(q24.3;q11.22q13.13), the male carries the insertion variant 46,XY, and there are three embryo samples to be tested.

[0213] 2. Sampling, library construction, sequencing, and typing were performed on the family according to the process of Example 1. For the female's Hi-C data, a contact matrix was generated using Juicer ( Figure 6), and according to the matrix analysis of the mutation breakpoints, the insertion position breakpoint chr17:70400000, the insertion sequence breakpoints chr20:32800000 and chr20:47300000 (the reference coordinate system is hg19), including one normal chain breakpoint (chr17:70400000) and two insertion mutation chain breakpoints (chr17:70400000-chr20:32800000; chr17:70400000-chr20:47300000).

[0214] 3. When screening for homozygous SNPs on the carrier side, take the paired sequence with one side located at chr17:60400000-70400000 and the other side located at chr17:70400000-80400000 as the normal chain for SNP calling.

[0215] 4. When screening for homozygous SNPs on the carrier side, take the paired sequence with one side located at chr17:60400000-70400000 and the other side located at chr20:32800000-42800000, or the paired sequence with one side located at chr20:37300000-47300000 and the other side located at chr17:70400000-80400000 as the variant chain for SNP calling.

[0216] 5. According to the haplotype carried by the embryo and the phenotypic information of its parents, the embryo is divided into those carrying structural variant haplotype and those carrying normal haplotype ( Figure 7 Embryo 1 carries the maternal mutation (inherited dark brown haplotype), while embryos 2 and 3 do not carry this haplotype and carry the normal haplotype (inherited light brown haplotype).

[0217] 6. The embryo amplification products were also screened for aneuploidy using CNV-seq, and no abnormal chromosome copy number was found (Table 5).

[0218] 7. Family linkage analysis confirmed that embryo No. 1 carried the maternal inversion, while embryos No. 2 and No. 3 did not carry the maternal inversion, which was consistent with the results of non-family linkage analysis (Table 5).

[0219] Table 5: Test results statistics in Example 3

[0220] Sample name Aneuploidy test results Structural variation detection results Embryo 1 46,XN Carry Insert Embryo 2 46,XN Do not carry Embryo 3 46,XN Do not carry

[0221] Example 5: Inference of deletion variants

[0222] 1. Since the insertion variation of copy number balance actually loses a sequence at the original position of the insertion fragment, it can be considered that the insertion variation of copy number balance also carries a deletion variation ( Figure 8 ). Under this premise, the deletion mutation carried by the female in Example 4 should be 46,XX,del(20)(q11.22q13.13).

[0223] 2. Sampling, library construction, sequencing, and typing were performed on the family according to the process of Example 4. For the female's Hi-C data, a contact matrix was generated using Juicer ( Figure 9 ), and according to the matrix analysis of mutation breakpoints, the deletion mutation breakpoints chr20:32800000 and chr20:47300000 (the reference coordinate system is hg19) were obtained, including 2 normal chain breakpoints (chr20:32800000; chr20:47300000) and 1 deletion mutation chain breakpoint (chr20:32800000-chr20:47300000).

[0224] 3. When screening for homozygous SNPs on the carrier side, take the paired sequence with one side located at chr20:22800000-32800000 and the other side located at chr20:32800000-42800000, or the paired sequence with one side located at chr20:37300000-47300000 and the other side located at chr20:47300000-57300000 as the normal chain for SNP calling.

[0225] 4. When screening for homozygous SNPs on the carrier side, take the paired sequence with one side located at chr20:22800000-32800000 and the other side located at chr20:47300000-57300000 as the variant chain for SNP calling.

[0226] 5. According to the haplotype carried by the embryo and the phenotypic information of its parents, the embryo is divided into those carrying structural variant haplotype and those carrying normal haplotype ( Figure 10 Embryo 1 carries the maternal mutation (inherited dark brown haplotype), while embryos 2 and 3 do not carry this haplotype and carry the normal haplotype (inherited light brown haplotype).

[0227] 6. Embryo amplification products were also screened for aneuploidy using CNV-seq, and no abnormal chromosome copy number was found (Table 6).

[0228] 7. Family linkage analysis confirmed that embryo No. 1 carried the maternal inversion, while embryos No. 2 and No. 3 did not carry the maternal inversion, which was consistent with the results of non-family linkage analysis (Table 6).

[0229] Table 6: Test results statistics in Example 3

[0230]

[0231]

Claims

1. A method for identifying structural variation linkage haplotypes, the method comprising the steps of: a1) Obtain biological samples from individuals carrying the structural variant and extract genomic DNA; a2) constructing a library from the extracted genomic DNA using chromosome conformation capture technology and sequencing; a3) aligning the genome sequence determined in step a2) to the human reference genome; a4) determining one or more structural variation breakpoints based on the sequence aligned to the reference genome in step a3) by combining the data characteristic signals obtained by sequencing using chromosome conformation capture technology; a5) for the sequence aligned to the reference genome in step a3), detecting SNPs upstream and downstream of the one or more breakpoints based on the one or more breakpoint information obtained in step a4), and retaining heterozygous SNP results; a6) For the sequences aligned to the reference genome in step a3), extract their paired-end sequencing information and screen them according to their genomic location. The screening criteria are: retaining paired-end sequences whose ends can be respectively aligned to the upstream and downstream breakpoint positions corresponding to normal chromosomes without mutation as normal chain characteristic sequences, and retaining paired-end sequences whose ends can be respectively aligned to the upstream and downstream breakpoints of newly generated chromosomes after mutation as variant chain characteristic sequences, and removing sequences that do not meet the above criteria; a7) detecting SNPs in the sequences retained in step a6), and retaining homozygous SNP results; a8) taking the intersection of the SNP sites retained in step a5) and the SNP sites retained in step a7), and comparing the effective depths of each site in step a5) and step a7), retaining sites with an effective depth in step a5) greater than that in step a7) as haplotype anchor points, wherein at this site, the homozygous base obtained in step a7) is the characteristic base representing the structural variation haplotype, and the other of the heterozygous bases obtained in step a5) is the characteristic base representing a normal chromosome; a9) Assembling and constructing locally linked SNP blocks for the sequences aligned to the reference genome in step a3); a10) In the locally linked SNP block obtained in step a9), search for the haplotype anchor position obtained in step a8). For the block containing the anchor SNP, use the base detected in step a8) as a reference and use the SNP allele linked to that base as the linked SNP allele representing the structural variant haplotype, and the other as the linked SNP allele representing the normal haplotype.

2. The method according to claim 1, wherein the chromosome conformation capture technology in step a2) is Hi-C technology.

3. The method of claim 1, wherein the reference genome in step a3) is selected from hg19, hg38, and T2T.

4. The method of claim 1, wherein the one or more structural variation breakpoints detected in step a4) include normal chain breaks and / or variant chain breaks.

5. The method of claim 4, wherein the variant chain breakpoints include variant position breakpoints and / or variant sequence breakpoints.

6. The method of claim 1, wherein: In step a3), the alignment is performed using genome alignment tool software, such as BWA, samtools, Bowtie2, HiC-Pro, and / or hisat2; In step a4), breakpoints are detected using Hi-C data analysis software, wherein the software is Juicer; The software used for detecting SNPs in step a5) is GATK or freebayes; and / or In step a9), haploid assembly analysis software is used to assemble and construct locally linked SNP blocks, and the software is HapCUT or HapCUT2.

7. The method according to claim 6, wherein in step a9), the locally linked SNP blocks are assembled and constructed using haploid assembly analysis software, and the software is HapCUT2.

8. The method of claim 1, wherein: In step a5), the upstream and downstream ranges of the breakpoints are 10 Mbp upstream and downstream of balanced translocation breakpoints, 10 Mbp downstream of Roche's ectopic breakpoints, 5 Mbp upstream and downstream of inversion breakpoints, and 10 Mbp upstream and downstream of insertion mutation breakpoints; and / or In step a6), the upstream and downstream ranges of the breakpoints are 10 Mbp upstream and downstream of balanced translocation breakpoints, 10 Mbp downstream of Robertson ectopic breakpoints, 5 Mbp upstream and downstream of inversion breakpoints, and 10 Mbp upstream and downstream of insertion mutation breakpoints.

9. The method of claim 1, wherein the structural variation is selected from the group consisting of structural rearrangements and copy number abnormalities (CNVs).

10. The method of claim 9, wherein the structural variation is a structural rearrangement.

11. The method of claim 10, wherein the structural rearrangement is a translocation.

12. The method of claim 11, wherein the translocation is selected from the group consisting of a balanced translocation, an unbalanced translocation, a reciprocal translocation, a nonreciprocal translocation, a Robertsonian translocation, and a complex translocation.

13. The method of claim 10, wherein the structural rearrangement is an inversion.

14. The method of claim 13, wherein the inversion is selected from the group consisting of an intra-arm inversion and an inter-arm inversion.

15. The method of claim 10, wherein the structural rearrangement is an insertion.

16. The method of claim 15, wherein the insertion is selected from the group consisting of a forward insertion and an inversion insertion.

17. The method of claim 1, wherein the structural variation comprises a copy number abnormality CNV, and the copy number abnormality structural variation is selected from a deletion and a duplication.

18. The method of claim 1, wherein the structural variation is selected from an isochromosome or a ring chromosome with a chromosome structure change having a sequence breakpoint.

19. The method of claim 1, wherein the sequencing is low-depth paired-end sequencing.

20. The method of claim 19, wherein the sequencing depth is no less than 2X and / or no greater than 10X.

21. The method of claim 20, wherein the sequencing depth is no less than 3X and / or no more than 5X.

22. The method of claim 21, wherein the sequencing depth is 30M reads to 40M reads.

23. The method of claim 19, wherein the paired-end sequencing is paired-end sequencing.

24. The method of claim 1, wherein the structural variation carrier has no genetic information of other family members.

25. A method for identifying the linkage haplotype status of a structural variation in an embryo, the method comprising the steps of: b1) Sampling: obtaining embryo biopsy cells and peripheral blood samples from both parents; b2) using the method of any one of claims 1 to 24 to identify the structural variation-linked haplotype of the structural variation carrier in one of the parents, and obtaining one or more breakpoint information related to the variation and two allele information: a linked SNP allele representing the structural variation haplotype and a linked SNP allele representing the normal haplotype; b3) DNA extraction, amplification, library construction, and sequencing were performed on peripheral blood samples and embryo biopsy cells from the non-carrier parent of both parents; b4) aligning the genome sequences of the two samples measured in step b3) to the same human reference genome; b5) for the sequence aligned to the reference genome in step b4), detecting SNPs upstream and downstream of the breakpoints based on the mutation-related breakpoint information of the carrier; b6) intersecting the variant haplotype-linked SNP result obtained in step a10) of the method of claims 1-24 with the SNP result obtained in step b5) according to their genomic positions to infer which allele of the carrier is inherited by the embryo at that position, thereby independently determining, for each SNP site in the embryo, whether its genotype is consistent with the variant haplotype-linked SNP allele or the normal haplotype SNP allele; b7) Counting the number of SNP sites determined in step b6) upstream and downstream of the structural variation-associated breakpoint; if the number of SNP sites consistent with the variant haplotype-linked SNP allele is greater than the number of sites consistent with the normal haplotype, the embryo is considered to carry the structural variation haplotype obtained in step a10) of the method of claims 1-24, i.e., is presumed to carry the variation; otherwise, it is presumed not to carry the variation. 26 . The method according to claim 25 , wherein the biopsied cells of the embryo in step b1) are single cells or several cells from the trophoblast of a blastocyst, and the blastocyst is a blastocyst cultured to day 5-8.

27. A detection product for implementing the method according to any one of claims 1 to 26, comprising one or more of the following modules: (1) gDNA extraction module: used to extract genomic DNA from samples; (2) Chromosome conformation capture module: used to fix the chromosome conformation in the sample, that is, to fix the interactions between spatially contacting or adjacent DNA fragments and / or with proteins so that they can be detected in subsequent steps; (3) Sequencing library construction module: used to generate whole genome DNA sequencing library; (4) Sequencing module: used for low-depth paired-end sequencing of sequencing libraries; (5) Alignment module: used to align the whole genome sequencing results generated by the sequencing module to the reference genome; (6) Structural variation assessment module: used to assess the chromosomal structural variation of sample genomic DNA from sequencing data; (7) SNP detection module: used to detect SNP sites and their genotypes from sequencing data; (8) Typing module: used to compare and screen related SNP sites and identify haplotype-related linked SNP alleles.

28. A detection system for implementing the method of any one of claims 1 to 26, the detection system comprising a computer and / or storable software, and the detection system comprising one or more of the following modules: (1) gDNA extraction module: used to extract genomic DNA from samples; (2) Chromosome conformation capture module: used to fix the chromosome conformation in the sample, that is, to fix the interactions between spatially contacting or adjacent DNA fragments and / or with proteins so that they can be detected in subsequent steps; (3) Sequencing library construction module: used to generate whole genome DNA sequencing library; (4) Sequencing module: used for low-depth paired-end sequencing of sequencing libraries; (5) Alignment module: used to align the whole genome sequencing results generated by the sequencing module to the reference genome; (6) Structural variation assessment module: used to assess the chromosomal structural variation of sample genomic DNA from sequencing data; (7) SNP detection module: used to detect SNP sites and their genotypes from sequencing data; (8) Typing module: used to compare and screen related SNP sites and identify haplotype-related linked SNP alleles.

29. The detection product of claim 27 or the detection system of claim 28, wherein the sample is a cell.

30. The detection product according to claim 27 or the detection system according to claim 28, wherein the chromosome conformation capture module fixes the chromosome conformation in the sample by covalent cross-linking.

Citation Information

Patent Citations

  • Method for variation detection before embryo implantation

    CN117925820A

  • Method for simultaneously completing gene locus, chromosome and linkage analysis

    WO2017084624A1