A tandem repeat variation typing detection device based on core family and its application
By using a core family-based tandem repeat variant typing detection device, combined with third-generation sequencing datasets and an improved Tandem Repeat Finder method, the problem of insufficient accuracy in detecting tandem repeat variants in existing technologies was solved, achieving full genome coverage and high-accuracy variant detection.
Patent Information
- Application Number
- CN202311109741.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-31
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2043-08-31
AI Technical Summary
Existing second-generation and third-generation sequencing technologies have problems such as insufficient read length, sequence bias, and insufficient detection accuracy when detecting tandem repeat variations, making it difficult to fully cover and accurately determine the variation information of the tandem repeat region.
A tandem repeat variation typing detection device based on a core family is used. Through the offspring typing and assembly unit, parent typing unit, and haplotype detection unit, combined with the improved Tandem Repeat Finder method, high-accuracy whole-genome coverage detection is performed using third-generation sequencing data sets to analyze tandem repeat variation information in the genetic process.
It achieves a comprehensive search of all tandem repeat regions on the reference genome, obtains copy number information for millions of sites, improves the accuracy and comprehensiveness of detection, and obtains more accurate haplotype variation information.
Smart Images

Figure CN117095742B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular to a core family-based tandem repeat variation typing detection device, and the application of the device in biology. Background Art
[0002] Variations in the human genome are closely related to human evolution, disease risk, and other aspects. With the development of second-generation short-read high-throughput sequencing technology (NGS), researchers have developed a series of methods to detect structural variations (SVs) in the genome. Without considering assembly, there are three main strategies for devices or methods to detect variations based on second-generation sequencing data: (1) ReadPair (RP), which classifies each read object as normal or SV based on the mapping distance and direction of pair reads on the reference genome, and then identifies regions with a high number of reads that meet the SV category and assigns a confidence score; (2) Split Read (SR), which means that only one of the two pair-end reads can be mapped to the reference genome, while the other cannot. The sites where such split reads are generated often have structural variations; (3) ReadDepth (RD), which is the read coverage depth, is mainly used to detect sequence loss or duplication. In particular, for tandem repeat regions, researchers have developed the following software to detect their variation information, including: (1) TSSV (https: / / pypi.org / project / tssv / ); (2) HipSTR (https: / / github.com / tfwillems / HipSTR); (3) STRSCan (http: / / darwin.informatics.indiana.edu / str / ), etc.
[0003] However, due to the short read lengths of second-generation sequencing, it is difficult to cover the entire region of long tandem repeats. Algorithms are mainly used to infer the repeats using information such as breakpoints and depth. In addition, NGS sequencing has sequence preferences during amplification, which can also affect TR detection.
[0004] Third-generation sequencing increases the read length from 150-200bp in the second generation to 15kb-4Mb. This long read length can directly cover most TR regions, providing a sequencing data foundation for accurately obtaining copy number information, making third-generation sequencing significantly superior to NGS technology in detecting tandem repeat variants. However, there are currently few TR detection methods developed based on third-generation sequencing data. The main ones are: (1) Straglr, which does not detect based on known TR regions and mainly detects newly occurring TR amplifications in the genome; (2) TriColor, which uses a parameter-free mode to detect amplified regions in the entire genome. These two methods detect a small number of variant sites. Compared with the approximately one million tandem repeat regions in the reference genome, these two methods can only obtain TR amplification sites in a few thousand to tens of thousands of regions; (3) TRGT, which detects known TR regions. This method can detect all tandem repeat regions in the reference genome. Existing methods mainly obtain TR typing information through clustering algorithms, and their accuracy remains to be evaluated.
[0005] In summary, variant detection in tandem repeat regions is a challenge for NGS technology, as its short read lengths cannot span repeat regions. Furthermore, it is difficult to determine breakpoint locations and depth information within tandem repeat regions during alignment. Furthermore, most current third-generation sequencing algorithms with long read lengths suffer from either a low number of effective sites or an inability to obtain accurate haplotype variant information. Summary of the Invention
[0006] To address the aforementioned issues with existing technologies, the present invention provides a core family-based tandem repeat variant typing detection device and its application. This device, developed for third-generation sequencing data, provides high-accuracy, genome-wide coverage of tandem repeat region variant detection. Furthermore, the device can obtain haplotype information for tandem repeat variants that occur during inheritance.
[0007] Specifically, the present invention relates to the following core family-based tandem repeat variation typing detection device and application.
[0008] In a first aspect, a tandem repeat variation typing detection device based on a core family comprises a progeny typing and assembly unit, a parent typing unit and a haplotype detection unit, wherein:
[0009] The offspring typing and assembly unit is used to perform a first typing and assembly on the offspring sequencing dataset based on the paternal sequencing dataset and the maternal sequencing dataset to obtain two sets of offspring haplotype genomes;
[0010] The parental typing unit is used to perform single nucleotide site variation detection on the paternal and maternal genomes respectively, using the obtained offspring haplotype genome as a reference, and perform a second typing on the paternal sequencing data set and the maternal sequencing data set respectively, to obtain two sets of paternal haplotype genomes and two sets of maternal haplotype genomes, and to clearly define the haplotype genome inherited from the father to the offspring and the haplotype genome inherited from the mother to the offspring;
[0011] The haplotype detection unit is used to perform tandem repeat variation detection on the haplotype genome inherited from the father to the offspring, the haplotype genome inherited from the mother to the offspring, and the two sets of offspring haplotype genomes.
[0012] Preferably, the method for detecting tandem repeat variations is an improved Tandem Repeat Finder method.
[0013] Preferably, the improvements include motif cycle characteristics and detection segment features.
[0014] Preferably, the sequencing data set includes long reads obtained by a third-generation sequencing method of the genome.
[0015] Preferably, the first typing method includes find-unique-kmers analysis based on the trio_binning method to obtain paternal and maternal specific kmer sequences.
[0016] Preferably, based on the classify_by_kmers method, the offspring sequencing dataset is judged as a sequencing dataset belonging to the father, a sequencing dataset belonging to the mother, or an untyped sequencing dataset according to the paternal and maternal specific kmer sequences.
[0017] Preferably, the assembly method is selected from at least one of hifiasm, wtdbg2, canu, and nextdenovo.
[0018] Preferably, the second typing method is based on the WhatsHap method.
[0019] Preferably, the device further comprises an analysis unit for analyzing the tandem repeat variation information of the two haploid genomes of the offspring during the inheritance process.
[0020] Secondly, the application of the above-mentioned core family-based tandem repeat variation typing detection device in biology.
[0021] Preferably, the application is in disease pathology research, paternity testing, crime investigation, and population genetics. DETAILED DESCRIPTION
[0022] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0023] Definition of terms
[0024] (1) Tandem repeats (TR): One or more nucleotides in DNA are repeated in a series of linked sequences. When the number of nucleotides in a repeating unit is less than 10, it is called a short tandem repeat (STR). When the number of nucleotides in a repeating unit is between 10 and 60, it is called a minisatellite. When the number of copies of a repeating unit varies among people, it is called a variable number tandem repeat (VNTR).
[0025] (2) Third-generation sequencing: The third-generation sequencing technology, single-molecule sequencing technology, does not require PCR amplification and can sequence each DNA molecule individually. It has no GC preference and has a faster data reading speed.
[0026] (3) Nuclear family: A family consisting of a child (offspring) and his or her parents.
[0027] (4) Genome typing: The process of detecting the DNA sequence of an individual through biological assays and comparing it with the genotypes or sequences of other individuals. It can be used to show the inheritance of the individual's alleles from his or her parents.
[0028] (5) Genome assembly: Genome assembly refers to the use of sequencing methods to generate sequence fragments (i.e., reads) from the genome of the species to be tested, and then splicing the fragments according to the overlapping regions between the reads, first splicing them into longer continuous sequences (contigs), and then splicing the contigs into longer scaffolds that allow the inclusion of blank sequences (gaps). By eliminating errors and gaps in the scaffolds, these scaffolds are located on the chromosomes, thereby obtaining a high-quality whole-genome sequence.
[0029] The first aspect of the present invention provides a tandem repeat variation typing detection device based on a core family, the device comprising a progeny typing and assembly unit, a parent typing unit and a haplotype detection unit, wherein:
[0030] The offspring typing and assembly unit is used to perform a first typing and assembly on the offspring sequencing dataset based on the paternal sequencing dataset and the maternal sequencing dataset to obtain two sets of offspring haplotype genomes;
[0031] The parental typing unit is used to perform single nucleotide site variation detection on the paternal and maternal genomes respectively, using the obtained offspring haplotype genome as a reference, and perform a second typing on the paternal sequencing data set and the maternal sequencing data set respectively, to obtain two sets of paternal haplotype genomes and two sets of maternal haplotype genomes, and to clearly define the haplotype genome inherited from the father to the offspring and the haplotype genome inherited from the mother to the offspring;
[0032] The haplotype detection unit is used to perform tandem repeat variation detection on the haplotype genome inherited from the father to the offspring, the haplotype genome inherited from the mother to the offspring, and the two sets of offspring haplotype genomes.
[0033] In the present invention, the method for detecting tandem repeat variations is an improved Tandem Repeat Finder method.
[0034] In the present invention, the improvements include motif loop characteristics and detection segment characteristics. Motif loop characteristics include, for example, motif base sequence. For example, but not limited to, considering the base sequence of the ATC motif, the TCA and CAT motif loops are also considered. Detection segment characteristics refer to considering the copy number of the entire segment between the start and end points when detecting a specific interval.
[0035] In the present invention, the sequencing dataset comprises long genomic reads; preferably, the sequencing dataset comprises long reads obtained using a third-generation sequencing method; more preferably, the third-generation sequencing method is selected from at least one of Pacbio and Nanopore; and more preferably, the sequencing dataset is HiFi reads. According to the method of the present invention, the average sequencing depth is preferably greater than 15x.
[0036] In the present invention, the method for described first typing is based on trio_binning method. trio_binning first uses the high-precision short read length data from two parental genomes to divide the long read length sequence of offspring into a haplotype-specific set, and then each haplotype is independently assembled to form a complete diploid reconstruction. Preferably, the method for described first typing includes find-unique-kmers analysis based on trio_binning method to obtain the specific kmer sequence of paternal and maternal parents. More preferably, based on classify_by_kmers method, and according to the specific kmer sequence of paternal and maternal parents, the offspring sequencing data set is judged to be the sequencing data set belonging to the paternal parent, the sequencing data set belonging to the maternal parent, or the sequencing data set for untyped. Preferably, the sequencing data set for untyped is not assembled.
[0037] In the present invention, the assembly method is selected from at least one of hifiasm, wtdbg2, canu, and nextdenovo.
[0038] In the present invention, the second typing method is based on the WhatsHap method.
[0039] In the present invention, the device further comprises: analyzing tandem repeat variation information of the two haploid genomes of the offspring during the inheritance process. The change in the copy number of the tandem repeat region during the inheritance process is calculated, and then analyzing the tandem repeat variation information of the two haploid genomes of the offspring during the inheritance process.
[0040] The second aspect of the present invention provides the application of the above-mentioned core family-based tandem repeat variation typing detection device in biology.
[0041] According to the application of the present invention, preferably, the application is application in disease pathology research, paternity testing, crime investigation, and population genetics.
[0042] The present invention has the following advantages:
[0043] 1. The device of the present invention combines assembly and typing technologies with variation detection. First, it accurately separates the two haplotype genomes of the offspring and the genome inherited from the parents from the sequencing data set. This is more accurate than the two haplotype genotypes obtained by clustering methods in existing algorithms.
[0044] 2. The device of the present invention adopts an improved TRF algorithm, which comprehensively considers the circular characteristics of the motif and the different characteristics of different parts of the segment, and the detected copy number is more accurate.
[0045] 3. The device of the present invention conducts a comprehensive search of all tandem repeat regions in the reference genome and obtains copy number information for millions of sites. Compared with the thousands to tens of thousands of sites obtained in the parameter-free mode of existing algorithms, the device of the present invention more comprehensively mines the tandem repeat variation information on the entire genome.
[0046] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is a flow chart of the core family-based tandem repeat variation typing detection device provided in Example 1 of the present invention.
[0048] Figure 2 This is a flow chart showing the principle of tandem repeat variant typing detection based on core family provided in Example 1 of the present invention.
[0049] Figure 3 This is the motif length distribution diagram provided in Example 1 of the present invention.
[0050] Figure 4 The motif length and frequency distribution diagram of Example 1 and Comparative Example 2.
[0051] The present invention will be described below with reference to specific examples. It should be noted that these examples are merely illustrative and are not to be construed as limiting the present invention.
[0052] [Example 1]
[0053] A device for detecting tandem repeat variant typing based on core family, the flow diagram of the device is as follows Figure 1 As shown, the principle flow chart is as follows Figure 2 shown. Figure 2 In the figure, p1 represents the sequencing data set of the offspring; Long reads represents long read length; P represents the offspring haploid genome from the father; M represents the offspring haploid genome from the mother; Fa represents the paternal sequencing data set; Mo represents the maternal sequencing data set; hap1 is the first set of haploid genomes, which is also the haploid genome inherited from parents to offspring; hap2 is the second set of haploid genomes.
[0054] The device includes a progeny typing and assembly unit, a parent typing unit, a haplotype detection unit, and an analysis unit, wherein:
[0055] The offspring typing and assembly unit is used to perform a first typing and assembly on the offspring sequencing dataset based on the paternal sequencing dataset and the maternal sequencing dataset to obtain two sets of offspring haplotype genomes;
[0056] The parental typing unit is used to perform single nucleotide site variation detection on the paternal and maternal genomes respectively, using the obtained offspring haplotype genome as a reference, and perform a second typing on the paternal sequencing data set and the maternal sequencing data set respectively, to obtain two sets of paternal haplotype genomes and two sets of maternal haplotype genomes, and to clearly define the haplotype genome inherited from the father to the offspring and the haplotype genome inherited from the mother to the offspring;
[0057] The haplotype detection unit is used to perform tandem repeat variation detection on the haplotype genome inherited from the father to the offspring, the haplotype genome inherited from the mother to the offspring, and the two sets of offspring haplotype genomes;
[0058] The analysis unit is used to analyze the tandem repeat variation information of the two haplotype genomes of the offspring during the inheritance process.
[0059] The device based on the present invention performs the following detection:
[0060] 1. Prepare a core family, sample the father, mother, and offspring, and perform third-generation HiFi sequencing on the samples. The sequencing data volume is 170G, and the average sequencing depth is 18.4×. Long reads of the father, mother, and offspring are obtained respectively (hereinafter referred to as long reads);
[0061] In the offspring typing and assembly unit: based on the paternal reads and maternal reads, the find-unique-kmers analysis of the offspring reads ( Figure 2 The P1 in the PCR product was typed to obtain the paternal and maternal specific kmer sequences. Based on the classify_by_kmers method, the offspring sequencing reads were judged as paternal reads, maternal reads, or untyped reads (not assembled) according to the paternal and maternal specific kmer sequences. Hifiasm was used for assembly to obtain two sets of offspring haplotype genomes ( Figure 2 P and M in );
[0062] In the parent typing unit: using the obtained offspring haplotype genome as a reference, single nucleotide site variation detection is performed on the paternal and maternal genomes, and the paternal reads and maternal reads are typed based on the WhatsHap method to obtain two sets of paternal haplotype genomes (hap1, hap2 of Fa) and two sets of maternal haplotype genomes (hap1, hap2 of Mo), and the haplotype genome inherited from the paternal to the offspring is determined ( Figure 2 hap1 in Fa), haploid genome inherited from mother to offspring ( Figure 2 The results are shown in Table 1 below.
[0063] Table 1
[0064] Sample Read_num Base_num(Gb) Phased_read_rate(%) hap1 hap2 Fa 4434073 62.44 98.72 2743164 1634016 Mo 3632160 53.06 99.80 2243035 1381715 P1 3251252 54.56 86.58 1420527 1394308
[0065] In Table 1, Sample refers to sample; Read-num refers to the number of reads; Base-num refers to the number of bases; Phased_read_rate refers to typing efficiency; hap1 refers to the number of reads of the first set of haplotype genomes; and hap2 refers to the number of reads of the second set of haplotype genomes.
[0066] In the haplotype detection unit, the improved Tandem Repeat Finder method is used to perform tandem repeat variation detection on the haplotype genome inherited from the father to the offspring, the haplotype genome inherited from the mother to the offspring, and the two sets of offspring haplotype genomes. Among them, the improved Tandem Repeat Finder method needs to take into account the motif cycle characteristics (motif base sequence) and the characteristics of the detection segment (when detecting a certain interval, consider the copy number of the entire segment between the start and end points of the segment), so as to obtain the copy number information of all known repeat intervals on the reference genome in the specific sample. As a result, a total of 1,025,347 tandem repeat (TR) site copy information was obtained, and the motif length distribution of these sites is as follows Figure 3 The tandem repeat typing copy number information of each sample was obtained statistically. Taking three intervals as an example, the results are shown in Table 2 below.
[0067] Table 2
[0068]
[0069]
[0070] In Table 2, "start" refers to the start position of the interval; "end" refers to the end position of the interval; "motif" refers to the base sequence of the motif; "motiflength" refers to the base length of the motif (in base pairs); "copy number" refers to the number of copies of the interval in the reference genome; "Father" refers to the tandem repeat variant detection result of the father; "Mother" refers to the tandem repeat variant detection result of the mother; and "p1" refers to the tandem repeat variant detection result of the offspring. In the variant detection results for the father, mother, and offspring, "GT" indicates the number of copies detected in the interval for the sample, "DP" indicates the sequencing read depth covering the interval in the sample, and "hap1" and "hap2" record the copy number information detected for each read.
[0071] Table 2 above shows the copy numbers of 1,025,347 tandem repeat intervals in parents and offspring.
[0072] Based on the core family, the variation information of tandem duplications in family inheritance was analyzed. Taking the three sites in Table 2 as examples, the results are shown in Table 3.
[0073] Table 3
[0074] Chr start end motif Motiflength Copynumber Diff_p1_fa Diff_p1_mo chrX 268059 268998 AGAGAG 6 155.2 -37.3 90 chrX 268071 269016 AGAGAGAC 8 116.6 -3.9 107.6 chrX 314102 314833 AGGA 4 174.5 -3.7 366.1
[0075] In Table 3, start refers to the starting position of the interval; end refers to the ending position of the interval; motif refers to the motif base sequence; Motiflength refers to the motif base length (bp); Copy number refers to the copy number (pieces); Diff_p1_fa refers to the number of copy number changes in the offspring compared with the father during the inheritance process (pieces); Diff_p1_mo refers to the number of copy number changes in the offspring compared with the mother during the inheritance process (pieces).
[0076] [Comparative Example 1]
[0077] A tandem repeat variation typing detection device based on the Straglr method.
[0078] The same nuclear family as in Example 1 was used to sample the father, mother, and offspring.
[0079] Tandem repeat variation typing detection was performed according to the method disclosed in Chiu, R., Rajan-Babu, I.-S., Friedman, JM, and Birol, I. (2021). Straglr: discovering and genotyping tandem repeat expansions using whole genome long-read sequences. Genome Biol 22, 224.10.1186 / s13059-021-02447-3.
[0080] According to this method, the number of effective sites detected in the three samples of the father, mother, and offspring is shown in Table 4. At the same time, the number of effective sites detected by the embodiment of the present invention is compared, and the results are shown in Table 4.
[0081] Table 4. Number of effective sites
[0082] Fa(piece) Mo(piece) P1 (pieces) Comparative Example 1 2959 2907 2962 Example 1 1025347 1025347 1025347
[0083] Comparing Example 1 using the device of the present invention with Comparative Example 1, it can be seen that the device of the present invention can generate more copy number information of effective sites, which is far more than the information that can be provided by the Straglr method.
[0084] [Comparative Example 2]
[0085] A tandem repeat variation typing detection device based on the TRGT method.
[0086] This method is officially recommended by PacBio. It is based on the known TR region and provides copy number information for two types.
[0087] The results of the two software for P1 samples were compared, and the number of sites in different copy number difference intervals was counted. The results are shown in Table 5 and Figure 4 shown.
[0088] Table 5
[0089] Copy number difference distribution (units) Count (frequency) ~-100 8339 -100~-50 412 -50~-20 1358 -20~-10 2779 -10~-5 6900 -5~0 160879 0~5 101920 5~10 5555 10~20 4180 20~50 1817 50~100 595 100~ 5528
[0090] In Table 5, distribution refers to the distribution of copy number differences obtained by the TRGT method and Example 1; count refers to the frequency under this distribution.
[0091] Through Table 5 and Figure 4 It can be seen that the distribution of the copy number difference between the TRGT method and the device of Example 1 is mainly concentrated between -5 to 0 and 0 to 5. It was found that 87.5% of the site copy number differences were within 5 copies, indicating that the accuracy of the device of the present invention is relatively high.
[0092] Compared with TRGT, the two sets of typing results provided by the device of the present invention are two sets of genomes obtained based on the family typing strategy, rather than simple clustering, so the variation information of TR in the genetic process finally obtained by the device of the present invention is more accurate.
[0093] In addition, through verification of local sites, as shown in Table 6, it was found that at some sites, TRGT would obtain abnormal amplification copy numbers that did not exist, that is, false positives, and at these sites, the accuracy of the device of the present invention was better than TRGT.
[0094] Table 6
[0095]
[0096]
[0097] As can be seen from Example 1, Comparative Example 1, and Comparative Example 2, the present invention obtains copy number variation information for 1,025,347 tandem repeat (TR) sites inherited from parents to offspring (Diff_p1_fa, Diff_p1_mo in Table 3). This information can be used for etiology analysis of genetic diseases, analysis of de novo mutations, and the like. However, the number of valid sites detected by Comparative Example 1 is far less than that detected by the device of the present invention (Table 4). Comparative Example 2 suffers from a high number of false positives at some sites (Table 6). Therefore, the device of the present invention has higher accuracy and comprehensiveness.
[0098] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0099] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
Claims
1. A tandem repeat variation typing detection device based on a core family, the device comprising a progeny typing and assembly unit, a parent typing unit and a haplotype detection unit, wherein: The offspring typing and assembly unit is used to perform a first typing and assembly on the offspring sequencing dataset based on the paternal sequencing dataset and the maternal sequencing dataset to obtain two sets of offspring haplotype genomes; The parental typing unit is used to perform single nucleotide site variation detection on the paternal and maternal genomes respectively, using the obtained offspring haplotype genome as a reference, and perform a second typing on the paternal sequencing data set and the maternal sequencing data set respectively, to obtain two sets of paternal haplotype genomes and two sets of maternal haplotype genomes, and to clearly define the haplotype genome inherited from the father to the offspring and the haplotype genome inherited from the mother to the offspring; The haplotype detection unit is used to perform tandem repeat variation detection on the haplotype genome inherited from the father to the offspring, the haplotype genome inherited from the mother to the offspring, and the two sets of offspring haplotype genomes; The method for detecting tandem repeat variations is an improved Tandem Repeat Finder method; The improvements include motif loop characteristics and detection segment features; The sequencing data set includes long reads obtained by the third-generation sequencing method of the genome; The first typing method includes a find-unique-kmers analysis based on the trio_binning method to obtain paternal and maternal specific kmer sequences; Based on the classify_by_kmers method and according to the paternal and maternal specific kmer sequences, the offspring sequencing dataset is judged as a sequencing dataset belonging to the paternal parent, a sequencing dataset belonging to the maternal parent, or an unclassified sequencing dataset; The second typing method is based on the WhatsHap method.
2. The device according to claim 1, characterized in that The assembly method is selected from at least one of hifiasm, wtdbg2, canu, and nextdenovo.
3. The device according to any one of claims 1 to 2, characterized in that The device also includes an analysis unit for analyzing the tandem repeat variation information of the two haploid genomes of the offspring during the inheritance process.
4. Application of the nuclear family-based tandem repeat variant typing detection device according to any one of claims 1 to 3 in biology; the application is application in disease pathology research, paternity testing, criminal investigation or population genetics.