Method for gene-level haploid typing based on PacBio HiFi three-generation DNA sequencing data
Gene-level haploid typing is performed through PacBio HiFi third-generation DNA sequencing data, which solves the problems of high resource consumption and high data error correction requirements in the existing technology, and realizes high-precision haploid variant detection, which is suitable for gene expression analysis of rare variants in medicine.
Patent Information
- Application Number
- CN202411994085.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
The existing haplotype construction methods rely on data across the entire genome, which consumes a lot of resources and requires high error correction requirements, and requires parental data to be difficult to obtain, resulting in inefficiency.
The third-generation DNA sequencing data of PacBio HiFi was used for gene-level haploid typing. By filtering short reads and long sequences, aligning them to the reference genome, variant detection and filtering, haploid typing and typing were constructed. The long reads and high accuracy characteristics of PacBio HiFi were used to accurately detect haploid mutation information of human genes.
It realizes high-precision gene-level haploid typing, saves time and resources, and can quickly and accurately obtain haploid information about individual genes, avoid false positive problems, and is suitable for variant detection of important genes in medicine.
Smart Images

Figure CN120015112A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular to a method for performing haploid typing at a gene level based on PacBio HiFi third-generation DNA sequencing data. Background Art
[0002] The research demand for individual haplotypes has promoted the continuous development and progress of haplotype construction methods and tools. Reconstructing haplotypes only from individual whole genome sequencing data without relying on parental data and population sequencing has become a research hotspot in this field. There are two main types of haplotype construction methods at present: one is a haplotype typing method based on alignment, and the other is a haplotype typing method based on assembly. Alignment-based haplotype analysis usually uses pre-identified heterozygous marker sites (mainly SNP and InDel markers) and sequencing fragments as input information, and converts the haplotype assembly process into minimum fragment removal, longest haplotype construction, minimum error correction, etc. The read length advantage of the third-generation sequencing sequence has great application potential in constructing haplotypes. In recent years, many genome typing tools based on third-generation sequences have been published, among which the whatHap method is representative. There are two main types of assembly-based haplotype typing methods. One is to have parental data, refer to the parental data for typing during assembly, and assemble two sets of haplotype genomes. The other method is to divide the assembled reads into two sets of haploid types based on the variant site information on the reference genome, and then assemble them separately. Since the above methods are based on haploid typing across the entire genome, the number of sequences and variant sites is relatively large, it takes a lot of time, consumes relatively large resources, and has higher requirements for data error correction. At the same time, some analysis methods need to rely on existing data, and parental data of offspring is usually difficult to obtain.
[0003] Based on this, the present invention provides a high-precision method for detecting haploid information at the gene level. Summary of the invention
[0004] Based on the above description, the present invention provides a method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data. This method combines the advantages of long read length, high precision and high accuracy of PacBio HiFi sequencing data, and can accurately obtain human gene haploid variation information based on human polymorphic site information.
[0005] The technical solution of the present invention to solve the above technical problems is as follows: A method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data. This method can perform haploid typing at the gene level and detect mutation information of two haplotypes of humans. The input files include: reads sequence of the third-generation PacBioHiFi and reference genome; the method includes the following steps: Filter short read length sequences: filter invalid short read length sequences less than 1000bp; Alignment: After filtering, use alignment software to align the third-generation sequencing data to the reference genome and generate a bam format file; Variant detection and filtering: Use the variation detection software DeepVariant to identify the original SNP sites in the sequenced genome data from the bam file, and filter these original SNP sites. The remaining mutation site set is used as the input for downstream operation processing. A sequencing fragment (read) contains multiple mutation sites and genotypes, which are regarded as a "local haplotype" of the haploid sequence; Constructing haplotypes: Using the overlapping relationship between local haplotypes on genome reads, construct haplotypes covering a large range of the genome; Typing: Align the reads to the reference genome, and type the sequenced reads according to the heterozygous variant sites on the sequenced reads to distinguish which sequenced fragments are from the father and which are from the mother. Then, assemble the fragments of the two genotypes separately to form two sets of genotypes. Construct the evaluation index: typing accuracy = 100% - the proportion of incorrect consecutive SNP pairs in the typing results. The larger the typing accuracy value, the higher the typing accuracy. Calculate the N50 value. The larger the N50 value, the better the continuity of the obtained haplotype.
[0006] Based on the above technical solution, the present invention can also be improved as follows.
[0007] Furthermore, in the above-mentioned method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data, the alignment step is: using minimap2 to align the PacBio HiFi data, aligning the HiFi data to the reference genome, and generating a bam format file.
[0008] Furthermore, in the above-mentioned method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data, in the variation detection and filtering steps, these original variation sites need to be filtered through the following steps: retaining PASS sites, filtering homozygous sites, deep filtering, and deviation processing.
[0009] Furthermore, in the above-mentioned method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data, the PASS sites are retained: only sites with a FILTER field value of "PASS" are retained, and two adjacent SNP sites with a distance of less than 3 bases are eliminated.
[0010] Furthermore, in the above-mentioned method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data, the homozygous sites are filtered: DeepVariant identifies the original SNP sites and their genotypes at the SNP sites contained in the sequencing reads data from the bam file, and stores them in the VCF file. In the VCF file, the sites with "GT" fields of "1 / 1", "0 / 0", ". / ." are filtered, and the heterozygous variation site "0 / 1" is retained; at the same time, the sites with a ratio of the number of reads of the two genotypes exceeding 80% or less than 20% are filtered out.
[0011] Furthermore, in the above-mentioned method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data, the depth filtering is as follows: if the number of reads under a site is greater than 2 times the average depth or less than half the average depth, the site is filtered, and the average depth is defined as the average number of reads under the SNP site.
[0012] Furthermore, the above-mentioned method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data, the deviation processing: the VCF file output by DeepVariant will record the position P of each SNP site, with position P as the center, extending 200bp before and after, the variant site information in the P+200bp and P-200bp regions is the amplified variant signal, and it is judged whether it is a genotype that meets the haploid typing, the genotype that does not meet the haploid typing is filtered out, and the genotype that meets the haploid typing is retained.
[0013] Furthermore, in the above-mentioned method for performing haplotype typing at the genetic level based on PacBio HiFi third-generation DNA sequencing data, in the step of constructing haplotypes, reads from the same haplotype can form a continuous path through overlapping SNP sites, and the paths formed in the graph by reads from different haplotypes do not cross each other. The higher the number of sequencing depths of reads supporting the path, the more reliable the path is; otherwise, it may be a noise path caused by errors and has lower reliability.
[0014] Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects: The third generation of PacBio HiFi data can cover the human gene sequence, and therefore can also fully sequence the human gene sequence. The sequencing length of PacBio HiFi data is up to 25kb, which can provide sufficient accurate coverage reads support for the vast majority of genes to obtain the corresponding structural variation information. Since the accuracy of PacBio HiFi data is as high as 99.9%, it can effectively avoid false positive problems. Haploids can be used to help detect and correct erroneous or missing sequencing data, which is crucial to interpreting individual genomes from a medical perspective, especially considering rare (or "private") variations in gene expression or function that have important medical effects.
[0015] This method combines the characteristics of the third-generation PacBio HiFi sequencing data. First, HiFi data is used to compare with the reference genome, and the corresponding mutation site information is found within the gene range to distinguish the haploid type information. Through longer gene fragments, the information of the overlap region on the reads fragments is then used to distinguish the haploid type information. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A flow chart of a method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data provided by the present invention; Figure 2 The method of step 3 of the present invention is to construct a local haplotype map based on the genotype information on PacBio HiFi sequencing reads; Figure 3 A diagram showing the relationship between overlapping regions obtained based on the long read length of PacBio HiFi sequencing reads in step 4 of the present invention; Figure 4 This is step five of the present invention, in which two genotype fragments are distinguished according to the heterozygous sites on the reads, and two sets of haploid sequence maps are obtained by assembling them separately. DETAILED DESCRIPTION
[0017] In order to facilitate understanding of the present application, the present application will be described more comprehensively below. The present application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0019] It should be noted that when an element is considered to be "connected" to another element, it can be directly connected to the other element, or connected to the other element through an intermediate element. The "connection" in the following embodiments should be understood as "electrical connection", "communication connection", etc. if the connected circuits, modules, units, etc. have electrical signals or data transmission between each other.
[0020] When used herein, the singular forms "a", "an", and "said / the" may also include plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "include / comprise" or "have" etc. specify the presence of stated features, wholes, steps, operations, components, parts or combinations thereof, but do not exclude the possibility of the presence or addition of one or more other features, wholes, steps, operations, components, parts or combinations thereof.
[0021] A method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data. This method can perform haploid typing at the gene level and detect mutation information of two haplotypes of humans. The input files include: reads sequence of the third-generation PacBioHiFi and reference genome. Sequencing genome: DNA reads sequence obtained by sequencing the sample. Reference genome: Published genome, such as hg19, hg38, etc.
[0022] The method includes the following steps: Filter short read length sequences: filter invalid short read length sequences less than 1000bp; Alignment: After filtering, use alignment software to align the third-generation sequencing data to the reference genome and generate a bam format file; Variant detection and filtering: Use the variation detection software DeepVariant to identify the original SNP sites in the sequenced genome data from the bam file, and filter these original SNP sites. The remaining mutation site set is used as the input for downstream operation processing. A sequencing fragment (read) contains multiple mutation sites and genotypes, which are regarded as a "local haplotype" of the haploid sequence; Constructing haplotypes: Using the overlapping relationship between local haplotypes on genome reads, construct haplotypes covering a large range of the genome; Typing: Align the reads to the reference genome, and type the sequenced reads according to the heterozygous variant sites on the sequenced reads to distinguish which sequenced fragments are from the father and which are from the mother. Then, assemble the fragments of the two genotypes separately to form two sets of genotypes. Construct the evaluation index: typing accuracy = 100% - the proportion of incorrect consecutive SNP pairs in the typing results. The larger the typing accuracy value, the higher the typing accuracy. Calculate the N50 value. The larger the N50 value, the better the continuity of the obtained haplotype.
[0023] Based on the above technical solution, the present invention can also be improved as follows.
[0024] Furthermore, in the above-mentioned method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data, the alignment step is: using minimap2 to align the PacBio HiFi data, aligning the HiFi data to the reference genome, and generating a bam format file.
[0025] Furthermore, in the above-mentioned method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data, in the variation detection and filtering steps, these original variation sites need to be filtered through the following steps: retaining PASS sites, filtering homozygous sites, deep filtering, and deviation processing.
[0026] Furthermore, in the above-mentioned method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data, the PASS sites are retained: only sites with a FILTER field value of "PASS" are retained, and two adjacent SNP sites with a distance of less than 3 bases are eliminated.
[0027] Furthermore, in the above-mentioned method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data, the homozygous sites are filtered: DeepVariant identifies the original SNP sites and their genotypes at the SNP sites contained in the sequencing reads data from the bam file, and stores them in the VCF file. In the VCF file, the sites with "GT" fields of "1 / 1", "0 / 0", ". / ." are filtered, and the heterozygous variation site "0 / 1" is retained; at the same time, the sites with a ratio of the number of reads of the two genotypes exceeding 80% or less than 20% are filtered out.
[0028] Furthermore, in the above-mentioned method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data, the depth filtering is as follows: if the number of reads under a site is greater than 2 times the average depth or less than half the average depth, the site is filtered, and the average depth is defined as the average number of reads under the SNP site.
[0029] Furthermore, the above-mentioned method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data, the deviation processing: the VCF file output by DeepVariant will record the position P of each SNP site, with position P as the center, extending 200bp before and after, the variant site information in the P+200bp and P-200bp regions is the amplified variant signal, and it is judged whether it is a genotype that meets the haploid typing, the genotype that does not meet the haploid typing is filtered out, and the genotype that meets the haploid typing is retained.
[0030] Furthermore, in the above-mentioned method for performing haplotype typing at the genetic level based on PacBio HiFi third-generation DNA sequencing data, in the step of constructing haplotypes, reads from the same haplotype can form a continuous path through overlapping SNP sites, and the paths formed in the graph by reads from different haplotypes do not cross each other. The higher the number of sequencing depths of reads supporting the path, the more reliable the path is; otherwise, it may be a noise path caused by errors and has lower reliability.
[0031] Example 1 The present invention provides a high-precision method for detecting haploid information based on the gene level, which detects mutation information of two haploid types of humans in combination with sequencing data of the third generation PacBioHiFi.
[0032] The present invention performs haploid typing at the gene level rather than the entire genome level, which helps save time and resources and obtain haploid information of genes of interest to customers more quickly and accurately.
[0033] The method includes the following steps: Input files: 1. Read sequences of the third generation PacBio HiFi; 2. Reference genome.
[0034] Step 1: Filter short reads: Filter invalid short reads less than 1000bp. The third-generation sequencing is characterized by long read length, and the average sequence length can reach 20kb, which facilitates the detection and determination of gene variation and genotype information. Compared with the second-generation data, the variation information of the third-generation data is more clustered, and the sites where the variation information is located within the same variation range are more concentrated. During the sequencing process, shorter sequencing sequences may be generated due to sequencing errors and other reasons. This sequence itself has variation signal errors, and may also be too short to affect the accurate classification of most sequences. Before comparison, it is first necessary to filter the shorter sequences in all sequencing sequences to avoid the impact of short and invalid sequences on subsequent results, so that subsequent sequence comparison results are more accurate.
[0035] Step 2: Alignment: After filtering, use alignment software to align the third-generation sequencing data to the reference genome and generate a bam file. Use minimap2 to align the PacBio HiFi data to the reference genome.
[0036] Step 3: Variation detection and filtering: Use the variation detection software DeepVariant to identify SNP sites in the sequenced genome data from the bam file. These original variation sites need to be filtered through the following 4 steps.
[0037] (1) DeepVariant identifies the original SNP sites and their genotypes contained in the sequencing reads data from the bam file and stores them in the VCF file. In the VCF file, only sites with a "PASS" field value in the FILTER are retained. In addition, sites with a distance of less than 3 bases between two adjacent SNP sites also need to be filtered out. This step is mainly to filter out densely populated heterozygous sites.
[0038] (2) Homozygous SNP sites, that is, sites with "1 / 1", "0 / 0", or ". / ." in the "GT" field in the VCF file need to be filtered out, leaving the heterozygous variant site "0 / 1". In addition, if the ratio of the number of reads of the two genotypes under a site exceeds 80% or is less than 20%, they also need to be filtered out. Identify the SNP sites contained in each read and its genotype at the SNP site from the bam file. The SNP genotype belongs to one of A, T, C, G, and - (missing). "-" means that the site has not been detected.
[0039] (3) If the number of reads under a site is greater than 2 times the average depth or less than half the average depth, such sites also need to be filtered out (the average depth is defined as the average number of reads under the SNP site).
[0040] (4) The VCF file output by DeepVariant will record the position P of each SNP site. Considering that the variation signals of multiple sequencing may have deviations, the position deviation is mostly a few bp or dozens of bp. For the deviation, if no processing is done, the range of variation signals traversed and counted is too narrow to fully consider all possible situations. Therefore, when defining the variation range, each of the current variation sites is expanded by 200bp (within the nearby 200bp area, there are DELs with the same sample, but the starting position is a few bp different. This type of variation can also be used for haploid typing to detect the mutation information of two haplotypes of a person), so as to include variation signals with partial position deviations. Specifically: with position P as the center, expand 200bp before and after, and the variation site information of P+200bp and P-200bp areas is the amplified variation signal. If the variation signals in the rest of the range are included in the variation range, the genotype information may be misjudged due to the wrong distribution characteristics of the variation signals. Therefore, it is necessary to screen all the amplified variation signals and determine whether they are valid signals. That is, whether the genotype is consistent with haploid typing, filter out the genotypes that do not meet the haploid typing, and retain the genotypes that meet the haploid typing.
[0041] After the above steps, the remaining mutation site set is filtered and used as input for downstream processing. A sequencing fragment (read) contains multiple mutation sites and genotypes, which can be regarded as a "local haplotype" of the haploid sequence, such as Figure 2 shown.
[0042] Step 4: Construct haplotype: Figure 3 As shown in the figure, due to its long read length, the third-generation sequencing data can easily span longer regions and obtain better overlap areas. By utilizing the overlapping relationship between local haplotypes on genome reads, haplotypes covering a large range of the genome are constructed. Under ideal conditions, reads from the same haplotype can form a continuous path through overlapping SNP sites, and the paths formed by reads from different haplotypes in the graph do not intersect each other. The higher the number of reads sequencing depth supporting the path, the more reliable the path. On the contrary, it is very likely to be a noise path caused by errors, and the reliability is relatively low. Since PacBio HiFireads do not require DNA amplification, the coverage preference brought by the PCR process is eliminated, and uniform coverage of the integrated genome can be achieved without introducing errors caused by amplification. Due to the long read length, it can cover areas with low diversity and can cross complex areas such as repetitive sequences. Due to the high accuracy of HiFi data, the genotype bases appearing at the SNP site are accurately read, so as to construct haploids at the gene level, reducing some complex structures such as branches, bubbles, and loop nesting.
[0043] Step 5 Typing: Typing the sequenced reads, distinguish which sequenced fragments are from the father and which are from the mother based on the heterozygous variant sites on the sequenced reads, and then assemble the fragments of the two genotypes separately to form two sets of genotypes. It is possible to distinguish the two types of F / M based on the variant sites. Align the reads to the reference genome, and distinguish the reads based on the variant site information on the sequenced reads. The red reads indicate that they are from the father, and the blue reads indicate that they are from the mother. After typing the reads, use the assembly tool Falcon to assemble them separately to obtain two sets of haploid sequences, such as Figure 4 shown.
[0044] Step 6: Construct evaluation indicators: Typing accuracy = 100% - the proportion of incorrect continuous SNP pairs in the typing results. For individuals, the sequences are long or short, and the number of incorrect typing is more or less. Assuming that there are 7 consecutive SNP pairs in the sequence, if 3 of them are incorrectly typed, the proportion of incorrect SNP pairs is 3 / 7=0.4285*100%, then the typing accuracy is 100%-42.85%=57.14%. The larger the typing accuracy value, the higher the typing accuracy. Calculate the N50 value. The larger the N50 value, the better the continuity of the obtained haplotype.
[0045] The third generation of PacBio HiFi data can cover the human gene sequence, and therefore can also fully sequence the human gene sequence. The sequencing length of PacBio HiFi data is up to 25kb, which can provide sufficient accurate coverage reads support for the vast majority of genes to obtain the corresponding structural variation information. Since the accuracy of PacBio HiFi data is as high as 99.9%, it can effectively avoid false positive problems. Haploids can be used to help detect and correct erroneous or missing sequencing data, which is crucial to interpreting individual genomes from a medical perspective, especially considering rare (or "private") variations in gene expression or function that have important medical effects.
[0046] This method first uses HiFi data to compare with the reference genome, and within the gene range, finds the corresponding mutation site information and distinguishes the haploid type information. Through longer gene fragments, the overlap region (the region formed by overlapping SNP sites of reads from the same haploid type) on the reads fragment is used to distinguish the haploid type information. This method can combine the advantages of long read length, high precision and high accuracy of the third-generation PacBio HiFi sequencing data, and accurately obtain the haploid variation information of human genes based on the polymorphic site information of humans.
[0047] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for performing haploid typing at the gene level based on PacBio HiFi third-generation DNA sequencing data, characterized in that: This method can perform haplotype typing at the gene level and detect mutation information of two haplotypes of humans. The input files include: reads sequence of the third generation PacBio HiFi and reference genome. The method includes the following steps: Filter short read length sequences: filter invalid short read length sequences less than 1000bp; Alignment: After filtering, use alignment software to align the third-generation sequencing data to the reference genome and generate a bam format file; Variant detection and filtering: Use the variation detection software DeepVariant to identify the original SNP sites in the sequenced genome data from the bam file, and filter these original SNP sites. The remaining mutation site set is used as the input for downstream operation processing. A sequencing fragment (read) contains multiple mutation sites and genotypes, which are regarded as a "local haplotype" of the haploid sequence; Constructing haplotypes: Using the overlapping relationship between local haplotypes on genome reads, construct haplotypes covering a large range of the genome; Typing: Align the reads to the reference genome, and type the sequenced reads according to the heterozygous variant sites on the sequenced reads to distinguish which sequenced fragments are from the father and which are from the mother. Then, assemble the fragments of the two genotypes separately to form two sets of haploid sequences; Construct the evaluation index: typing accuracy = 100% - the proportion of incorrect consecutive SNP pairs in the typing results. The larger the typing accuracy value, the higher the typing accuracy. Calculate the N50 value. The larger the N50 value, the better the continuity of the obtained haplotype.
2. The method for performing gene-level haplotype typing based on PacBio HiFi third-generation DNA sequencing data according to claim 1, characterized in that: The alignment step is: using minimap2 to align the PacBio HiFi data, aligning the HiFi data to the reference genome, and generating a bam format file.
3. The method for performing gene-level haplotype typing based on PacBio HiFi third-generation DNA sequencing data according to claim 1, characterized in that: In the variation detection and filtering step, these original variation sites need to be filtered through the following steps: retaining PASS sites, filtering homozygous sites, deep filtering, and deviation processing.
4. The method for performing gene-level haplotype typing based on PacBio HiFi third-generation DNA sequencing data according to claim 3, characterized in that: The retained PASS sites: only the sites with the FILTER field value of "PASS" are retained, and two adjacent SNP sites with a distance of less than 3 bases are eliminated.
5. The method for performing gene-level haplotype typing based on PacBio HiFi third-generation DNA sequencing data according to claim 3, characterized in that: The filtering of homozygous sites: DeepVariant identifies the original SNP sites and their genotypes at the SNP sites contained in the sequencing reads data from the bam file and stores them in the VCF file. In the VCF file, the sites with "GT" fields of "1 / 1", "0 / 0", and ". / ." are filtered, and the sites of heterozygous variation "0 / 1" are retained; at the same time, the sites with the ratio of the number of reads of the two genotypes exceeding 80% or less than 20% are filtered out.
6. The method for performing gene-level haplotype typing based on PacBio HiFi third-generation DNA sequencing data according to claim 3, characterized in that: The depth filtering: if the number of reads under a site is greater than 2 times the average depth or less than half the average depth, the site is filtered. The average depth is defined as the average number of reads under the SNP site.
7. The method for performing gene-level haplotype typing based on PacBio HiFi third-generation DNA sequencing data according to claim 3, characterized in that: The deviation processing: The VCF file output by DeepVariant will record the position P of each SNP site, expand 200bp before and after the position P, and determine whether it is a genotype that conforms to haplotype typing, filter out genotypes that do not conform to haplotype typing, and retain genotypes that conform to haplotype typing.
8. The method for performing gene-level haplotype typing based on PacBio HiFi third-generation DNA sequencing data according to claim 1, characterized in that: In the step of constructing haplotypes, reads from the same haplotype can form a continuous path through overlapping SNP sites, and the paths formed by reads from different haplotypes in the graph do not cross each other. The higher the number of sequencing depths of reads supporting the path, the more reliable the path is; otherwise, it may be a noise path caused by errors and has low reliability.
Citation Information
Cited By
Haploid horizontal transposon insertion identification method and system based on physical typing
CN120526846A
Method and system for haplotype level transposon insertion identification based on physical typing
CN120526846B