Method, device, medium, and program product for determining a target fragment haplotype
By using single-molecule long-read sequencing and recursive partitioning of heterozygous SNPs, the method addresses inefficiencies in haplotype phasing for multi-copy genes, achieving enhanced accuracy and efficiency in resolving gene copy differences.
Patent Information
- Application Number
- CN202510541354.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-27
AI Technical Summary
When processing haplotype phasing of multi-copy genes, the prior art has insufficient resolution, low computational efficiency and low reliability of the results, making it difficult to accurately distinguish the difference in haplotypes between different gene copies.
By acquiring the sequencing sequences and reference genome sequences, the information of heterozygous single nucleotide polymorphisms is determined, and multiple recursive divisions are performed based on the contribution degree, and the contribution of each heterozygous single nucleotide polymorphism site to haplotype phase is systematically evaluated. The sequencing sequence is obtained using single-molecule long read and long sequencing technology.
The haplotype phasing efficiency and accuracy of multi-copy genes are improved, and the haplotype typing can be efficiently and accurately, which improves the haplotype phasing ability of genes with complex structures.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention generally relates to bioinformatics processing, and specifically, to a method, a computing device, a computer storage medium, and a computer program product for determining the haplotype of a target fragment. Background Art
[0002] A haplotype refers to a combination of a series of genetic variant sites coexisting on a single chromosome, which is a combination of allelic sequences from the same parent and describes the sequence differences on a single chromosome segment. The combination of different genetic loci between homologous chromosomes has an important impact on biological phenotypes, such as human genetic diseases. Haplotype phasing aims to obtain a haplotype-resolved assembled sequence, that is, to obtain a complete haplotype genome sequence, so as to deeply understand the combination and inheritance of different genetic loci on a single chromosome or a specific single-chromosome region, which is helpful for understanding and diagnosing the pathogenesis of genetic diseases.
[0003] Currently, with the development of sequencing technology, the high-resolution sequencing data of genomes is increasing day by day, and great progress has been made in haplotype phasing technology. However, haplotype phasing for multi-copy genes remains a challenging problem. Traditional analysis methods often have difficulty accurately distinguishing haplotype differences between highly similar different gene copies. Existing technologies often have problems such as insufficient resolution, low computational efficiency, and low reliability of results when dealing with such complex situations.
[0004] Currently, software for haplotype phasing detection, such as Whatshap, has obvious limitations. First, these software usually can only handle the classical situation of two alleles and cannot effectively process the presence of multi-copy genes. Second, compared with traditional methods, the detection performance of these software is lower and prone to a higher error rate, which may lead to inaccurate or even misleading analysis results.
[0005] Therefore, there is still an urgent need for a new efficient and accurate algorithm for haplotype phasing of multi-copy genes. Summary of the Invention
[0006] The present invention provides a method, a computing device, a computer storage medium, and a computer program product for determining the haplotype of a target fragment, which can significantly improve the efficiency and accuracy of haplotype phasing for multi-copy genes.
[0007] According to a first aspect of the present invention, there is provided a method for determining the haplotype of a target fragment. The method includes: obtaining a set of input sequencing sequences of the target fragment from a sample to be tested, the set of input sequencing sequences including multiple sequencing sequences of the target fragment; aligning the sequencing sequences with a reference genome sequence to determine heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment; and performing multiple recursive partitions on the set of input sequencing sequences based on the contribution degree of each heterozygous single nucleotide polymorphism locus to haplotype phasing, so as to determine the haplotype of the target fragment.
[0008] In one embodiment, the sequencing sequence of the target fragment is a single molecule long read sequencing sequence of the target fragment obtained by a single molecule long read sequencing technology.
[0009] In one embodiment, the determining the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment includes: aligning the sequencing sequence with the reference genome sequence to determine the single nucleotide polymorphism locus corresponding to each sequencing sequence within the target fragment; retaining the single nucleotide polymorphism locus with the maximum base frequency within a preset base frequency range as the heterozygous single nucleotide polymorphism locus, and obtaining the base types of the heterozygous single nucleotide polymorphism locus corresponding to each sequencing sequence.
[0010] In one embodiment, the determining the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment further includes filtering the heterozygous single nucleotide polymorphism loci according to a preset filtering rule, and the preset filtering rule includes at least one of the following: removing the heterozygous single nucleotide polymorphism loci located in a region with a continuous number of identical bases exceeding a preset threshold; removing the heterozygous single nucleotide polymorphism loci located in a tandem repeat region; and removing the heterozygous single nucleotide polymorphism loci located in a preset "blacklist" region.
[0011] In one embodiment, the performing multiple recursive partitions on the set of input sequencing sequences based on the contribution degree of each heterozygous single nucleotide polymorphism locus to haplotype phasing, so as to determine the haplotype of the target fragment includes: in each recursion, based on the current set of input sequencing sequences, evaluating the contribution degree of each heterozygous single nucleotide polymorphism locus to haplotype phasing; based on the contribution degree, dividing the sequencing sequences in the current set of input sequencing sequences into multiple subsets; in response to any one of the subsets satisfying a termination condition, independently outputting the subset satisfying the termination condition as a haplotype of the target fragment; and in response to any one of the subsets not satisfying the termination condition, independently updating the current set of input sequencing sequences based on the subset not satisfying the termination condition, and advancing to the next recursion until the termination condition is satisfied.
[0012] In one embodiment, the method for evaluating the contribution degree of a heterozygous single nucleotide polymorphism locus to haplotype phasing includes: for the heterozygous single nucleotide polymorphism locus to be evaluated, dividing the sequencing sequences in the currently input sequencing sequence set into multiple subsets according to the base types of each sequencing sequence at the heterozygous single nucleotide polymorphism locus to be evaluated; traversing all heterozygous single nucleotide polymorphism loci of the target fragment, and for each of them: in response to the base types of each sequencing sequence at the heterozygous single nucleotide polymorphism locus being unique in each subset, adding 1 point to the contribution degree of the heterozygous single nucleotide polymorphism locus to be evaluated; in response to the base types of each sequencing sequence at the heterozygous single nucleotide polymorphism locus not being unique in any subset, not increasing the contribution degree of the heterozygous single nucleotide polymorphism locus to be evaluated; accumulating the total score as the contribution degree of the heterozygous single nucleotide polymorphism locus to be evaluated to haplotype phasing.
[0013] In one embodiment, dividing the sequencing sequences in the currently input sequencing sequence set into multiple subsets based on the contribution degree includes: selecting the heterozygous single nucleotide polymorphism locus with the highest contribution degree, and dividing the sequencing sequences in the currently input sequencing sequence set into multiple subsets according to the base types of each sequencing sequence at the heterozygous single nucleotide polymorphism locus with the highest contribution degree.
[0014] In one embodiment, when multiple heterozygous single nucleotide polymorphism loci have the same highest contribution degree, any one of them is arbitrarily selected as the basis for set division.
[0015] In one embodiment, the method for determining the haplotype of the target fragment further includes, in each recursion, processing the multiple subsets in parallel until all subsets meet the termination condition.
[0016] In one embodiment, the termination condition is that all sequencing sequences in the subset are compatible.
[0017] According to the second aspect of the present invention, there is also provided a computing device, which includes: a memory configured to store one or more computer programs; and a processor coupled to the memory and configured to execute one or more programs to cause the device to execute the method of the first aspect of the present invention.
[0018] According to the third aspect of the present invention, there is also provided a non-transitory computer-readable storage medium. Machine-executable instructions are stored on the non-transitory computer-readable storage medium, and when the machine-executable instructions are executed, the machine is caused to execute the method of the first aspect of the present invention.
[0019] According to a fourth aspect of the present invention, there is also provided a computer program product. The computer program product includes a computer program which, when executed by a machine, performs the method according to the first aspect of the present invention.
[0020] The Summary of the Invention section is provided to introduce, in a simplified form, a selection of concepts that are further described below in the Detailed Description. The Summary of the Invention section is not intended to identify key features or essential features of the present invention, nor is it intended to limit the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A schematic diagram of a system for implementing a method for determining a target fragment haplotype according to an embodiment of the present invention is shown.
[0022] Figure 2 A flowchart of a method for determining a target fragment haplotype according to an embodiment of the present invention is shown.
[0023] Figure 3 A flowchart of a method for performing multiple recursive partitions on an input sequencing sequence set to determine the haplotype of a target fragment according to an embodiment of the present invention is shown.
[0024] Figure 4 A flowchart of a method for evaluating the contribution of a heterozygous single nucleotide polymorphism site to haplotype phasing according to an embodiment of the present invention is shown.
[0025] Figure 5 A block diagram of an electronic device suitable for implementing an embodiment of the present invention is schematically shown.
[0026] In the respective drawings, the same or corresponding reference numerals denote the same or corresponding parts. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] Preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0028] As used herein, the term "comprising" and its variations mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.
[0029] As described above, for the haplotype typing of multi-copy genes, traditional analysis methods often have difficulty in accurately distinguishing haplotype differences between different gene copies. Existing software for haplotype typing detection also often has problems such as insufficient resolution, low computational efficiency, and low result reliability when dealing with such complex situations of multi-copy genes.
[0030] To at least partially solve one or more of the above problems and other potential problems, example embodiments of the present invention propose a method for determining the haplotype of a target fragment for multi-copy genes. This method systematically evaluates the contribution degree of each heterozygous single nucleotide polymorphism site to the accurate division of haplotypes, and based on this, performs multiple recursive divisions on multiple sequencing sequences of the target fragment, and determines the haplotype of the target fragment according to the compatibility of the sequencing sequences. The present invention can efficiently and accurately perform haplotype typing on multi-copy genes.
[0031] Figure 1 FIG. shows a schematic diagram of a system 100 for implementing the method for determining the haplotype of a target fragment according to an embodiment of the present invention. As Figure 1 shown, the system 100 includes: a computing device 110, a sequencing device 130. In some embodiments, the system 100 further includes: a server 140, a network 120. In some embodiments, the computing device 110, the server 140, and the sequencing device 130 can perform data interaction in a wired or wireless manner (such as the network 120).
[0032] Regarding the sequencing device 130, it is, for example, used to sequence a sample to be tested to generate a sequencing sequence of a target fragment; and send the generated sequencing sequence data to the computing device 110. In some embodiments, the server 140 can also send the sequencing sequence of the sample to be tested to the computing device 110. In some embodiments, the sequencing device 130 is, for example, a sequencing device based on single molecule long read length sequencing technology. The sequencing device 130 is, for example, but not limited to, a sequencer developed by Pacific Biosciences (PacBio) using single molecule real-time (SMRT) sequencing technology.
[0033] Regarding the server 140, it is, for example, used to provide sequence information of a reference genome. For example, provide human reference genome sequence information, and the human reference genome is, for example, but not limited to, the human reference genome GRCh37 or the human reference genome GRCh38.
[0034] Regarding computing device 110, which is used, for example, to determine a target fragment haplotype. Specifically, computing device 110 is used to obtain a set of input sequencing sequences of a target fragment from a sample to be tested, the set of input sequencing sequences including multiple sequencing sequences of the target fragment; align the sequencing sequences with a reference genomic sequence to determine heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment; and based on the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing, perform multiple recursive partitions on the set of input sequencing sequences, thereby determining the haplotype of the target fragment.
[0035] In some embodiments, computing device 110 may have one or more processing units, including dedicated processing units such as GPUs, FPGAs, and ASICs, as well as general-purpose processing units such as CPUs. Computing device 110 may be an integration of multiple physical servers, an integration of multiple processing units, etc. Additionally, one or more virtual machines may also be running on each computing device 110. In some embodiments, computing device 110 and server 130 may be integrated together or may be separately provided from each other.
[0036] In some embodiments, computing device 110 includes, for example: a sequencing sequence data acquisition unit 112, a heterozygous single nucleotide polymorphism site determination unit 114, and a haplotype phasing unit 116. The above-mentioned sequencing sequence data acquisition unit 112, heterozygous single nucleotide polymorphism site determination unit 114, and haplotype phasing unit 116 may be configured on one or more computing devices 110.
[0037] Regarding the sequencing sequence data acquisition unit 112, it is used to obtain a set of input sequencing sequences of a target fragment from a sample to be tested, the set of input sequencing sequences including multiple sequencing sequences of the target fragment.
[0038] Regarding the heterozygous single nucleotide polymorphism site determination unit 114, it is used to align the sequencing sequences with a reference genomic sequence to determine heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment.
[0039] Regarding the haplotype phasing unit 116, it is used to perform multiple recursive partitions on the set of input sequencing sequences based on the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing, thereby determining the haplotype of the target fragment.
[0040] The following will be combined with Figure 2 Describe a method for determining a target fragment haplotype according to an embodiment of the present invention. Figure 2 The flowchart of a method 200 for determining a target fragment haplotype according to an embodiment of the present invention is shown. It should be understood that method 200 may, for example, be inFigure 5 performed at the described electronic device 500. It can also be performed at Figure 1 the described computing device 110. It should be understood that method 200 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the present invention is not limited in this regard.
[0041] At step 202, the computing device 110 obtains a set of input sequencing sequences of a target fragment from a sample to be tested, and the set of input sequencing sequences includes multiple sequencing sequences of the target fragment.
[0042] Regarding the sequencing sequences of the target fragment from the sample to be tested, for example but not limited to, single-molecule long-read sequencing sequences of the target fragment obtained by single-molecule long-read sequencing technology. For example, the sequencing sequences can be single-molecule long-read sequencing sequences obtained by using a single-molecule sequencing platform (such as the SMRT sequencing platform or the nanopore sequencing platform) after specific amplification or capture of the target fragment by one or more pairs of primers. Compared with short-read sequencing or whole-genome sequencing, the single-molecule long-read sequencing obtained based on specific amplification or capture can more completely cover gene regions with complex structures or variable copy numbers, thereby facilitating the improvement of the accuracy of subsequent haplotype phasing.
[0043] After sequencing is completed, the sequencing sequences can be subjected to preliminary quality control and filtering, such as including but not limited to processes such as low-quality read removal, sequencing error correction, and redundant sequence removal. Those skilled in the art should understand that the above quality control and filtering can be implemented using any suitable quality control and filtering tools, and their specific parameters can be flexibly set according to the sequencing platform, sample characteristics, and experimental purposes.
[0044] Regarding single-molecule long-read sequencing technologies, there are multiple types. For example, but not limited to, Pacific Biosciences (PacBio) has developed sequencers that use single-molecule real-time (SMRT) sequencing technology (such as the RS I, RS II, and Sequel machines). Compared with second-generation sequencing technologies, taking the single-molecule long-read sequencing platform PacBio Sequel II as an example, it uses SMRT sequencing technology to achieve single-molecule real-time sequencing. The principle of SMRT sequencing is based on SMRT Cells. Each SMRT Cell is covered with millions of zero-mode waveguide holes (ZMWs). During sequencing, DNA polymerase and a template molecule are anchored at the bottom of the ZMWs for reaction. The excitation light at the bottom of the small holes can excite the fluorescent labels on the nucleotide substrates, and then the fluorescence signals are recorded by the monitoring system to obtain base information. Another example is the nanopore sequencing device developed by Oxford Nanopore Technologies (ONT). It directly reads sequence information by measuring the current changes caused when DNA molecules pass through the nanopores, supporting real-time analysis of ultra-long reads (>100 kb).
[0045] Regarding the target fragment, for example, it is a target fragment regarding the gene to be tested. It should be understood that the gene to be tested can be a multi-copy gene or other complex genomic regions that require haplotype structure analysis. In some embodiments, the gene to be tested, for example, includes SMN1 gene and SMN2 gene.
[0046] At step 204, the computing device 110 aligns the sequencing sequences with the reference genome sequence to determine the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment.
[0047] Regarding the method for determining the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment, for example, it includes: the computing device 110 aligns the sequencing sequences with the reference genome sequence to determine the single nucleotide polymorphism sites corresponding to each sequencing sequence within the target fragment; retains the single nucleotide polymorphism sites where the maximum base frequency is within a preset base frequency range as heterozygous single nucleotide polymorphism sites. At the same time, obtains the base types of the heterozygous single nucleotide polymorphism sites corresponding to each sequencing sequence.
[0048] In some embodiments, the alignment operation can be implemented by an alignment algorithm applicable to long-read sequencing data, such as tools like BWA-MEM, Minimap2, or BLASR. Specifically, the computing device 110 can align the sequencing sequences (e.g., in FASTQ format) with a reference genome sequence (such as the human reference genome in GRCh37 or GRCh38 version), generating an alignment result file (e.g., a BAM or SAM file). Subsequently, the computing device 110 can use variant detection tools such as GATK HaplotypeCaller, Samtools mpileup, DeepVariant, etc. to analyze the alignment results within the target fragment region, generating a list of candidate single nucleotide polymorphism (SNP) sites. Those skilled in the art should understand that the above alignment and variant detection can be implemented using any suitable alignment and variant detection tools and algorithms, and their specific parameters can be adjusted according to actual needs.
[0049] After obtaining the candidate single nucleotide polymorphism sites, since homozygous single nucleotide polymorphism sites cannot provide useful information for haplotype phasing, the computing device 110 can further identify the heterozygous single nucleotide polymorphism sites among them. It should be understood that "homozygous" means that the base types of multiple sequencing sequences at a certain single nucleotide polymorphism site are the same; if the base types of any two sequencing sequences are different at a certain single nucleotide polymorphism site, it is called "heterozygous". Considering reasons such as sequencing errors, the computing device 110 can identify the heterozygous single nucleotide polymorphism sites according to whether the maximum base frequency at the single nucleotide polymorphism site is within a preset base frequency range. Among them, the maximum base frequency refers to the proportion of the number of sequencing sequences supporting the base with the highest occurrence frequency at this site to the total effective sequencing depth of this site. It can usually be obtained by dividing the number of sequencing sequences supporting the base with the highest occurrence frequency by the total effective sequencing depth of this site.
[0050] Regarding the preset base frequency range, it is, for example, but not limited to, greater than or equal to a first preset base frequency threshold and less than or equal to a second preset base frequency threshold. Regarding the first preset base frequency threshold, it is, for example, 10%, and in some embodiments, the first preset base frequency threshold is, for example, 20%. Regarding the second preset base frequency threshold, it is, for example, 90%. In some embodiments, the second preset base frequency threshold is, for example, 80%. It should be understood that the preset base frequency range can be adjusted according to the experimental needs of those skilled in the art. The single nucleotide polymorphism sites within this preset base frequency range can be identified as heterozygous single nucleotide polymorphism sites for subsequent analysis.
[0051] In addition, to support subsequent haplotype phasing algorithms, computing device 110 may further extract the base type information of each heterozygous single nucleotide polymorphism site in different sequencing sequences. This process can be achieved by parsing the alignment result file and the list of heterozygous single nucleotide polymorphism sites, and extracting the base type of each sequencing sequence at each heterozygous single nucleotide polymorphism site one by one. If a certain sequencing sequence does not cover a single nucleotide polymorphism site, it is recorded as "missing".
[0052] In some embodiments, to improve the quality of heterozygous single nucleotide polymorphism sites that support haplotype phasing algorithms, computing device 110 may further filter the heterozygous single nucleotide polymorphism sites according to preset filtering rules to improve the accuracy and reliability of subsequent haplotype phasing analysis.
[0053] Regarding the preset filtering rules, they include, for example, but are not limited to one or more of the following: removing heterozygous single nucleotide polymorphism sites located in regions where the number of consecutive identical bases exceeds a preset threshold. For example, when the number of consecutive occurrences of the same base in a certain segment reaches or exceeds 5, the heterozygous single nucleotide polymorphism sites appearing in this region will be excluded; removing heterozygous single nucleotide polymorphism sites located in tandem repeat regions. For example, when a sequence of two or more bases that repeats multiple times (e.g., more than 5 times) in a head-to-tail manner appears in a certain segment, the heterozygous single nucleotide polymorphism sites appearing in this region will be excluded; and, by comparing with a preset "blacklist" region, removing heterozygous single nucleotide polymorphism sites located in "blacklist" regions such as known low-quality regions or regions with inaccurate sequencing. The preset "blacklist" region can be determined by manual review or reference to public databases. It should be understood that various thresholds in the preset filtering rules (such as the threshold for the number of consecutive identical bases, the length of tandem repeat units, the number of tandem repeat unit repetitions, and the preset "blacklist" region, etc.) can be flexibly adjusted according to factors such as the experimental design of those skilled in the art, the characteristics of the target region, and the characteristics of the sequencing platform.
[0054] In some embodiments, in order to systematically record the information of heterozygous single nucleotide polymorphism sites for subsequent analysis and processing, computing device 110 can, for example, arrange the above-mentioned heterozygous single nucleotide polymorphism sites according to their coordinates on the reference genome and their corresponding base types for each sequencing fragment to form a heterozygous single nucleotide polymorphism matrix, and the heterozygous single nucleotide polymorphism matrix indicates the base type of each heterozygous single nucleotide polymorphism site in each sequencing sequence.
[0055] Specifically, let the input sequencing sequence set containing sequencing sequences be:
[0056]
[0057] Among them, represents a sequencing sequence regarding the target fragment.
[0058] Suppose the target fragment contains pre-processed and filtered heterozygous single nucleotide polymorphism sites, denoted as:
[0059]
[0060] Among them, represents a heterozygous single nucleotide polymorphism site regarding the target fragment and is arranged in the order of its coordinates in the reference genome.
[0061] Accordingly, a heterozygous single nucleotide polymorphism matrix of size is defined as follows:
[0062]
[0063] Among them, the matrix element indicates the base type of the sequencing sequence at the heterozygous single nucleotide polymorphism site . If , it means that the sequencing sequence does not cover the heterozygous single nucleotide polymorphism site .
[0064] Through this matrix, the base types of each sequencing sequence at each heterozygous single nucleotide polymorphism site can be systematically described, thus supporting a series of subsequent processes based on the method of the present invention. In addition, in some prior arts, the haplotype phasing algorithm usually constructs a heterozygous single nucleotide polymorphism matrix based on the dimorphism of single nucleotide polymorphism sites. However, when dealing with multi-copy genes or regions with multiple allelic base types, this method reduces the ability of the algorithm to describe polymorphisms, thereby affecting the accuracy of haplotype phasing for genes with complex structures such as multi-copy genes. In contrast, the present application completely retains the original base information and improves the ability to accurately phase haplotypes for genes with complex structures.
[0065] At step 206, the computing device 110 recursively partitions the input sequencing sequence set multiple times based on the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing, so as to determine the haplotypes of the target fragment.
[0066] A method for determining the haplotype of a target fragment by performing multiple recursive partitions on the input sequencing sequence set based on the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing, for example, includes: in each recursion, the computing device 110 evaluates the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing based on the current input sequencing sequence set; divides the sequencing sequences in the current input sequencing sequence set into multiple subsets based on the contribution degree; in response to any one of the subsets satisfying the termination condition, independently outputs the subset satisfying the termination condition as a haplotype of the target fragment; and, in response to any one of the subsets not satisfying the termination condition, independently updates the current input sequencing sequence set based on the subset not satisfying the termination condition and advances to the next recursion until the termination condition is satisfied. The following will be combined with Figure 3 Specifically illustrate the method for performing multiple recursive partitions on the input sequencing sequence set to determine the haplotype of the target fragment, which will not be elaborated here.
[0067] In the above solution, by determining the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment based on the alignment result of the sequencing sequence of the target fragment and the reference genome sequence; performing multiple recursive partitions on the input sequencing sequence set based on the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing, thereby determining the haplotype of the target fragment. The method of the present invention systematically evaluates the contribution degree of each heterozygous single nucleotide polymorphism site to accurately partitioning the haplotype, and based on this, performs multiple recursive partitions on multiple sequencing sequences of the target fragment, and determines the haplotype of the target fragment according to the compatibility of the sequencing sequences. The present invention can efficiently and accurately perform haplotype typing on multi-copy genes.
[0068] The following will be combined with Figure 3 Describe method 300 for performing multiple recursive partitions on an input sequencing sequence set to determine the haplotype of a target fragment according to an embodiment of the present invention. Figure 3 FIG. shows a flowchart of method 300 for performing multiple recursive partitions on an input sequencing sequence set to determine the haplotype of a target fragment according to an embodiment of the present invention. It should be understood that method 300 can be executed, for example, at Figure 5 the described electronic device 500. It can also be executed at Figure 1 the described computing device 110. It should be understood that method 300 may further include additional actions not shown and / or may omit the shown actions, and the scope of the present invention is not limited in this regard.
[0069] In method 300, the computing device 110 initiates a recursive process for determining the haplotype of the target fragment.
[0070] Specifically, to start the process, the computing device 110 initially evaluates the contribution of each heterozygous single nucleotide polymorphism site to haplotype phasing based on the current input sequencing sequence set at step 302 .
[0071] In some embodiments, due to multiple rounds of recursion, at least some of the single nucleotide polymorphism sites have reached homozygosity in the current input sequencing sequence set. Therefore, in order to at least partially save computing resources, the computing device 110 can re-identify heterozygous single nucleotide polymorphism sites based on the current input sequencing sequence set.
[0072] A method for re-identifying heterozygous single nucleotide polymorphism sites based on a current input sequencing sequence set, for example, comprises: for each single nucleotide polymorphism site The computing device 110 compares the base types of each sequencing sequence in the current input sequencing sequence set at the site. If two or more different non-missing bases are observed at the site, the site is retained as a heterozygous site in the current recursive stage, and all sites that meet this condition are Add to collection .
[0073] In some embodiments, a method for evaluating the contribution of a heterozygous single nucleotide polymorphism site to haplotype phasing includes: for the heterozygous single nucleotide polymorphism site to be evaluated, , according to the current input sequencing sequence set The base type of each sequencing sequence at the heterozygous single nucleotide polymorphism site to be evaluated , divide the sequencing sequences in the current input sequencing sequence set into multiple subsets ; Traverse all heterozygous single nucleotide polymorphism sites about the target fragment , for each heterozygous single nucleotide polymorphism site : In response to each subset In the example, the base types of each sequencing sequence at the heterozygous single nucleotide polymorphism site are all unique, and the heterozygous single nucleotide polymorphism site to be evaluated is 1 point is added for the contribution degree; in response to the fact that in any subset, the base type of each sequencing sequence at the heterozygous single nucleotide polymorphism site is not unique, the heterozygous single nucleotide polymorphism site to be evaluated The contribution of does not increase; the total score is accumulated as the heterozygous single nucleotide polymorphism site to be evaluated Contribution to haplotype phasing. The higher the contribution of a single nucleotide polymorphism site, the more helpful the site is for accurate haplotype phasing. Figure 4 The method 400 for evaluating the contribution of heterozygous single nucleotide polymorphism sites to haplotype phasing is specifically described and will not be repeated here.
[0074] Once the contribution of each heterozygous single nucleotide polymorphism site to haplotype phasing is obtained, computing device 110 divides the sequencing sequences in the current input sequencing sequence set into multiple subsets at step 304 based on the contribution for further recursive processing.
[0075] In one embodiment, computing device 110 selects the heterozygous single nucleotide polymorphism site with the highest contribution as the basis for dividing the current set, and divides the current input sequencing sequence set into multiple subsets . Specifically, if multiple different bases (such as A, T, C or G) are observed at the site , then the sequencing sequences in the current input sequencing sequence set are grouped according to the base type, and a set of subsets is constructed, where each subset contains the sequencing sequences with the base at the site . This strategy helps to preferentially select the heterozygous single nucleotide polymorphism sites that are most discriminative for different haplotypes, thereby improving the overall typing efficiency and accuracy.
[0076] In some cases, when there are multiple heterozygous single nucleotide polymorphism sites with the same and highest contribution, it means that these sites are equivalent in differentiating the target fragment haplotypes at the current stage. Therefore, computing device 110 can arbitrarily select one of the heterozygous single nucleotide polymorphism sites as the basis for set division without distinction.
[0077] Subsequently, at step 306, computing device 110 determines whether each subset obtained from the above division satisfies the termination condition. In one embodiment, the termination condition is that all the sequencing sequences in the subset are compatible. It should be understood that for two sequencing sequences, if there is a single nucleotide polymorphism site such that the two sequencing sequences have different and non-missing bases at this single nucleotide polymorphism site, then these two sequencing sequences are said to be "in conflict"; conversely, if the two sequencing sequences do not conflict at all single nucleotide polymorphism sites, then these two sequencing sequences are said to be "compatible". If all the sequencing sequences in a sequencing sequence set are compatible, it indicates that this set can uniquely determine a haplotype.
[0078] In response to the judgment result in step 306, if there is any subset that does not satisfy the termination condition, then computing device 110 updates the current input sequencing sequence set , and independently proceed to the next round of recursion, repeating the above steps 302 to 306 until all subsets meet the termination condition. Since the sequence conflicts in the subsets are strictly reduced in each round of recursion, the recursive processing of the method of the present application will eventually converge and return a set of non-overlapping and internally compatible subsets of sequencing sequences, corresponding to the complete haplotype set of the target fragment.
[0079] In response to the judgment result in step 306, if a subset meets the termination condition, that is, all the sequencing sequences in the subset are compatible, it indicates that the subset can uniquely determine a haplotype. Therefore, the computing device 110 outputs the subset independently as a haplotype of the target fragment at step 310.
[0080] In one embodiment, to improve the overall processing efficiency, the computing device 110 can process multiple subsets in parallel during each recursive process, thereby shortening the total calculation time of the phasing algorithm of the present invention.
[0081] Furthermore, the recursive phasing process in the method 300 of the present invention is further described below in conjunction with mathematical formulas.
[0082] First, optionally, for the current input set of sequencing sequences , re-identify the set of heterozygous single nucleotide polymorphism sites for the target fragment that can be used for phasing:
[0083]
[0084] Among them, represents that in the current input set of sequencing sequences , at least two different non-missing bases are observed at the locus , so the locus is considered heterozygous in the current input set of sequencing sequences and is included in the set .
[0085] For the heterozygous single nucleotide polymorphism locus to be evaluated, evaluate the contribution of this locus to haplotype phasing:
[0086]
[0087] The function is used to evaluate the contribution of the heterozygous single nucleotide locus to haplotype phasing, which will be specifically described below in conjunction with Figure 4 , and will not be elaborated here.
[0088] The greater the contribution of a locus to be evaluated, the more this locus The stronger the ability to accurately distinguish different haplotypes, the more helpful it is to divide the sequencing sequence set into homozygous subsets. Therefore, select the site with the highest contribution as the basis for dividing the current set:
[0089]
[0090] According to the base type at the site , divide the current input sequencing sequence set into multiple subsets :
[0091]
[0092] Among them, represents the set of all bases that have appeared at the site in the set :
[0093]
[0094] For each newly generated subset , check whether all the sequencing sequences in it are compatible:
[0095]
[0096] Similarly, means that in the subset , only one non-missing base is observed at the site . If only one non-missing base is observed at all sites , then it is considered that all the sequencing sequences in this subset are compatible. Therefore, the subset can uniquely determine a haplotype.
[0097] For subsets that do not meet the above conditions, they are used as the new current input sequencing sequence set , and enter the next round of recursion:
[0098]
[0099] The function takes the union of all the returned subset processing results and finally returns a set of non-overlapping and internally compatible sequencing sequence subsets, corresponding to the complete haplotype combination of the target fragment.
[0100] By adopting the above recursive phasing strategy, the present invention achieves an efficient partitioning of the sequencing sequence set, thereby gradually converging to the complete haplotype combination of the target fragment. This strategy can effectively improve the efficiency and accuracy of haplotype typing for multi-copy genes, providing an important data basis for further research and applications.
[0101] The following will be combined with Figure 4 Describe a method for evaluating the contribution degree of heterozygous single nucleotide polymorphism sites to haplotype phasing according to an embodiment of the present invention. Figure 4 FIG. shows a flowchart of a method 400 for evaluating the contribution degree of heterozygous single nucleotide polymorphism sites to haplotype phasing according to an embodiment of the present invention. It should be understood that the method 400 can be executed, for example, at Figure 5 The electronic device 500 described. It can also be executed at Figure 1 The computing device 110 described. It should be understood that the method 400 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this regard.
[0102] At step 402, the computing device 110 divides the current input sequencing sequence set for the heterozygous single nucleotide polymorphism site to be evaluated into multiple subsets according to the base types of each sequencing sequence at this site in the current input sequencing sequence set . Specifically, if multiple different bases (such as A, T, C, or G) are observed at the site , then the sequencing sequences in the current input sequencing sequence set are grouped according to the base types to construct a set of subsets , where each subset contains the sequencing sequences with the base being at the site.
[0103] At step 404, the computing device 110 traverses each heterozygous single nucleotide polymorphism site regarding the target fragment, and determines whether all the sequencing sequences in each subset divided according to have a unique base type at the site . It should be understood that the base types at the site in different subsets can be different.
[0104] In response to the determination at step 404, if the base types of all the sequencing sequences at the site are unique within each subset, it indicates that the site to be evaluated Clearly distinguish different haplotypes at the locus This has a positive effect on accurately performing haplotype phasing. Therefore, the computing device 110 adds 1 point to the contribution degree of the heterozygous single nucleotide polymorphism locus to be evaluated at step 406; on the contrary, if at any subset, at the locus there are multiple different bases, it indicates that the locus cannot be distinguished according to the locus to be evaluated , and it has no positive effect on accurate phasing. Therefore, the computing device 110 does not add points to the contribution degree of the locus at step 408.
[0105] At step 410, the computing device 110 determines whether it has completed traversing all the heterozygous single nucleotide polymorphism loci regarding the target fragment. If the traversal has not been completed, it returns to step 404 to continue the determination; if all loci have been traversed, then at step 412, the total score is accumulated to obtain the contribution degree of the locus to be evaluated to haplotype phasing. The higher the contribution degree of a certain single nucleotide polymorphism locus, the more heterozygous single nucleotide polymorphism loci can be distinguished according to this locus, and the more helpful it is for accurate haplotype phasing.
[0106] Furthermore, the following further illustrates the method for evaluating the contribution degree of heterozygous single nucleotide polymorphism loci to haplotype phasing in combination with mathematical formulas.
[0107] First, for the current input sequencing sequence set , and the heterozygous single nucleotide polymorphism locus to be evaluated, the set is divided into multiple subsets according to the base type at the locus :
[0108]
[0109] Among them, represents the set of all bases that have appeared at the locus in the set :
[0110]
[0111] Define the function to determine whether the base type at the locus in a certain subset is unique:
[0112]
[0113] Accordingly, the contribution degree of the heterozygous single nucleotide polymorphism site to be evaluated is defined as:
[0114]
[0115] wherein, used to determine whether, among all subsets partitioned according to the site the bases of all sequencing sequences at the site are unique. Therefore, the function indicates that: if, among all subsets at a certain site the condition that the base types are unique is satisfied, it is regarded that the site to be evaluated can correctly distinguish the site Therefore, the contribution degree of the site to be evaluated is increased by one point. Traverse all Each time this condition is met, the contribution degree of the site is increased by 1 point.
[0116] In the above manner, the present invention quantitatively scores the heterozygous single nucleotide polymorphism sites through the contribution degree, so as to objectively evaluate the discrimination ability of the heterozygous single nucleotide polymorphism sites in the process of phasing the target fragment. Therefore, the method of the present invention can effectively improve the haplotype phasing accuracy in the background of multi-copy genes.
[0117] Figure 5 FIG. schematically shows a block diagram of an electronic device 500 suitable for implementing the embodiments of the present invention. The electronic device 500 may be a device for implementing the methods 200 to 400 shown in Figures 2 to 4 As shown in Figure 5 the electronic device 500 includes a central processing unit (i.e., CPU 501), which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (i.e., ROM 502) or computer program instructions loaded from a storage unit 508 into a random access memory (i.e., RAM 503). In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The CPU 501, ROM 502, and RAM 503 are connected to each other through a bus 504. An input / output interface (i.e., I / O interface 505) is also connected to the bus 504.
[0118] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508. The CPU 501 executes the various methods and processes described above, such as executing methods 200 to 400. For example, in some embodiments, methods 200 to 400 may be implemented as computer software programs, which are stored in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the CPU 501, one or more operations of methods 200 to 400 described above can be executed. Alternatively, in other embodiments, the CPU 501 may be configured to execute one or more actions of methods 200 to 400 by any other suitable means (e.g., by means of firmware).
[0119] It should be further noted that the present invention can be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0120] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being an instantaneous signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0121] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0122] The computer program instructions for carrying out operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present invention.
[0123] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0124] These computer-readable program instructions can be provided to a processor in a voice interaction device, a general-purpose computer, a special-purpose computer, or other programmable data processing device's processing unit, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0125] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing device, or other device, such that a series of operation steps are executed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing device, or other device to implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and this module, segment of a program, or part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the boxes may occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as combinations of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that executes the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0127] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art in the technical field to understand the disclosed embodiments.
[0128] The above are only alternative embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for determining the haplotype of a target fragment, comprising: Obtaining a set of input sequencing sequences of the target fragment from a sample to be tested, the set of input sequencing sequences comprising multiple sequencing sequences of the target fragment; Aligning the sequencing sequences with a reference genomic sequence to determine heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment; Based on the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing, performing multiple recursive partitions on the set of input sequencing sequences, thereby determining the haplotype of the target fragment; Wherein the evaluation method for the contribution degree of the heterozygous single nucleotide polymorphism site to haplotype phasing includes: For a heterozygous single nucleotide polymorphism site to be evaluated, according to the base types of each sequencing sequence in the current set of input sequencing sequences at the heterozygous single nucleotide polymorphism site to be evaluated, partitioning the sequencing sequences in the current set of input sequencing sequences into multiple subsets; Traversing all heterozygous single nucleotide polymorphism sites of the target fragment, for each of the heterozygous single nucleotide polymorphism sites: In response to the base types of all sequencing sequences in each subset being unique at the heterozygous single nucleotide polymorphism site, adding 1 point to the contribution degree of the heterozygous single nucleotide polymorphism site to be evaluated; In response to the base types of the sequencing sequences in any one subset not being unique at the heterozygous single nucleotide polymorphism site, not increasing the contribution degree of the heterozygous single nucleotide polymorphism site to be evaluated; Accumulating the total score as the contribution degree of the heterozygous single nucleotide polymorphism site to be evaluated to haplotype phasing.
2. The method according to claim 1, wherein the sequencing sequences of the target fragment are single molecule long read sequencing sequences of the target fragment obtained by single molecule long read sequencing technology.
3. The method according to claim 1, wherein determining the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment includes: Aligning the sequencing sequences with a reference genomic sequence to determine single nucleotide polymorphism sites corresponding to each sequencing sequence within the target fragment; Retaining single nucleotide polymorphism sites with the maximum base frequency within a preset base frequency range as heterozygous single nucleotide polymorphism sites, and obtaining the base types of the heterozygous single nucleotide polymorphism sites corresponding to each sequencing sequence.
4. The method according to claim 3, the method further comprising filtering the heterozygous single nucleotide polymorphism sites according to preset filtering rules, the preset filtering rules including at least one of the following: Removing heterozygous single nucleotide polymorphism sites located within a region with a continuous number of identical bases exceeding a preset threshold; Removing heterozygous single nucleotide polymorphism sites located within a tandem repeat region; and, Removing heterozygous single nucleotide polymorphism sites located within a preset "blacklist" region.
5. The method according to claim 1, wherein based on the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing, performing multiple recursive partitions on the set of input sequencing sequences, thereby determining the haplotype of the target fragment includes: In each recursion, based on the current input set of sequencing sequences, the contribution of each heterozygous single nucleotide polymorphism (SNP) site to haplotype phasing is evaluated; Based on the contribution, the sequencing sequences in the current input set of sequencing sequences are divided into multiple subsets; In response to any one of the subsets satisfying the termination condition, the subset that satisfies the termination condition is independently output as a haplotype of the target fragment; and, In response to any one of the subsets not satisfying the termination condition, based on the subset that does not satisfy the termination condition independently, the current input set of sequencing sequences is updated and advanced to the next recursion until the termination condition is satisfied.
6. The method according to claim 5, wherein the dividing the sequencing sequences in the current input set of sequencing sequences into multiple subsets based on the contribution includes: Selecting the heterozygous SNP site with the highest contribution, and dividing the sequencing sequences in the current input set of sequencing sequences into multiple subsets according to the base types of the sequencing sequences at the heterozygous SNP site with the highest contribution.
7. The method according to claim 6, when multiple heterozygous SNP sites have the same highest contribution, any one of the heterozygous SNP sites is arbitrarily selected as the basis for set division.
8. The method according to claim 5, the method further includes, in each recursion, processing the multiple subsets in parallel until all subsets satisfy the termination condition.
9. The method according to claim 5, wherein the termination condition is that all the sequencing sequences in the subset are compatible.
10. A computing device, comprising: At least one processing unit; And, At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the device to perform the steps of the method according to any one of claims 1 to 9.
11. A computer-readable storage medium, having a computer program stored thereon, the computer program when executed by a machine implements the method according to any one of claims 1 to 9.
12. A computer program product, which includes a computer program, the computer program when executed by a machine performs the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method for rebuilding individual single somatotype based on optimizing solution aggregate
CN101256602A
Haploid phasing and variation detection methods and devices of diploid genome read
CN110021355A