Method, device, medium and program product for determining haplotype of target segment
By sequencing sequence alignment and recursive division of the target fragments of multi-copy genes, their haplotypes are determined, which solves the problems of insufficient resolution and low computational efficiency of haplotype phasing of multi-copy genes in the prior art, and achieves efficient and accurate haplotype phasing.
Patent Information
- Application Number
- CN202510541354.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The prior art has problems such as insufficient resolution, low computational efficiency and low reliability of results when dealing with haplotype phasing of multi-copy genes.
By obtaining the input sequencing sequence set of the target fragment, aligning with the reference genomic sequences, heterozygous single nucleotide polymorphism information is determined, and multiple recursive divisions are made based on the contribution of each heterozygous single nucleotide polymorphism site to determine the haplotype of the target fragment.
The efficiency and accuracy of multi-copy genes in haplotype phasing are significantly improved, and the problems of insufficient resolution and low computational efficiency in the prior art are solved.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention generally relates to bioinformatics processing, and specifically, to a method, a computing device, a computer storage medium, and a computer program product for determining the haplotype of a target fragment. Background Art
[0002] A haplotype refers to a combination of a series of genetic variant sites co-existing on a single chromosome, which is a combination of allelic sequences from the same parent and describes the sequence differences on a single chromosome segment. The combination of different genetic loci between homologous chromosomes has an important impact on biological phenotypes, such as human genetic diseases. Haplotype phasing aims to obtain a haplotype-resolved assembled sequence, that is, to obtain a complete haplotype genomic sequence, so as to deeply understand the combination and inheritance of different genetic loci on a single chromosome or a specific single-chromosome region, which is helpful for understanding and diagnosing the pathogenesis of genetic diseases.
[0003] Currently, with the development of sequencing technology, the high-resolution sequencing data of genomes is increasing day by day, and haplotype phasing technology has made great progress. However, haplotype phasing for multi-copy genes remains a challenging problem. Traditional analysis methods often have difficulty accurately distinguishing haplotype differences between highly similar different gene copies. Existing technologies often have problems such as insufficient resolution, low computational efficiency, and low reliability of results when dealing with such complex situations.
[0004] Currently, software for haplotype phasing detection, such as Whatshap, has obvious limitations. First of all, these software usually can only handle the classical situation of two alleles, and cannot effectively handle the presence of multi-copy genes. Secondly, compared with traditional methods, the detection performance of these software is lower, and it is easy to generate a higher error rate, which may lead to inaccurate or even misleading analysis results.
[0005] Therefore, there is still an urgent need for a new efficient and accurate algorithm for haplotype phasing of multi-copy genes. Summary of the Invention
[0006] The present invention provides a method, a computing device, a computer storage medium, and a computer program product for determining the haplotype of a target fragment, which can significantly improve the efficiency and accuracy of haplotype phasing for multi-copy genes.
[0007] According to a first aspect of the present invention, a method for determining the haplotype of a target fragment is provided. The method includes: obtaining a set of input sequencing sequences regarding the target fragment from a sample to be tested, the set of input sequencing sequences including multiple sequencing sequences regarding the target fragment; aligning the sequencing sequences with a reference genome sequence to determine the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment; and performing multiple recursive partitions on the set of input sequencing sequences based on the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing, so as to determine the haplotype of the target fragment.
[0008] In one embodiment, the sequencing sequence regarding the target fragment is a single molecule long read sequencing sequence regarding the target fragment obtained by a single molecule long read sequencing technology.
[0009] In one embodiment, the determining the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment includes: aligning the sequencing sequence with a reference genome sequence to determine the single nucleotide polymorphism sites corresponding to each sequencing sequence within the target fragment; retaining the single nucleotide polymorphism sites where the maximum base frequency is within a preset base frequency range as heterozygous single nucleotide polymorphism sites, and obtaining the base types of the heterozygous single nucleotide polymorphism sites corresponding to each sequencing sequence.
[0010] In one embodiment, the determining the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment further includes filtering the heterozygous single nucleotide polymorphism sites according to a preset filtering rule, and the preset filtering rule includes at least one of the following: removing the heterozygous single nucleotide polymorphism sites located in a region where the number of consecutive identical bases exceeds a preset threshold; removing the heterozygous single nucleotide polymorphism sites located in a tandem repeat region; and removing the heterozygous single nucleotide polymorphism sites located in a preset "blacklist" region.
[0011] In one embodiment, the performing multiple recursive partitions on the set of input sequencing sequences based on the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing, so as to determine the haplotype of the target fragment includes: in each recursion, evaluating the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing based on the current set of input sequencing sequences; dividing the sequencing sequences in the current set of input sequencing sequences into multiple subsets based on the contribution degree; in response to any one of the subsets satisfying a termination condition, independently outputting the subset satisfying the termination condition as a haplotype of the target fragment; and in response to any one of the subsets not satisfying the termination condition, independently updating the current set of input sequencing sequences based on the subset not satisfying the termination condition, and advancing to the next recursion until the termination condition is satisfied.
[0012] In one embodiment, the method for evaluating the contribution degree of a heterozygous single nucleotide polymorphism locus to haplotype phasing includes: for a heterozygous single nucleotide polymorphism locus to be evaluated, dividing the sequencing sequences in the currently input sequencing sequence set into multiple subsets according to the base types of each sequencing sequence at the heterozygous single nucleotide polymorphism locus to be evaluated; traversing all heterozygous single nucleotide polymorphism loci of a target fragment, and for each of them: in response to the base types of each sequencing sequence at the heterozygous single nucleotide polymorphism locus being unique in each subset, adding 1 point to the contribution degree of the heterozygous single nucleotide polymorphism locus to be evaluated; in response to the base types of each sequencing sequence at the heterozygous single nucleotide polymorphism locus not being unique in any subset, not increasing the contribution degree of the heterozygous single nucleotide polymorphism locus to be evaluated; accumulating the total score as the contribution degree of the heterozygous single nucleotide polymorphism locus to be evaluated to haplotype phasing.
[0013] In one embodiment, dividing the sequencing sequences in the currently input sequencing sequence set into multiple subsets based on the contribution degree includes: selecting the heterozygous single nucleotide polymorphism locus with the highest contribution degree, and dividing the sequencing sequences in the currently input sequencing sequence set into multiple subsets according to the base types of each sequencing sequence at the heterozygous single nucleotide polymorphism locus with the highest contribution degree.
[0014] In one embodiment, when multiple heterozygous single nucleotide polymorphism loci have the same highest contribution degree, any one of them is arbitrarily selected as the basis for set division.
[0015] In one embodiment, the method for determining the haplotype of a target fragment further includes, in each recursion, processing the multiple subsets in parallel until all subsets meet the termination condition.
[0016] In one embodiment, the termination condition is that all sequencing sequences in the subset are compatible.
[0017] According to the second aspect of the present invention, there is also provided a computing device, which includes: a memory configured to store one or more computer programs; and a processor coupled to the memory and configured to execute one or more programs to cause the device to execute the method of the first aspect of the present invention.
[0018] According to the third aspect of the present invention, there is also provided a non-transitory computer-readable storage medium. Machine-executable instructions are stored on the non-transitory computer-readable storage medium, and when the machine-executable instructions are executed, the machine is caused to execute the method of the first aspect of the present invention.
[0019] According to a fourth aspect of the present invention, there is also provided a computer program product. The computer program product includes a computer program which, when executed by a machine, performs the method according to the first aspect of the present invention.
[0020] The Summary of the Invention section is provided to introduce, in a simplified form, a selection of concepts that are further described below in the Detailed Description. The Summary of the Invention section is not intended to identify key features or essential features of the present invention, nor is it intended to limit the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A schematic diagram of a system for implementing a method for determining a target fragment haplotype according to an embodiment of the present invention is shown.
[0022] Figure 2 A flowchart of a method for determining a target fragment haplotype according to an embodiment of the present invention is shown.
[0023] Figure 3 A flowchart of a method for performing multiple recursive partitions on an input sequencing sequence set to determine the haplotype of a target fragment according to an embodiment of the present invention is shown.
[0024] Figure 4 A flowchart of a method for evaluating the contribution of heterozygous single nucleotide polymorphism sites to haplotype phasing according to an embodiment of the present invention is shown.
[0025] Figure 5 A block diagram of an electronic device suitable for implementing an embodiment of the present invention is schematically shown.
[0026] In the respective drawings, the same or corresponding reference numerals indicate the same or corresponding parts. DETAILED DESCRIPTION
[0027] Preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention will be more thorough and complete, and will fully convey the scope of the present invention to those skilled in the art.
[0028] As used herein, the term "comprising" and its variations mean open-ended inclusion, i.e., "including but not limited to". Unless specifically stated otherwise, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.
[0029] As described above, for haplotype typing of multi-copy genes, traditional analysis methods often have difficulty in accurately distinguishing haplotype differences between different gene copies. Existing software for haplotype typing detection also often has problems such as insufficient resolution, low computational efficiency, and low result reliability when dealing with such complex situations of multi-copy genes.
[0030] To at least partially solve one or more of the above problems and other potential problems, example embodiments of the present invention propose a method for determining the haplotype of a target fragment for multi-copy genes. This method systematically evaluates the contribution degree of each heterozygous single nucleotide polymorphism site to the accurate division of haplotypes, and based on this, recursively divides multiple sequencing sequences of the target fragment multiple times, and determines the haplotype of the target fragment according to the compatibility of the sequencing sequences. The present invention can efficiently and accurately perform haplotype typing on multi-copy genes.
[0031] Figure 1 FIG. shows a schematic diagram of a system 100 for implementing the method for determining the haplotype of a target fragment according to an embodiment of the present invention. As Figure 1 shown, the system 100 includes: a computing device 110, a sequencing device 130. In some embodiments, the system 100 further includes: a server 140, a network 120. In some embodiments, the computing device 110, the server 140, and the sequencing device 130 can perform data interaction in a wired or wireless manner (such as the network 120).
[0032] Regarding the sequencing device 130, it is used, for example, to sequence a sample to be tested to generate a sequencing sequence of the target fragment; and send the generated sequencing sequence data to the computing device 110. In some embodiments, the server 140 can also send the sequencing sequence of the sample to be tested to the computing device 110. In some embodiments, the sequencing device 130 is, for example, a sequencing device based on single molecule long read length sequencing technology. The sequencing device 130 is, for example, but not limited to, a sequencer developed by Pacific Biosciences (PacBio) using single molecule real-time (SMRT) sequencing technology.
[0033] Regarding the server 140, it is used, for example, to provide sequence information of a reference genome. For example, provide human reference genome sequence information, and the human reference genome is, for example, but not limited to, the human reference genome GRCh37 or the human reference genome GRCh38.
[0034] Regarding computing device 110, which is used, for example, to determine the target fragment haplotype. Specifically, computing device 110 is used to obtain a set of input sequencing sequences of a target fragment from a sample to be tested, the set of input sequencing sequences including multiple sequencing sequences of the target fragment; align the sequencing sequences with a reference genomic sequence to determine heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment; and based on the contribution degree of each heterozygous single nucleotide polymorphism locus to haplotype phasing, perform multiple recursive partitions on the set of input sequencing sequences to thereby determine the haplotype of the target fragment.
[0035] In some embodiments, computing device 110 may have one or more processing units, including dedicated processing units such as GPUs, FPGAs, and ASICs, as well as general-purpose processing units such as CPUs. Computing device 110 may be an integration of multiple physical servers, an integration of multiple processing units, etc. Additionally, one or more virtual machines may also be running on each computing device 110. In some embodiments, computing device 110 and server 130 may be integrated together or may be separately provided from each other.
[0036] In some embodiments, computing device 110 includes, for example: a sequencing sequence data acquisition unit 112, a heterozygous single nucleotide polymorphism locus determination unit 114, and a haplotype phasing unit 116. The above-mentioned sequencing sequence data acquisition unit 112, heterozygous single nucleotide polymorphism locus determination unit 114, and haplotype phasing unit 116 may be configured on one or more computing devices 110.
[0037] Regarding the sequencing sequence data acquisition unit 112, it is used to obtain a set of input sequencing sequences of a target fragment from a sample to be tested, the set of input sequencing sequences including multiple sequencing sequences of the target fragment.
[0038] Regarding the heterozygous single nucleotide polymorphism locus determination unit 114, it is used to align the sequencing sequences with a reference genomic sequence to determine heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment.
[0039] Regarding the haplotype phasing unit 116, it is used to perform multiple recursive partitions on the set of input sequencing sequences based on the contribution degree of each heterozygous single nucleotide polymorphism locus to haplotype phasing, thereby determining the haplotype of the target fragment.
[0040] The following will be combined with Figure 2 Describe a method for determining the haplotype of a target fragment according to an embodiment of the present invention. Figure 2 A flowchart of a method 200 for determining the haplotype of a target fragment according to an embodiment of the present invention is shown. It should be understood that method 200 may, for example, be inFigure 5 performed at the described electronic device 500. It may also be performed at Figure 1 the described computing device 110. It should be understood that method 200 may also include additional actions not shown and / or may omit the shown actions, and the scope of the present invention is not limited in this regard.
[0041] At step 202, the computing device 110 obtains a set of input sequencing sequences of a target fragment from a sample to be tested, and the set of input sequencing sequences includes multiple sequencing sequences of the target fragment.
[0042] Regarding the sequencing sequences of the target fragment from the sample to be tested, for example but not limited to, single-molecule long-read sequencing sequences of the target fragment obtained by single-molecule long-read sequencing technology. For example, the sequencing sequences may be single-molecule long-read sequencing sequences obtained by using a single-molecule sequencing platform (such as the SMRT sequencing platform or the nanopore sequencing platform) after specifically amplifying or capturing the target fragment via one or more pairs of primers. Compared with short-read sequencing or whole-genome sequencing, the single-molecule long-read sequencing obtained based on specific amplification or capture can more completely cover gene regions with complex structures or variable copy numbers, thus facilitating the improvement of the accuracy of subsequent haplotype phasing.
[0043] After sequencing is completed, the sequencing sequences may undergo preliminary quality control and filtering, for example including but not limited to processes such as low-quality read removal, sequencing error correction, and redundant sequence removal. Those skilled in the art should understand that the above quality control and filtering can be implemented using any suitable quality control and filtering tools, and their specific parameters can be flexibly set according to the sequencing platform, sample characteristics, and experimental purposes.
[0044] Regarding single-molecule long-read sequencing technology, it includes many types. For example, but not limited to, Pacific Biosciences (PacBio) has developed sequencers (such as RS I, RS II and Sequel machines) using single-molecule real-time (SMRT) sequencing technology. Compared with the second-generation sequencing technology, taking the single-molecule long-read sequencing platform PacBio Sequel II as an example, it uses SMRT sequencing technology to achieve single-molecule real-time sequencing. The principle of SMRT sequencing is to use SMRT Cell as a carrier. Each SMRT Cell is covered with millions of zero-mode waveguide holes (ZMW). During sequencing, DNA polymerase and a template molecule are targeted at the bottom of the ZMW hole for reaction. The excitation light at the bottom of the small hole can excite the fluorescent marker on the nucleotide substrate, and then the fluorescent signal is recorded through the monitoring system to obtain base information. For another example, the nanopore sequencing device developed by Oxford Nanopore Technology (ONT) directly reads sequence information by measuring the current changes caused by the DNA molecule passing through the nanopore, supporting real-time analysis of ultra-long reads (>100 kb).
[0045] Regarding the target fragment, it is, for example, a target fragment about a gene to be tested. It should be understood that the gene to be tested can be a multi-copy gene or other complex genomic region that needs to resolve the haplotype structure. In some embodiments, the gene to be tested, for example, includes SMN1 Genes and SMN2 Gene.
[0046] At step 204 , the computing device 110 compares the sequencing sequence with the reference genome sequence to determine the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence in the target fragment.
[0047] The method for determining the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence in the target fragment includes, for example: the computing device 110 compares the sequencing sequence with the reference genome sequence to determine the single nucleotide polymorphism site corresponding to each sequencing sequence in the target fragment; retaining the single nucleotide polymorphism site whose maximum base frequency is within the preset base frequency range as the heterozygous single nucleotide polymorphism site. At the same time, obtaining the base type corresponding to each sequencing sequence of the heterozygous single nucleotide polymorphism site.
[0048] In some embodiments, the alignment operation can be implemented by an alignment algorithm suitable for long-read sequencing data, such as tools like BWA-MEM, Minimap2, or BLASR. Specifically, the computing device 110 can align the sequencing sequences (e.g., in FASTQ format) with a reference genomic sequence (such as the human reference genome in GRCh37 or GRCh38 version) to generate an alignment result file (e.g., BAM or SAM file). Subsequently, the computing device 110 can use variant detection tools such as GATK HaplotypeCaller, Samtools mpileup, DeepVariant, etc. to analyze the alignment results within the target fragment region to generate a list of candidate single nucleotide polymorphism (SNP) sites. Those skilled in the art should understand that the above alignment and variant detection can be implemented using any suitable alignment and variant detection tools and algorithms, and their specific parameters can be adjusted according to actual needs.
[0049] After obtaining the candidate single nucleotide polymorphism sites, since homozygous single nucleotide polymorphism sites cannot provide useful information for haplotype phasing, the computing device 110 can further identify the heterozygous single nucleotide polymorphism sites among them. It should be understood that "homozygous" means that the base types of multiple sequencing sequences at a single nucleotide polymorphism site are the same; if the base types of any two sequencing sequences are different at a single nucleotide polymorphism site, it is called "heterozygous". Considering reasons such as sequencing errors, the computing device 110 can identify the heterozygous single nucleotide polymorphism sites according to whether the maximum base frequency at the single nucleotide polymorphism site is within a preset base frequency range. Among them, the maximum base frequency refers to the proportion of the number of sequencing sequences supporting the base with the highest occurrence frequency at this site to the total effective sequencing depth of this site. It can usually be obtained by dividing the number of sequencing sequences supporting the base with the highest occurrence frequency by the total effective sequencing depth of this site.
[0050] Regarding the preset base frequency range, for example, but not limited to, it is greater than or equal to a first preset base frequency threshold and less than or equal to a second preset base frequency threshold. Regarding the first preset base frequency threshold, for example, it is 10%, and in some embodiments, the first preset base frequency threshold is, for example, 20%. Regarding the second preset base frequency threshold, for example, it is 90%. In some embodiments, the second preset base frequency threshold is, for example, 80%. It should be understood that the preset base frequency range can be adjusted according to the experimental needs of those skilled in the art. The single nucleotide polymorphism sites within this preset base frequency range can be identified as heterozygous single nucleotide polymorphism sites for subsequent analysis.
[0051] In addition, to support subsequent haplotype phasing algorithms, the computing device 110 may further extract the base type information of each heterozygous single nucleotide polymorphism site in different sequencing sequences. This process can be achieved by parsing the alignment result file and the list of heterozygous single nucleotide polymorphism sites, and extracting the base type of each sequencing sequence at each heterozygous single nucleotide polymorphism site one by one. If a certain sequencing sequence does not cover a single nucleotide polymorphism site, it is recorded as "missing".
[0052] In some embodiments, to improve the quality of heterozygous single nucleotide polymorphism sites that support haplotype phasing algorithms, the computing device 110 may further filter the heterozygous single nucleotide polymorphism sites according to preset filtering rules, so as to improve the accuracy and reliability of subsequent haplotype phasing analysis.
[0053] Regarding the preset filtering rules, they include, for example, but are not limited to one or more of the following: removing heterozygous single nucleotide polymorphism sites located within a region where the number of consecutive identical bases exceeds a preset threshold. For example, when the number of consecutive occurrences of the same base in a certain segment reaches or exceeds 5, the heterozygous single nucleotide polymorphism sites appearing in this region will be excluded; removing heterozygous single nucleotide polymorphism sites located within a tandem repeat region. For example, when a sequence of two or more bases that repeats multiple times (for example, more than 5 times) in a head-to-tail manner appears in a certain segment, the heterozygous single nucleotide polymorphism sites appearing in this region will be excluded; and, by comparing with a preset "blacklist" region, removing heterozygous single nucleotide polymorphism sites located within known low-quality regions or regions with inaccurate sequencing such as the "blacklist" region. The preset "blacklist" region can be determined by manual review or reference to public databases. It should be understood that various thresholds in the preset filtering rules (such as the threshold for the number of consecutive identical bases, the length of tandem repeat units, the number of tandem repeat unit repetitions, and the preset "blacklist" region, etc.) can be flexibly adjusted according to factors such as the experimental design of those skilled in the art, the characteristics of the target region, and the characteristics of the sequencing platform.
[0054] In some embodiments, in order to systematically record the information of heterozygous single nucleotide polymorphism sites for subsequent analysis and processing, the computing device 110 can, for example, arrange the above-mentioned heterozygous single nucleotide polymorphism sites according to their coordinates on the reference genome and their corresponding base types for each sequencing fragment to form a heterozygous single nucleotide polymorphism matrix, and the heterozygous single nucleotide polymorphism matrix indicates the base type of each heterozygous single nucleotide polymorphism site in each sequencing sequence.
[0055] Specifically, let the input sequencing sequence set containing sequencing sequences be:
[0056] where Represents a sequencing sequence for a target fragment.
[0057] Suppose the target fragment contains pre - processed and filtered heterozygous single - nucleotide polymorphism (SNP) sites, denoted as:
[0058] Among them, represents a heterozygous single - nucleotide polymorphism site for the target fragment and is arranged in the order of its coordinates in the reference genome.
[0059] Accordingly, define a heterozygous single - nucleotide polymorphism matrix of size :
[0060] Among them, the matrix element indicates the base type of the sequencing sequence at the heterozygous single - nucleotide polymorphism site . If , it means that the sequencing sequence does not cover the heterozygous single - nucleotide polymorphism site .
[0061] Through this matrix, the base types of each sequencing sequence at each heterozygous single - nucleotide polymorphism site can be systematically described, thus supporting a series of subsequent processes of the method of the present invention on this basis. In addition, in some prior arts, haplotype phasing algorithms usually construct a heterozygous single - nucleotide polymorphism matrix based on the dimorphism of single - nucleotide polymorphism sites. However, when dealing with multi - copy genes or regions with multiple allelic base types, this method reduces the ability of the algorithm to describe polymorphisms, and thus affects the accuracy of haplotype phasing for genes with complex structures such as multi - copy genes. In contrast, this application completely retains the original base information and improves the ability to accurately phase haplotypes for genes with complex structures.
[0062] At step 206, the computing device 110 recursively partitions the input sequencing sequence set multiple times based on the contribution degree of each heterozygous single - nucleotide polymorphism site to haplotype phasing, so as to determine the haplotype of the target fragment.
[0063] A method for determining the haplotype of a target fragment by performing multiple recursive partitions on the input sequencing sequence set based on the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing, for example, includes: in each recursion, the computing device 110 evaluates the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing based on the current input sequencing sequence set; divides the sequencing sequences in the current input sequencing sequence set into multiple subsets based on the contribution degree; in response to any one of the subsets satisfying the termination condition, independently outputs the subset satisfying the termination condition as a haplotype of the target fragment; and, in response to any one of the subsets not satisfying the termination condition, independently updates the current input sequencing sequence set based on the subset not satisfying the termination condition and advances to the next recursion until the termination condition is satisfied. The following will be combined with Figure 3 Specifically illustrate the method for performing multiple recursive partitions on the input sequencing sequence set to determine the haplotype of the target fragment, which will not be elaborated here.
[0064] In the above solution, by determining the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence within the target fragment based on the alignment result of the sequencing sequence of the target fragment with the reference genomic sequence; performing multiple recursive partitions on the input sequencing sequence set based on the contribution degree of each heterozygous single nucleotide polymorphism site to haplotype phasing, thereby determining the haplotype of the target fragment. The method of the present invention systematically evaluates the contribution degree of each heterozygous single nucleotide polymorphism site to accurately partitioning the haplotype, and based on this, performs multiple recursive partitions on multiple sequencing sequences of the target fragment, and determines the haplotype of the target fragment according to the compatibility of the sequencing sequences. The present invention can efficiently and accurately perform haplotype typing on multi-copy genes.
[0065] The following will be combined with Figure 3 Describe method 300 for performing multiple recursive partitions on an input sequencing sequence set to determine the haplotype of a target fragment according to an embodiment of the present invention. Figure 3 FIG. shows a flowchart of method 300 for performing multiple recursive partitions on an input sequencing sequence set to determine the haplotype of a target fragment according to an embodiment of the present invention. It should be understood that method 300 can be executed, for example, at Figure 5 the described electronic device 500. It can also be executed at Figure 1 the described computing device 110. It should be understood that method 300 may further include additional actions not shown and / or may omit the shown actions, and the scope of the present invention is not limited in this regard.
[0066] In method 300, the computing device 110 starts a recursive process for determining the haplotype of the target fragment.
[0067] Specifically, to start the process, computing device 110 initially, at step 302, evaluates the contribution of each heterozygous single nucleotide polymorphism (SNP) site to haplotype phasing based on the current input set of sequencing sequences.
[0068] In some embodiments, due to multiple rounds of recursion, at least some SNP sites have become homozygous in the current input set of sequencing sequences. Thus, to at least partially save computing resources, computing device 110 may re-identify heterozygous SNP sites based on the current input set of sequencing sequences.
[0069] Regarding the method of re-identifying heterozygous SNP sites based on the current input set of sequencing sequences, for example, it includes: for each SNP site , computing device 110 compares the base types of each sequencing sequence in the current input set of sequencing sequences at this site. If two or more different non-deleted bases are observed at this site, this site is retained as a heterozygous site in the current recursion stage, and all sites that meet this condition are included in the set .
[0070] In some embodiments, regarding the method of evaluating the contribution of a heterozygous SNP site to haplotype phasing, for example, it includes: for the heterozygous SNP site to be evaluated , according to the base types of each sequencing sequence in the current input set of sequencing sequences at the heterozygous SNP site to be evaluated , the sequencing sequences in the current input set of sequencing sequences are divided into multiple subsets ; traverse all heterozygous SNP sites regarding the target fragment , for each of the heterozygous SNP sites among them : in response to the base types of each sequencing sequence in each subset being unique at this heterozygous SNP site, the contribution of the heterozygous SNP site to be evaluated is incremented by 1 point; in response to the base types of each sequencing sequence in any subset not being unique at this heterozygous SNP site, the contribution of the heterozygous SNP site to be evaluated is not increased; accumulate the total score as the contribution of the heterozygous SNP site to be evaluated to haplotype phasing. The higher the contribution of a certain SNP site, the more it helps to accurately perform haplotype phasing. Method 400 for evaluating the contribution of a heterozygous SNP site to haplotype phasing will be specifically described below in conjunction with Figure 4 and will not be elaborated here.
[0071] Once the contribution of each heterozygous single nucleotide polymorphism (SNP) site to haplotype phasing is obtained, computing device 110 divides the sequencing sequences in the current input sequencing sequence set into multiple subsets based on the contribution at step 304 for further recursive processing.
[0072] In one embodiment, computing device 110 selects the heterozygous SNP site with the highest contribution as the basis for dividing the current set, and divides the current input sequencing sequence set into multiple subsets based on the base types of the sequencing sequences at this site. Specifically, if multiple different bases (such as A, T, C, or G) are observed at the site , then the sequencing sequences in the current input sequencing sequence set are grouped according to the base types to construct a set of subsets , where each subset contains the sequencing sequences with the base at the site . This strategy helps to preferentially select the heterozygous SNP sites that are most discriminative for different haplotypes, thereby improving the overall typing efficiency and accuracy.
[0073] In some cases, when there are multiple heterozygous SNP sites with the same and highest contribution, it means that these sites have equivalent discrimination ability for the target fragment haplotypes at the current stage. Therefore, computing device 110 can arbitrarily select one of the heterozygous SNP sites as the basis for set division without distinction.
[0074] Subsequently, at step 306, computing device 110 determines whether each of the subsets obtained from the above division satisfies the termination condition. In one embodiment, the termination condition is that all the sequencing sequences in the subset are compatible. It should be understood that for two sequencing sequences, if there exists a single nucleotide polymorphism site such that the two sequencing sequences have different and non-missing bases at this single nucleotide polymorphism site, then the two sequencing sequences are said to be "in conflict"; conversely, if the two sequencing sequences do not conflict at all single nucleotide polymorphism sites, then the two sequencing sequences are said to be "compatible". If all the sequencing sequences in a sequencing sequence set are compatible, it indicates that this set can uniquely determine a haplotype.
[0075] In response to the judgment result in step 306, if there is any subset that does not satisfy the termination condition, then computing device 110 updates the current input sequencing sequence set , and independently proceed to the next round of recursion, repeating the above steps 302 to 306 until all subsets meet the termination condition. Since each round of recursion strictly reduces the sequence conflicts in the subset, the recursive processing of the method of the present application will eventually converge and return a set of non-overlapping and internally compatible subsets of sequencing sequences, corresponding to the complete haplotype set of the target fragment.
[0076] In response to the judgment result in step 306, if a subset meets the termination condition, that is, all the sequencing sequences in the subset are compatible, it indicates that this subset can uniquely determine a haplotype. Therefore, the computing device 110 outputs this subset independently as a haplotype of the target fragment at step 310.
[0077] In one embodiment, to improve the overall processing efficiency, the computing device 110 can process multiple subsets in parallel during each recursive process, thereby shortening the total calculation time of the phasing algorithm of the present invention.
[0078] Furthermore, the following further illustrates the recursive phasing process in the method 300 of the present invention in combination with mathematical formulas.
[0079] First, optionally, for the current input sequencing sequence set , re-identify the set of heterozygous single nucleotide polymorphism sites on the target fragment that can be used for phasing:
[0080] Among them, represents that in the current input sequencing sequence set , at least two different non-missing bases are observed at the locus . Therefore, the locus is considered heterozygous in the current input sequencing sequence set and is included in the set .
[0081] For the heterozygous single nucleotide polymorphism site to be evaluated, evaluate the contribution degree of this site to haplotype phasing:
[0082] The function is used to evaluate the contribution degree of the heterozygous single nucleotide site to haplotype phasing. It will be specifically described below in combination with Figure 4 , and will not be elaborated here.
[0083] The greater the contribution degree of a certain site to be evaluated, the stronger the ability of this site to accurately distinguish different haplotypes, and the more helpful it is to divide the sequencing sequence set Divided into homozygous subsets. Therefore, select the locus with the highest contribution as the basis for dividing the current set:
[0084] According to the locus and the base type , divide the current input sequencing sequence set into multiple subsets :
[0085] Among them, represents all bases that have appeared at the locus in the set :
[0086] For each newly generated subset , check whether all the sequencing sequences in it are compatible:
[0087] Similarly, means that in the subset , only one non-missing base is observed at the locus . If only one non-missing base is observed at all loci , then it is considered that all the sequencing sequences in this subset are compatible. Therefore, the subset can uniquely determine a haplotype.
[0088] For subsets that do not meet the above conditions , take them as the new current input sequencing sequence set , and enter the next round of recursion:
[0089] The function takes the union of the processing results of all returned subsets, and finally returns a set of non-overlapping and internally compatible sequencing sequence subsets, corresponding to the complete haplotype combination of the target fragment.
[0090] By adopting the above recursive phasing strategy, the present invention realizes the efficient division of the sequencing sequence set, thereby gradually converging to the complete haplotype combination of the target fragment. This strategy can effectively improve the efficiency and accuracy of haplotyping of multi-copy genes, providing an important data basis for further research and application.
[0091] The following will be combined with Figure 4A method for evaluating the contribution degree of a heterozygous single nucleotide polymorphism site to haplotype phasing according to an embodiment of the present invention is described. Figure 4 FIG. 400 is a flowchart of a method for evaluating the contribution degree of a heterozygous single nucleotide polymorphism site to haplotype phasing according to an embodiment of the present invention. It should be understood that the method 400 can be executed, for example, at the Figure 5 electronic device 500 described. It can also be executed at the Figure 1 computing device 110 described. It should be understood that the method 400 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this regard.
[0092] At step 402, the computing device 110 divides the current input sequencing sequence set for the heterozygous single nucleotide polymorphism site to be evaluated into multiple subsets according to the base types of each sequencing sequence at this site in the current input sequencing sequence set . Specifically, if multiple different bases (such as A, T, C, or G) are observed at the site , the sequencing sequences in the current input sequencing sequence set are grouped according to the base types to construct a set of subsets , where each subset contains the sequencing sequences with the base being at the site.
[0093] At step 404, the computing device 110 traverses each heterozygous single nucleotide polymorphism site of the target fragment and determines whether all the sequencing sequences in each subset obtained according to have a unique base type at the site . It should be understood that the base types at the site in different subsets can be different.
[0094] In response to the determination at step 404, if in each subset, all the sequencing sequences have a unique base type at the site , it indicates that different haplotypes can be clearly distinguished at the site according to the site to be evaluated, which has a positive effect on accurate haplotype phasing. Therefore, the computing device 110 adds 1 point to the contribution degree of the heterozygous single nucleotide polymorphism site to be evaluated at step 406; on the contrary, if multiple different bases appear at the site in any subset, it indicates that the site to be evaluated cannot be used to distinguish different haplotypes clearly. Differentiating site has no positive effect on accurate phasing. Therefore, the computing device 110 does not add points to the contribution degree of the site at step 408.
[0095] At step 410, the computing device 110 determines whether all heterozygous single nucleotide polymorphism sites regarding the target fragment have been traversed. If the traversal is not completed, it returns to step 404 to continue the determination; if all sites have been traversed, the total score is accumulated at step 412 to obtain the contribution degree of the site to haplotype phasing. The higher the contribution degree of a certain single nucleotide polymorphism site, the more heterozygous single nucleotide polymorphism sites can be differentiated based on this site, which is more conducive to accurate haplotype phasing.
[0096] Furthermore, the following further illustrates the method for evaluating the contribution degree of heterozygous single nucleotide polymorphism sites to haplotype phasing in combination with mathematical formulas.
[0097] First, for the current input sequencing sequence set , and the heterozygous single nucleotide polymorphism site to be evaluated , the set is divided into multiple subsets according to the base type at the site :
[0098] Among them, represents the set of all bases that have appeared at the site in the set :
[0099] Define the function to determine whether the base type at the site in a certain subset is unique:
[0100] Accordingly, the contribution degree of the heterozygous single nucleotide polymorphism site to be evaluated is defined as:
[0101] Among them, is used to determine that in all subsets divided according to the site , all sequencing sequences at the site Whether the base at a certain position is unique. Therefore, the function indicates that: if in all subsets at a certain position the condition that the base type is unique is satisfied, it is considered that the site to be evaluated can correctly distinguish the site , so one point is added to the contribution degree of the site to be evaluated . Traverse all . Each time this condition is met, one point is added to the contribution degree of the site .
[0102] By the above method, the present invention quantitatively scores the heterozygous single nucleotide polymorphism sites through the contribution degree, so as to objectively evaluate the discrimination ability of the heterozygous single nucleotide polymorphism sites in the process of phasing the target fragment. Therefore, the method of the present invention can effectively improve the haplotype phasing accuracy in the context of multi-copy genes.
[0103] Figure 5 Schematically shows a block diagram of an electronic device 500 suitable for implementing the embodiments of the present invention. The electronic device 500 may be a device for implementing the methods 200 to 400 shown in Figures 2 to 4 . As shown in Figure 5 , the electronic device 500 includes a central processing unit (i.e., CPU 501), which can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (i.e., ROM 502) or the computer program instructions loaded from the storage unit 508 into the random access memory (i.e., RAM 503). In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The CPU 501, ROM 502, and RAM 503 are connected to each other through a bus 504. The input / output interface (i.e., I / O interface 505) is also connected to the bus 504.
[0104] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508. The CPU 501 executes the various methods and processes described above, such as executing methods 200 to 400. For example, in some embodiments, methods 200 to 400 may be implemented as computer software programs, which are stored in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the CPU 501, one or more operations of methods 200 to 400 described above may be performed. Alternatively, in other embodiments, the CPU 501 may be configured to execute one or more actions of methods 200 to 400 by any other suitable means (e.g., by means of firmware).
[0105] It should be further noted that the present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0106] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0107] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0108] The computer program instructions for carrying out operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages - such as Smalltalk, C++, etc., and conventional procedural programming languages - such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present invention.
[0109] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0110] These computer-readable program instructions can be provided to a processing unit of a processor in a voice interaction device, a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that, when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which causes a computer, a programmable data processing device, and / or other devices to operate in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0111] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing device, or other device, so that a series of operation steps are executed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable data processing device, or other device implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, a program segment, or a part of an instruction, and the module, program segment, or part of an instruction includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the boxes may occur in a different order than marked in the accompanying drawings. For example, two consecutive boxes may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that executes the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0113] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the technical field to understand the embodiments disclosed herein.
[0114] The above are only alternative embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for determining a target fragment haplotype, comprising: Obtaining an input sequencing sequence set for a target fragment from a sample to be tested, wherein the input sequencing sequence set comprises a plurality of sequencing sequences for the target fragment; Comparing the sequencing sequence with the reference genome sequence to determine the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence in the target fragment; Based on the contribution of each heterozygous single nucleotide polymorphism site to haplotype phasing, the input sequencing sequence set is recursively partitioned multiple times to determine the haplotype of the target fragment.
2. The method according to claim 1, wherein the sequencing sequence of the target fragment is a single molecule long read sequencing sequence of the target fragment obtained by single molecule long read sequencing technology.
3. The method according to claim 1, wherein the step of determining the heterozygous single nucleotide polymorphism information corresponding to each sequencing sequence in the target fragment comprises: Compare the sequencing sequence with the reference genome sequence to determine the single nucleotide polymorphism site corresponding to each sequencing sequence in the target fragment; The single nucleotide polymorphism sites whose maximum base frequencies are within the preset base frequency range are retained as heterozygous single nucleotide polymorphism sites, and the base types of the heterozygous single nucleotide polymorphism sites corresponding to each sequencing sequence are obtained.
4. The method according to claim 3, further comprising filtering the heterozygous single nucleotide polymorphism sites according to a preset filtering rule, wherein the preset filtering rule comprises at least one of the following: Removing heterozygous single nucleotide polymorphism sites located in regions where the number of consecutive identical bases exceeds a preset threshold; Removing heterozygous single nucleotide polymorphism sites located within tandem repeat regions; and, Remove heterozygous single nucleotide polymorphism sites located in the preset "blacklist" region.
5. The method according to claim 1, wherein the step of recursively partitioning the input sequencing sequence set multiple times based on the contribution of each heterozygous single nucleotide polymorphism site to haplotype phasing, thereby determining the haplotype of the target fragment comprises: In each recursion, based on the current input sequencing sequence set, the contribution of each heterozygous single nucleotide polymorphism site to haplotype phasing is evaluated; Based on the contribution, the sequencing sequences in the current input sequencing sequence set are divided into a plurality of subsets; In response to any of the subsets satisfying a termination condition, independently outputting the subset satisfying the termination condition as a haplotype of the target fragment; and, In response to any of the subsets not satisfying the termination condition, the current input sequencing sequence set is updated independently based on the subset that does not satisfy the termination condition, and the process proceeds to the next recursion until the termination condition is satisfied.
6. according to the method described in claim 1 or 5, the assessment method of the contribution of the heterozygous single nucleotide polymorphism site to haplotype phasing comprises: For the heterozygous single nucleotide polymorphism site to be evaluated, the sequencing sequences in the current input sequencing sequence set are divided into a plurality of subsets according to the base type of each sequencing sequence in the current input sequencing sequence set at the heterozygous single nucleotide polymorphism site to be evaluated; Traverse all heterozygous single nucleotide polymorphism sites about the target fragment, for each heterozygous single nucleotide polymorphism site: In response to the fact that in each subset, the base types of each sequencing sequence at the heterozygous single nucleotide polymorphism site are all unique, the contribution of the heterozygous single nucleotide polymorphism site to be evaluated is increased by 1 point; In response to the fact that in any subset, the base type of each sequencing sequence at the heterozygous single nucleotide polymorphism site is not unique, the contribution of the heterozygous single nucleotide polymorphism site to be evaluated does not increase; The total score is accumulated as the contribution of the heterozygous single nucleotide polymorphism site to be evaluated to the haplotype phasing.
7. The method according to claim 5, wherein said dividing the sequencing sequences in the current input sequencing sequence set into a plurality of subsets based on the contribution comprises: The heterozygous single nucleotide polymorphism site with the highest contribution is selected, and the sequencing sequences in the current input sequencing sequence set are divided into multiple subsets according to the base type of each sequencing sequence in the current input sequencing sequence set at the heterozygous single nucleotide polymorphism site with the highest contribution.
8. The method according to claim 7, wherein when a plurality of heterozygous single nucleotide polymorphism sites have the same highest contribution, one of the heterozygous single nucleotide polymorphism sites is arbitrarily selected as the basis for set division.
9. The method according to claim 5, further comprising processing the multiple subsets in parallel in each recursion until all subsets satisfy the termination condition.
10. The method according to claim 5, wherein the termination condition is that all sequencing sequences in the subset are compatible.
11. A computing device comprising: at least one processing unit; as well as, At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the device to perform the steps of the method according to any one of claims 1 to 10.
12. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 10 when executed by a machine.
13. A computer program product comprising a computer program, which, when executed by a machine, performs the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Method for rebuilding individual single somatotype based on optimizing solution aggregate
CN101256602A
Haploid phasing and variation detection methods and devices of diploid genome read
CN110021355A
Method and device for haplotype phasing of diploid genome based on third generation capture sequencing
CN110621785A
Structural genetic variation detection method and device and computer equipment
CN118782139A