Systems and methods for related error event mitigation for variant recognition

By considering the related error events of mapping errors and sequencing errors, a probability model is used to determine the possibility of true variants, which solves the genotyping error problem of existing variant recognizers when identifying variants, and improves recognition accuracy.

CN120015114APending Publication Date: 2025-05-16ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510083750.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-02-16
Filing Date
2019-02-19
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing variant recognizers are susceptible to mapping errors and sequencing errors when identifying variants, resulting in genotyping errors.

Method used

By considering the indication of related error events, a method action is adopted, which includes accessing from the memory accumulating the stacking of multiple sequence reads aligned with the reference genome, obtaining characteristic information of the reads, and inputting it into the probability model based on this information to determine the possibility of a true variant.

Benefits of technology

It improves the accuracy of variant recognition, reduces the occurrence of genotyping errors, and can more accurately distinguish between errors and true variants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015114A_ABST
    Figure CN120015114A_ABST
Patent Text Reader

Abstract

The present invention relates to methods, systems, and apparatus, including computer programs, for improving the accuracy of variant identification by taking into account indications of related error events. In one aspect, a method may comprise the actions of accessing a stack of sequence reads aligned to a first region of a reference genome; obtaining information describing one or more characteristics of each of the plurality of reads of the stack; providing one or more inputs describing the one or more characteristics of the plurality of reads of the stack to a probabilistic model, wherein the probabilistic model is configured to determine, for each of one or more hypotheses selected based on the one or more inputs, a score indicating whether each hypothesis is true; obtaining output information of each hypothesis in the one or more hypotheses; and determining a likelihood of presence of a true variant at the first location based on the obtained output information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a Chinese invention patent application with an application date of February 19, 2019, application number 201980002932.0, and name “System and method for mitigating related error events for variant identification”.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims the benefit of U.S. Provisional Application No. 62 / 710,348, filed on February 16, 2018, and entitled “Methods, Devices, and Systems for Performing Foreign Read Detection and Burst Error Detection,” the disclosure of which is incorporated herein by reference in its entirety. Background Art

[0004] A nucleic acid sequencer is an instrument configured to automate the process of sequencing nucleic acids, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Nucleic acid sequencing is the process of determining the order of nucleotides in a gene sequence.

[0005] The nucleic acid sequencer is configured to receive a nucleic acid sample and generate output data, which is called one or more "reads" that represent the sequence of nucleotides in the nucleic acid sample. The nucleotides in the DNA sample can include one or more of guanine (G), cytosine (C), adenine (A), and thymine (T) in any combination. The nucleotides in the RNA sample can include one or more of G, C, A, and uracil (U) in any combination.

[0006] The DNA reads generated by the DNA sequencer can be mapped to and aligned with the known DNA sequence of a reference genome. Once mapped to and aligned with the reference genome, the mapped and aligned read sequences can be analyzed in light of the reference genome to identify potential variations between the mapped and aligned read sequences and the reference genome. Summary of the invention

[0007] Aspects of the present disclosure relate to methods, systems and apparatus, including computer programs, for accounting for relevant error events that create problems for a variant caller, wherein the variant caller is used to determine whether an alternative (hereinafter referred to as "alt") allele identified in a pileup of mapped and aligned sequence reads is a true variant.

[0008] According to one innovative aspect of the present disclosure, a method for improving the accuracy of variant identification by considering indications of correlated error events is disclosed. In some embodiments, the method may include method actions, the method actions including: accessing, by one or more computers and from one or more memory devices, a stack of multiple sequence reads aligned to a first region of a reference genome; obtaining, by the one or more computers, information describing one or more characteristics of each of the multiple reads of the stack corresponding to a first position of the reference genome; providing, by the one or more computers and based on the obtained information, one or more inputs describing the one or more characteristics of the multiple reads of the stack to a probability model, wherein the probability model is configured to determine, for each of one or more hypotheses selected based on the one or more inputs, a score indicating whether the hypothesis is true; obtaining, by the one or more computers, output information for each of the one or more hypotheses, wherein the output information for each of the one or more hypotheses (i) is generated by the probability model based on the one or more inputs describing the one or more characteristics of the corresponding reads of the stack processed to the probability model by the probability model, and (ii) indicates a score indicating whether the hypothesis is true; and determining, by the one or more computers and based on the obtained output information generated by the probability model for each of the multiple hypotheses, a likelihood that a true variant exists at the first position.

[0009] Other versions include corresponding systems, apparatus, and computer programs for performing the method actions defined by instructions encoded on a computer-readable storage device.

[0010] These and other versions may optionally include one or more of the following features. For example, in some embodiments, determining, by the one or more computers and based on the obtained output information generated by the probability model for each of the multiple hypotheses, the likelihood that a true variant exists at the first position may include: determining, by the one or more computers, a total score based on the output information generated by the probability model for each of the multiple hypotheses, wherein the total score indicates the likelihood that the true variant exists; determining, by the one or more computers, whether the score generated by the total score satisfies a predetermined threshold; and adding information indicating that a true variant exists at the first position to the VCF file based on determining, by the one or more computers, that the total score satisfies the predetermined threshold.

[0011] In some embodiments, the information indicating the presence of a true variant at the first position can include information identifying: (i) the first position, (ii) a candidate alt allele at the first position, and (iii) the total score.

[0012] In some embodiments, the information describing the one or more characteristics of the corresponding reads may include information describing: (i) a mapping quality score for each sequence read in the stack at the first position; and (ii) a read allele score for each sequence read in the stack at the first position for each candidate allele at the first position.

[0013] In some embodiments, the stacked read allele score for each read at the first position is based on an output generated by a P-HMM model for each of the reads at the first position, the output indicating the probability that a particular candidate allele G is present given the candidate allele G. m,φ The observed read segment r i probability.

[0014] In some embodiments, the output information may include: first output information for a first hypothesis among the one or more hypotheses, the first output information including the possibility that the sequence read at the first position indicates the occurrence of a homozygous reference with an alien allele matching alt; and second output information for a second hypothesis among the one or more hypotheses, the second output information including the possibility that the sequence read at the first position indicates the occurrence of a homozygous alt with an alien allele matching a reference allele.

[0015] In some embodiments, the information describing the one or more characteristics of the corresponding sequence reads may include information describing: (i) the read orientation of each sequence read at the first position of the stack, (ii) the position of each base at the first position within each sequence read at the first position of the stack with reference to the 5′ end of the sequence read, (iii) the read allele score of each sequence read in the multiple sequence reads for each candidate allele at the reference position, and (iv) the base quality score of each read of the base at the first position.

[0016] In some embodiments, the read allele score for each of the sequence reads at the first position of the stack is based on an output generated by a P-HMM model for each of the sequence reads at the first position, the output indicating the probability that a particular candidate allele G is present given the candidate allele G. m,φ The observed sequence reads r iprobability.

[0017] In some embodiments, the information describing the one or more characteristics of the corresponding sequence reads may include information describing: (i) the read orientation of each sequence read at the first position of the stack, (ii) the position of each base at the first position within each sequence read at the first position of the stack with reference to the 5′ end of the sequence read, and (iii) the read allele score of each sequence read in the multiple sequence reads for each candidate allele at the reference position.

[0018] In some embodiments, the output information may include: first output information for a first hypothesis among the one or more hypotheses, the first output information may include the possibility that the sequence read at the first position indicates the occurrence of a homozygous reference with a sequencing error that matches the alt allele; and second output information for a second hypothesis among the one or more hypotheses, the second output information includes the possibility that the sequence read at the first position indicates the occurrence of a homozygous alt with a sequencing error that matches the reference allele.

[0019] In some embodiments, the information describing the one or more characteristics of the corresponding sequence reads may include information describing: (i) the read orientation of each sequence read at the first position of the stack, (ii) the position of each base at the first position within each sequence read, e.g., position "0" 142, with reference to the 5' end of the sequence read, (iii) the mapping quality score of each sequence read at the first position of the stack, (iv) the read allele score of each sequence read in the multiple reads for each candidate allele at the reference position, and (v) the base quality score of each read of the base aligned at the first position.

[0020] In some embodiments, the read allele score for each of the sequence reads at the first position of the stack is based on an output generated by a P-HMM model for each of the sequence reads at the first position, the output indicating the probability that a particular candidate allele G is present given the candidate allele G. m,φ The observed read segment r i probability.

[0021] In some embodiments, the output information may include: the sequence read at the first position indicates a first possibility of the occurrence of a homozygous reference with an alien allele matching alt; the sequence read at the first position indicates a second possibility of the occurrence of a homozygous alt with an alien allele matching a reference allele; third output information of a first hypothesis among the one or more hypotheses, the third output information including the possibility that the sequence read at the first position indicates the occurrence of a homozygous reference with a sequencing error matching the alt allele; and fourth output information of a second hypothesis among the one or more hypotheses, the fourth output information including the possibility that the sequence read at the first position indicates the occurrence of a homozygous alt with a sequencing error matching the reference allele.

[0022] In some embodiments, the one or more memory devices receive the stack of aligned sequence reads from a field programmable gate array (FPGA) device, wherein the FPGA comprises one or more configurable digital logic gates that have been configured as a mapping and alignment unit to perform read mapping and alignment.

[0023] In some embodiments, the computer is configured to access the one or more memory devices using one or more wired or wireless networks, wherein a field programmable gate array (FPGA) device and the one or more memory devices are housed in an expansion card that has been coupled to a circuit board of a sequencer, wherein the sequencer is configured to generate sequence reads based on input samples and store the generated sequence reads in the one or more memory devices, and wherein the mapping and alignment unit of the FPGA is configured to access the one or more memory devices to obtain the generated sequence reads.

[0024] In some embodiments, the computer and the sequencer are each configured to access the one or more memory devices using one or more wired or wireless networks, wherein the field programmable gate array (FPGA) device and the one or more memory devices are housed in an expansion card that has been coupled to a circuit board of a server located remote from the computer and the sequencer, wherein the sequencer is configured to generate sequence reads based on input samples, provide the generated sequence reads to the server using the one or more wired or wireless networks to store the generated sequence reads in the one or more memory devices, and wherein the mapping and alignment unit of the FPGA is configured to access the one or more memory devices to obtain the generated sequence reads.

[0025] These and other aspects of the disclosure are discussed in more detail below in the detailed description with reference to the accompanying figures. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a context diagram of an example of a system for associated error mitigation for variant identification.

[0027] Figure 2 is a flow chart of an example of a process for associated error mitigation for variant identification.

[0028] Figure 3 is a flow chart of an example of a process for mapping error mitigation for variant identification.

[0029] Figure 4 is a line graph of an example of a function for converting an example of an output value from a mapping and alignment unit representing a first mapping confidence score to a second mapping confidence score (μ).

[0030] Figure 5 is a flow chart of an example of a process for sequencing error mitigation for variant identification.

[0031] Figure 6 is another flow chart of an example of a process for related error mitigation for variant identification.

[0032] Figure 7 is another flow chart of an example of a process for related error mitigation for variant identification.

[0033] Figure 8 is a block diagram of system components that may be used to implement a system for correlated error mitigation for variant identification.

[0034] Fig. 9A is an example of a summary of experimental results from execution of a process for relevant error mitigation for variant calling that has been performed on a pile of sequencing reads whose results indicate examples of true positive results.

[0035] Fig. 9B yes Fig. 9A An instance of a stacked set of modified probability outcomes.

[0036] Fig. 10A is an example of a summary of experimental results from the execution of a process for relevant error mitigation of variant calling that has been performed on a pile of sequencing reads whose results indicate a low probability of the occurrence of mapping errors.

[0037] Fig. 10B yes Fig. 10A An instance of a stacked set of modified probability outcomes.

[0038] Fig.11Ais an example of a summary of experimental results from the execution of a process for associated error mitigation of variant identification that has been performed on a pile of sequencing reads whose results indicate a high likelihood that a candidate alt allele is unlikely to be a true variant due to sequencing error.

[0039] Fig. 11B yes Fig.11A An instance of a stacked set of modified probability outcomes.

[0040] Fig. 12A is an example of a summary of experimental results from the execution of a process for associated error mitigation for variant identification, which has been performed on a pile of sequencing reads whose results indicate a high probability that the candidate alt allele is unlikely to be a true variant due to sequencing errors in both read orientations.

[0041] Fig. 12B is an example of a set of modified probability results for the stack of FIG. 12 . DETAILED DESCRIPTION

[0042] Aspects of the present disclosure relate to methods, systems, and apparatus, including computer programs, for accounting for errors that create problems for variant callers.

[0043] Related error events can cause conventional variant identifiers to produce genotyping errors, because the internal probability mathematics used by conventional variant identifiers is usually based on the assumption that errors are unrelated. There are two phenomena that often produce highly correlated errors in variant identification: (1) mapping errors and (2) sequence-specific errors. When the read is mapped to a specific location of the reference genome other than the true source of the read, a mapping error will occur. The appearance of sequence-specific errors, which are referred to as sequencing errors or system errors in this manual, is because some base sequences often have a high probability of producing sequencing errors. Two types of errors may cause high-confidence false positives and other genotyping errors in conventional variant identifiers.

[0044] Some variant callers attempt to mitigate these problems using ad hoc rules for screening reads and variants, but such rules do not approach the performance limits that could be achieved using more complex algorithms. Other variant callers use machine learning to identify and suppress such errors, but machine learning has other drawbacks, such as being "fragile" due to difficulty handling situations not represented in the training data, or "opaque" because the output of such situations cannot be explained. The present disclosure describes new methods for handling both types of errors. Specifically, the present disclosure addresses the possible fact that certain errors are associated with probability calculations in read stacking, rather than relying on a series of ad hoc rules or machine learning to account for the associated properties of these types of errors.

[0045] For purposes of this disclosure, a "correlated error event" is an error class that refers to two or more mapping errors or two or more sequencing errors. The processes described herein can be applied to consider a single type of correlated error event, such as one or more mapping errors or one or more sequencing errors. Alternatively, the processes described herein can be applied to consider multiple types of correlated errors, such as one or more mapping errors and one or more sequencing errors.

[0046] Related Error Event Mitigation System

[0047] Figure 1 1 is a context diagram of an example of a system 100 for variant identification-related error event mitigation. The system 100 includes a nucleic acid sequencer 110 and an auxiliary analysis unit 120.

[0048] The nucleic acid sequencer 110 is configured to perform a primary analysis on the biological sample 105 to convert the raw physical signal detected by the sequencing instrument into a series of ordered nucleotide base recognitions with associated quality scores. The primary analysis is specific to the nature of the sequencing technology employed. In some embodiments, for example, nucleotides can be detected by sensing changes in fluorescence, charge, current, or radiated light, or any combination thereof. In some embodiments, the biological sample includes a chimeric or hybrid form of DNA, RNA, PNA, LNA, nucleic acid.

[0049] The nucleic acid sample can be a purified sample, or a crude DNA sample containing, for example, a lysate from an oral swab, paper, fabric, or other substrate that may be impregnated with saliva, blood, or other body fluids. In some embodiments, the nucleic acid sample may comprise a small amount of DNA or a fragmented portion of DNA, such as genomic DNA. In some embodiments, the target sequence may be present in one or more body fluids, including but not limited to blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence may comprise a nucleic acid obtained from non-human DNA such as microbial, plant, or insect DNA.

[0050] The primary analysis may include receiving a biological sample 105 and generating output data 112, referred to as one or more base calls, each of which has a quality score, compiled into a plurality of "reads," each of which represents a set of ordered nucleotides in a sequence fragment prepared from the received biological sample 105. In some embodiments, the biological sample 105 may include a DNA sample, and the sequencer 110 may perform the primary analysis to output a plurality of reads comprising a series of ordered nucleotides or bases from the DNA sample. In such embodiments, the sequenced nucleotide sequence comprises one or more of guanine (G), cytosine (C), adenine (A), and thymine (T) in any combination. In other embodiments, the biological sample 105 may include an RNA sample. In such embodiments, the sequenced nucleotide sequence comprises one or more of G, C, A, and uracil (U) in any combination. Therefore, although Figure 1 The example of describes a DNA sequencer 110 that generates output reads based on an input DNA sample, but other embodiments may include a sequencer 110 that generates output reads based on an RNA sample. Depending on the sequencing method used, the length of the read comprising an ordered sequence of contiguous base pairs sequenced from one or more fragments of the biological sample 105 can vary in the range of about 30 base pairs to more than 10,000 base pairs. For example, in some embodiments, the read length of the sequenced fragments can be between about 150 base pairs and 500 base pairs. About 150 base pairs, about 250 base pairs, or about 300 base pairs. The read can be a single read or a double-end read from a fragment prepared from the biological sample 105.

[0051] In some embodiments, the nucleic acid sequencer 110 comprises a next generation sequencer (NGS) configured to generate sequence reads 112 for a given sample 105 in a manner that achieves ultra-high throughput, scalability, and speed by using massively parallel sequencing technology. In various examples, NGS enables rapid sequencing of entire genomes, can amplify target regions for deep sequencing, discover novel RNA variants and splice sites using RNA sequencing (RNA-seq) or quantify mRNA for gene expression analysis, analyze epigenetic factors such as genome-wide DNA methylation and DNA-protein interactions, sequence cancer samples to study rare somatic variants and tumor subclones, and study microbial diversity in the human body or environment.

[0052] The nucleic acid sequencer 110 is configured to generate sequence reads 112 and provide the generated sequence reads 112 to an auxiliary analysis unit 120. The auxiliary analysis unit 120 may include one or more memory devices 122 and one or more computers, such as a field programmable gate array 124 and a variant identification unit 130. The one or more computers may include one or more devices configured to perform one or more operations. The one or more computers may include only hardware, only software, or any combination thereof.

[0053] In some embodiments, the auxiliary analysis unit 120 may be integrated with the nucleic acid sequencer 110. In such embodiments, for example, each of the one or more components of the auxiliary analysis unit 120 may be housed in an expansion card such as a peripheral component interconnect (PCI) expansion card and installed in the nucleic acid sequencer 110. In other embodiments, for example, each of the one or more components of the auxiliary analysis unit 120 may be part of another computer different from the nucleic acid sequencer 110 and directly connected to the nucleic acid sequencer 110 using an Ethernet cable, a USB cable, a USB-C cable, etc. In still other embodiments, for example, each of the components of the auxiliary analysis unit 120 is integrated into a cloud-based server that can be remotely accessed by the nucleic acid sequencer 110 using one or more wired or wireless networks, such as a local area network (LAN), a wide area network (WAN), a cellular network, the Internet, or a combination thereof. In still other embodiments, for example, one or more of the components of the auxiliary analysis unit 120 are integrated into the nucleic acid sequencer 110, and one or more of the components of the auxiliary analysis unit 120 are integrated into another computer, such as a cloud-based server. In such embodiments, for example, the FPGA 124 for implementing the mapping and alignment unit 126 is integrated into the nucleic acid sequencer 110 and the memory 122, and the variant identification unit 130 is integrated into another computer, such as a cloud-based server.

[0054] refer to Figure 1Each of the components discussed, including the nucleic acid sequencer 110, the auxiliary analysis unit 120, and one or more other computers, such as one or more cloud-based servers, may alternatively or additionally be enabled to communicate via one or more wired or wireless networks, including one or more of a LAN, a WAN, a cellular network, the Internet, or a combination thereof, if not enabled to communicate with a direct connection. Similarly, each of the components of the auxiliary analysis unit 120 may be configured to communicate with each other or with components external to the auxiliary analysis unit 120 using one or more buses, one or more direct connections, or one or more networks to enable interactions between the respective components described herein.

[0055] The auxiliary analysis unit 120 is configured to receive the read segment 112 and store the read segment 112 in the first part 122a of the memory 122. The field programmable gate array (FPGA) 124 can be dynamically configured to implement one or more modules of the genome data analysis pipeline. For example, the FPGA 124 can be dynamically configured to implement the mapping and comparison unit 126, the hidden Markov model (P-HMM) unit 128 or both. In some embodiments, the mapping and comparison unit 126 is a single functional module. In other embodiments, the mapping and comparison unit 126 is divided into two discrete functional modules, including a dedicated mapping unit 126a and a dedicated comparison unit 126b. In some embodiments, the FPGA 124 is configured to implement both the mapping and comparison unit 126 and the P-HMM unit 128 at a specific time.

[0056] However, in other embodiments, FPGA 124 can be dynamically reconfigured as needed to implement a specific genome analysis module or any other computer described herein at any given time. For example, FPGA 124 can first be configured to include mapping and alignment unit 126, and then, once FPGA 124 has performed mapping and alignment operations on the read segments obtained from memory 122, FPGA 124 can then be dynamically reconfigured as P-HMM unit 128. FPGA 124 can be dynamically reconfigured from one genome analysis module to another genome analysis module or other computer as required by any specific genome analysis workflow. For the purposes of this disclosure, the terms unit and module are used interchangeably to mean one or more hardware components, one or more software components, or any combination thereof that are configured to perform one or more specific operations.

[0057] With reference to the FPGA unit 124, the implementation of the functional operations of the corresponding mapping and comparison unit 126 and the P-HMM unit 128 can be implemented in hardware by programming the programmable digital logic gates using a hardware description programming language such as a Very High Speed ​​Integrated Circuit (VHSIC) Hardware Description Language (VHDL) to dynamically configure or reconfigure the programmable logic gates. Alternatively, although the FPGA 124 can be used to implement the mapping and comparison unit 126 and the P-HMM unit 128, the present disclosure need not be so limited. For example, in other embodiments, one or more of the mapping and comparison unit 126 and the P-HMM unit 128 can also be implemented using software on one or more computers local to the DNA sequencer, or remote from the DNA sequencer. In yet other embodiments, the FPGA 124 can also be configured to perform the functions of the variant identification unit 130 to perform analysis of the modified probability results, thereby determining whether such probability results should be included in the VCF file 170 generated by the variant identification unit 130.

[0058] However, the present disclosure is not limited to using a dynamically reconfigurable FPGA 124 to implement one or more of the following: the mapping and alignment unit 126, the P-HMM module 128 or other computers of the auxiliary analysis unit 120, such as the variant identification unit 130 described herein. Alternatively, other types of programmable or non-programmable integrated circuits can be used. For example, one or more application-specific integrated circuits (Application-Specific Integrated Circuit, ASIC) can be programmed to perform one or more functions of the corresponding genome analysis module or other computers described herein. ASIC includes an integrated circuit, and the integrated circuit includes one or more programmable logic circuits similar to the FPGA described herein, and the similarity is that the digital logic gates of ASIC can be programmed using a hardware description language such as VHDL. However, the difference between ASIC and FPGA is that ASIC can be programmed only once, and once programmed, it cannot be dynamically reconfigured. In addition, aspects of the present disclosure are not limited to using FPGA or ASIC to implement the genome analysis module or other computers of the auxiliary analysis unit 120. Alternatively, one or more central processing units (CPUs), graphical processing units (GPUs), or any combination thereof may be used to implement any of the genomic analysis modules or other computers of the auxiliary analysis unit 120, thereby implementing the genomic analysis module or other computer of the auxiliary analysis unit 120 by executing software instructions.

[0059] In some embodiments, the mapping and alignment unit 126 may be implemented using an FPGA 124 configured to map the generated reads 112 stored in the first portion 122a of the memory 122 to a reference genome stored in another portion 122b of the memory 122 and align the generated reads 112 to the reference genome. However, the present disclosure is not limited to storing reads in the memory 122, accessing reads from the memory 122, storing a reference genome in the memory 122, or accessing a reference genome in the memory 122. Instead, in some embodiments, the generated reads 112, the reference genome, or both may be stored in a memory device in a cloud-based server accessible via one or more networks.

[0060] The mapped and aligned reads may be output by the mapping and alignment unit 126 to be stored in the memory 122 at the third portion 122c of the memory 122. In some embodiments, the write to the memory 122 containing the mapped and aligned reads from the FPGA 124 is referred to as being stored in the third portion 122c, but may actually be stored in the memory 122 in a manner that overwrites the initially generated reads 112 output by the nucleic acid sequencer 110 and stored in the first portion 122a. Therefore, although information at multiple stages is shown as being stored in the memory 122 at the first portion 122a, the second portion 122b, and the third portion 122c, respectively, the present disclosure does not require that all data described by the present disclosure as being stored in one of these respective portions of the memory 122 be present in the memory 122 at any particular time during the execution of the process disclosed by the present disclosure, but there may be a situation where all data described in this specification is stored in the memory 122 at the same time.

[0061] In some embodiments, memory 122 may include a single memory device or multiple memory devices. Using additional memory devices can reduce latency in accessing data and increase throughput by implementing writes and reads to fast memory devices such as flash memory, compared to requesting reads or writes to one or more disk storage devices using multiple levels of access to a memory hierarchy to create a fast memory effect.

[0062] Similarly, in some embodiments, each genome analysis module or other computer implementing the auxiliary analysis unit 120 using integrated circuits such as FPGA 124, ASIC, CPU, GPU, or a combination thereof may include a single FPGA 124, a single ASIC, a single CPU, a single GPU, or any combination thereof. Alternatively or in addition, each genome analysis module or other computer implementing the auxiliary analysis unit 120 using integrated circuits such as FPGA 124, ASIC, CPU, GPU, or a combination thereof may include multiple FPGAs 124, multiple ASICs, multiple CPUs, or multiple GPUs, or any combination thereof. Using additional integrated circuits such as multiple FPGAs to implement the genome analysis unit or other computer of the auxiliary analysis unit 120 can reduce the amount of time spent performing auxiliary analysis operations such as mapping, alignment, P-HMM probability calculation, and variant identification. In some embodiments, using FPGAs to implement these auxiliary analysis operations can reduce the time spent to complete these auxiliary analysis operations from 24 hours or more to 30 minutes or less. In some embodiments, using multiple FPGAs to perform these auxiliary analysis operations can enable these auxiliary analysis operations to be completed in as little as 5 minutes.

[0063] The output of mapping and comparison unit 126 comprises the accumulation of the reads mapped to the reference genome and compared with it. The accumulation can comprise a text-based format, which is used to summarize the base recognition of the compared reads from the DNA sample to the reference genome or a part of the reference genome. This output can be stored in one or more parts 122b, 122c of memory 122 in a computer-readable binary format, and the computer-readable binary format can be accessed and analyzed by one or more in FPGA 124, P-HMM unit 128 and variant identification unit 130. Alternatively, the output of mapping and comparison unit 126 can be stored in the memory of one or more remote computers using a computer-readable binary format, and the remote computer is a remote cloud-based server such as using one or more network accesses. The mapping of FPGA 124 and the output of comparison unit 126 can be described on the graphical user interface of the user device. The example of this graphical user interface is illustrated with reference to interface 140.

[0064] The interface 140 may be provided for display on a user interface of a user device that may access the memory 122 of the auxiliary analysis unit 120. For example, in some embodiments, the auxiliary analysis unit 120 has an attached display device. Alternatively or in addition, other devices having a display such as a smartphone or tablet may be connected to the same network as the auxiliary analysis unit 120, access the memory 122, and then display an accumulation, such as the accumulation 141 of the interface 140. In such embodiments, the variant identification unit 130 may perform the following operations: (i) access the output of the FPGA 124, which maps the obtained reads 112 to and aligns with the reference genome stored in the memory; and (ii) generate rendering data that, when rendered on the display device, causes the data output representing the reads that have been mapped to and aligned with the reference genome by the FPGA 124 and stored in the memory 122 to be displayed on the user device in a human-friendly manner, which data may be read by a person using the interface 140.

[0065] Interface 140 shows that the output of FPGA 124 includes a pile 141 of mapped and aligned reads. In this example, interface 140 shows fourteen reads with fourteen corresponding horizontal lines. These reads are grouped based on the DNA strand from which the reads were generated. For example, Figure 1In the example, the stack 141 of mapped and aligned reads includes a first group of reverse aligned reads extending from the 5′1 end of the read to the left toward the 3′1 end of the read in a first direction, and a second group of forward aligned reads extending from the 5′2 end of the read to the right toward the 3′2 end of the read in a second direction. Therefore, in this example of the interface 140 of fourteen mapped and aligned reads output from the FPGA 124, the seven reads at the bottom represent the first group of mapped and aligned reads, and the seven reads at the top represent the second group of mapped and aligned reads. The corresponding 5′ or 3′ ends of the two middle reads not shown in the interface 140 appear outside the window 140. Although this example used to illustrate the concepts of the present disclosure depicts a stack of only fourteen reads, the present disclosure is not limited thereto. Furthermore, although the stack 141 displays reverse aligned reads at the bottom of the stack and forward aligned reads at the top of the stack, other alternatives may exist. For example, as shown in reference Figure 7 to 1 0, forward aligned reads may be presented at the bottom of the stack 141, and reverse aligned reads may be presented at the top of the stack.

[0066] Although an example of an interface 140 that can display information describing characteristics of a sequence read is provided, no embodiment of the present disclosure is required to output information describing characteristics of a sequencer read on a display, or to require the variant caller unit 130 or other components described herein to access information from the interface 140. Instead, the interface 140 is provided merely as an example of the type of information that describes characteristics that make up a sequence read.

[0067] The present disclosure can use FPGA 124 to generate the accumulation of mapped and aligned reads, and the reads include tens of reads, hundreds of reads, thousands of reads or more depending on the needs of a specific genome sequence analysis workflow. By way of example, the high-throughput next-generation sequencing of nucleic acid samples can generate hundreds of thousands of short reads, which need to be mapped to one or more regions or parts thereof of a reference genome sequence and aligned with it. Such a large number of reads are mapped and aligned to generate a large number of overlapping or repeated short sequence nucleic acid reads. In some embodiments, for example, the number of overlapping or repeated short sequence reads can include 1x, 5x, 10x, 30x, 100x or greater coverage of one or more corresponding reference positions of a reference genome sequence. "30x coverage" refers to the mapping and alignment of the short reads of the accumulation of 30 or more overlapping reads, for example, comprising one or more reference positions of a reference genome sequence. By way of another example, "5x coverage" refers to the mapping and alignment of the short reads of the accumulation of 5 or more overlapping reads, for example, comprising one or more reference positions of a reference genome sequence.

[0068] Methods for mitigating related errors

[0069] Therefore, new read processing methods need to be designed to accurately and efficiently map and align large numbers of reads to reference genome sequences or portions thereof. For example, data generated by human genome sequencing can generate hundreds of millions of short reads, which often need to be mapped to and aligned to locations in a complete reference genome, which can then be further analyzed for potential variants to determine their biological, diagnostic, and / or therapeutic relevance.

[0070] The stacking of overlapping reads enables each read in different reads to be compared at a specific reference position of the reference genome sequence. The analysis of multiple overlapping reads at a specific reference position allows accurate determination of the following: whether there is a true variation, variant or deviation in the reads mapped to a specific position of the reference genome and aligned therewith, or whether there may be an error in any one of the reads at the discussed position in the stacking. For example, if only one or two of the 30 reads detect a specific nucleotide at position "X" of the reference genome sequence, and each of the 28 or 29 other reads supports the determination that another nucleotide is present at position "X", then two unrelated reads can be excluded as being wrong for position "X".

[0071] Relative to the method that does not assess the similarity or difference of overlapping reads, the analysis of the accumulation of overlapping reads can realize a more accurate analysis of reads, to determine how the genome of the subject is different from the reference genome, such as a model genome.For example, the analysis of the accumulation of overlapping reads can more accurately identify errors, such as chemical errors, machine errors or read errors, and distinguish such errors from true variants.More specifically, in the case where the subject has a true variant at the position "X" of the reference genome, by, for example, most of the reads include true variants, most of the reads in the accumulation should support the existence of true variants.Statistical models, such as those described herein, can then be implemented to determine the true gene sequence of the subject, wherein all true variants of the subject are from the reference genome.

[0072] Therefore, in various cases, once reads have been generated for nucleic acid samples, the sequential order of the reads compared and the generated reads have been mapped to a reference genome or its portion, the true gene sequence of the genome of the experimenter can be determined. Once the true sample genome is determined, one or more true variations can be determined based on the comparison of the true sample genome with a reference genome or its portion. Once one or more variations between the true sample genome and a reference genome or its portion are determined, the list of all true variants or deviations between the sample genome and the reference genome can be determined and called out. This variation may be due to various reasons, and may have biological, diagnostic and / or therapeutic relevance.

[0073] The exemplary stack depicted by interface 140 shows that the output of mapping and alignment unit 126 of FPGA 124 includes a base quality score 143 and a mapping confidence score 144 for each read in the stack. Base quality score 143 includes a value indicating the confidence level that the base identified for a read at a particular position of interest, such as position "0" 142, is accurate. Figure 1 In the example of , the base quality score is represented by a range of values ​​defined by: a high base quality score of "41", which indicates that the base call of the read at position "0" 142 is accurate with a high confidence level; and a low base quality score of "2", which indicates that the base call of the read at position "0" 142 is accurate with a low confidence level. In some embodiments, the base quality score is an output of the nuclear sequencer 110, which is received by the auxiliary analysis unit 120 and can be used to identify the base call error Q phred-base The phred-scale probability is determined by phred-base ,=-10*log10(P e-base ). In this example, P e-base is the probability of a base call error for a particular read. In some embodiments, a low base quality score can be a factor indicative of a sequencing error.

[0074] The mapping confidence score 144 indicates the confidence level that the obtained read 112 is accurately mapped to the reference genome 145 by the mapping and alignment unit 126 at a particular position of interest, such as position "0" (indicated by reference numeral 142). Figure 1 In the example of , the mapping confidence score is represented by a range of values ​​defined by: a high mapping confidence score of "250", which indicates a high confidence level that the read is accurately mapped to the reference genome 145 at position "0" 142; and a low mapping confidence score of "0", which indicates a low confidence level that the read is accurately mapped to the reference genome 145 at position "0" 142. In some embodiments, the mapping confidence score is an output of the mapping and alignment unit 126 and can be used to calculate the mapping error Q. phred-mapping The phred level probability is determined by Q phred-mapping ,=-10*log10(P e-mapping ). In this example, P e-mapping is the probability of a mapping error for a particular read. The value of the mapping confidence score 144 can be proportional to the difference between the best alignment score from an alignment algorithm such as a Smith-Waterman aligner and the second best score of the aligner. In some embodiments, this method can be adjusted to take into account the number of auxiliary alignments. In some embodiments, a low mapping score can be a factor indicating a mapping error.

[0075] Interface 140 also shows that the output of the FPGA includes the base nucleotide identified at position "0" 142. Figure 1 In the example of , the top twelve reads in the stack 141 are determined to have the same base recognition as the reference genome at position "0" 142. The example interface 140 depicts this determination by not depicting the letters A, C, G, or T representing the nucleotide of each of the top twelve reads at position "0" 142. Therefore, based on a review of the information depicted in the interface 140, it can be determined that the top twelve reads have a base recognition of "A" (adenine) at position "0" 142. The interface 140 also shows that the alt allele of G (guanine) is determined as the base recognition of the last two reads in the stack. G (guanine) is the alt allele because it is different from the nucleotide base of the reference genome at position "0".

[0076] The output of the FPGA also includes additional information that can be determined based on the analysis of the stacking 141 shown in the interface 140. First, information describing the chain direction of each read segment, also known as the read segment orientation of each read segment, can be determined. For example, the interface 140 shows that each alt allele appears in the same chain direction or the same read segment orientation. In this example, this information is demonstrated by the fact that the alt allele (e.g., "G") appears in the first group of read segments in the reverse alignment direction. By way of another example, the information describing the chain can also include the chain direction of each read segment in the stacking, and the chain direction can be determined with reference to the corresponding 3' and 5' ends. Second, information describing the position of the alt allele within each read segment can also be determined. For example, the proximity between the alt allele of each read segment at position "0" 142 and the 5' end of the read segment can be determined. In this example, the determined alt allele "G" appears away from the 5' end of the corresponding read segment of the alt allele because each alt allele is near the opposite 3' end. The further the alt allele appears from the 5′ end, the more likely the alt allele is to be associated with a sequencing error. Third, the base quality of each read at position “0” 142 can be determined. By way of example, the base qualities of the reads with the alt allele “G” are 6 and 2, respectively. Fourth, a mapping confidence score for each read can be determined for each read at the reference position “0”. By way of example, the reads with the alt allele “G” at position “0” 142 have mapping confidence scores of 45 and 3, respectively. The information shown in interface 140 for the purpose of this example is merely an example to illustrate the features of the present disclosure. However, reference Figure 6 9 show real-world examples of such interfaces.

[0077] Each type of information displayed by interface 140 may be referred to as a characteristic of one or more DNA reads. The characteristics include read-specific characteristics, such as base quality scores 143 or mapping confidence scores 144. Additional examples of read-specific characteristics are set forth with reference to interface 140 below.

[0078] The information displayed by the interface 140 is also stored in the memory 122. By way of example, the interface used to generate the interface 140 uses information describing characteristics of the DNA reads to generate the interface 140, such as information indicating a pile 141 of mapped and aligned reads, information indicating a reference genome 145 (or a portion thereof), information indicating a base quality score for each read 143, information indicating a mapping confidence score 144 for each read, information indicating a base call for each read at a reference position (e.g., position "0" 142), information indicating whether each read contains an alt allele, information indicating whether each read contains a nucleotide sequence ... Information indicating the identity of reads having alt alleles, information indicating the position of each alt allele with reference to the 5′ end of the reads containing the alt allele, information indicating the direction (or orientation) of each read (e.g., via forward alignment or reverse alignment), information indicating the direction (or orientation) of each read having the alt allele (e.g., via forward alignment or reverse alignment), and information indicating whether the alt allele is determined to occur in a first set of reads in the forward alignment direction or in a second set of reads in the reverse alignment direction. Information describing, indicating, or otherwise representing each of these characteristics can be stored in the memory 122, for example, at positions 122b, 122c. By way of example, these characteristics can be stored in the memory 122 in a machine-readable binary form.

[0079] The variant identification unit 130 can obtain information describing the characteristics of the mapped and aligned reads from the memory 122 and displayed in the interface 140. For some inputs, the variant identification unit 130 can use the information describing the characteristics of the mapped and aligned reads from the memory 122 to generate inputs to one or more probability models 131 of the variant identification unit 130. The variant identification unit 130 can provide the generated input as input to the one or more probability models 131. In some embodiments, the variant identification unit 130 can also request the P-HMM unit 128 to generate one or more probabilities to input to the one or more probability models 131 used by the variant identification unit 130. By way of example, the variant identification unit 130 can request the P-HMM unit 128 to determine, for each read, a probability that a particular candidate allele G is considered at a reference position, such as reference position "0" 142. m,φ, The observed read segment r iIn such embodiments, for example, the variant identification unit 130 can use the probability value returned by the P-HMM unit 128 as the read allele score. The variant identification unit 130 can then provide information describing the characteristics of the mapped and aligned reads, including (i) information describing the characteristics obtained from the memory 112 and / or remote memory, and (ii) information probability of one or more characteristics of the pile described by the interface 140 calculated by the P-HMM unit 128, or any combination thereof. In some embodiments, the variant identification unit 130 can calculate or receive a calculation result from another computer of the auxiliary analysis unit 120, which represents an alternative form of the read allele score. These different forms of the read allele score will be described in more detail below.

[0080] It should be noted that the memory 122 stores information that describes, supports, or describes and supports one or more characteristics of the reads described by the interface 140. In some cases, it may be necessary to derive the type of information described with reference to the interface 140 from the information actually stored in the memory 122. For example, in some embodiments, the memory 122 may store the position of a candidate alt allele, such as "G", in the interface 140, and store the position of the 5' end of the read containing the candidate alt allele, such as "G", in the interface 140. Then, based on the information stored in the memory 122, the variant identification unit 130, or other components of the auxiliary analysis unit 120, the distance of the candidate alt allele "G" from the 5' end of the read can be determined. In this case, it is not necessary for the memory 122 to store the distance of the candidate alt allele "G" from the 5' end, because such information can be derived from the information stored in the memory 122. Other types of information describing characteristics of sequence reads can be derived in a similar manner from the actually stored information describing the characteristics.

[0081] The probability models 131 described herein are used by the variant identification unit 130 described herein to determine the probability scores of various candidate genotypes as true. The improvements presented in the present disclosure improve the accuracy of the variant identifier in a way that conventional probability models cannot achieve, and the method enables the present probability model to consider the occurrence of related error events such as two or more mapping errors and / or two or more sequencing errors. These new probability models 131 include mapping error probability models 132 and sequencing error probability models 134. These probability models provide technical advantages because they are not limited by rule-based decision making or the characteristics of a predetermined training data set.

[0082] Mapping Error Probability Model

[0083] The mapping error probability model 132 has been designed to take into account mapping errors, such as those errors that occur when multiple regions of a reference genome, such as a first region and a second region, contain similar and in some cases almost identical base sequences. In this case, due to the similarity of the base sequences in each corresponding region, the mapping and alignment unit may incorrectly map a set of reads to the first region instead of the second region. The possibility of mapping errors may increase when there are naturally occurring variants at one or more positions in the first region, which may make the sequence identical to the second region. In order to take into account such errors, the mapping error model receives a set of information obtained by the variant identification unit 130 from the memory 122 as input to the probability model from the P-HMM unit 128 or a combination thereof, the information describing the characteristics of the mapped and aligned reads.

[0084] To determine the probability that a mapping error has occurred, mapping error probability model 132 receives input obtained by variant identification unit 130, the input comprising: (i) for each candidate allele at a reference position, such as reference position "0" 142, the read allele score of each read stored in memory 122; and (ii) the mapping quality score of each read stored in memory 122. In some embodiments, the read allele score may include a value indicating that the sequencing process will consider a sequence containing allele G. m,φ DNA molecules to generate read segments r i The probability value P(r i |G m,φ There are various ways to calculate or estimate this value P(r i |G m,φ ).

[0085] When using, for example, the Genome Analysis Tool Kit (GATK) or In embodiments of haplotype-based identifiers such as the . k , where a haplotype represents a base sequence extending beyond a reference position, such as reference position "0" 142, in one or both directions. A hidden Markov model (HMM), such as the P-HMM unit 128, can then be used to calculate the read haplotype score P (r i |H k ), which means that the sequencing process will take into account the presence of haplotype H k DNA molecules to generate read segments r i probability.

[0086] The HMM calculation can take into account possible uncertainty in the alignment, summing the probabilities of multiple possible alignments rather than assuming that the alignment returned by the mapper / aligner is correct. Next, in some embodiments, the read allele score can then be assigned a score that is higher than the best score for the haplotype containing the allele:

[0087]

[0088] A detailed description of using variant identification unit 130 to calculate such probabilities using HMM units is described in more detail in U.S. Publication No. 2016 / 0306922, the entirety of which is incorporated herein by reference.

[0089] Variant callers other than haplotype-based callers may use simpler estimates in order to reduce computational complexity. For example, a variant caller may assume that the alignment from the mapper / aligner is correct and estimate the score accordingly. To detect SNPs, such a variant caller may call bases based on alignment at a reference position such as reference position "0" 142. i and base quality q i The read allele fraction for each read of the pileup stored in memory 122 is estimated as follows:

[0090]

[0091] For indels, this variant caller can assign scores based on the length of the indel, whether the indel is an insertion or a deletion, and the surrounding sequence context (e.g., the period and length of short tandem repeats).

[0092] Regardless of whether a column-by-column embodiment related to SNP detection or a more general embodiment related to SNP and indel detection is used, the output of the mapping error probability model 132 includes one or more probabilities, including scores indicating the likelihood of one or more mapping errors occurring at a reference position, such as position "0" 142. In some embodiments, the mapping error probability model 132 is configured to output two probability scores for two different hypotheses, including (i) the likelihood that a read at the reference position 142 indicates the occurrence of a homozygous ref with an alien allele that matches alt, and (ii) the likelihood that a read at the reference position 142 indicates the occurrence of a homozygous alt with an alien allele that matches the reference allele (165).

[0093] Sequencing error probability model

[0094] The sequencing error probability model 134 is a probability model that takes into account sequencing errors, which may occur because certain combinations of nucleotides may confuse the sequencing algorithm to generate an incorrect sequence. As with the mapping error model 132 described above, in some embodiments, the input provided to the sequencing error probability model 134 may vary based on the complexity of the variant recognition unit used. More complex haplotype variant identifiers may use the read allele fractions calculated using equation (1) above, while other variant identifiers other than haplotype variant identifiers may use the read alleles calculated using equation (2). In some embodiments, the type of read allele fraction used can be determined based on whether the variant recognition unit 130 will be used to detect only SNPs or SNPs and indels.

[0095] Regardless of the type of read allele fractions used, the sequencing error probability model 134 receives as input from the variant identification unit 130 a set of information describing the characteristics of the mapped and aligned reads retrieved by the variant identification unit 130. To determine the probability that a sequencing error has occurred, the sequencing error probability model 134 receives input generated or otherwise obtained by the variant identification unit 130, the input comprising: (i) the read orientation of each read stored in the memory 122, which is stacked; (ii) the position of the 5′ end of each base at a reference position within each read, such as position “0” 142, the reference read; (iii) the read allele fraction of each read stored in the memory 122 for each candidate allele at a reference position, such as reference position “0” 142; and (iv) the base quality score of each read of the base at the reference position, such as position “0”.

[0096] Other variations of the sequencing error probability model can be configured to receive other inputs. For example, in some embodiments, the base quality score of each read of the base at the reference position of, for example, position "0" 142 is not required as an input. An example of the case of the base quality score of each read of the base at the reference position "0" 142 is the case of determining the read allele score using equation (2). In this case, the base quality score of each read of the base at the reference position of, for example, reference position "0" 142 can be derived from the read allele score determined using equation (2). However, there may be other embodiments in which the base quality score of each read of the base at the reference position of, for example, position "0" 142 is not required as a dedicated input because the base quality score can be derived from another received input instead.

[0097] In yet other embodiments, another fourth input may be provided to the sequencing error probability model. For example, the sequencing error probability model may be configured to receive an input describing a plurality of homopolymers of the reference genome 145, the reference genome 145 comprising, for example, Figure 1 Before the reference position, for example, position "0" 142, in the same direction (or read orientation) of the read of the candidate alt allele "G" at position "0" of the interface 140 in the reference genome 145. It should be noted that, for example, three reference alleles "G" appear in the reference genome 145 before the reference position 142, and the reference alleles are the same as the candidate alt allele "G". This number of homopolymers can be input as another input into the sequencing error probability model. This input describing the number of homopolymers can be added to improve the accuracy of the model.

[0098] The output of the sequencing error probability model 134 includes one or more probabilities, including a score indicating the likelihood of one or more sequencing errors occurring at a reference position, such as position "0" 142. In some embodiments, the sequencing error probability model 134 is configured to output two probability scores for two different hypotheses, including (i) the likelihood that a read at the reference position 142 indicates the occurrence of a homozygous reference with a sequencing error that matches an alt allele (166), and (ii) the likelihood that a read at the reference position 142 indicates the occurrence of a homozygous alt with a sequencing error that matches a reference allele (167).

[0099] The variant identification unit 130 generates a set of modified probability results 150 for each of the plurality of hypotheses using the conventional probability model and one or more. The plurality of hypotheses may be determined based on the inputs provided to the one or more probability models 131, the specific one or more probability models 131 used, or a combination thereof. The generated set of modified probability results 150 may include a probability value for each of the plurality of hypotheses, the probability value indicating the likelihood that each respective hypothesis is true.

[0100] By way of example, in some embodiments, if the variant identification unit 130 only generates or otherwise provides input to the mapping error model 132, then a set of modified probability results 150 will include conventional probabilities for hypotheses 161, 162, 163 and unconventional probabilities for hypotheses 164, 165, each of which is described in more detail below, and the probability of one or more mapping errors occurring at the reference position 142 is considered. By way of another example, if the variant identification unit 130 only generates or otherwise provides input to the sequencing error model 134, then a set of modified probability results 150 will include conventional probabilities for hypotheses 161, 162, 163 and unconventional probabilities for hypotheses 166, 167, each of which is described in more detail below, and the probability of one or more sequencing errors occurring at the reference position 142 is considered. However, if the variant identification unit 130 generates or otherwise provides input to both the mapping error model 132 and the sequencing error model 134, then the set of modified probability results 150 generated by the variant identification unit 130 will include conventional probabilities 161, 162, 163 and unconventional probabilities 164, 165, 166, 167, each of which is described in more detail below. The set of modified probability results 150 can be generated and provided in a computer-readable binary format. The set of modified probability results 150 improves the results of the conventional probability calculations typically performed by the variant identification unit 129 because the modified probability results 150 can also include additional probability scores for one or more additional hypotheses 164, 165, 166, 167, which the variant identification unit 130 can use to use the modified probability calculations to account for the occurrence of one or more mapping errors, one or more sequencing errors, or a combination of the two.

[0101] A human-readable version of the modified set of probability results 150 may be shown using a graphical user interface on a display of a user device that has access to the set of probability results 150 generated by the variant identification unit 130. Figure 1 An example of such a graphical user interface is shown in the example of FIG. In some embodiments, the display may include, for example, a display device coupled to the auxiliary analysis unit 120. Each probability in the set of modified probabilities 150 is discussed below with reference to the display of probabilities in the interface 160. However, these probabilities may also be obtained and analyzed by the variant identification unit 130 in a machine-readable format.

[0102] The conventional probability calculated by variant identifier 130 can include a set of probabilities determined using conventional probability models. These conventional probability models are configured to determine probability scores for each of the three hypotheses, the hypothesis comprising (i) the read at reference position 142 indicates the possibility 161 of the occurrence of a homozygous reference, (ii) the read at reference position 142 indicates the possibility 162 of the occurrence of a heterozygous alt, and (iii) the read at reference position 142 indicates the possibility of the occurrence of a homozygous alt. When the two alleles at reference position 142 are identical, a homozygous reference will occur. In such examples, alt alleles do not occur at reference position 142. When one allele at reference position 142 is an alt allele and another allele at reference position 142 is a reference allele, heterozygous alt will occur. When the two alleles at reference position 142 are alt alleles, homozygous alt will occur. The conventional probability calculations that generate these three hypotheses assume that all reads in the stack are correctly mapped and that sequencing errors throughout the reads are unrelated.

[0103] However, mapping errors often occur, and mapping errors and sequencing errors throughout the reads are often highly correlated. In order to take into account the occurrence of these errors, the present invention uses a variant identification unit 130 configured to generate a set of modified probability results 150 using one or more modified probability models 131, and a set of modified probability results 150 includes probability scores for four or more additional unconventional hypotheses. In some embodiments, for example, for a single alt allele (e.g., a reference allele and a first alt allele), a set of modified probability results 150 may include probability scores for four additional unconventional hypotheses 164, 165, 166, 167 described herein. However, in other embodiments where there are more than a single alt allele (e.g., a reference allele, a first alt allele, and a second alt allele), a set of modified probability results 150 may include more than four unconventional hypotheses described herein. In such a case, each respective combination of the first alt allele, the second alt allele, and the reference allele will have a probability score generated for a set of unconventional hypotheses corresponding to the four unconventional hypotheses 164, 165, 166, 167 described herein. Compared to conventional variant identifiers that do not themselves utilize probability scores output by one or more probability models 131, a set of modified probability results 150 improves the variant identification unit 130 by allowing the variant identification unit 130 to consider the likelihood of one or more correlated error events occurring at a reference position, such as a reference position 142 of the stack 141.

[0104] A set of modified probability results 150 includes corresponding unconventional probability scores for different hypotheses that take into account the potential occurrence of one or more mapping errors. These additional probabilities include the probability (164) that the read at the reference position 142 indicates the occurrence of a homozygous ref with an alien allele that matches alt, and (ii) the probability (165) that the read at the reference position 142 indicates the occurrence of a homozygous alt with an alien allele that matches the reference allele. The alien allele may include an allele that is mapped to a base nucleotide of the reference genome 145 at the reference position 142, and the allele is the result of a mapping error. The alien allele may be incorrectly mapped to a first region of the reference genome, which has a nucleotide base sequence that is substantially similar or identical to one or more second regions of the reference genome 145.

[0105] A set of modified probability results 150 includes corresponding unconventional probabilities that take into account the potential occurrence of one or more sequencing errors. These additional probabilities include (i) the probability that the read at reference position 142 indicates the occurrence of a homozygous reference with a sequencing error that matches the alt allele (166), and (ii) the probability that the read at reference position 142 indicates the occurrence of a homozygous alt with a sequencing error that matches the reference allele (167).

[0106] A set of modified probability results 150 collectively represent the probability of the presence of a true variant of the base nucleotide at the reference position 142. In addition, each corresponding probability in the set of modified probability results 150, including probabilities 161, 162, 163, 164, 165, 166, 167, provides a specific probability score for the presence of a specific genotype represented by each hypothesis. Here, the specific genotype can include homozygous reference, heterozygous alt, homozygous alt, homozygous ref with an alien allele matching alt, homozygous alt with an alien allele matching a reference allele, homozygous reference with a sequencing error matching alt allele, or homozygous alt with a sequencing error matching a reference allele.

[0107] The variant identification unit 130 can use a set of modified probability results 150 to determine whether a candidate alt allele of one or more sample reads is a true variant of a reference allele at a reference position of interest. Figure 1In an example, the variant identification unit 130 can use a set of modified probability results 150 to determine whether the candidate alt allele "G" shown in the interface 140 is a true variant of the reference allele "A" at the reference position. For example, the variant identification unit 130 can process the set of modified probability results 150 and determine 136 whether data identifying a true variant at the reference position should be included in a variant call format (VCF) file 170 generated by the variant identification unit 130.

[0108] The variant identification unit 130 can determine whether a candidate alt allele of one or more reads of a sample, such as data describing a read having a candidate alt allele "G" shown in the interface 140, should represent a true variant by determining a total score based on the modified probability result 150 and evaluating the total score using one or more predetermined thresholds. The total score can indicate the likelihood that a true variant exists. In some embodiments, the total score can be determined using equation (14). However, other methods of determining the total score are within the scope of the present disclosure. If the computer determines 136 that the total score meets the predetermined threshold, the computer can add information indicating the presence of a true variant at the first position to the VCF file. The information added to the VCF file can include, for example, the position of the candidate alt allele, an identifier of the alt allele, a genotype of the alt allele, and data indicating the total score. Alternatively, if the computer determines 136 that the total score does not meet the predetermined threshold, the computer can discard the output data and determine that the output data indicates a false positive at the reference position.

[0109] exist Figure 1 In the example shown, the variant identification unit 130 can determine 136 that the total score based on the probabilities in the set of modified probabilities 150 fails to meet the predetermined threshold and does not contain information identifying the candidate alt allele "G" in the VCF file 170 as a true variant. This is because the total score based on the set of modified probabilities 140 will indicate a high probability that one or more sequencing errors exist, because the candidate alt allele "G" only appears in a single strand in the forward alignment position and has low base quality in a position away from the 5' end of the corresponding read of the candidate alt allele "G". Therefore, the variant identification unit 130 can determine that the candidate alt allele "G" at the reference position 142 is not a true variant, but a false positive, and that there is no true variant at the reference position 142 based on the evaluation of the set of modified probabilities 150.

[0110] The probabilities shown in interface 160 are merely examples and are provided for the purpose of illustrating examples of the present disclosure. The probabilities shown in interface 160 are not intended to be Figure 1The result of putting the actual information depicted in the actual probability model described in this specification.

[0111] Figure 2 is a flow chart of an example of a process 200 for variant identification related error event mitigation. The process 200 is described below as being performed by, for example Figure 1 In some embodiments, the variant identification unit can be implemented using hardware such as one or more FPGAs, ASICs, CPUs, GPUs, or any combination thereof, using software, or any combination thereof.

[0112] The computer may start execution process 200 by accessing a stack of aligned sequence reads stored in one or more memory devices (step 210). Mapping and alignment units implemented with configurable logic gates of FGPA devices may have been used to generate aligned sequence reads. In some embodiments, the accessed reads are stored in one or more memory devices and may include a first set of sequence reads aligned in the forward direction and a second set of sequence reads aligned in the reverse direction. Each set of corresponding sequence reads corresponds to a specific read orientation or direction.

[0113] The computer may continue to perform process 200 by obtaining information describing one or more characteristics of the corresponding read at the first position of the sequence read pile (step 220). The one or more characteristics may include attributes of the read at the first position in the pile that can be used to consider the probability of occurrence of one or more related error events.

[0114] The computer may continue to perform process 200 by providing one or more inputs describing one or more characteristics of the reads stacked at the first position to the probability model (step 230). The one or more characteristics associated with the reads stacked in the first position may include (i) information describing one or more characteristics obtained from one or more memory devices, (ii) information generated by one or more models, such as P-HMM models, based on one or more models processing information describing one or more characteristics obtained from one or more memory devices, or a combination thereof. In some embodiments, the probability model is configured to determine, for each of the one or more hypotheses selected based on the one or more inputs, a score indicating that the hypothesis is true.

[0115] The computer may continue to perform process 200 by obtaining output information for each of the one or more hypotheses based on the one or more inputs. The output information for each hypothesis may be (i) generated by the probability model based on one or more inputs processed by the probability model to the probability model describing one or more characteristics of the corresponding read segments of the stack, and (ii) indicating the probability that the hypothesis is true (step 240). In some embodiments, the computer may obtain such output information for each of the one or more hypotheses or a subset of the one or more hypotheses. It may be determined whether a particular hypothesis will be included in the output information based on the input provided to the probability model.

[0116] In some embodiments, based on one or more inputs to the probability model, as described above, one or more hypotheses can include (i) the likelihood that the read at the reference position indicates the occurrence of a homozygous reference, (ii) the likelihood that the read at the reference position indicates the occurrence of a heterozygous alt, (iii) the likelihood that the read at the reference position indicates the occurrence of a homozygous alt, (iv) the likelihood that the read at the reference position indicates the occurrence of a homozygous ref with an alien allele matching alt, (v) the likelihood that the read at the reference position indicates the occurrence of a homozygous alt with an alien allele matching the reference allele, (vi) the likelihood that the read at the reference position indicates the occurrence of a homozygous reference with a sequencing error matching the alt allele, (vii) the likelihood that the read at the reference position indicates the occurrence of a homozygous alt with a sequencing error matching the reference allele, or any combination thereof.

[0117] The computer may continue to perform process 200 by determining the likelihood of the presence of a true variant at the first position based on the obtained output data generated by the probability model for each of the plurality of hypotheses (step 250). In some embodiments, this determination may be made, for example, based on evaluating the probability scores determined for each of the one or more hypotheses individually or collectively against one or more predetermined thresholds.

[0118] For example, the computer can determine a total score based on the output data generated by the probability model for each of the multiple hypotheses. The total score can indicate the possibility of the presence of a true variant. In some embodiments, the total score can be determined using equation (14). However, other methods for determining the total score are within the scope of the present disclosure. If the computer determines that the total score meets a predetermined threshold, the computer can add information indicating the presence of a true variant at the first position to the VCF file. Alternatively, if the computer determines that the total score does not meet the predetermined threshold, the computer can discard the output data and determine that the output data indicates a false positive at the reference position.

[0119] Figure 3is a flow chart of an example of a process 300 for mapping error mitigation for variant identification. The process 300 is described below as being performed by, for example Figure 1 In some embodiments, the variant identification unit can be implemented using hardware such as one or more FPGAs, ASICs, CPUs, GPUs, or any combination thereof, using software, or any combination thereof.

[0120] The computer can start execution process 300 by accessing the accumulation of aligned sequence reads stored in one or more memory devices (step 310). The mapped and aligned sequence reads may have been generated using a mapping and alignment unit implemented with a configurable logic gate of an FGPA device. In some embodiments, the accessed reads are stored in one or more memory devices and may include a first group of sequence reads aligned in the forward direction and a second group of sequence reads aligned in the reverse direction. Each group of corresponding sequence reads corresponds to a specific read orientation or direction.

[0121] The computer can continue executing process 300 by obtaining information describing: (i) a mapping confidence score for each read in a stack of aligned sequence reads stored in one or more memories, and (ii) a read allele score for each candidate allele at a reference position in a stack of aligned sequence reads stored in one or more memories (320).

[0122] In some embodiments, the mapping confidence score may comprise the output of the mapping and alignment unit 126 and may be calculated using the mapping error Q phred-mapping The phred level probability is determined by Q phred-mapping ,=-10*log10(P e-mapping ). In this example, P e-mapping is the probability that a particular read is mapped incorrectly. The value of the mapping confidence score 144 may be proportional to the difference between the best alignment score from an alignment algorithm, such as a Smith-Waterman aligner, and the second best score from the aligner.

[0123] The read allele fraction can be determined in many different ways. For example, a more complex haplotype variant identifier can use the read allele fraction calculated using the above equation (1). Alternatively, other variant identifiers other than the haplotype variant identifier can use the read allele fraction calculated using the above equation (2). In some embodiments, the type of read allele fraction used can be determined based on whether the variant identification unit 130 will be used to detect only SNPs or SNPs and insertions and deletions. For example, in some embodiments where the variant identification unit will be used to detect SNPs and insertions and deletions, the read allele fraction calculated using the above equation (1) can be used. By way of another example, in other embodiments where the variant identification unit will be used to detect only SNPs, the read allele fraction calculated using equation (2) can be used.

[0124] The computer can continue to perform process 300 by providing one or more inputs describing the obtained information to the probability model (step 330), wherein the obtained information describes (i) the mapping confidence score of each read in the pile of aligned sequence reads stored in one or more memories, and (ii) the read allele score of each read in the pile of aligned sequence reads stored in one or more memories for each candidate allele at a reference position.

[0125] The computer can continue to perform process 300 by obtaining output information (340) for each of the one or more hypotheses based on the one or more inputs obtained at 320. In an example of process 300, the inputs obtained include inputs to a mapping error probability model, including (i) a mapping confidence score for each read, and (ii) a read allele score for each read of each candidate allele at a reference position. Thus, based on receiving these inputs to the mapping error probability model, the computer will generate output information for one or more hypotheses that take into account the occurrence of one or more mapping errors. Such hypotheses include (i) the probability that the read at the reference position indicates the occurrence of a homozygous ref with an alien allele that matches alt, and (ii) the probability that the read at the reference position indicates the occurrence of a homozygous alt with an alien allele that matches the reference allele.

[0126] For each of these hypotheses, the output information includes (i) information generated by the mapping error probability model based on the probability model processing the input to the probability model, the input describing (i) the mapping confidence score of each read in the pile of aligned sequence reads stored in one or more memories, and (ii) the read allele score of each read in the pile of aligned sequence reads stored in one or more memories for each candidate allele at the reference position. In addition, the output information obtained includes a score in each specific hypothesis that takes into account the occurrence of one or more mapping errors, the score indicating the likelihood that the hypothesis is true.

[0127] The computer can continue to perform process 300 by determining the likelihood that a true variant exists at the first position based on the obtained output data generated by the probability model for each of the plurality of hypotheses (step 350). In some embodiments, this determination can be made, for example, based on evaluating the probability scores determined for each of the one or more hypotheses individually or collectively against one or more predetermined thresholds.

[0128] For example, the computer can determine a total score based on the output data generated by the probability model for each of the multiple hypotheses. The total score can indicate the possibility of the presence of a true variant. In some embodiments, the total score can be determined using equation (14). However, other methods for determining the total score are within the scope of the present disclosure. If the computer determines that the total score meets a predetermined threshold, the computer can add information indicating the presence of a true variant at the first position to the VCF file. Alternatively, if the computer determines that the total score does not meet the predetermined threshold, the computer can discard the output data and determine that the output data indicates a false positive at the reference position.

[0129] The above process 300 describes a method for using a probabilistic model that can be used to account for the likelihood of mapping errors. Examples of probabilistic models that can be used to account for the likelihood of one or more mapping errors are described in more detail below.

[0130] In one embodiment, the probability model is modified to incorporate the probability that a pile of sequence reads stored in a memory device contains multiple mismapped reads, resulting in one or more mapping errors. The probability model is applicable to the following situations: (1) each read r i The mapping quality μ is accompanied by the phred-level confidence that it is correctly mapped. i Therefore, R = {r i ,μ i :i=1…N R}; (2) auxiliary alignments (i.e., other loci in the genome to which the reads align well) may be unknown and / or too numerous to tabulate; and (3) as input, for each read r i and candidate allele G m,φ Given the base mass P(r i |G m,φ ).

[0131] In one embodiment, the present disclosure modifies a conventional variant calling probability model that supplements a list of candidate genotypes with an expanded candidate genotype that represents the hypothesis that a pile of sequence reads stored in a memory device contains a mixture of local alleles and foreign alleles.

[0132] For a diploid genome with one foreign allele, the method of the present invention defines an expanded candidate genotype G′ m =[G m,1 G m,2 F m ], where F m It is a foreign allele. The local allele G m,1 and G m,2 Each is assumed to have an allele frequency of (1-β) / 2, and the foreign allele F m with allele frequency β, where β is unknown.

[0133] For each expanded candidate genotype G m , model calculation:

[0134]

[0135] where U(t) is the Heavyside unit step function,

[0136]

[0137] And P0(F m ) is the genotype present at the locus of interest [ρF m ], where ρ is the reference allele at the locus of interest. The value Ψ m is the joint probability P(G′ m ,R) estimation.

[0138] The probabilistic model used to determine the likelihood of the occurrence of one or more mapping errors is based on the mapping quality μ of the mismapped reads. iand the prior probability of a variant (or cluster of variants) causing read i to be mismapped. In a probabilistic model for determining the likelihood of one or more mapping errors, the quantity p F Indicates that sufficient number of variants occur at another position to cause mismapping to the read. Mapping quality μ = -10log 10 (p F / P0(F m )) is the prior probability. i The above term #1 is non-zero only when it does not exceed this threshold, and the threshold increases with p F How many variants may be present at a remote location is usually unknown, and scanning p F To find m The value of maximizing , thereby testing each hypothesis with respect to the number of remote variants, which may cause foreign reads to end up there.

[0139] In some embodiments, the complexity of the probability model used to determine the likelihood of the occurrence of one or more mapping errors is optimized. m may seem to have high computational complexity, because, as shown in (Equation #3), β and p are scanned independently F Both, and evaluate the result over a continuous range of values. Depending on the resolution, this can be computationally very complex. However, only i A set of discrete values ​​of the scan p F ,in:

[0140] P(r i |F m )>(P(r i |G m,1 )+P(r i |G m,2 )) / 2, equation (5)

[0141] It is typically a small value. In practice, there is no need to scan term β at all; instead the optimal value can be estimated as the fraction of reads where term #1 exceeds term #2, and the effect of this estimate on the result is typically negligible. Finally, the computational cost of this operation may typically be insignificant relative to other parts of the system (e.g., Hidden Markov Model (HMM) computations).

[0142] In some embodiments, it may be necessary to adjust the mapping confidence score. It should be noted that in some embodiments, the mapping confidence score reported by the mapping and alignment unit may represent an estimate of the phred-level confidence that the read is correctly mapped, but in reality, this estimate may not be accurate. Thus, depending on the mapper, it may be beneficial to adjust the mapping confidence score to better match the true probability of a mismap. In some embodiments, the mapping confidence score may be used. Figure 4 The function 400 shown converts an output value of a first mapping confidence score of a mapping and alignment unit, which represents a MAPQ score, for example, into a mapping confidence score μ i .

[0143] In some embodiments, the term β may be limited to the range [0, 0.5]. A value of β = 0.5 corresponds to a situation where all reads of the alternative reference position are mapped to the reference position of interest. Although it is conceivable that there may be more than one source of foreign reads, which suggests a possible higher value of β, limiting β to 0.5 can improve overall accuracy, thereby recovering some true positives that would otherwise be suppressed.

[0144] In some embodiments, the number of candidates can be reduced to only G m,1 =G m,2 =G m In this case, equation (3) is simplified to:

[0145]

[0146] In some embodiments, the above expression may assume that only inputs that include reads that overlap with genotyping events, i.e., those that favor one allele over another. However, in some embodiments, there may be ambiguity as to which reads overlap with events. In such cases, the following expression avoids the complications associated with including non-overlapping reads in the calculation:

[0147]

[0148] In some embodiments, there is another variation of the probability model used to determine the likelihood of the occurrence of one or more mapping errors. In such embodiments, the probability model can take into account data describing the knowledge of the strand direction of each read. Typically, mismapped reads are often limited to single-stranded directions, and this can be useful information to support the hypothesis that these reads are foreign. To account for this potential error, the indicator in the probability model used to determine the likelihood of the occurrence of one or more mapping errors is θ iindicates the strand direction of read i, where values ​​0 and 1 indicate the forward and reverse strand directions, respectively. In the case of a strand-aware modified probabilistic mapping model for determining the likelihood of the presence of one or more mapping errors, three hypotheses can be evaluated. The three hypotheses include the presence of foreign reads only on the forward strand, only on the reverse strand, or on both strands. The scheme includes using Ψ m The assumption of maximization. Starting with equation (5), this changes as follows:

[0149]

[0150] where Λ0={i:θ i =0}, Λ1={i:θ i =1}, Λ2={i:θ i =0 or 1}.

[0151] Figure 5 5 is a flow chart of an example of a process for sequencing error mitigation for variant identification. Process 500 is described below as being performed by, for example Figure 1 In some embodiments, the variant identification unit can be implemented using hardware such as one or more FPGAs, ASICs, CPUs, GPUs, or any combination thereof, using software, or any combination thereof.

[0152] The computer can start execution process 500 by accessing the accumulation (step 510) of aligned sequence reads stored in one or more memory devices. The mapping and alignment unit implemented with the configurable logic gate of the FGPA device may have been used to generate the aligned sequence reads. In some embodiments, the accessed reads are stored in one or more memory devices and may include a first group of sequence reads aligned in the forward direction and a second group of sequence reads aligned in the reverse direction. Each group of corresponding sequence reads corresponds to a specific read orientation or direction.

[0153] The computer can continue to execute process 500 by obtaining information describing: (i) the read orientation of each read in the stack of aligned sequence reads stored in one or more memories; (ii) for each read in the stack of aligned sequence reads stored in one or more memories, the position of the 5′ end of the reference read for each base at a reference position; (iii) the read allele score of each read in the stack of aligned sequence reads stored in one or more memories (520) for each candidate allele at the reference position; and (iv) the base quality score of each read of the base at the reference position, such as position “0”.

[0154] The read allele fraction can be determined in many different ways. For example, a more complex haplotype variant identifier can use the read allele fraction calculated using the above equation (1). Alternatively, other variant identifiers other than the haplotype variant identifier can use the read allele fraction calculated using the above equation (2). In some embodiments, the type of read allele fraction used can be determined based on whether the variant identification unit 130 will be used to detect only SNPs or SNPs and insertions and deletions. For example, in some embodiments where the variant identification unit will be used to detect SNPs and insertions and deletions, the read allele fraction calculated using the above equation (1) can be used. By way of another example, in other embodiments where the variant identification unit will be used to detect only SNPs, the read allele fraction calculated using equation (2) can be used.

[0155] The computer may continue to perform process 500 by providing one or more inputs describing the obtained information to the probabilistic model, wherein the obtained information describes: (i) the read orientation of each read in the stack of aligned sequence reads stored in one or more memories; (ii) for each read in the stack of aligned sequence reads stored in one or more memories, the position of the 5′ end of the reference read for each base at a reference position; (iii) the read allele score of each read in the stack of aligned sequence reads stored in one or more memories (step 530) for each candidate allele at the reference position; and (iv) the base quality score of each read of the base at the reference position, for example, position “0”.

[0156] The computer may continue to perform process 500 by obtaining output information (540) for each of the one or more hypotheses based on the one or more inputs obtained at 520. In an example of process 500, the obtained inputs include inputs to a sequencing error probability model, the inputs including: (i) a read orientation of each read in a pile of aligned sequence reads stored in one or more memories; (ii) for each read in a pile of aligned sequence reads stored in one or more memories, the position of the 5′ end of the reference read for each base at a reference position; (iii) a read allele score for each read in a pile of aligned sequence reads stored in one or more memories for each candidate allele at a reference position; and (iv) a base quality score for each read of a base at a reference position, such as position “0”. Thus, based on receiving these inputs to the sequencing error probability model, the computer will generate output information for one or more hypotheses that take into account the occurrence of one or more sequencing errors. This hypothesis includes the probability that (i) the read at the reference position indicates the occurrence of a homozygous reference with a sequencing error that matches the alt allele, and (ii) the probability that the read at the reference position indicates the occurrence of a homozygous alt with a sequencing error that matches the reference allele.

[0157] For each of these hypotheses, the output information obtained includes (i) information generated by the sequencing error probability model based on the probability model processing the input to the probability model, the input describing: (i) the read orientation of each read in the pile of aligned sequence reads stored in one or more memories; (ii) for each read in the pile of aligned sequence reads stored in one or more memories, the position of the 5' end of the reference read for each base at the reference position; (iii) the read allele score of each read in the pile of aligned sequence reads stored in one or more memories for each candidate allele at the reference position; and (iv) the base quality score of each read of the base at the reference position, such as position "0". In addition, the output information obtained includes a score for each hypothesis in a specific hypothesis that takes into account the occurrence of one or more sequencing errors, the score indicating the likelihood that the hypothesis is true.

[0158] The computer can continue to perform process 500 by determining the likelihood that a true variant exists at the first position based on the obtained output data generated by the probability model for each of the plurality of hypotheses (step 550). In some embodiments, this determination can be made, for example, based on evaluating the probability scores determined for each of the one or more hypotheses individually or collectively against one or more predetermined thresholds.

[0159] For example, the computer can determine a total score based on the output data generated by the probability model for each of the multiple hypotheses. The total score can indicate the possibility of the presence of a true variant. In some embodiments, the total score can be determined using equation (14). However, other methods for determining the total score are within the scope of the present disclosure. If the computer determines that the total score meets a predetermined threshold, the computer can add information indicating the presence of a true variant at the first position to the VCF file. Alternatively, if the computer determines that the total score does not meet the predetermined threshold, the computer can discard the output data and determine that the output data indicates a false positive at the reference position.

[0160] The process 500 described above describes a method for using a probabilistic model that can be used to account for the likelihood of sequencing errors. Examples of probabilistic models that can be used to account for the likelihood of one or more sequencing errors, also referred to as systematic errors, are described in more detail below.

[0161] In one embodiment, the probability model is modified to take into account the observation that certain base sequences tend to have a high probability of producing base calling errors, and this probability is not used to calculate P(r i |G m,φ ) is represented by the base quality.

[0162] Probabilistic models for determining the likelihood of a sequencing error occurring are applicable to the following situations:

[0163] (1) Errors appear to occur independently in each chain direction because one chain direction is often damaged while the other is error-free;

[0164] (2) Errors are more likely to occur farther from the 5' end of the read, and the error rate often decreases abruptly at a certain distance from the 5' end. Therefore, when the reads in a given strand direction are listed in order of decreasing distance from the 5' end, all errors are included in the subset at the beginning of the list;

[0165] (3) errors are often accompanied by a decrease in the average base quality in a subset of reads containing errors, rather than every erroneous read having a low base quality, and the average base quality is often not low enough to reflect the true error rate associated with these error events; and

[0166] (4) Errors are often preceded by homopolymers that match the error, for example, a T ==> G error is often preceded by a sequence of G.

[0167] Thus, the present disclosure provides a probability model for determining the likelihood of the occurrence of a sequencing error that takes into account the four characteristics described above.

[0168] In one embodiment, a probabilistic model for determining the likelihood of occurrence of a sequencing error taking into account the four characteristics described above can be implemented by starting with the following definitions:

[0169] Let θ indicate the chain direction, θ=0,1.

[0170] Make the read segment r θ,i Sort by strand direction and distance from the locus of interest to the 5' end, with i=1 being farthest from the 5' end and most likely to be affected by error events.

[0171] Make q θ,i Instructions and Reads θ,i The phred-level base quality of the aligned bases at the locus of interest.

[0172] make is the ordered read segment i=1…n in the chain direction θ θ Average base quality in a subset of :

[0173] where n θ ≥1.

[0174] Define the extended candidate genotype G′ m =[G m,1 G m,2 E m,0 E m,1 ], where E m,θ is the wrong allele in the strand direction θ.

[0175] Make L E,θ is the base E immediately preceding the error in the strand direction θ m,θ Match the length of the homopolymer.

[0176] For each expanded candidate genotype G′ m ,calculate:

[0177]

[0178] in Indicates the a priori probability of error events affecting a subset of reads as a function of the average base quality in the subset and the length of the homopolymer matched to the error.

[0179] Quantity m Denote the joint probability P(G′) under the following assumptions m, R) are estimated as follows: (1) when sorting by decreasing distance from the 5' end, an error event affects a contiguous subset of reads, starting with the first read, and does not affect reads beyond the subset; (2) the error event occurs independently for each strand direction; and (3) the a priori probability of such an error event is the average base quality in the subset of reads affected by the error event and the average base quality of the base immediately preceding the error in strand direction θ that is identical to base E m,θ Function of the length of the matched homopolymer.

[0180] In some embodiments, the number of candidates can be reduced to only G m,1 =G m,2 =G m In this case, the expression simplifies to

[0181]

[0182] In some embodiments, α θ The value of can be fixed to α θ =0.5. Therefore, (7) can be rewritten as:

[0183]

[0184] In some embodiments, the prior probability function in (6) is expressed as a general function that can have a wide range of shapes. In theory, this function can be trained on real data. However, given the limitations of using this function to train real data, some embodiments of current probability models for determining sequencing errors may use

[0185] In some embodiments, the above equation can accommodate the possibility of different erroneous alleles in each strand direction. However, in some embodiments, the probability model used to determine the probability of one or more sequencing errors can only consider the probability that E m,0 =E m,1 In such embodiments, the θ subscript can be dropped, and the error allele can then be represented as E m .

[0186] Figure 6 600 is another flow chart of an example of a process 600 for variant identification related error mitigation. The process 600 is described below as being performed by, for example, Figure 1 In some embodiments, the variant identification unit can be implemented using hardware such as one or more FPGAs, ASICs, CPUs, GPUs, or any combination thereof, using software, or any combination thereof.

[0187] The computer can start execution process 600 by accessing the accumulation (step 610) of aligned sequence reads stored in one or more memory devices. The mapping and alignment unit implemented with the configurable logic gate of the FGPA device may have been used to generate the aligned sequence reads. In some embodiments, the accessed reads are stored in one or more memory devices and may include a first group of sequence reads aligned in the forward direction and a second group of sequence reads aligned in the reverse direction. Each group of corresponding sequence reads corresponds to a specific read orientation or direction.

[0188] The computer can continue to execute process 600 by obtaining information describing: (i) the read orientation of each read in the stack of aligned sequence reads stored in one or more memories; (ii) for each read in the stack of aligned sequence reads stored in one or more memories, the position of the 5′ end of the reference read for each base at a reference position; (iii) a mapping confidence score for each read; (iv) a read allele score for each read in the stack of aligned sequence reads stored in one or more memories for each candidate allele at the reference position; and (v) a base quality score for each read of the base at the reference position, e.g., position “0” (620).

[0189] In some embodiments, the mapping confidence score may comprise the output of the mapping and alignment unit 126 and may be calculated using the mapping error Q phred-mapping The phred level probability is determined by Q phred-mapping ,=-10*log10(P e-mapping ). In this example, P e-mapping is the probability that a particular read is mapped incorrectly. The value of the mapping confidence score 144 may be proportional to the difference between the best alignment score from an alignment algorithm, such as a Smith-Waterman aligner, and the second best score from the aligner.

[0190] The read allele fraction can be determined in many different ways. For example, a more complex haplotype variant identifier can use the read allele fraction calculated using the above equation (1). Alternatively, other variant identifiers other than the haplotype variant identifier can use the read allele fraction calculated using the above equation (2). In some embodiments, the type of read allele fraction used can be determined based on whether the variant identification unit 130 will be used to detect only SNPs or SNPs and insertions and deletions. For example, in some embodiments where the variant identification unit will be used to detect SNPs and insertions and deletions, the read allele fraction calculated using the above equation (1) can be used. By way of another example, in other embodiments where the variant identification unit will be used to detect only SNPs, the read allele fraction calculated using equation (2) can be used.

[0191] The computer can continue to perform process 600 by providing one or more inputs describing the obtained information to the probability model, wherein the obtained information describes: (i) the read orientation of each read in the stack of aligned sequence reads stored in one or more memories; (ii) for each read in the stack of aligned sequence reads stored in one or more memories, the position of the 5′ end of the reference read for each base at a reference position; (iii) a mapping confidence score for each read; (iv) the read allele score for each read in the stack of aligned sequence reads stored in one or more memories (step 630) for each candidate allele at the reference position; and (v) a base quality score for each read of the base at the reference position, such as position “0” (630).

[0192] The computer may continue to perform process 600 by obtaining output information (640) for each of the one or more hypotheses based on the one or more inputs obtained at 620. In an example of process 600, the obtained inputs include inputs to a sequencing error probability model, the inputs including: (i) a read orientation of each read in the pile of aligned sequence reads stored in one or more memories; (ii) for each read in the pile of aligned sequence reads stored in one or more memories, the position of the 5′ end of the reference read for each base at a reference position; (iii) a mapping confidence score for each read; (iv) a read allele score for each read in the pile of aligned sequence reads stored in one or more memories for each candidate allele at the reference position; and (v) a base quality score for each read of the base at the reference position, such as position “0”. Thus, based on receiving these inputs to the mapping error probability model and the sequencing error probability model, the computer will generate output information for one or more hypotheses that take into account the occurrence of one or more mapping errors and one or more sequencing errors. This hypothesis includes (i) the probability that the read at the reference position indicates the occurrence of homozygous ref with an alien allele matching alt, (ii) the probability that the read at the reference position indicates the occurrence of homozygous alt with an alien allele matching the reference allele, (iii) the probability that the read at the reference position indicates the occurrence of homozygous ref with a sequencing error matching the alt allele, and (iv) the probability that the read at the reference position indicates the occurrence of homozygous alt with a sequencing error matching the reference allele.

[0193] For each of these hypotheses, the output information obtained includes information generated by the mapping error probability model and the sequencing error probability model based on the inputs processed by each corresponding probability model to the probability model, the input describing: (i) the read orientation of each read in the stack of aligned sequence reads stored in one or more memories; (ii) for each read in the stack of aligned sequence reads stored in one or more memories, the position of the 5′ end of the reference read for each base at a reference position; (iii) a mapping confidence score for each read; (iv) a read allele score for each read in the stack of aligned sequence reads stored in one or more memories for each candidate allele at the reference position; and (v) a base quality score for each read of the base at the reference position, such as position “0” (620). In addition, the output information obtained includes a score for each of the specific hypotheses that takes into account the occurrence of one or more mapping errors or one or more sequencing errors, the score indicating the likelihood that the hypothesis is true.

[0194] The computer can continue to perform process 600 by determining the likelihood that a true variant exists at the first position based on the obtained output data generated by the probability model for each of the plurality of hypotheses (step 650). In some embodiments, this determination can be made, for example, based on evaluating the probability scores determined for each of the one or more hypotheses individually or collectively against one or more predetermined thresholds.

[0195] For example, the computer can determine a total score based on the output data generated by the probability model for each of the multiple hypotheses. The total score can indicate the possibility of the presence of a true variant. In some embodiments, the total score can be determined using equation (14). However, other methods for determining the total score are within the scope of the present disclosure. If the computer determines that the total score meets a predetermined threshold, the computer can add information indicating the presence of a true variant at the first position to the VCF file. Alternatively, if the computer determines that the total score does not meet the predetermined threshold, the computer can discard the output data and determine that the output data indicates a false positive at the reference position.

[0196] refer to Figure 2 , 3 The processes described in the flowcharts of and 5 generally describe processes for variant identification related error event mitigation. These corresponding processes describe the processes used for reference Figure 2 The overall process of mitigating related error events is for reference Figure 3 The process of mapping error mitigation and reference for Figure 5The process of sequencing error mitigation. However, when sufficient input for considering one or more mapping errors and one or more sequencing errors is obtained by the computer, the present disclosure does not need to perform separate, uncombined probability calculations for mapping errors and sequencing errors, respectively. Instead, in some embodiments, a complete probability model can be used in the process of, for example, process 600 to consider mapping errors and sequencing errors. Although a combined probability model for considering mapping errors and sequencing errors will be described below, it is not required to use a combined probability model. Instead, as described with reference to other embodiments, a separate, uncombined probability model can be relied upon.

[0197] The following description derives the full probability calculation in the case of a diploid genome with at most one foreign allele and one systematic error allele. However, given the following description, this can be directly extended to more alleles. The expanded candidate genotype can be defined as G′ m =[G m,1 G m,2 F m E m In the following expressions, the notation 0 indicates the REF allele, 1 indicates the first ALT allele, and 2 indicates the second ALT allele, etc. m or E m The dashes indicate the absence of foreign alleles or systematic error alleles.

[0198] When the candidate G′ m In the absence of foreign or mis-alleles, the following expression is produced:

[0199]

[0200] When G′ m When foreign alleles are present, equation (4) described above or one of its variations can be used. m When error alleles are present, equation (10) or one of its variations may be used. Generally, it is not necessary to test for both foreign alleles and error alleles, since it is extremely difficult to find a pileup that is affected by both types of errors.

[0201] As an example of a candidate list, for the common case of a single alt allele (REF=0, ALT=1), the following expanded candidate genotypes can be tested:

[0202]

[0203] For each (non-expanded) candidate G m , the joint probability P(G m ,R) is related to G mThe maximum value among the matching candidates:

[0204]

[0205] And the posterior probability is just

[0206]

[0207] in

[0208] References below Figures 9A to 12B The examples described illustrate the reference Figure 3 , 5 and 6 are examples of the calculations described for various pile-ups.

[0209] Figure 7 700 is another flow chart of an example of a process 700 for variant identification related error mitigation. The process 700 is described below as being performed by, for example, Figure 1 In some embodiments, the variant identification unit can be implemented using hardware such as one or more FPGAs, ASICs, CPUs, GPUs, or any combination thereof, using software, or any combination thereof.

[0210] A computer storing one or more probability models may begin executing process 700 by receiving input data (710) containing information describing one or more characteristics of a read from a stack of aligned sequence reads. In some embodiments, the one or more characteristics may include: (i) information describing the read orientation of each read in the stack of aligned sequence reads; (ii) information describing the position of the 5′ end of each base at a reference position for each read in the stack of aligned sequence reads; (iii) a mapping quality score for each read in the stack of aligned sequence reads; (iv) a read allele score for each read in the stack of aligned sequence reads for each candidate allele at a reference position; and (v) a base quality score for each read of a base at a reference position, such as position “0”. In some embodiments, all of these inputs or a subset of these inputs may be provided as input to a probability model, as with reference to other probability models herein. In some embodiments, one or more probability models stored by a computer may include a mapping error probability model and a sequencing error probability model.

[0211] In some embodiments, the mapping confidence score may comprise the output of the mapping and alignment unit 126 and may be calculated using the mapping error Q phred-mapping The phred level probability is determined by Q phred-mapping ,=-10*log10(P e-mapping ). In this example, Pe-mapping is the probability that a particular read is mapped incorrectly. The value of the mapping confidence score 144 may be proportional to the difference between the best alignment score from an alignment algorithm, such as a Smith-Waterman aligner, and the second best score from the aligner.

[0212] The read allele fraction can be determined in many different ways. For example, a more complex haplotype variant identifier can use the read allele fraction calculated using the above equation (1). Alternatively, other variant identifiers other than the haplotype variant identifier can use the read allele fraction calculated using the above equation (2). In some embodiments, the type of read allele fraction used can be determined based on whether the variant identification unit 130 will be used to detect only SNPs or SNPs and insertions and deletions. For example, in some embodiments where the variant identification unit will be used to detect SNPs and insertions and deletions, the read allele fraction calculated using the above equation (1) can be used. By way of another example, in other embodiments where the variant identification unit will be used to detect only SNPs, the read allele fraction calculated using equation (2) can be used.

[0213] The computer can continue to perform process 700 by determining a set of one or more hypotheses based on the received inputs (720). For example, if the one or more inputs include information describing read characteristics associated with a probability model that accounts for one or more mapping errors, the computer can determine a set of one or more hypotheses that account for one or more mapping errors. Such hypotheses can include, for example, (i) the likelihood that a read at a reference position indicates the occurrence of a homozygous ref with an alien allele that matches alt, and (ii) the likelihood that a read at a reference position indicates the occurrence of a homozygous alt with an alien allele that matches a reference allele.

[0214] Alternatively or additionally, for example, if one or more inputs include information describing read characteristics associated with a probabilistic model that accounts for one or more sequencing errors, then the computer can determine a set of one or more hypotheses that account for one or more sequencing errors. Such hypotheses can include (i) the likelihood that a read at a reference position indicates the occurrence of a homozygous reference with a sequencing error that matches an alt allele, and (ii) the likelihood that a read at a reference position indicates the occurrence of a homozygous alt with a sequencing error that matches a reference allele.

[0215] In some embodiments, one or more received inputs can cause the computer to determine a set of hypotheses that consider both mapping errors and sequencing errors. When, for example, one or more received inputs contain information describing the read characteristics associated with a probability model that considers both one or more mapping errors and one or more sequencing errors, such determination can be performed by a computer. A set of hypotheses that consider both one or more mapping errors and one or more sequencing errors can include, for example, (i) the read at the reference position indicates the possibility of the occurrence of a homozygous ref with an alien allele that matches alt, (ii) the read at the reference position indicates the possibility of the occurrence of a homozygous alt with an alien allele that matches the reference allele, (iii) the read at the reference position indicates the possibility of the occurrence of a homozygous reference with a sequencing error that matches the alt allele, and (iv) the read at the reference position indicates the possibility of the occurrence of a homozygous alt with a sequencing error that matches the reference allele.

[0216] The computer can continue to perform process 700 by determining a score 730 indicating the probability that each corresponding hypothesis determined in stage 720 is true for each of the one or more hypotheses. Determining the score of each hypothesis by the computer can include, for example, using one or more probability models to determine the probability score of each of the one or more hypotheses. An example of a probability model that can be used to determine the score of each mapping error-related hypothesis is described with reference to equation (3) described above. However, other variations of other probability models for calculating the score of each mapping error hypothesis are also described above. An example of a probability model that can be used to determine the score of each sequencing error-related hypothesis is described with reference to equation (9). However, other variations of other probability models for calculating the score of each sequencing error hypothesis are also described above.

[0217] The computer can continue to perform process 700 by providing output data (740) of the score of each hypothesis in one or more hypotheses determined at stage 720 generated by the probability model. In some embodiments, the output data provided can be provided to a second computer, which is configured to determine the possibility of true variant existence based on the output data. In some embodiments, the second computer can be a computer that provides input at stage 710. In some embodiments, the second computer can be another genome analysis model or other computer module of an auxiliary analysis unit, and the computer is a part of the auxiliary analysis unit. In some embodiments, the second computer can be a remote computer with an auxiliary analysis unit, and the auxiliary analysis unit has a mapper and an alignment module, but does not have a variant identification module, and the auxiliary analysis unit can use one or more networks to communicate with the computer.

[0218] However, in other embodiments, the computer is not required to provide input to the second computer. Instead, the output information provided at 740, including a set of scores for each of the one or more hypotheses, can be used by a computer, such as a variant caller, to determine whether a candidate allele at a reference position is a true variant or a false positive, such as a reference position. Figure 1 described.

[0219] System Components

[0220] Figure 8 is a block diagram of system components that may be used to implement a system for correlated error mitigation for variant identification.

[0221] Computing device 800 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 850 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, and other similar computing devices. In addition, computing device 800 or 850 may include a Universal Serial Bus (USB) flash drive. A USB flash drive may store an operating system and other applications. A USB flash drive may include an input / output component, such as a wireless transmitter or a USB connector that may be inserted into a USB port of another computing device. The components shown here, their connections and relationships, and their functions are intended to be exemplary only, and are not intended to limit the embodiments of the invention described and / or claimed in this document.

[0222] The computing device 800 includes a processor 802, a memory 804, a storage device 808, a high-speed interface 808 connected to the memory 804 and a high-speed expansion port 810, and a low-speed interface 812 connected to a low-speed bus 814 and the storage device 808. Each of the components 802, 804, 808, 808, 810 and 812 is interconnected using various buses and can be installed on a common motherboard or otherwise installed when appropriate. The processor 802 can process instructions for execution within the computing device 800, including instructions stored in the memory 804 or on the storage device 808, and the instructions are used to display graphical information of a GUI on an external input / output device, such as a display 816 coupled to the high-speed interface 808. In other embodiments, multiple processors and / or multiple buses can be used together with multiple memories and memory types when appropriate. In addition, multiple computing devices 800 can be connected, each of which provides a portion of the necessary operations, such as a server group, a group of blade servers, or a multi-processor system.

[0223] The memory 804 stores information within the computing device 800. In one embodiment, the memory 804 is one or more volatile memory units. In another embodiment, the memory 804 is one or more non-volatile memory units. The memory 804 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.

[0224] Storage device 808 can provide mass storage for computing device 800. In one embodiment, storage device 808 can be or contain a computer readable medium, such as a floppy disk device, hard disk device, optical disk device or tape device, flash memory or other similar solid-state memory device or device array, including devices in a storage area network or other configuration. A computer program product can be tangibly embodied in an information carrier. A computer program product can also contain instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer readable medium or a machine readable medium, such as memory 804, storage device 808, or memory on processor 802.

[0225] High-speed controller 808 manages bandwidth-intensive operations of computing device 800, while low-speed controller 812 manages lower bandwidth-intensive operations. This functional allocation is exemplary only. In one embodiment, high-speed controller 808 is coupled to memory 804, such as being coupled to display 816 through a graphics processor or accelerator, and is coupled to a high-speed expansion port 810 that can accept various expansion cards (not shown). In the embodiment, low-speed controller 812 is coupled to storage device 808 and low-speed expansion port 814. The low-speed expansion port can include various communication ports, such as USB, Bluetooth, Ethernet, wireless Ethernet, and the low-speed expansion port can be coupled to one or more input / output devices, such as keyboards, pointing devices, microphone / speaker pairs, scanners or networking devices, such as switches or routers, for example, through a network adapter. Computing device 800 can be implemented in many different forms, as shown in the figure. For example, computing device can be implemented as a standard server 820, or implemented multiple times in a group of such servers. Computing device can also be implemented as a part of rack server system 824. In addition, computing device can be implemented in personal computers such as notebook computers 822. Alternatively, components from computing device 800 may be combined with other components in a mobile device (not shown), such as device 850. Each of such devices may contain one or more of computing devices 800, 850, and the entire system may be comprised of multiple computing devices 800, 850 in communication with each other.

[0226] The computing device 800 can be implemented in many different forms, as shown. For example, the computing device can be implemented as a standard server 820, or multiple times in a group of such servers. The computing device can also be implemented as part of a rack server system 824. In addition, the computing device can be implemented in a personal computer such as a laptop 822. Alternatively, components from the computing device 800 can be combined with other components in a mobile device (not shown) such as device 850. Each of such devices can contain one or more of the computing devices 800, 850, and the entire system can be composed of multiple computing devices 800, 850 that communicate with each other.

[0227] The computing device 850 includes a processor 852, a memory 864, and input / output devices such as a display 854, a communication interface 866, and a transceiver 868, among other components. The device 850 may also be provided with a storage device such as a micro drive or other device to provide additional storage. Each of the components 850, 852, 864, 854, 866, and 868 are interconnected using various buses, and several components may be mounted on a common motherboard or otherwise when appropriate.

[0228] The processor 852 can execute instructions within the computing device 850, including instructions stored in the memory 864. The processor can be implemented as a chipset of chips, including single and multiple analog and digital processors. In addition, the processor can be implemented using any of several architectures. For example, the processor 810 can be a Complex Instruction Set Computer (CISC) processor, a Reduced Instruction Set Computer (RISC) processor, or a Minimal Instruction Set Computer (MISC) processor. The processor can, for example, implement coordination of other components of the device 850, such as controlling a user interface, applications run by the device 850, and wireless communications performed by the device 850.

[0229] The processor 852 can communicate with the user through a control interface 858 and a display interface 856 coupled to the display 854. The display 854 can be, for example, a thin film transistor liquid crystal display (TFT) display or an organic light emitting diode (OLED) display, or other appropriate display technology. The display interface 856 may include an appropriate circuit system for driving the display 854 to present graphics and other information to the user. The control interface 858 can receive commands from the user and convert the commands to submit to the processor 852. In addition, an external interface 862 that communicates with the processor 852 can be provided so that the device 850 can communicate with other devices in a near area. The external interface 862 can, for example, implement wired communication in some embodiments or wireless communication in other embodiments, and multiple interfaces can also be used.

[0230] The memory 864 stores information within the computing device 850. The memory 864 may be implemented as one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. An expansion memory 874 may also be provided and connected to the device 850 via an expansion interface 872, which may include, for example, a single in-line memory module (SIMM) card interface. Such an expansion memory 874 may provide additional storage space for the device 850, or may also store applications or other information for the device 850. Specifically, the expansion memory 874 may include instructions for executing or supplementing the processes described above, and may also include security information. Thus, for example, the expansion memory 874 may be provided as a security module for the device 850, and may be programmed with instructions for permitting secure use of the device 850. In addition, a security application may be provided through a SIMM card together with additional information, such as placing identification information on the SIMM card in an inviolable manner.

[0231] The memory may include, for example, flash memory and / or NVRAM memory, as discussed below. In one embodiment, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable medium or a machine-readable medium, such as memory 864, expansion memory 874, or on-processor memory 852, which may be received, for example, via transceiver 868 or external interface 862.

[0232] The device 850 may communicate wirelessly via a communication interface 866, which may include digital signal processing circuitry, if necessary. The communication interface 866 may enable communication in various modes or protocols, such as GSM voice calls, SMS, EMS or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, etc. Such communication may occur, for example, via a radio frequency transceiver 868. In addition, short-range communication may occur, such as using Bluetooth, Wi-Fi, or other such transceivers (not shown). In addition, a global positioning system (GPS) receiver module 870 may provide additional navigation-related wireless data and location-related wireless data to the device 850, which may be used by applications running on the device 850 as appropriate.

[0233] Device 850 may also communicate audibly using audio codec 860, which may receive voice information from a user and convert the voice information into usable digital information. Audio codec 860 may also generate sounds audible to the user, such as through a speaker, such as a microphone in device 850. Such sounds may include sounds from voice phone calls, may include recorded sounds, such as voice messages, music files, etc., and may also include sounds generated by applications running on device 850.

[0234] The computing device 850 can be implemented in many different forms, as shown. For example, the computing device can be implemented as a cellular phone 880. The computing device can also be implemented as part of a smart phone 882, a personal digital assistant, or other similar mobile device.

[0235] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, specially designed application specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations of such embodiments. These various embodiments can include embodiments in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor that can be special-purpose or general-purpose and is coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to the storage system, at least one input device, and at least one output device.

[0236] These computer programs (also referred to as programs, software, software applications or code) contain machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages ​​and / or in assembly / machine languages. As used herein, the terms "machine-readable medium", "computer-readable medium" refer to any computer program product, device and / or means for providing machine instructions and / or data to a programmable processor, such as a disk, an optical disk, a memory, a programmable logic device (PLD), including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0237] To enable interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device, such as a cathode ray tube (CRT) or liquid crystal display (LCD) monitor, for displaying information to the user; and a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other types of devices can also be used to enable interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and the input from the user can be received in any form, including acoustic, voice, or tactile input.

[0238] The systems and techniques described herein can be implemented in a computing system that includes a back-end component, such as a data server, or includes a middleware component, such as an application server, or includes a front-end component, such as a client computer with a graphical user interface or a web browser, through which a user can interact with an embodiment of the systems and techniques described herein, or includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by digital data communication of any form or medium, such as through a communication network. Examples of communication networks include local area networks ("local area networks, LANs"), wide area networks ("wide area networks, WANs"), and the Internet.

[0239] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0240] Several embodiments have been described. However, it should be understood that various modifications may be made without departing from the spirit and scope of the present invention. In addition, the logic flow depicted in the figure does not need to be in the specific order or sequential order shown to achieve the desired results. In addition, other steps may be provided, or steps may be removed from the described flow, and other components may be added to or removed from the described system. Therefore, other embodiments are within the scope of the following claims.

[0241] Examples

[0242] The subject matter of the present disclosure is further described with reference to the following examples, which are not intended to limit the scope of the present disclosure in any way.

[0243] The examples provided in this section show real-world examples illustrating computations using the probabilistic model described herein on real-world pileups. Each pileup graph contains a MAPQ mapping confidence score for each read, a base quality at a reference position such as position "0", and an average base quality for each strand (blue = forward, red = reverse).

[0244] Example 1 - Typical true variant

[0245] Fig. 9A is an example of a summary of experimental results from execution of a process for relevant error mitigation for variant calling that has been performed on a pile of sequencing reads whose results indicate examples of true positive results.

[0246] In more detail, Fig. 9A An image 940 of a pile 941 of aligned sequence reads showing a typical true positive example is shown. A true positive is an example of a pile 941 of reads having a true variant at a reference position 942. In this example, it can be seen that the read characteristics indicate the presence of a candidate alt allele "C" at the reference position 942 that is different from the reference "T" of the reference genome to which the pile 941 of reads maps and aligns.

[0247] In this example, and with reference to Fig. 9AFrom the image 940 of the stack 941, it can be seen that the alt allele frequency of "C" is balanced between the first forward read direction (or orientation) and the second reverse read direction (or orientation), and for each read at the reference position, the MAPQ mapping confidence score is high, with most mapping confidence scores 944 being the maximum value "250", and the average base quality in both strand directions (or orientations) is high, with most base quality scores 943 being "35" or higher. Therefore, the probability score of candidate [10|-|-] 962 is high, as shown in the probability score result column 980 and the normalized probability score result column 980. Therefore, information identifying the true alt "C" or otherwise associated with it can be included in the VCF file.

[0248] Fig. 9B A complete set of modified probability results 960 is shown in FIG. Candidate [10|-|-] 962 corresponds to the above-mentioned hypothesis 162, which contains the possibility that the read segment at the reference position 142 indicates the occurrence of the heterozygous alt allele. Probability score result column 980 and normalized probability score result column 990 respectively show the probability score results and normalized score results of each of the other conventional hypotheses 961, 962, 963 and unconventional hypotheses 964, 965, 966, 967. For clarity, conventional hypothesis 961 corresponds to the above-mentioned conventional hypothesis 161, conventional hypothesis 962 corresponds to the above-mentioned conventional hypothesis 162, and conventional hypothesis 963 corresponds to the above-mentioned conventional hypothesis 163. In addition, unconventional hypothesis 964 corresponds to the above-mentioned unconventional hypothesis 164, unconventional hypothesis 965 corresponds to the above-mentioned unconventional hypothesis 165, unconventional hypothesis 966 corresponds to the above-mentioned unconventional hypothesis 166, and unconventional hypothesis 967 corresponds to the above-mentioned unconventional hypothesis 167.

[0249] Example 2 – Low probability of mapping error

[0250] Fig. 10A is an example of a summary of experimental results from the execution of a process for relevant error mitigation of variant calling that has been performed on a pile of sequencing reads whose results indicate a low probability of the occurrence of mapping errors.

[0251] In more detail, Fig. 10AAn image 1040 of a pile 1041 of aligned sequence reads showing possible examples of foreign reads or possible examples of mapping errors is shown. In this example, for reads with alt alleles, the pile has a low alt allele frequency and a moderately low MAPQ mapping confidence score 1044. Both of these factors will be utilized by the mapping error probability model provided by the present disclosure. This means that the resulting likelihood of a true variant is therefore reduced. Review of the probability scores for each corresponding hypothesis indicates that it is not possible to conclude with high confidence that the "C" allele is a foreign read, nor is it possible to make a decision in heterozygous identification with high confidence. Therefore, the output information can be discarded without adding any information identifying the candidate alt allele to the VCF file.

[0252] Fig. 10B 1060. The probability score result column 1080 and the normalized probability score result column 1090 respectively show the probability score result and the normalized score result for each of the other conventional hypotheses 1061, 1062, 1063 and the unconventional hypotheses 1064, 1065, 1066, 1067. For clarity, the conventional hypothesis 1061 corresponds to the conventional hypothesis 161 described above, the conventional hypothesis 1062 corresponds to the conventional hypothesis 162 described above, and the conventional hypothesis 1063 corresponds to the conventional hypothesis 163 described above. In addition, the unconventional hypothesis 1064 corresponds to the unconventional hypothesis 164 described above, the unconventional hypothesis 1065 corresponds to the unconventional hypothesis 165 described above, the unconventional hypothesis 1066 corresponds to the unconventional hypothesis 166 described above, and the unconventional hypothesis 1067 corresponds to the unconventional hypothesis 167 described above.

[0253] Example 3 - Impossible true variant due to sequencing error

[0254] Fig.11A is an example of a summary of experimental results from the execution of a process for associated error mitigation of variant identification that has been performed on a pile of sequencing reads whose results indicate a high likelihood that a candidate alt allele is unlikely to be a true variant due to sequencing error.

[0255] In more detail, Fig.11A An image 1140 of a pile 1141 of aligned sequence reads showing systematic or sequencing errors is shown.

[0256] In this example, all "G" alt alleles appear in a single read orientation in the forward alignment direction. In addition, all "G" alt alleles appear on a subset of forward oriented reads that are farthest from the 5' end of the read. The read subset with the "G" alt allele has an extremely low base quality, as demonstrated by the base quality score 1143. In addition, the "G" alt allele matches the two base alleles presenting the base allele at the current reference position 1145 of the reference genome. The sequencing error probability model takes into account the aforementioned characteristics of the reads, and the probability scores output for each of the seven hypotheses 1161, 1162, 1163, 1164, 1165, 1166, 1167 and shown in the probability score results 1180 and the normalized probability score results column 1190 support with high confidence that the "G" alt allele is unlikely to support the true variant. Therefore, the output information can be discarded without adding any information identifying the candidate alt alleles to the VCF file.

[0257] Fig. 11B 160. The probability score result column 1180 and the normalized probability score result column 1190 respectively show the probability score result and the normalized score result for each of the other conventional hypotheses 1161, 1162, 1163 and the unconventional hypotheses 1164, 1165, 1166, 1167. For clarity, the conventional hypothesis 1161 corresponds to the conventional hypothesis 161 described above, the conventional hypothesis 1162 corresponds to the conventional hypothesis 162 described above, and the conventional hypothesis 1163 corresponds to the conventional hypothesis 163 described above. In addition, the unconventional hypothesis 1164 corresponds to the unconventional hypothesis 164 described above, the unconventional hypothesis 1165 corresponds to the unconventional hypothesis 165 described above, the unconventional hypothesis 1166 corresponds to the unconventional hypothesis 166 described above, and the unconventional hypothesis 1167 corresponds to the unconventional hypothesis 167 described above.

[0258] Example 4—Unlikely deformation due to sequencing errors in both read orientations

[0259] Fig. 12A is an example of a summary of experimental results from the execution of a process for associated error mitigation for variant identification, which has been performed on a pile of sequencing reads whose results indicate a high probability that the candidate alt allele is unlikely to be a true variant due to sequencing errors in both read orientations.

[0260] In more detail, Fig. 12A An image 1240 of a pile-up 1241 of aligned sequence reads showing systematic errors or sequencing errors in the orientation of two reads is shown.

[0261] In this example, there is an extreme drop in base quality, as evidenced by base quality score 1243. On the forward alignment reads, the "G" alt allele is limited to a subset of reads that are further away from the 5' end of the locus of interest. In addition, the "G" alt allele matches the first three reference alleles of the reference genome before the reference allele at reference position 1242. The sequencing error probability model takes into account the aforementioned characteristics of the reads, and the probability scores output for each of the seven hypotheses 1261, 1262, 1263, 1264, 1265, 1266, 1267 and shown in the probability score results 1280 and the normalized probability score results column 1290 support with high confidence that the "G" alt allele is unlikely to support a true variant. Therefore, the output information can be discarded without adding any information identifying the candidate alt allele to the VCF file.

[0262] Fig. 12B 1260. The probability score result column 1280 and the normalized probability score result column 1290 respectively show the probability score result and the normalized score result for each of the other conventional hypotheses 1161, 1262, 1263 and the unconventional hypotheses 1264, 1265, 1266, 1267. For clarity, the conventional hypothesis 1261 corresponds to the conventional hypothesis 161 described above, the conventional hypothesis 1262 corresponds to the conventional hypothesis 162 described above, and the conventional hypothesis 1263 corresponds to the conventional hypothesis 163 described above. In addition, the unconventional hypothesis 1264 corresponds to the unconventional hypothesis 164 described above, the unconventional hypothesis 1265 corresponds to the unconventional hypothesis 165 described above, the unconventional hypothesis 1266 corresponds to the unconventional hypothesis 166 described above, and the unconventional hypothesis 1267 corresponds to the unconventional hypothesis 167 described above.

[0263] Other embodiments

[0264] It should be understood that although the invention has been described in conjunction with its drawings and detailed description, the foregoing description is intended to illustrate rather than limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages and modifications are within the scope of the following claims.

Claims

1. A method for determining whether a candidate variant allele is not a true variant allele, the method comprising: Obtaining a pile of reads aligned to a reference sequence; identifying a plurality of alt alleles corresponding to the candidate variant allele, wherein each alt allele corresponds to a base call of at least one read strand in the pile of reads at a particular reference position of the reference sequence; obtaining first data indicating whether each of the plurality of alt alleles is present in a single read strand having a read orientation in a forward alignment direction; obtaining second data indicating the positions of the plurality of alt alleles relative to the 5' end of each corresponding read strand comprising the alt allele; acquiring third data indicating a base quality score of each read having a read chain including an alt allele; Providing the first data, the second data, and the third data as inputs of a sequencing error probability model; Obtaining output data generated by the sequencing error probability model, wherein the obtained output data indicates a likelihood of whether the candidate variant allele is a true variant; and A determination is made based on the output data obtained whether the candidate variant allele is a true variant.

2. The method according to claim 1, further comprising: Based on determining that the candidate variant allele is not a true variant allele, perform the following steps: discarding the obtained output data without adding any additional information to a variant call format (VCF) file identifying candidate variant alleles; and The VCF that does not contain data identifying the candidate variant allele file is stored in a storage device of the computer system.

3. The method according to claim 1, further comprising: Based on the determination that the candidate variant allele is likely to be a true variant allele, the following steps are performed: Add information to the VCF file identifying candidate variant alleles; and The VCF file including information identifying the candidate variant alleles is stored in a storage device of the computer system.

4. The method according to claim 3, wherein: The information added to the VCF file for a candidate variant allele includes one or more of the following: the position of the candidate variant allele, an identifier of the candidate variant allele, the genotype of the candidate variant allele, or data indicating the likelihood that the candidate variant allele is a true variant.

5. The method according to claim 1, wherein: The output data indicating the likelihood of whether a candidate variant allele is a true variant is a value corresponding to the probability indicating whether the candidate variant allele is a true variant.

6. The method according to claim 5, wherein: Determining whether a candidate variant allele is a true variant based on the obtained output data includes: The probability indicating whether the candidate variant allele is a true variant is compared to a predetermined threshold.

7. The method according to claim 6, further comprising: Based on the comparing, determining that a probability indicating whether the candidate variant allele is a true variant does not meet the predetermined threshold; and Based on determining that the probability indicating whether the candidate variant allele is a true variant does not meet the predetermined threshold, it is determined that the candidate variant allele is not a true variant.

8. The method according to claim 6, further comprising: Based on the comparison, determining whether a probability indicating whether the candidate variant allele is a true variant satisfies the predetermined threshold; and Based on determining that the probability indicating whether the candidate variant allele is a true variant satisfies the predetermined threshold, it is determined that the candidate variant allele is likely to be a true variant.

9. A system comprising: one or more computers; and One or more memory devices storing instructions, which, when executed by the one or more computers, cause the one or more computers to perform the operations described in any one of claims 1-8.

10. One or more computer-readable storage media storing instructions which, when executed by one or more computers, cause the one or more computers to perform the operations of any one of claims 1-8.