Methods for detecting nucleic acid variants
By using flow cycle sequential sequencing technology in DNA samples, selecting target short genetic variants and determining their existence through matching scores, the accuracy and efficiency of detecting short genetic variants in the prior art are solved, and efficient and accurate variant judgment is achieved.
Patent Information
- Application Number
- CN202080048530.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-07
- Filing Date
- 2020-05-01
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-05-01
AI Technical Summary
The prior art has problems with accuracy and efficiency in detecting short genetic variants in DNA samples, especially in high-deep sequencing.
By selecting the target short genetic variant, the target sequence is sequenced using non-terminal nucleotides provided in separate nucleotide streams according to the flow cycle sequence, the target sequencing data set and the reference sequencing data set are obtained, and the presence or absence of the target short genetic variant in the test sample is determined by match scores.
Improve the accuracy and efficiency of base and variant judgments, reduce the cost and time consumption of high-deep sequencing, and can determine the existence of genetic variants with confidence.
Smart Images

Figure CN114072523B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of priority to U.S. Provisional Patent Application Serial No. 62 / 842,534, filed on May 3, 2019; and U.S. Provisional Patent Application Serial No. 62 / 971,530, filed on February 7, 2020; the contents of each of which are incorporated herein by reference in their entirety.
[0003] Submit sequence listing as ASCII text file
[0004] The contents of the following submitted ASCII text file are incorporated herein by reference in their entirety: Sequence Listing in Computer Readable Form (CRF) (File Name: 165272000540SEQLIST.TXT, Record Date: April 27, 2020, Size: 5KB). Technical Field
[0005] Described herein are methods of sequencing polynucleotides, including methods of generating and / or analyzing sequencing data, including detecting genetic variants. Background Art
[0006] Genetic variants in a DNA sample can be detected by sequencing the DNA in the sample, aligning the sequence to a reference sequence, and assessing the differences. High-confidence differences between the sequenced DNA and the reference sequence are called variants of the organism from which the DNA sample was derived. Next-generation sequencing provides research and clinical laboratories with the tools needed to sequence many different nucleic acid molecules in a single sample simultaneously, generating large amounts of data for analysis.
[0007] In addition, reversible terminator sequencing by synthesis (e.g., sequencing methods with reversible termination dye labels) provides a single difference signal for each base, so single signal sequencing errors can lead to incorrect variant determinations. In some cases, this can be overcome by high-depth sequencing, effectively overwhelming the false determinations with true positive signals, but sequencing at such a high depth is expensive and time-consuming.
[0008] There remains a need in the art for efficient and accurate base calling and variant calling schemes.
[0009] Brief description of the invention
[0010] Methods for detecting short genetic variants in a test sample containing nucleic acid molecules are described herein, and in certain embodiments, the methods can be computer-implemented methods. Systems for performing such methods are also described herein. Methods for sequencing nucleic acid molecules are further described.
[0011] In some embodiments, a method for detecting short genetic variants in a test sample includes (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing a target sequence using a non-termination nucleotide provided in a separate nucleotide flow according to a flow cycle order, the target sequencing dataset associated with a target sequence comprising the target short genetic variant is different from a reference sequencing dataset associated with a reference sequence at more than two flow positions; (b) obtaining one or more test sequencing datasets, each test sequencing dataset being associated with a test nucleic acid molecule, each test nucleic acid molecule at least partially overlapping a locus associated with the target short genetic variant and derived from the test sample, wherein the one or more test sequencing datasets are determined by sequencing the test nucleic acid molecules using a non-termination nucleotide provided in a separate nucleotide flow according to a flow cycle order, and wherein the test sequencing dataset includes flow signals at multiple flow positions; (c) for each test nucleic acid molecule associated with the test sequencing dataset, determining a match score indicating the likelihood that the test sequencing dataset associated with the nucleic acid molecule matches the target sequence. score), or a match score indicating the likelihood that a test sequencing data set associated with a nucleic acid molecule matches a reference sequence; and (d) using one or more of the determined match scores to determine the presence or absence of a target short genetic variant in the test sample.
[0012] In some embodiments of the above methods, the obtaining step comprises sequencing the test nucleic acid molecule using non-terminating nucleotides provided in separate nucleotide streams according to a flow cycle order.
[0013] In some embodiments of the above methods, before determining the presence or absence of a target short genetic variant in a test sample, a target short genetic variant is preselected. In some embodiments, after determining the presence or absence of a target short genetic variant in a test sample based on a determination confidence, a target short genetic variant is selected. In some embodiments, the method further comprises generating a personalized biomarker panel for a subject associated with the test sample. The biomarker panel includes a target short genetic variant.
[0014] In some embodiments of the above method, the method further comprises selecting a flow cycle sequence.
[0015] In some embodiments, the target sequencing data set is an expected target sequencing data set, or the reference sequencing data set is an expected reference sequencing data set. In some embodiments, the target sequence and the reference sequence are sequenced in silico to obtain an expected target sequencing data set and an expected reference sequencing data set.
[0016] In some embodiments of the above methods, the target sequencing data set is different from the reference sequencing data at more than two non-continuous flow positions. In some embodiments, the target sequencing data set is different from the reference sequencing data at more than two continuous flow positions. In some embodiments, the target sequence is different from the reference sequence at X base positions, and wherein the target sequencing data set is different from the reference sequencing data at (X+2) or more continuous flow positions. In some embodiments, the (X+2) flow position differences include the difference between a value substantially equal to zero and a value substantially greater than zero. In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles. In some embodiments, the flow signal includes a base count indicating the number of bases of the test nucleic acid molecules sequenced at each flow position.
[0017] In some embodiments of the above methods, the flow signal includes a statistical parameter indicating the likelihood of at least one base count at each flow position, wherein the base count indicates the number of bases of the test nucleic acid molecule sequenced at the flow position. In some embodiments, the flow signal includes a statistical parameter indicating the likelihood of multiple base counts at each flow position, wherein each base count indicates the number of bases of the test nucleic acid molecule sequenced at the flow position.
[0018] In some embodiments of the above method, step (c) includes (i) selecting a statistical parameter of the base counts of the target sequence at each flow position in the test sequencing data set corresponding to the flow position, and determining a match score indicating the possibility that the test sequencing data set matches the target sequence; or (ii) selecting a statistical parameter of the base counts of the reference sequence at each flow position in the test sequencing data set corresponding to the flow position, and determining a match score indicating the possibility that the test sequencing data set matches the reference sequence. In some embodiments, the match score determined in step (c) is a combined value of the selected statistical parameters across the flow positions in the test sequencing data set. In some embodiments, step (c) includes determining a match score indicating the possibility that the test sequencing data set matches the target sequence. In some embodiments, step (c) includes determining a match score indicating the possibility that the test sequencing data set matches the reference sequence.
[0019] In some embodiments of the above methods, the one or more test sequencing data sets include multiple test sequencing data sets. In some embodiments, the presence or absence of the target short genetic variant is determined separately for each of the one or more test sequencing data sets. In some embodiments, at least a portion of the multiple test sequencing data sets are associated with different test nucleic acid molecules having different sequencing start positions.
[0020] In some embodiments of the above methods, the flow cycle sequence comprises 4 individually separate streams repeated in the same order. In some embodiments, the flow cycle sequence comprises 5 or more individually separate streams.
[0021] In some embodiments of the above methods, the method is a computer-implemented method. For example, in some embodiments, the computer-implemented method includes selecting a target short genetic variant using one or more processors; obtaining one or more test sequencing data sets by receiving one or more test sequencing data sets at one or more processors; determining one or more match scores using one or more processors; and determining the presence or absence of the target short genetic variant in the test sample using one or more processors.
[0022] The present invention also provides a system, including: one or more processors; and a non-transitory computer-readable medium storing one or more programs including instructions for implementing the above method.
[0023] In some embodiments, a method for detecting short genetic variants in a test sample includes: (a) obtaining one or more first test sequencing data sets, each first test sequencing data set being associated with a different test nucleic acid molecule derived from the test sample, wherein the first test sequencing data set is determined by sequencing the one or more test nucleic acid molecules according to a first flow cycle order using non-termination nucleotides provided in a separate nucleotide flow, and wherein the one or more first test sequencing data sets include flow signals at flow positions corresponding to the nucleotide flow; (b) obtaining one or more second test sequencing data sets, each second test sequencing data set being associated with the same test nucleic acid molecule as the first test sequencing data set, wherein the second test sequencing data set is determined by sequencing the one or more test nucleic acid molecules according to a second flow cycle order using non-termination nucleotides provided in a separate nucleotide flow, wherein the first flow cycle order and the second flow cycle order are different, and wherein the test sequencing data sets include flow signals at flow positions corresponding to the nucleotide flow; (c) for each of the first sequencing data set and the second sequencing data set, determining a match score for one or more candidate sequences, wherein the match score indicates a likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from the one or more candidate sequences; and (d) determining the presence or absence of the short genetic variant in the test sample using the determined match score.
[0024] In some embodiments of the above method, the method includes sequencing the test nucleic acid molecules according to a first flow cycle order using non-termination nucleotides provided in a separate nucleotide flow, and sequencing the test nucleic acid molecules according to a second flow cycle order using non-termination nucleotides provided in a separate nucleotide flow.
[0025] In some embodiments of the above methods, the match score indicates the likelihood that the first test sequencing data set matches the candidate sequence, or the likelihood that the second test sequencing data set matches the candidate sequence. In some embodiments, the match score indicates the likelihood that both the first test sequencing data set and the second sequencing data set match the candidate sequence.
[0026] In some embodiments of the above method, the one or more candidate sequences include two or more different candidate sequences, and for each nucleic acid molecule associated with the first sequencing data set and the second sequencing data set, the method includes: selecting a candidate sequence from the two or more different candidate sequences, wherein the selected candidate sequence has the highest probability of matching the first test sequencing data set, the second test sequencing data set, or both; and using the selected candidate sequence to determine the presence or absence of a short genetic variant in the test sample. In some embodiments, at least one unselected candidate sequence from the two or more different candidate sequences is different from the selected candidate sequence at two or more flow positions according to the first flow cycle order or the second flow cycle order. In some embodiments, at least one unselected candidate sequence from the two or more different candidate sequences is different from the selected candidate sequence at two or more flow positions according to both the first flow cycle order and the second flow cycle order. In some embodiments, at least one unselected candidate sequence from the two or more different candidate sequences is different from the selected candidate sequence at two or more non-continuous flow positions according to the first flow cycle order or the second flow cycle order. In some embodiments, at least one unselected candidate sequence from two or more different candidate sequences is different from the selected candidate sequence at two or more non-continuous flow positions according to both the first flow cycle order and the second flow cycle order. In some embodiments, at least one unselected candidate sequence from two or more different candidate sequences is different from the selected candidate sequence at two or more continuous flow positions according to the first flow cycle order or the second flow cycle order. In some embodiments, at least one unselected candidate sequence from two or more different candidate sequences is different from the selected candidate sequence at two or more continuous flow positions according to both the first flow cycle order and the second flow cycle order. In some embodiments, at least one unselected candidate sequence from two or more different candidate sequences is different from the selected candidate sequence at 3 or more flow positions according to the first flow cycle order or the second flow cycle order. In some embodiments, at least one unselected candidate sequence from two or more candidate sequences is different from the selected candidate sequence at 3 or more flow positions according to both the first flow cycle order and the second flow cycle order. In some embodiments, at least one unselected candidate sequence from the two or more different candidate sequences differs from the selected candidate sequence at X base positions, and wherein the test sequencing data set associated with the test nucleic acid molecule differs from at least one unselected candidate sequence from the two or more different candidate sequences at (X+2) or more flow positions according to the first flow cycle order or the second flow cycle order.In some embodiments, at least one unselected candidate sequence from two or more different candidate sequences is different from the selected candidate sequence at X base positions, and wherein the test sequencing data set associated with the test nucleic acid molecule is different from at least one unselected candidate sequence from two or more different candidate sequences at (X+2) or more flow positions according to both the first flow cycle order and the second flow cycle order. In some embodiments, the (X+2) flow position differences include the difference between a value substantially equal to zero and a value substantially greater than zero. In some embodiments, at least one unselected candidate sequence from two or more different candidate sequences is different from the selected candidate sequence across one or more flow cycles according to the first flow cycle order or the second flow cycle order. In some embodiments, at least one unselected candidate sequence from two or more different candidate sequences is different from the selected candidate sequence across one or more flow cycles according to both the first flow cycle order and the second flow cycle order.
[0027] In some embodiments of the above method, the flow signal includes a base count indicating the number of bases of the test nucleic acid molecule sequenced at each flow position. In some embodiments, the flow signal includes a statistical parameter indicating the possibility of at least one base count at each flow position, wherein the base count indicates the number of bases of the test nucleic acid molecule sequenced at the flow position. In some embodiments, the flow signal includes a statistical parameter indicating the possibility of multiple base counts at each flow position, wherein each base count indicates the number of bases of the test nucleic acid molecule sequenced at the flow position. In some embodiments, determining the match score includes, for each of one or more different candidate sequences, selecting a statistical parameter corresponding to the base count of the candidate sequence at the flow position at each flow position in the first test sequencing data set and the second test sequencing data set. In some embodiments of the above method, the method includes: for one or more different candidate sequences, generating a candidate sequencing data set including the base count of the candidate sequence at each flow position. In some embodiments, the candidate sequencing data set is generated on a computer. In some embodiments, the match score is a combined value of the selected statistical parameters across the flow positions in the first test sequencing data set and the second test sequencing data set.
[0028] In some embodiments of the above methods, at least a portion of the test nucleic acid molecules have different sequencing start positions.
[0029] In some embodiments of the above method, the method further includes selecting a target short genetic variant, wherein when a target sequencing data set and a reference sequencing data set are obtained by sequencing the target sequence according to a first flow cycle order or a second flow cycle order using a non-terminal nucleotide provided in a separate nucleotide flow, the target sequencing data set associated with the target sequence containing the target short genetic variant is different from the reference sequencing data set associated with the reference sequence at two or more flow positions, wherein the first flow cycle order is different from the second flow cycle order, and wherein the flow position corresponds to the nucleotide flow; wherein one or more candidate sequences include the target sequence and the reference sequence. In some embodiments, the target short genetic variant is pre-selected before determining the presence or absence of the target short genetic variant in the test sample. In some embodiments, the target short genetic variant is selected after determining the presence or absence of the target short genetic variant in the test sample based on the confidence of the determination. In some embodiments, the method further includes generating a personalized biomarker panel for a subject associated with the test sample, the biomarker panel including the target short genetic variant present in the test sample. In some embodiments, if the reference sequence is sequenced using non-terminal nucleotides provided in a separate flow according to the first flow cycle order or the second flow cycle order, the reference sequencing data set is obtained by determining the expected reference sequencing data set. In some embodiments, if the reference sequence is sequenced using non-terminal nucleotides provided in a separate flow according to both the first flow cycle order and the second flow cycle order, the reference sequencing data set is obtained by determining the expected reference sequencing data set. In some embodiments, according to both the first flow cycle order and the second flow cycle order, the target sequence is different from the reference sequence at two or more flow positions. In some embodiments, according to the first flow cycle order or the second flow cycle order, the target sequence is different from the reference sequence at two or more non-continuous flow positions. In some embodiments, according to both the first flow cycle order and the second flow cycle order, the target sequence is different from the reference sequence at two or more non-continuous flow positions. In some embodiments, according to the first flow cycle order or the second flow cycle order, the target sequence is different from the reference sequence at two or more continuous flow positions. In some embodiments, according to both the first flow cycle order and the second flow cycle order, the target sequence is different from the reference sequence at two or more continuous flow positions. In some embodiments, according to both the first flow cycle order and the second flow cycle order, the target sequence is different from the reference sequence at two or more continuous flow positions. In some embodiments, the target sequence differs from the reference sequence at three or more flow positions according to the first flow cycle sequence or the second flow cycle sequence. In some embodiments, the target sequence differs from the reference sequence at three or more flow positions according to both the first flow cycle sequence and the second flow cycle sequence. In some embodiments, the target sequence differs from the reference sequence across one or more flow cycles according to the first flow cycle sequence or the second flow cycle sequence.In some embodiments, the target sequence differs from the reference sequence across one or more flow cycles according to both the first flow cycle sequence and the second flow cycle sequence.
[0030] In some embodiments of the above method, the first flow cycle sequence or the second flow cycle sequence includes 4 separate flows repeated in the same order. In some embodiments, the first flow cycle sequence or the second flow cycle sequence includes 5 or more separate flows repeated in the same order.
[0031] In some embodiments of the above method, the method includes sequencing a test nucleic acid molecule, including providing a non-terminating nucleotide in a separate nucleotide stream according to a first flow cycle order, extending a sequencing primer, and detecting the presence or absence of nucleotide integration into the sequencing primer after each nucleotide stream to generate a first test sequencing data set; removing the extended sequencing primer; and sequencing the same test nucleic acid molecule, including providing a non-terminating nucleotide in a separate nucleotide stream according to a second flow cycle order, extending a sequencing primer, and detecting the presence or absence of nucleotide integration into the sequencing primer after each nucleotide stream to generate a second test sequencing data set.
[0032] In some embodiments of the above methods, the methods are computer-implemented methods. For example, in some embodiments, the computer-implemented methods include receiving one or more first sequencing data sets at one or more processors; receiving one or more first sequencing data sets at the one or more processors; determining a match score using the one or more processors; and determining the presence or absence of a target short genetic variant in a test sample using the one or more processors.
[0033] Also described herein is a system comprising one or more processors; and a non-transitory computer-readable medium storing one or more programs comprising instructions for implementing any of the methods described above.
[0034] In some embodiments of any of the methods or systems described above, the separate flows include a single base type.
[0035] In some embodiments of any of the methods or systems described above, at least one of the individually separated flows includes 2 or 3 different base types.
[0036] In some embodiments of any of the methods or systems described above, the method includes generating or updating a variant call file indicating the presence, identity, or absence of short genetic variants in the test sample.
[0037] In some embodiments of any of the methods or systems described above, the method includes generating a report indicating the presence, identity, or absence of the short genetic variant in the test sample. In some embodiments, the report includes a text, probability, numeric, or graphical output indicating the presence, identity, or absence of the short genetic variant in the test sample. In some embodiments, the method includes providing the report to the patient or a health care representative of the patient.
[0038] In some embodiments of any of the methods or systems described above, the short genetic variants comprise single nucleotide polymorphisms.
[0039] In some embodiments of any of the methods or systems described above, the short genetic variants include insertions and deletions (indels).
[0040] In some embodiments of any of the methods or systems described above, the test sample comprises fragmented DNA.
[0041] In some embodiments of any of the methods or systems described above, the test sample comprises cell-free DNA. In some embodiments, the cell-free DNA comprises circulating tumor DNA (ctDNA).
[0042] In some embodiments, the sequencing method of nucleic acid molecules includes hybridizing nucleic acid molecules with primers to form a hybridization template; according to a repeated flow cycle sequence including five or more separate nucleotide streams, using a labeled non-terminated nucleotide extension primer provided in a separate nucleotide stream; and when the primer is extended by the nucleotide stream, detecting a signal from the integrated labeled nucleotide or the absence of a signal. In some embodiments, the method includes detecting the absence of a signal or a signal after each nucleotide stream. In some embodiments, the method includes sequencing a plurality of nucleic acid molecules. In some embodiments, the nucleic acid molecules in the plurality of nucleic acid molecules have different sequencing start positions relative to the locus. In some embodiments, the test sample is cell-free DNA. In some embodiments, the cell-free DNA includes circulating tumor DNA (ctDNA). In some embodiments, for 50% or more of the possible SNP arrangements at at least 5% of the random sequencing start positions, the flow cycle sequence induces a signal change at more than two flow positions. In some embodiments, the induced signal change is a change in signal intensity, or a new substantially zero (or new zero) or a new substantially non-zero (or new non-zero) signal. In some embodiments, the induced signal change is a new substantially zero (or new zero) or a new substantially non-zero (or new non-zero) signal. In some embodiments, the flow cycling sequence has an efficiency of 0.6 or more bases per flow integration. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1AShown is the sequencing data obtained by extending the primer with the sequence TATGGTCGTCGA (SEQ ID NO: 1) by a repeated flow cycle sequence using TACG. The sequencing data represents the extended primer strand, and it can be easily determined that the sequencing information of the complementary template strand is effectively equivalent.
[0044] Figure 1B Shows Figure 1A Sequencing data shown. Given the sequencing data, the most likely sequence was selected based on the highest likelihood at each flow position (as indicated by an asterisk).
[0045] Figure 1C Shows Figure 1A Sequencing data shown, where the traces represent two different candidate sequences: TATGGTCATCGA (SEQ ID NO: 2) (filled circles) and TATGGTCGTCGA (SEQ ID NO: 1) (open circles). The likelihood that the sequencing data matches a given sequence can be determined by multiplying the likelihood that each flow position matches the candidate sequence.
[0046] Figure 2A Shown are the alignment results of sequencing reads R1 (SEQ ID NO:1), R2 (SEQ ID NO:3) and R3 (SEQ ID NO:4) (each represented by the sequence of an extension primer) aligned with two candidate sequences H1 (SEQ ID NO:5) and H2 (SEQ ID NO:6) (each represented by their complementary sequences). Figure 2B Sequencing data corresponding to R1 are shown, with traces representing H1 (filled circles) and H2 (open circles). Figure 2C Sequencing data corresponding to R2 are shown, with traces representing H1 (filled circles) and H2 (open circles). Figure 2D Sequencing data corresponding to R3 are shown, with traces representing H1 (filled circles) and H2 (open circles).
[0047] Figure 3 A flow chart of an exemplary method for detecting short genetic variants in a test sample is shown.
[0048] Figure 4A shows sequencing data of a nucleic acid molecule having an extended primer sequence of TATGGTCGTCGA (SEQ ID NO: 1) obtained by sequencing the nucleic acid molecule using a first flow cycle order (TACG), and Figure 4B The sequencing data obtained by sequencing the same nucleic acid molecule using the second flow cycle order (AGCT) are shown. In addition, Figure 4A and Figure 4BThe traces from the first candidate sequence TATGGTCGTCGA (SEQ ID NO: 1) (filled circles) and the second candidate sequence TATGGTCATCGA (SEQ ID NO: 2) (open circles) are shown respectively. Figure 4A and Figure 4B As shown, differences in the order of flow cycles can drastically change the signal detected at a given flow position, and for variant content, more significant signal differences can be detected when better flow cycles are used.
[0049] Figure 5 Another exemplary method for detecting the presence or absence of short genetic variants in a test sample is shown.
[0050] Figure 6 Another exemplary method for detecting the presence or absence of short genetic variants in a test sample is shown.
[0051] Figure 7 An example of a computing device according to one embodiment is illustrated, which may be used to implement the methods described herein.
[0052] Figure 8 Sequencing data of a putative nucleic acid molecule sequenced using ATGC flow cycle sequence is shown. Traces can be generated using potential haplotype sequences TATGGTCG-TCGA (SEQ ID NO: 7) (H1) and TATGGTCGATCG (SEQ ID NO: 8) (H2), where H1 has a 1 base deletion relative to H2. The sequencing data has a better match with the H2 candidate sequence, and no indels were determined in this sequence.
[0053] Fig. 9 Shown are the sensitivity of detected SNP permutations given random sequencing start positions for four exemplary flow cycle sequences, including three of which are extended flow cycle sequences. Fig. 9 In , the x-axis indicates the fraction of mobile phase (or fragmentation starting position), while the y-axis indicates the fraction of SNP arrangements with induced signal changes at more than two mobile positions. DETAILED DESCRIPTION OF THE INVENTION
[0055] Described herein are methods for detecting one or more short genetic variants, such as single nucleotide polymorphisms (SNPs), multinucleotide polymorphisms (MNPs), or indels, in a test sample derived from a subject. Test sequencing data associated with a test nucleic acid molecule from the test sample is analyzed to determine a match between the test sequencing data and another sequence, such as a test sequence, a candidate sequence (or a candidate haplotype sequence and / or a reference sequence), which can be reflected by determining a match score indicating the proximity of the match (e.g., given the test sequencing data, the likelihood that the test sequencing data is from a nucleic acid molecule of a comparison sequence). The match score can then be used to determine the presence or identity or absence of a short genetic variant in the test sample.
[0056] The test sequencing data set is uniquely structured to provide computationally efficient analysis. For example, the test sequencing data set can be generated by sequencing the test nucleic acid molecule using the non-terminated nucleotides provided in the nucleotide stream separated according to the flow cycle order. Then, the test sequencing data set of the nucleic acid molecule includes the flow signal at the flow position, and each flow position corresponds to the flow of a specific nucleotide. Using this uniquely structured data set, a nucleic acid molecule (or multiple nucleic acid molecules) can be analyzed in "flow space" rather than "base space" (also referred to as "nucleotide space" or "sequence space"). The flow space data depends on the additional information related to the flow cycle order, which is not carried by the basic space data. The analysis of the data collected in the flow space provides at least two advantages that are better than the analysis of the data collected in the base space or in the base space. First, when compared with the reference sequence in the flow space, the most common variant type (replacing SNP) in the test nucleic acid molecule will produce two or more different flow signals (which can propagate the entire flow cycle or more), and when analyzing the sequence in the base space, only one data signal is available. That is to say, in base space, each base position is associated with a single signal, and the variant base only affects the signal of the variant base without affecting adjacent signals. In flow space, variants can affect multiple flow positions, and for some variants, variants can induce the displacement of subsequent flow graph signals relative to the reference sequence, thereby forming a continuous enhancement of effective variant detection. Secondly, flow space data can be analyzed to determine the match with one or more candidate flow space sequences without the need to test the direct comparison between the sequence of the nucleic acid molecule and one or more candidate sequences. Sequence alignment is computationally expensive and can be simplified using matching analysis as described herein.
[0057] The multi-signal indicator in the flow space of a given genetic variant improves the variant determination accuracy of a single signal indicator that can be identified in base space analysis. In addition, a greater number of flow signal differences increase the possibility of detecting a variant determination. As further discussed herein, in some cases, it is desirable to determine a pre-selected variant with high confidence, and those variants and / or flow orders can be selected to determine genetic variants confidently to ensure that the flow signal differences of the desired number are generated. The sequencing data set of nucleic acid molecules can be compared with the candidate sequence to determine the matching score of the possibility of indicating the test sequencing data set to match the candidate sequence.
[0058] The comparison of the sequence determined in base space and the candidate sequence (such as candidate haplotype sequence) is computationally expensive, and is the most intensive step in the current genome analysis tool kit (GATK) HaplotypeCaller. In HaplotypeCaller, PairHMM compares each sequencing read with each haplotype, and uses base quality as the estimation of error to determine the possibility of the haplotype of a given sequencing read. However, the structure of the data set used together with the method described herein retains the possibility of error pattern, which makes variant determination more efficient computationally. For example, a given genotype possibility can be simply determined as the product of the possibility in each flow position with the sequence comparison of genotype. The possibility determined in flow space can replace the PairHMM module of HaplotypeCaller for more efficient variant determination computationally.
[0059] The flow signal of any flow position in the sequencing data set is flow order-related, because the flow order for sequencing the nucleic acid molecules at any base position can affect the flow signal at this position.As further described herein, this discovery can be utilized in one or more ways.First, the random fragmentation (fragmentation in vivo, such as cell-free DNA, or in vitro fragmentation, such as by ultrasonic treatment or enzymatic digestion) of the nucleic acid molecules overlapping at the same locus causes a plurality of different sequencing start sites (relative to the locus) of nucleic acid molecules.In some cases, different flow contents are available at the locus (for example, when reordering with different flow orders, or when using quasi-periodic flow orders).Therefore, even if other nucleic acid molecules cause lower confidence signals (for example, single flow signal changes), variants at the locus can also be accurately detected based on single nucleic acid molecules, which have a high sensitivity flow signal for variants (for example, compared with reference or unselected candidate sequences, there are two or more flow signal differences).Secondly, the first flow order can be used to order given nucleic acid molecules, and the second (different) flow order is used to reorder, thereby different flow order contents are provided on nucleic acid molecules. If the possibility that a nucleic acid molecule with a variant matches a candidate sequence with a variant is low using one flow order, the possibility that a nucleic acid molecule matches a candidate sequence may be high using a second flow order. Third, the flow order can be an extended flow cycle (e.g., having more than four base types in the cycle), which means that the four flows of more than four base types A, C, T and G are repeated periodically. In some cases, the recombinant unit is longer than four bases, such as including all possible two-base flow sequences (i.e., all XY pairs are within the repeating unit, where X is all four bases, and Y is each non-X base) or three-base flow sequences (i.e., all possible XYZ arrangements are within the repeating unit). Fourth, the flow sequencing order can be selected to target specific genetic variants.
[0060] In some embodiments, a method for detecting short genetic variants in a test sample includes: (a) obtaining one or more test sequencing data sets, each test sequencing data set being associated with a test nucleic acid molecule derived from the test sample, wherein the test sequencing data set is generated by sequencing the test nucleic acid molecule using non-terminal nucleotides provided in separate nucleotide flows according to a flow order, and wherein the test sequencing data set includes flow signals at flow positions corresponding to the flow of the nucleic acid; (b) for each test nucleic acid molecule associated with the test sequencing data set, determining a match score indicating the likelihood that the test sequencing data set matches one or more candidate sequences; and (c) using the one or more determined match scores to determine the presence or absence of the target short genetic variant in the test sample.
[0061] In some embodiments, a method for detecting short genetic variants in a test sample includes: (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order, the target sequencing dataset associated with the target sequence including the target short genetic variant is different from the reference sequencing dataset associated with the reference sequence, wherein the flow position corresponds to the nucleotide flow; (b) obtaining one or more test sequencing datasets, each test sequencing dataset associated with a test nucleic acid molecule, each test nucleic acid molecule at least partially overlapping with a locus associated with the target short genetic variant and derived from the test sample, wherein the one or more test sequencing datasets are determined by sequencing the test nucleic acid molecules using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order, and wherein the test sequencing dataset includes flow signals at multiple flow positions; (c) for each test nucleic acid molecule associated with the test sequencing dataset, determining a match score indicating the likelihood that the test sequencing dataset associated with the nucleic acid molecule matches the target sequence, or a match score indicating the likelihood that the test sequencing dataset associated with the nucleic acid molecule matches the reference sequence; and (d) using one or more determined match scores to determine the presence or absence of the target short genetic variant in the test sample.
[0062] In some embodiments, a method for detecting short genetic variants in a test sample includes: (a) obtaining one or more first test sequencing data sets, each first test sequencing data set associated with a different test nucleic acid molecule derived from the test sample, wherein the first test sequencing data set is determined by sequencing the one or more test nucleic acid molecules according to a first flow cycle order using non-termination nucleotides provided in separate nucleotide flows, and wherein the one or more first test sequencing data sets include flow signals at flow positions corresponding to the nucleotide flows; (b) obtaining one or more second test sequencing data sets, each second test sequencing data set is associated with the same test nucleic acid molecule as the first test sequencing data set, wherein the second test sequencing data set is determined by sequencing the one or more test nucleic acid molecules according to a second flow cycle order using non-termination nucleotides provided in separate nucleotide flows, wherein the first flow cycle order and the second flow cycle order are different, and wherein the test sequencing data sets include flow signals at flow positions corresponding to the nucleotide flows; (c) for each of the first sequencing data set and the second sequencing data set, determining a match score for one or more candidate sequences, wherein the match score indicates a likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from the one or more candidate sequences; and (d) using the determined match score to determine the presence or absence of the short genetic variant in the test sample.
[0063] The method described herein may be a computer-implemented method, and one or more steps of the method may be performed, for example, using one or more computer processors.
[0064] Also provided herein is a non-transitory computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform any one or more methods described herein.
[0065] This document further describes an electronic device, which includes one or more processors, a memory, and one or more programs stored in the memory, the one or more programs configured to be executed by the one or more processors. The one or more programs may include instructions for executing any one or more methods described herein.
[0066] Methods for sequencing nucleic acid molecules are also described herein. For example, a method for sequencing a nucleic acid molecule may include: hybridizing a nucleic acid molecule with a primer to form a hybridization template; extending the primer using a labeled non-terminator nucleotide provided in a separate nucleotide stream according to a repeated flow cycle sequence including five or more separate nucleotide streams; and detecting a signal or the absence of a signal from an incorporated labeled nucleotide as the primer is extended along the nucleotide stream.
[0067] definition
[0068] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0069] Reference herein to "about" a value or parameter includes (and describes) variations with respect to that value or parameter itself. For example, description referring to "about X" includes description of "X".
[0070] "Expected sequencing data" or "expected sequencing data set" for a given sequence refers to the calculated sequencing data that will be generated if the sequence is sequenced according to the flow order using non-terminal nucleotides provided in a separate nucleotide stream. For example, the expected sequencing data or expected sequencing data set can be determined by computer modeling (i.e., in silico).
[0071] "Flow order" refers to the order of individual separate nucleotide flows used to sequence nucleic acid molecules using non-terminator nucleotides. The flow order can be divided into cycles of repeating units, and the flow order of repeating units is called the "flow cycle order". "Flow position" refers to the sequential position of a given individual nucleotide flow during the sequencing process.
[0072] The terms "individual," "patient," and "subject" are used synonymously and refer to animals, including humans.
[0073] As used herein, the term "label" refers to a detectable portion that is or can be coupled to another portion (e.g., a nucleotide or nucleotide analog). The label can emit a signal or change the signal transmitted to the label so that the presence or absence of the label can be detected. In some cases, the coupling can be carried out by a linker, which can be cleavable, such as photocleavable (e.g., cleavable under ultraviolet light), chemically cleavable (e.g., by a reducing agent, such as dithiothreitol (DTT), tris (2-carboxyethyl) phosphine (TCEP)) or enzymatically cleavable (e.g., by an esterase, lipase, peptidase, or protease). In some embodiments, the label is a fluorophore.
[0074] A "non-terminal nucleotide" is a nucleic acid moiety that can be attached to the 3' end of a polynucleotide using a polymerase or transcriptase and to which another non-terminal nucleic acid can be attached using a polymerase or transcriptase without removing a protecting group or reversible terminator from the nucleotide. Naturally occurring nucleic acids are a class of non-terminal nucleic acids. Non-terminal nucleic acids can be labeled or unlabeled.
[0075] "Nucleotide flow" refers to a set of one or more non-terminating nucleotides (which may be labeled or a portion of which may be labeled).
[0076] "Short genetic variant" is used herein to describe a genetic polymorphism (i.e., mutation) of 10 consecutive bases or less (i.e., 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base in length). The term includes single nucleotide polymorphisms (SNPs), multiple nucleotide polymorphisms (MNPs), and indels of 10 consecutive bases or less in length.
[0077] It should be understood that the aspects and variations of the present invention described herein include "consisting of" and / or "consisting essentially of" the aspects and variations.
[0078] When providing a range of values, it will be understood that each intermediate value between the upper and lower limits of the range, as well as any other stated or intermediate values in the state range, are included within the scope of the present disclosure. Where the stated range includes an upper or lower limit, ranges excluding any of those included limits are also included in the present disclosure.
[0079] Some analysis methods described herein include mapping a sequence to a reference sequence, determining sequence information, and / or analyzing sequence information. It is well understood in the art that complementary sequences can be easily determined and / or analyzed, and the description provided herein encompasses analysis methods performed with reference to complementary sequences.
[0080] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described. This description is presented to enable one of ordinary skill in the art to make and use the invention, and is provided in the context of a patent application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the embodiments shown, but to conform to the widest scope consistent with the principles and features described herein.
[0081] The accompanying drawings show processes according to various embodiments. In the exemplary process, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some examples, additional steps can be combined with the exemplary process. Therefore, the operations shown (and described in more detail below) are exemplary in nature and should not be considered as limiting.
[0082] The disclosures of all publications, patents, and patent applications mentioned herein are each incorporated herein by reference in their entirety. In the event of a conflict between any reference incorporated by reference and the present disclosure, the present disclosure shall control.
[0083] Flow Sequencing Method
[0084] Sequencing data can be generated using a flow sequencing method, the method comprising extending a primer bound to a template polynucleotide molecule according to a predetermined flow cycle, in which a single type of nucleotide is accessible to the primer being extended at any given flow position. In some embodiments, at least some specific types of nucleotides include a label, and when the labeled nucleotide is integrated into the extension primer, the label produces a detectable signal. The resulting sequence integrated into the extension primer by such nucleotides should be the reverse complementary sequence of the template polynucleotide molecule sequence. In some embodiments, for example, sequencing data is generated using a flow sequencing method, the method comprising extending the primer using labeled nucleotides, and detecting the presence or absence of the labeled nucleotides integrated into the extension primer. The flow sequencing method may also be referred to as "natural sequencing by synthesis" or "sequencing by synthesis" method. Exemplary methods are described in U.S. Patent No. 8,772,473, which is incorporated herein by reference as a whole. Although the following description is provided with reference to the flow sequencing method, it should be understood that all or part of the sequencing region can be sequenced using other sequencing methods. For example, the sequencing data discussed herein can be generated using a pyrosequencing method.
[0085] Flow sequencing includes using nucleotides to extend primers hybridized with polynucleotides. If there are complementary bases in the template strand, nucleotides of a given base type (e.g., A, C, G, T, U, etc.) can be mixed with the hybridized template to extend the primer. The nucleotides can be, for example, non-terminated nucleotides. When the nucleotide is non-terminated, if there are more than one continuous complementary base in the template strand, more than one continuous base can be integrated into the primer strand being extended. In contrast to non-terminated nucleotides are nucleotides with 3' reversible terminators, in which the blocking group is usually removed before connecting continuous nucleotides. If there are no complementary bases in the template strand, primer extension stops until a nucleotide complementary to the next base in the template strand is introduced. At least a portion of the nucleotides can be labeled so that their integration can be detected. Most commonly, only a single nucleotide type (i.e., discrete addition) is introduced at a time, although two or three different types of nucleotides can be introduced simultaneously in certain embodiments. In contrast to this methodology is a sequencing method using a reversible terminator, in which primer extension stops after each single base extension, and the terminator is reversed to allow the integration of the next subsequent base.
[0086] Nucleotides can be introduced in a flow order during primer extension, which can be further divided into flow cycles. Flow cycles are the repetitive order of nucleotide flows, and can have any length. Nucleotides are added stepwise, which allows the added nucleotides to be integrated into the ends of the sequencing primers with complementary bases in the template strand. As an example only, the flow order of the flow cycle can be ATGC, or the flow cycle order can be ATCG. Those skilled in the art can easily envision alternative sequences. The flow cycle order can be any length, although the flow cycle containing four unique base types (A, T, C and G in any order) is the most common. In some embodiments, the flow cycle includes 5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20 or more nucleotide flows separated in the flow cycle order. For example only, the flow cycle order can be TCACGATGCATGCTAG, wherein these 16 nucleotides provided separately are provided for several cycles with this flow cycle order. Between the introduction of different nucleotides, unintegrated nucleotides can be removed, for example, by washing the sequencing platform with washing solution.
[0087] By integrating one or more nucleotides at the end of the primer in a template-dependent manner, a polymerase can be used to extend the sequencing primer. In some embodiments, the polymerase is a DNA polymerase. The polymerase can be a naturally occurring polymerase or a synthetic (e.g., mutant) polymerase. A polymerase can be added in the initial step of primer extension, but a supplementary polymerase can be optionally added during sequencing, such as with the gradual addition of nucleotides or after multiple flow cycles. Exemplary polymerases include DNA polymerases, RNA polymerases, thermostable polymerases, wild-type polymerases, modified polymerases, Bst DNA polymerases, Bst 2.0 DNA polymerases, Bst 3.0 DNA polymerases, Bsu DNA polymerases, Escherichia coli DNA polymerase I, T7 DNA polymerases, bacteriophage T4 DNA polymerases, Φ29 (phi29) DNA polymerases, Taq polymerases, Tth polymerases, Tli polymerases, Pfu polymerases, and SeqAmp DNA polymerases.
[0088] When determining the sequence of the template strand, the introduced nucleotides may include labeled nucleotides, and the presence or absence of the integrated labeled nucleic acid may be detected to determine the sequence. The label may be, for example, an optically active label (e.g., a fluorescent label) or a radioactive label, and a detector may be used to detect the signal emitted or changed by the label. The presence or absence of the labeled nucleotides integrated into the primers hybridized with the template polynucleotide may be detected, which allows the determination of the sequence (e.g., by generating a flow graph). In some embodiments, the labeled nucleotides are labeled with fluorescent, luminescent or other luminescent moieties. In some embodiments, the label is connected to the nucleotides via a linker. In some embodiments, the linker is cleavable, for example, by photochemical or chemical cleavage reactions. For example, the label may be cut after detection and before the integration of continuous nucleotides. In some embodiments, the label (or linker) connects the nucleotide bases, or connects another site on the nucleotide, without interfering with the extension of the nascent chain of the DNA. In some embodiments, the linker includes a disulfide or a PEG-containing part.
[0089] In some embodiments, the introduced nucleotides include only unlabeled nucleotides, and in some embodiments, the nucleotides include a mixture of labeled and unlabeled nucleotides. For example, in some embodiments, the labeled nucleotide portion is about 90% or less, about 80% or less, about 70% or less, about 60% or less, about 50% or less, about 40% or less, about 30% or less, about 20% or less, about 10% or less, about 5% or less, about 4% or less, about 3% or less, about 2.5% or less, about 2% or less, about 1.5% or less, about 1% or less, about 0.5% or less, about 0.25% or less, about 0.1% or less, about 0.05% or less, about 0.025% or less, or about 0.01% or less compared to total nucleotides. In some embodiments, the portion of labeled nucleotides compared to the total nucleotides is about 100%, about 95% or more, about 90% or more, about 80% or more, about 70% or more, about 60% or more, about 50% or more, about 40% or more, about 30% or more, about 20% or more, about 10% or more, about 5% or more, about 4% or more, about 3% or more, about 2.5% or more, about 2% or more, about 1.5% or more, about 1% or more, about 0.5% or more, about 0.25% or more, about 0.1% or more, about 0.05% or more, about 0.025% or more, or about 0.01% or more. In some embodiments, the portion of labeled nucleotides compared to the total nucleotides is from about 0.01% to about 100%, such as from about 0.01% to about 0.025%, from about 0.025% to about 0.05%, from about 0.05% to about 0.1%, from about 0.1% to about 0.25%, from about 0.25% to about 0.5%, from about 0.5% to about 1%, from about 1% to about 1.5%, from about 1.5% to about 2%, from about 2% to about 1.5 ... About 2.5%, about 2.5% to about 3%, about 3% to about 4%, about 4% to about 5%, about 5% to about 10%, about 10% to about 20%, about 20% to about 30%, about 30% to about 40%, about 40% to about 50%, about 50% to about 60%, about 60% to about 70%, about 70% to about 80%, about 80% to about 90%, about 90% to less than 100%, or about 90% to about 100%.
[0090] Before generating sequencing data, polynucleotides are hybridized with sequencing primers to produce hybridization templates. Polynucleotides can be connected with joints (adapters) during sequencing library preparation. Joints can include hybridization sequences that hybridize with sequencing primers. For example, the hybridization sequence of joints can be a uniform sequence across multiple different polynucleotides, and the sequencing primer can be a uniform sequencing primer. This allows multiple sequencing of different polynucleotides in sequencing libraries.
[0091] Polynucleotides can be attached to a surface (e.g., a solid support) for sequencing. Polynucleotides can be amplified (e.g., by bridge amplification or other amplification techniques) to generate polynucleotide sequencing colonies. The polynucleotides amplified in the cluster are substantially identical or complementary (some errors may be introduced during the amplification process so that a part of the polynucleotides may not necessarily be identical to the original polynucleotides). Colony formation allows signal amplification, so that the detector can accurately detect the integration of the labeled nucleotides of each colony. In some cases, colonies are formed on beads using emulsion PCR, and beads are distributed on the sequencing surface. Examples of systems and methods for sequencing can be found in U.S. Patent No. 10,344,328, which is incorporated herein by reference in its entirety.
[0092] Primers hybridized to the polynucleotide are extended through the nucleic acid molecules using separate separate streams of nucleotides according to a flow order (which may be cyclic according to a flow cycle order), and incorporation of the nucleotides may be detected as described above, thereby generating a sequencing data set for the nucleic acid molecules.
[0093] Primer extension using flow sequencing allows length to be hundreds or even thousands of base magnitude long range sequencing. The number of flow steps or cycles can be increased or decreased to obtain the required sequencing length. The extension of primers in the first region or the third region can include one or more flow steps of primers that are progressively extended using nucleotides with one or more different base types. In some embodiments, the extension of primers includes 1 to approximately 1000 flow steps, such as 1 to approximately 10 flow steps, approximately 10 to approximately 20 flow steps, approximately 20 to approximately 50 flow steps, approximately 50 to approximately 100 flow steps, approximately 100 to approximately 250 flow steps, approximately 250 to approximately 500 flow steps or approximately 500 to approximately 1000 flow steps. The flow steps can be divided into identical or different flow cycles. The number of bases integrated into the primer depends on the sequence of the sequencing region, and the flow order for extending the primer. In some embodiments, the sequencing region is about 1 base to about 4000 bases long, such as about 1 base to about 10 bases long, about 10 bases to about 20 bases long, about 20 bases to about 50 bases long, about 50 bases to about 100 bases long, about 100 bases to about 250 bases long, about 250 bases to about 500 bases long, about 500 bases to about 1000 bases long, about 1000 bases to about 2000 bases long, or about 2000 bases to about 4000 bases long.
[0094] The polynucleotides used in the methods described herein can be obtained from any suitable biological source, such as tissue samples, blood samples, plasma samples, saliva samples, fecal samples or urine samples. The polynucleotides can be DNA or RNA polynucleotides. In some embodiments, before the polynucleotides are hybridized with sequencing primers, the RNA polynucleotides are reverse transcribed into DNA polynucleotides. In some embodiments, the polynucleotides are cell-free DNA (cfDNA), such as circulating tumor DNA (ctDNA) or fetal cell-free DNA. Nucleic acid molecules can be randomly fragmented, for example, in vivo (e.g., as in cfDNA) or in vitro (e.g., by ultrasound or enzyme fragmentation).
[0095] The polynucleotide library can be prepared by known methods. In some embodiments, the polynucleotide can be connected to a linker sequence. The linker sequence can include a hybridization sequence that hybridizes to a primer extended during the generation of coupled sequencing read pairs.
[0096] In some embodiments, sequencing data is obtained without amplifying nucleic acid molecules before establishing sequencing colonies (also referred to as sequencing clusters). Methods for generating sequencing colonies include bridge amplification or emulsion PCR. Methods that rely on shotgun sequencing and determine consensus sequences typically use unique molecular identifiers (UMIs) to label nucleic acid molecules and amplify nucleic acid molecules to generate many copies of the same independently sequenced nucleic acid molecules. The amplified nucleic acid molecules can then be connected to a surface and bridge amplified to generate independently sequenced sequencing clusters. UMIs can then be used to associate independently sequenced nucleic acid molecules. However, the amplification process may introduce errors into nucleic acid molecules, for example due to the limited fidelity of DNA polymerases. In some embodiments, nucleic acid molecules are not amplified before amplification to generate colonies for obtaining sequencing data. In some embodiments, nucleic acid sequencing data is obtained without using unique molecular identifiers (UMIs).
[0097] Sequencing datasets and variant detection
[0098] Sequencing data can be generated based on the detection of the nucleotides integrated and the order in which the nucleotides are introduced. For example, take the extended sequence of flow (i.e., each reverse complementary sequence of the corresponding template sequence): CTG, CAG, CCG, CGT and CAT (assuming that there is no preceding sequence or in the following sequence for sequencing method), and the repeated flow cycle of TACG (i.e., sequentially adding T, A, C and G nucleotides in repeated cycles). Only when there are complementary bases in the template polynucleotide, the nucleotides of the specific type at the given flow position will be integrated into the primer. An exemplary resulting flow diagram is shown in Table 1, wherein 1 indicates the integration of the nucleotides introduced and 0 indicates that the nucleotides introduced are not incorporated. The flow diagram can be used to derive the sequence of the template chain. For example, the sequencing data discussed herein (e.g., flow diagram) represents the sequence of the primer chain extended, and its reverse complementary sequence can be easily determined as the sequence representing the template chain. The asterisk (*) in Table 1 indicates that if other nucleotides are integrated in the extended sequencing chain (e.g., a longer template chain), there may be a signal in the sequencing data.
[0099] Table 1
[0100]
[0101]
[0102] Flow graphs can be binary or non-binary. Binary flow graphs detect the presence (1) or absence (0) of integrated nucleotides. Non-binary flow graphs can more quantitatively determine the number of nucleotides that are gradually introduced into integration each time. For example, the extension sequence of CCG will include the integration of two C bases in the extension primer within the same C flow (for example, at flow position 3), and the signal emitted by the base of the mark will have an intensity greater than the intensity level corresponding to the integration of a single base. This is shown in Table 1. Non-binary flow graphs also indicate the presence or absence of bases, and can provide additional information, including the number of bases that may be integrated into each extension primer at a given flow position. These values do not need to be integers. In some cases, these values can reflect the uncertainty and / or possibility of the number of bases integrated at a given flow position.
[0103] In some embodiments, the sequencing data set includes a flow signal representing a base count, which indicates the number of bases in the sequenced nucleic acid molecule integrated at each flow position. For example, as shown in Table 1, a primer extended with a CTG sequence using a TACG flow cycle sequence has a value of 1 at position 3, indicating that the base count at this position is 1 (1 base is C, which is complementary to the G in the template strand being sequenced). Also in Table 1, a primer extended with a CCG sequence using a TACG flow cycle sequence has a value of 2 at position 3, indicating that the base count of the extended primer at this position during this flow position is 2. Here, 2 bases refer to the CC sequence at the start of the CCG sequence in the extension primer sequence, and it is complementary to the GG sequence in the template strand.
[0104] The flow signal in the sequencing data set may include one or more statistical parameters, which indicate the possibility or confidence interval of one or more base counts at each flow position. In some embodiments, the flow signal is determined by the analog signal detected during the sequencing process, such as the fluorescent signal of one or more bases integrated into the sequencing primer during sequencing. In some cases, the analog signal can be processed to generate statistical parameters. For example, a machine learning algorithm can be used to correct the contextual effect of the analog sequencing signal, as described in the disclosed international patent application WO2019084158A1, which is incorporated herein by reference as a whole. Although zero or more bases of an integer are integrated at any given flow position, a given analog signal may not be fully matched with the analog signal. Therefore, given a detected signal, it is possible to determine the statistical parameters indicating the possibility of the number of bases integrated at the flow position. For example only, for the CCG sequence in Table 1, the flow signal indicates that the possibility of integrating 2 bases at flow position 3 can be 0.999, and the flow signal indicates that the possibility of integrating 1 base at flow position 3 can be 0.001. The sequencing data set can be formatted as a sparse matrix, where the flow signal includes statistical parameters indicating the likelihood of multiple base counts at each flow position. By way of example only, a primer extended with the following sequence using repeated flow cycles of TACG: TATGGTCGTCGA (SEQ ID NO: 1) can generate Figure 1A . The statistical parameter or likelihood value may vary, for example, based on noise or other artifacts present during detection of the analog signal during sequencing. In some embodiments, if the statistical parameter or likelihood is below a predetermined threshold, the parameter may be set to a predetermined non-zero value (i.e., some very small or negligible value) that is substantially zero to assist in the statistical analysis discussed further herein, wherein a true zero value may cause computational errors or insufficiently distinguish levels of improbability, e.g., very unlikely (0.0001) and incredible (0).
[0105] A value indicating the likelihood of a given sequence in a sequencing data set can be determined from a sequencing data set without sequence alignment. For example, given the data, the most likely sequence can be determined by selecting the base count with the highest likelihood at each flow position, such as Figure 1B As shown by the star in Figure 1A ). Therefore, the sequence for primer extension can be determined based on the most likely base count at each flow position: TATGGTCGTCGA (SEQ ID NO: 1). From this, the reverse complementary sequence (i.e., the template strand) can be easily determined. Furthermore, given the TATGGTCGTCGA (SEQ ID NO: 1) sequence (or the reverse complementary sequence), the likelihood of this sequencing data set can be determined as the product of the likelihoods selected at each flow position.
[0106] The sequencing data set associated with the nucleic acid molecule can be compared with one or more (e.g., 2, 3, 4, 5, 6 or more) possible candidate sequences. The close match between the sequencing data set and the candidate sequence (based on the matching score, as discussed below) indicates that the sequencing data set may come from a nucleic acid molecule having the same sequence as the candidate sequence of the close match. In some embodiments, the sequence of the nucleic acid molecule sequenced can be mapped to a reference sequence (e.g., using Burrows-Wheeler alignment (BWA) algorithm or other suitable alignment algorithms) to determine the locus (or one or more loci) of the sequence. As described above, the sequencing data set in the flow space can be easily converted to base space (or vice versa if the flow order is known), and can be mapped in flow space or base space. The locus (or multiple loci) corresponding to the mapping sequence can be associated with one or more variant sequences, and the variant sequence can be operated as a candidate sequence (or haplotype sequence) of the analytical method described herein. An advantage of the method described herein is that in some cases, the sequence of the nucleic acid molecule sequenced does not need to be aligned with each candidate sequence using an alignment algorithm, which is usually expensive in calculation. Instead, the matching score for each candidate sequence can be determined using sequencing data in flow space, which is a more computationally efficient operation.
[0107] The match score indicates the degree to which the sequencing data set supports the candidate sequence. For example, given the expected sequencing data for the candidate sequence, a match score indicating the likelihood that the sequencing data set matches the candidate sequence can be determined by selecting a statistical parameter (e.g., likelihood) at each flow position that corresponds to the base count at that flow position. The product of the selected statistical parameters can provide the match score. For example, assuming that Figure 1A Sequencing datasets of extended primers shown in , and TATGGTC ACandidate primer extension sequence of TCGA (SEQ ID NO: 2). Figure 1C (show Figure 1A The same sequencing dataset in Figure 2 shows the traces of candidate sequences (solid circles). For comparison, TATGGTC G TCGA (SEQ ID NO: 1) sequence (see Figure 1B ) trace in Figure 1C The match score indicating the likelihood that the sequencing data matches the first candidate sequence TATGGTCATCGA (SEQ ID NO: 2) is substantially different from the match score indicating the likelihood that the sequencing data matches the second candidate sequence TATGGTCGTCGA (SEQ ID NO: 1), even though the sequences vary by only a single base change. Figure 1C As shown in , the difference between the traces is observed at flow position 12 and propagates for at least 9 flow positions (and possibly longer if the sequencing data extends across additional flow positions). This continued propagation across one or more flow cycles may be referred to as a "flow shift" or "cycle shift" and is generally a very unlikely event if the sequencing data set matches the candidate sequence.
[0108] A match score between each sequencing data set and the candidate sequence (or each candidate sequence) can then be determined. For example, the likelihood (e.g., the product) of the selected base counts at each flow position of the given candidate sequence can be used to determine the likelihood L(R) of the sequencing data set matching the given candidate sequence. j |H i ).
[0109] Matching score can be used for classifying test sequencing data and / or nucleic acid molecules related to test sequencing data.Classifier can indicate that nucleic acid molecules include variants (for example, variants included in candidate sequence), nucleic acid molecules do not include variants, or can indicate empty judgment.Empty judgment neither indicates the presence or absence of variants in nucleic acid molecules related to test sequencing data, but indicates that matching score can not be used for determination with required statistical confidence.For example, if matching score is higher than expected confidence threshold, test sequencing data or nucleic acid molecules can be classified as having variants.On the contrary, for example, if matching score is lower than expected confidence threshold, test sequencing data or nucleic acid molecules can be classified as not having variants.
[0110] The above analysis can be applied to select a candidate sequence from two or more different candidate sequences. A match score indicating the likelihood that a sequencing data set matches each candidate sequence can be determined. For example, a statistical parameter corresponding to the base count of the candidate sequence at each flow position in the sequencing data set can be selected for each candidate sequence. In some embodiments, such analysis includes generating expected sequencing data for candidate sequencing, assuming that the candidate sequence is sequenced using the same flow order used to generate a sequencing data set of a test nucleic acid molecule for sequencing. This can be generated by sequencing a nucleic acid molecule with a candidate sequence, or by generating a candidate sequencing data set on a computer based on the candidate sequence and the flow order. An exemplary candidate sequencing data set is shown in Figure 1C , where the first candidate sequence (TATGGTCATCGA (SEQ ID NO: 2)) corresponds to the solid circle trace and the second candidate sequence (TATGGTCGTCGA (SEQ ID NO: 1)) corresponds to the hollow circle trace. In some embodiments, for example, if the match scores of two or more different candidate sequences are determined, the test sequencing data or nucleic acid molecules can be classified as having a variant of one of the two or more candidate sequences, not having a variant of one of the two or more candidate sequences, or a null call can be made between the two or more candidate sequences (e.g., if no call can be made for any candidate sequence, or if the match scores indicate two or more different variants at the same locus).
[0111] Once the match score of the sequencing data set of the candidate sequence is determined, the candidate sequence with short genetic variants (e.g., the candidate sequence with the highest possible match score generated from two or more candidate sequences) can be selected based on the match score. Short genetic variants can be variants or mutations found, for example, in individual subgroups, or variants or mutations unique to a single or specific individual. Short genetic variants can be germline variants or somatic variants. The sequencing data produced by the sequence nucleic acid molecules with short genetic variants will match the candidate sequence with short genetic variants, and the candidate sequence can be selected, while the rejected (or unselected) candidate sequence does not include short genetic variants, as indicated by the less likely match (based on the determined match score of those candidate sequences). Unselected candidate sequences can be different from the selected candidate sequence (which best matches the nucleic acid molecule sequencing data set for sequencing) at two or more flow positions, and the two or more flow positions can be two or more continuous flow positions or two or more non-continuous flow positions. In some embodiments, unselected candidate sequence is different from selected candidate sequence at 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more flow positions.In some embodiments, unselected candidate sequence is different from selected candidate sequence in 1 or more, 2 or more, 3 or more, 4 or more or 5 or more flow cycles.In some embodiments, unselected candidate sequence is different from selected candidate sequence at X base positions, and the sequencing data set related to sequence nucleic acid molecule is different from unselected candidate sequence at (X+2) or more flow positions.The increase of the quantity of different flow positions between selected candidate sequence and unselected candidate sequence (wherein the nucleic acid molecule sequencing data set of sequencing and selected candidate sequence best match) reduces the possibility of the nucleic acid molecule sequencing data set of sequencing obtained from using unselected candidate sequence to sequence nucleic acid molecule.
[0112] The probability that the sequencing data set of the sequenced nucleic acid molecule matches the unselected candidate sequence is preferably low, such as less than 0.05, less than 0.04, less than 0.03, less than 0.02, less than 0.01, less than 0.005, less than 0.001, less than 0.0005 or less than 0.0001. The probability that the sequencing data set of the sequenced nucleic acid molecule matches the selected candidate sequence is preferably high, such as greater than 0.95, greater than 0.96, greater than 0.97, greater than 0.98, greater than 0.99, greater than 0.995 or greater than 0.999.
[0113] In some embodiments, a method for detecting short genetic variants in a test sample can include analyzing multiple test sequencing data sets, wherein each test sequencing data set is associated with a separate test nucleic acid molecule in the test sample. For example, if the sequence of the nucleic acid molecule is aligned with a reference sequence, the nucleic acid molecules at least partially overlap at the locus. At least a portion of the nucleic acid molecules can have different sequencing start positions (relative to the locus), which results in different flow positions and / or different flow order environments for a given base within the sequence. In this way, the same candidate sequence can be used to analyze multiple test sequencing data sets. For each candidate sequence, a match score indicating the likelihood that multiple test sequencing data sets match the candidate sequence can be determined, and the candidate sequence with the highest likelihood of matching (and therefore, including the short genetic variant) can be selected. An exemplary analysis of detecting short genetic variants using multiple test sequencing data sets is shown in FIG. Figures 2A-2D In. Figure 2A In , sequences corresponding to three sequenced test nucleic acid molecules (R1, R2 and R3, each represented by the sequence of an extended primer) are aligned to a reference sequence at overlapping loci associated with two candidate sequences (HI and H2). Figure 2B , Figure 2C and Figure 2D Exemplary sequencing data sets for R1, R2, and R3 are shown, respectively, along with selected statistical parameters at each flow position in the sequencing data set corresponding to a base of H1 (solid circles) or H2 (open circles).
[0114] One or more determined match scores can be used to determine the presence (or identity) or absence of short genetic variants in a test sample. In some embodiments, for example, a single nucleic acid molecule (or a related test sequencing data set) classified as having a variant may be sufficient to determine the presence, consistency or absence of a variant, for example, if the match score indicates that it matches the candidate sequence with an expected or preset confidence. In some embodiments, a predetermined number (e.g., 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, etc.) of nucleic acid molecules (or a test sequencing data set associated with a nucleic acid molecule) is classified as having a variant before determining a variant for a test sample. In some embodiments, the number of nucleic acid molecules (or a test sequencing data set associated with a nucleic acid molecule) is dynamically selected according to the match score; for example, a single nucleic acid molecule classified as a variant with a high confidence match score can be used to determine a variant, or two or more nucleic acid molecules classified as variants with a lower confidence match score can be used to determine a variant.
[0115] Optionally, the individual match scores of the sequencing data sets are analyzed together to determine the match scores of the multiple test sequencing data sets. For example, once the match score of each test sequencing data set for each candidate sequence is determined using the methods described herein, a known Bayesian method can be used, for example, using the HaplotypeCaller algorithm included in the Genome Analysis Toolkit (GATK), to determine a match score indicating the likelihood that the multiple test sequencing data sets match the candidate sequence, and the candidate sequence with the highest likelihood of matching can be selected. See, e.g., Depristo et al., A framework for variation discovery and genotyping using next-generation DNA sequencing data, Nature Genetics 43, 491-498 (2011); and Poplin et al., Scaling Accurate Genetic Variant Discovery to Tens of Kestomen of Samples, BioRxiv, www.bioRxiv.org / content / 10.1101 / 201178v3 (July 24, 2018); Hwang et al., Systematic Comparison of Variant Calling Pipelines Using Gold Standard Personal Exome Variants, Scientific Reports, Vol. 5, No. 17875 (2015); the contents of each of which are incorporated herein.
[0116] Selection of target variants and / or flow cycle sequence
[0117] Target short genetic variants can be selected, for example, as a basis for selecting a flow order and / or candidate sequence (i.e., by preselecting target short genetic variants), or as a basis for downstream analysis. Downstream analysis can include, for example, assembling a biomarker panel including the identified short genetic variants. The biomarker panel can be personalized for an individual subject associated with a test sample. For example, a biomarker panel can include one or more short genetic variants associated with a disease (e.g., cancer), such as a variant signature. In another example, the biomarker panel is personalized for a subject, including one or more short genetic variants previously detected in a sample from a subject, which can be attributed to a disease (e.g., cancer) of the subject.
[0118] The method for identifying short genetic variants as described herein may be particularly useful when one or more target short genetic variants are selected in advance. The limit of detection (LOD) for a given short genetic variant may depend on the sequence background of the short genetic variant (e.g., the sequence of the nucleic acid molecule flanking the target short genetic variant locus) and the flow order (or flow cycle order) for sequencing the nucleic acid molecule and generating the sequencing data set of the nucleic acid molecule. That is, using a given flow order, short genetic variant and short genetic variant environment, the number of flow position changes in the flow space of nucleic acid molecules with short genetic variants and nucleic acid molecules without short genetic variants (e.g., reference sequences) can be determined. This allows the selection of particularly sensitive variants or the selection of flow orders that can detect specific variants with high sensitivity. The target sequencing data set associated with the target sequence including the target short genetic variant can be compared with the reference sequencing data set associated with the reference sequence without the target short genetic variant to determine the number of flow position differences between the target sequence and the reference sequence. That is, except for the target short genetic variant, the reference sequence is identical to the target sequence. A larger number of flow position differences indicates a higher sensitivity (i.e., detection limit) of the variant. The target and reference sequencing data sets can be determined by actually sequencing nucleic acid molecules having the target sequence and / or nucleic acid molecules having the reference sequence, or the data sets can be expected sequencing data sets (e.g., as determined in silico).
[0119] In one example, a genetic fingerprint of a particular subject or cancer may be desired, but it is not necessary to detect every short genetic variant in the subject or cancer genome. Instead, one or more short genetic variants that have a particularly high sensitivity to a given flow sequence can be pre-selected. By pre-selecting sensitive variants, a lower sequencing depth of the test sample can be used to confidently call variants.
[0120] In some embodiments, a method for detecting a target short genetic variant in a test sample may include selecting a target short genetic variant, wherein when a target sequencing data set is obtained by sequencing a target sequence using non-terminal nucleotides provided in separate nucleotide flows according to a flow cycle order, the target sequencing data set associated with the target sequence including the target short genetic variant is different from a reference sequencing data set associated with a reference sequence at two or more flow positions, wherein the flow position corresponds to a nucleotide flow. In some embodiments, the target sequencing data set is different from the reference sequencing data at two or more non-continuous flow positions. In some embodiments, the target sequencing data set is different from the reference sequencing data at two or more continuous flow positions. In some embodiments, the target sequencing data set is different from the reference sequencing data at three or more flow positions, and the flow positions may be continuous or non-continuous. In some embodiments, the target sequence is different from the reference sequence at X base positions, and wherein the target sequencing data set is different from the reference sequencing data at (X+2) or more continuous flow positions. In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0121] In some embodiments, a method for detecting a target short genetic variant in a test sample may include selecting a target short genetic variant, wherein when a target sequencing data set and a reference sequencing data set are obtained by sequencing a target sequence and a reference sequence using non-terminal nucleotides provided in separate nucleotide flows according to a flow cycle order, the target sequencing data set associated with the target sequence including the target short genetic variant is different from the reference sequencing data set associated with the reference sequence at two or more flow positions, wherein the flow position corresponds to the nucleotide flow. In some embodiments, the target sequencing data set is different from the reference sequencing data at two or more non-continuous flow positions. In some embodiments, the target sequencing data set is different from the reference sequencing data at two or more continuous flow positions. In some embodiments, the target sequencing data set is different from the reference sequencing data at three or more flow positions, and the flow positions may be continuous or non-continuous. In some embodiments, the target sequence is different from the reference sequence at X base positions, and wherein the target sequencing data set may be different from the reference sequencing set at (X+2) or more continuous flow positions. In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0122] The detection of the targeted short genetic variant selected can usually be performed as described above. For example, in some embodiments, a test sequencing data set related to a test nucleic acid molecule of a locus with a target short genetic variant can be obtained. Sequencing data is generated by sequencing the test nucleic acid molecule using the non-terminated nucleotides provided in a separate nucleotide stream according to the same flow cycle sequence for generating the target and reference sequencing data sets. A match score indicating the possibility of the test sequencing data set matching the target sequence with a short genetic variant (or, alternatively or additionally, a match score indicating the possibility of the test sequencing data set matching the reference sequence) is determined, and the determined match score can be used to determine the presence or absence of the target short genetic variant in the test sample.
[0123] In some embodiments, multiple test sequencing data sets are used to detect target short genetic variants in a test sample, wherein each test sequencing data set is associated with a different test nucleic acid molecule in the test sample. The analyzed test nucleic acid molecules overlap at the target short genetic variant locus, and the test nucleic acid molecules are sequenced using the same flow cycle order for selecting the target short genetic variant to generate a data set. A match score indicating the likelihood of a multiple test sequencing data set matching a target sequence with a short genetic variant is determined (or, alternatively or additionally, a match score indicating the likelihood of a multiple test sequencing data set matching a reference sequence), and the determined match score can be used to determine the presence or absence of the target short genetic variant in the test sample.
[0124] In some embodiments, the flow order or flow cycle order for generating sequencing data is preselected. As discussed herein, the content of the variant in the flow order can affect the signal difference between the variant sequence and the (e.g., reference) sequence of comparison. In order to increase the possibility of detecting the selected target variant, the flow order or flow cycle order can be preselected.
[0125] Figure 3A flow chart of an exemplary method for detecting short genetic variants in a test sample is shown. In step 302, a target short genetic variant is selected. The target short genetic variant is selected so that when a target sequencing data set and a reference sequencing data set are obtained by sequencing the target sequence using non-terminal nucleotides provided in separate nucleotide streams according to a flow cycle order, the target sequencing data associated with the target sequence including the target short genetic variant is different from the sequencing data set associated with the reference sequence at more than two flow positions, wherein the flow position corresponds to the nucleotide stream. In step 304, one or more test sequencing data sets are obtained, such as by sequencing one or more test nucleic acid molecules to obtain one or more test sequencing data sets, or by receiving one or more test sequencing data sets. Each of the test sequencing data sets is associated with a test nucleic acid molecule derived from a test sample. In order to analyze the selected target short genetic variant, the test nucleic acid molecule at least partially overlaps with the locus associated with the target short genetic variant. The sequencing data set can be determined (or can have been determined before) by sequencing the test nucleic acid molecule using non-terminal nucleotides provided in separate nucleotide streams according to a flow cycle order, wherein the test sequencing data set includes flow signals at multiple flow positions. In step 306, for each test nucleic acid molecule associated with the test sequencing data set, a match score is determined. The match score indicates the likelihood that the test sequencing data set associated with the nucleic acid molecule matches the target sequence. Alternatively, the match score can indicate the likelihood that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence. In step 308, one or more determined match scores are used to determine the presence or absence of the target short genetic variant in the test sample.
[0126] In some embodiments, a method for detecting a short genetic variant in a test sample comprises: (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing a target sequence using non-termination nucleotides provided in separate nucleotide flows according to a flow cycle order, the target sequencing dataset associated with the target sequence including the target short genetic variant differs from a reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to nucleotide flows; (b) obtaining one or more test sequencing datasets, each test sequencing dataset being associated with a test nucleic acid molecule, each test nucleic acid molecule at least partially overlapping a target sequence associated with the target short genetic variant and derived from the test sequence; The method comprises the steps of: (a) determining a test nucleic acid molecule for a test sample, wherein one or more test sequencing data sets are determined by sequencing the test nucleic acid molecule using non-terminating nucleotides provided in separate nucleotide flows according to a flow cycle order, and wherein the test sequencing data set includes flow signals at the multiple flow positions; (c) for each test nucleic acid molecule associated with the test sequencing data set, determining a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the target sequence, or indicating a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence; and (d) using one or more determined match scores to determine the presence or absence of a target short genetic variant in the test sample. In some embodiments, the method further comprises generating a personalized biomarker group for a subject associated with the test sample, the biomarker group comprising a target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0127] In some embodiments, a method for detecting short genetic variants in a test sample comprises: (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order, the target sequencing dataset associated with the target sequence including the target short genetic variant is different from the reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to the nucleotide streams; (b) sequencing one or more test nucleic acid molecules using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order to obtain one or more test sequencing datasets including flow signals at multiple flow positions, each test sequencing dataset being associated with a test nucleic acid molecule, and each test nucleic acid molecule at least partially overlapping with a locus associated with the target short genetic variant and derived from the test sample; (c) for each test nucleic acid molecule associated with the test sequencing dataset, determining a match score indicating the likelihood that the test sequencing dataset associated with the nucleic acid molecule matches the target sequence, or a match score indicating the likelihood that the test sequencing dataset associated with the nucleic acid molecule matches the reference sequence; and (d) using one or more determined match scores to determine the presence or absence of the target short genetic variant in the test sample. In some embodiments, the method further comprises generating a personalized biomarker panel for a subject associated with the test sample, the biomarker panel comprising a target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0128] In some embodiments, a method for detecting short genetic variants in a test sample comprises: (a) preselecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing a target sequence using non-termination nucleotides provided in separate nucleotide flows according to a flow cycle order, the target sequencing dataset associated with the target sequence including the preselected target short genetic variant is different from the reference sequencing dataset associated with the reference sequence, wherein the flow position corresponds to the nucleotide flow; (b) obtaining one or more test sequencing datasets, each test sequencing dataset being associated with a test nucleic acid molecule, each test nucleic acid molecule at least partially overlapping with a target sequence associated with the preselected target short genetic variant and derived from The loci of the test sample, wherein one or more test sequencing data sets are determined by sequencing the test nucleic acid molecules using non-terminal nucleotides provided in separate nucleotide flows according to the flow cycle order, and wherein the test sequencing data set includes flow signals at multiple flow positions; (c) for each test nucleic acid molecule associated with the test sequencing data set, a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the target sequence is determined, or a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence is determined; and (d) using one or more determined match scores to determine the presence or absence of a pre-selected target short genetic variant in the test sample. In some embodiments, the method also includes generating a personalized biomarker group for a subject associated with the test sample, the biomarker group including a target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0129] In some embodiments, a method for detecting short genetic variants in a test sample includes: (a) pre-selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order, the target sequencing dataset associated with the target sequence including the pre-selected target short genetic variant is different from the reference sequencing dataset associated with the reference sequence, wherein the flow position corresponds to the nucleotide flow; (b) sequencing one or more test nucleic acid molecules using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order to obtain one or more test sequencing datasets including flow signals at multiple flow positions, each test sequencing dataset being associated with a test nucleic acid molecule, and each test nucleic acid molecule at least partially overlapping with a locus associated with the target short genetic variant and derived from the test sample; (c) for each test nucleic acid molecule associated with the test sequencing dataset, determining a match score indicating the likelihood that the test sequencing dataset associated with the nucleic acid molecule matches the target sequence, or a match score indicating the likelihood that the test sequencing dataset associated with the nucleic acid molecule matches the reference sequence; and (d) using one or more determined match scores to determine the presence or absence of the pre-selected target short genetic variant in the test sample. In some embodiments, the method further comprises generating a personalized biomarker panel for a subject associated with the test sample, the biomarker panel comprising a target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0130] In some embodiments, the method for detecting short genetic variants in a test sample comprises: (a) preselecting a target short genetic variant and a flow cycle sequence, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing a target sequence using non-termination nucleotides provided in separate nucleotide flows according to the preselected flow cycle sequence, the target sequencing dataset associated with the target sequence including the preselected target short genetic variant differs from a reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to nucleotide flows; (b) obtaining one or more test sequencing datasets, each test sequencing dataset being associated with a test nucleic acid molecule, each test nucleic acid molecule at least partially overlapping a region corresponding to the preselected target short genetic variant. Variant-related and derived from a locus of a test sample, wherein one or more test sequencing data sets are determined by sequencing the test nucleic acid molecule using non-terminal nucleotides provided in a separate nucleotide flow according to a pre-selected flow cycle order, and wherein the test sequencing data set includes flow signals at multiple flow positions; (c) for each test nucleic acid molecule associated with the test sequencing data set, a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the target sequence is determined, or a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence is determined; and (d) using one or more determined match scores to determine the presence or absence of a pre-selected target short genetic variant in the test sample. In some embodiments, the method also includes generating a personalized biomarker group for a subject associated with the test sample, the biomarker group including a target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0131] In some embodiments, a method for detecting short genetic variants in a test sample includes: (a) preselecting a target short genetic variant and a flow cycle order, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide flows according to the preselected flow cycle order, the target sequencing dataset associated with the target sequence including the preselected target short genetic variant is different from the reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to the nucleotide flows; (b) sequencing one or more test nucleic acid molecules using non-termination nucleotides provided in separate nucleotide flows according to the preselected flow cycle order to obtain one or more test sequencing datasets, each test sequencing dataset being associated with a test nucleic acid molecule, and each test nucleic acid molecule at least partially overlapping with a locus associated with the target short genetic variant and derived from the test sample; (c) for each test nucleic acid molecule associated with the test sequencing dataset, determining a match score indicating the likelihood that the test sequencing dataset associated with the nucleic acid molecule matches the target sequence, or a match score indicating the likelihood that the test sequencing dataset associated with the nucleic acid molecule matches the reference sequence; and (d) using one or more determined match scores to determine the presence or absence of the preselected target short genetic variant in the test sample. In some embodiments, the method further comprises generating a personalized biomarker panel for a subject associated with the test sample, the biomarker panel comprising the target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0132] Selection of target variants and / or flow cycle sequence
[0133] The flow cycle sequence need not be limited to four basic flow cycles (e.g., each of A, G, C, and T, in any repeating order), and can be an extended flow cycle with more than four basic types in the cycle. The extended cycle sequence can be repeated to obtain a desired number of cycles to extend the sequencing primer. For example, in some embodiments, the extended flow sequence includes 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more individual nucleotide streams in the flow cycle sequence. The cycle can include at least one of each of A, G, C, and T, but one or more base types are repeated within the cycle before the cycle is repeated.
[0134] Compared with the flow cycle order with four repeated bases, the extended flow cycle order can be used to detect the small genome variants (such as SNP) of a larger proportion. For example, there are 192 kinds of substitution SNPs with the effective configuration of XYZ→XQZ form, wherein Q≠Y (and Q, X, Y and Z are each any one of A, C, G and T). Among these, 168 can generate new signals (that is, new non-zero signals or new zero signals) in sequencing data sets (for example, flow graphs). In view of the identical tail sequence relative to the reference in the variant, the new zero or non-zero signal combined with the sensitive flow order can generate the signal (for example, flow shift or cyclic shift, which can extend more than the length of the cycle) propagated for multiple flow positions. It should be noted that the insertion or deletion of homopolymers, rather than homopolymer length changes, can cause signal difference propagation. The remaining 24 kinds of variants cause homopolymer length changes at the affected flow positions, but this change does not cause the propagation signal to change. Therefore, the theoretical maximum of 87.5% of SNPs can generate new signals for more than two flow positions that are different from the reference (or candidate) sequence. As described above, the propagated signal differences increase the probability of differences between the test sequencing data set and the incorrectly matched candidate sequence. In addition, the propagated signal changes depend on the flow order across the variant.
[0135] When using flow order to extend sequencing primers, nucleic acid molecules that have been randomly fragmented in the test sample are sequenced to cause random displacement of the content of the variant flow order. That is, the flow position of the variant can be changed according to the starting position of the nucleic acid molecules sequenced. For all 87.5% of SNPs, not all flow cycle combinations can detect signal changes at more than two flow positions, even if all sequencing starting positions in the nucleic acid molecule sequence are utilized. For example, for 41.7% of SNPs, the four-base flow cycle order TACG can cause the test sequencing data set to be different from the reference sequencing data set at more than two flow positions. As further discussed herein, in view of sufficiently high sequencing depth (i.e., sampling sufficiently large starting positions), the extended flow cycle order has been designed so that all theoretical maximum values of SNP (i.e., 87.5% of possible SNPs, or all SNPs except those causing homopolymer length changes) can produce differences at more than two flow positions between the test sequencing data set and the reference sequencing data set.
[0136] The expanded sequencing flow order can have different efficiencies (i.e., the average number of integrations per flow when used to sequence a human reference genome). In some embodiments, the flow order has an efficiency of about 0.6 or greater (e.g., about 0.62 or greater, about 0.64 or greater, about 0.65 or greater, about 0.66 or greater, or about 0.67 or greater). In some embodiments, the flow order has an efficiency of about 0.6 to about 0.7. Examples of flow cycle orders and corresponding estimated efficiencies are shown in Table 2.
[0137] In some embodiments, an extended sequencing flow order is selected to produce a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), and the two sequencing data sets are associated with nucleic acid molecules with different SNPs in about 50% to 87.5% of the SNP arrangements of at least 5% of the random sequencing starting positions. In some embodiments, an extended sequencing flow order is selected to produce a signal difference at more than two flow positions between two sequencing data sets associated with nucleic acid molecules (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), and the two sequencing data sets are associated with nucleic acid molecules with different SNPs in about 60% to 87.5% of the SNP arrangements of at least 5% of the random sequencing starting positions (i.e., "flow phase"). In some embodiments, an extended sequencing flow order is selected to produce a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), and the two sequencing data sets are associated with nucleic acid molecules with different SNPs in about 70% to 87.5% of the SNP arrangements at at least 5% of the random sequencing starting positions. In some embodiments, an extended sequencing flow order is selected to produce a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), and the two sequencing data sets are associated with nucleic acid molecules with different SNPs in about 80% to 87.5% of the SNP arrangements at at least 5% of the random sequencing starting positions.
[0138] In some embodiments, the extended sequencing flow order is selected to produce a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), and the two sequencing data sets are associated with nucleic acid molecules with different SNPs in about 50% to 87.5% of the SNP arrangements of at least 10% of the random sequencing starting positions. In some embodiments, the extended sequencing flow order is selected to produce a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), and the two sequencing data sets are associated with nucleic acid molecules with different SNPs in about 60% to 87.5% of the SNP arrangements of at least 10% of the random sequencing starting positions. The extended sequencing flow order is selected to produce a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), and the two sequencing data sets are associated with nucleic acid molecules with different SNPs in about 70% to 87.5% of the SNP arrangements of at least 10% of the random sequencing starting positions. The extended sequencing flow order is selected to produce a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), wherein the two sequencing data sets are associated with nucleic acid molecules that differ in SNPs in approximately 80% to 87.5% of the SNP arrangements at at least 10% of the random sequencing start positions.
[0139] In some embodiments, the extended sequencing flow order is selected to produce a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), and the two sequencing data sets are associated with nucleic acid molecules with different SNPs in about 50% to 87.5% of the SNP arrangements of at least 20% of the random sequencing starting positions. In some embodiments, the extended sequencing flow order is selected to produce a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), and the two sequencing data sets are associated with nucleic acid molecules with different SNPs in about 60% to 87.5% of the SNP arrangements of at least 20% of the random sequencing starting positions. The extended sequencing flow order is selected to produce a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), and the two sequencing data sets are associated with nucleic acid molecules with different SNPs in about 70% to 87.5% of the SNP arrangements of at least 20% of the random sequencing starting positions. The extended sequencing flow order is selected to produce a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), wherein the two sequencing data sets are associated with nucleic acid molecules that differ in SNPs in approximately 80% to 87.5% of the SNP arrangements at at least 20% of the random sequencing start positions.
[0140] In some embodiments, an extended sequencing flow order is selected to generate a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), and the two sequencing data sets are associated with nucleic acid molecules with different SNPs in about 50% to 87.5% of the SNP arrangements of at least 30% of the random sequencing starting positions. In some embodiments, an extended sequencing flow order is selected to generate a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), and the two sequencing data sets are associated with nucleic acid molecules with different SNPs in about 60% to 87.5% of the SNP arrangements of at least 30% of the random sequencing starting positions. An extended sequencing flow order is selected to generate a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), and the two sequencing data sets are associated with nucleic acid molecules with different SNPs in about 70% to 87.5% of the SNP arrangements of at least 30% of the random sequencing starting positions. The extended sequencing flow order is selected to produce a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), wherein the two sequencing data sets are associated with nucleic acid molecules that differ in SNPs in approximately 80% to 87.5% of the SNP arrangements at at least 30% of the random sequencing start positions.
[0141] In some embodiments, the expanded sequencing flow order is any one of the expanded sequencing flow orders in Table 2. "Shift sensitivity" refers to the maximum sensitivity that produces a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set) over all possible SNP arrangements. "Maximum shift sensitivity" refers to the maximum sensitivity that produces a signal difference at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set) over all possible SNP arrangements under the mobile phase that maintains the highest score for the sensitivity.
[0142]
[0143]
[0144]
[0145] In some embodiments, a nucleic acid sequencing method comprises (a) hybridizing a nucleic acid molecule with a primer to form a hybridization template; (b) extending the primer using a labeled non-terminator nucleotide provided in a separate nucleotide stream according to a repeated flow cycle sequence comprising five or more separate nucleotide streams; and (c) detecting a signal from an integrated labeled nucleotide or the absence of a signal as the primer is extended through the nucleotide stream. In some embodiments, for 50% or more of the possible SNP arrangements at a 5% random sequencing start position, the flow cycle sequence induces a signal change at more than two flow positions. In some embodiments, the induced signal change is a change in signal intensity, or a new substantially zero (or new zero) or a new substantially non-zero (or new non-zero) signal. In some embodiments, the induced signal change is a new substantially zero (or new zero) or a new substantially non-zero (or new non-zero) signal. In some embodiments, the flow cycle sequence has an efficiency of 0.6 or more bases per flow integration. In some embodiments, the flow cycle is any one of the flow cycle sequences listed in Table 2.
[0146] Resequencing with a different flow order
[0147] Since the sensitivity of the detected short genetic variants depends on the flow cycle order used to sequence the nucleic acid molecules, the methods described herein can be applied to analyze test nucleic acid molecules (or multiple nucleic acid molecules with overlapping loci) sequenced using two or more different flow cycle orders. A match score can be determined based on the match of two or more different sequencing data sets (generated by different flow cycle orders) with one or more candidate sequences. The presence or absence of a variant can be determined and / or a candidate sequence can be selected based on the match score as described above.
[0148] The method may include obtaining a first test sequencing data set associated with a test nucleic acid molecule derived from a test sample sequenced using a first flow cycle order, and a second test sequencing data set associated with the same test nucleic acid molecule sequenced using a second flow cycle order. For example, the test nucleic acid molecule may be sequenced by providing a non-terminating nucleic acid molecule in a separate nucleotide stream according to the first flow cycle order, extending a sequencing primer, and detecting the presence or absence of a nucleotide incorporated into the sequencing primer after each nucleotide stream to generate the first test sequencing data set; removing the extended sequencing primer; and sequencing the same test nucleic acid molecule by providing a non-terminating nucleotide in a separate nucleotide stream according to the second flow cycle order, extending the sequencing primer, and detecting the presence or absence of a nucleotide incorporated into the sequencing primer after each nucleotide stream to generate the second test sequencing data set.
[0149] The sequencing data sets are different because different flow cycle orders are used to sequence the nucleic acid molecules. Figure 4A and Figure 4B shows the use of the first flow cycle sequence (TACG) ( Figure 4A ) and the second flow cycle sequence (AGCT) ( Figure 4B ) determined by the extended primer sequence TATGGTCGTCGA (SEQ ID NO: 1). As can be seen, Figure 4A and Figure 4B The sequencing data sets in the example are different due to differences in the order of the flow cycles, even if the nucleic acid molecule sequence does not change. Within the sequencing data set, statistical parameters corresponding to the base counts of the first candidate extension primer sequence TATGGTCGTCGA (SEQ ID NO: 1) (filled circles) and the second candidate extension primer sequence TATGGTCATCGA (SEQ ID NO: 2) (open circles) at each flow position can be selected. Figure 4A and Figure 4B The results demonstrate that the flow cycle order significantly changes the sensitivity of variant detection. For example, the difference between the first candidate sequence and the second candidate sequence using the first flow cycle order is obvious at flow positions 12-20 ( Figure 4A ), while the difference between the first candidate sequence and the second candidate sequence using the first flow cycle order is only apparent at positions 17 and 18 ( Figure 4B ).
[0150] A match score indicating the likelihood that the first sequencing data set and the second sequencing data set match one or more candidate sequences (e.g., a target sequence having a pre-selected target short genetic variant, a reference sequence having a sequence without a pre-selected target short genetic variant, or other possible candidate sequences (such as a haplotype)) can be determined, and the presence or absence of the target short genetic variant can be determined or a candidate sequence can be selected.
[0151] As discussed herein, this method can be used when sequencing multiple different test nucleic acid molecules overlapping at a common locus. For example, multiple first test sequencing data sets can be obtained, wherein each test sequencing data set is associated with a test nucleic acid molecule sequenced using a first flow cycle sequence, and multiple second test sequencing data sets can be obtained, wherein each test sequencing data set is associated with the same nucleic acid molecule sequenced using a second flow cycle sequence. The first flow cycle sequence and the second flow cycle sequence are different. A match score indicating the possibility of matching multiple first sequencing data sets and multiple second sequencing data sets with one or more candidate sequences (e.g., a target sequence with a pre-selected target short genetic variant, a reference sequence with a sequence without a pre-selected target short genetic variant, or other possible candidate sequences (such as haplotypes)) can be determined, and the presence or absence of a target short genetic variant can be determined or a candidate sequence can be selected.
[0152] Figure 5 An exemplary method for detecting the presence or absence of short genetic variants in a test sample is shown. In step 502, one or more first test sequencing data sets are obtained. One or more first test sequencing data sets can be obtained, for example, by receiving one or more first test sequencing data sets or by sequencing one or more nucleic acid molecules. Each of the first test sequencing data sets is associated with a different nucleic acid molecule derived from the test sample. The first sequencing data set is determined by sequencing one or more test nucleic acid molecules using non-termination nucleotides provided in a separate nucleotide stream according to a first flow cycle order. The resulting one or more first test sequencing data sets each include a flow signal at a flow position corresponding to the nucleotide stream. In step 504, one or more second test sequencing data sets are obtained. One or more second test sequencing data sets can be obtained, for example, by receiving one or more second test sequencing data sets or by sequencing one or more nucleic acid molecules. Each of the second test sequencing data sets is associated with the same nucleic acid molecule as the first test sequencing data set. That is, the nucleic acid molecule is associated with both the first sequencing data set and the second sequencing data set. The second sequencing data set is determined by sequencing one or more test nucleic acid molecules using non-termination nucleotides provided in a separate nucleotide stream according to a second flow cycle order different from the first flow cycle order. The resulting one or more second test sequencing data sets each include a flow signal at a flow position corresponding to the nucleotide stream. At step 506, for each of the first sequencing data set and the second sequencing data set, a match score is determined. The match score indicates that the first test sequencing data set, the sequencing data set, or both match a candidate sequence from one or more candidate sequences. At step 508, the determined match score is used to determine the presence or absence of the short genetic variant in the test sample.
[0153] Figure 6Another exemplary method for detecting the presence or absence of short genetic variants in a test sample is shown. In step 602, a target short genetic variant is selected. The target short genetic variant is selected so that when the target sequencing data set and the reference sequencing data set are obtained by sequencing the target sequence using non-terminal nucleotides provided in a separate nucleotide flow according to a first flow cycle order or a second flow cycle order or both, the target sequencing data associated with the target sequence including the target short genetic variant is different from the sequencing data set associated with the reference sequence at more than two flow positions, wherein the first flow cycle order and the second flow cycle order are different, and wherein the flow position corresponds to the nucleotide flow. In step 604, one or more first test sequencing data sets are obtained. One or more first test sequencing data sets can be obtained, for example, by receiving one or more first test sequencing data sets or by sequencing one or more nucleic acid molecules. Each of the first test sequencing data sets is associated with a different nucleic acid molecule derived from a test sample. The first sequencing data set is determined by sequencing one or more test nucleic acid molecules using non-terminal nucleotides provided in a separate nucleotide flow according to a first flow cycle order. The resulting one or more first test sequencing data sets each include a flow signal at a flow position corresponding to a nucleotide flow. In step 606, one or more second test sequencing data sets are obtained. For example, one or more second test sequencing data sets can be obtained by receiving one or more second test sequencing data sets or by sequencing one or more nucleic acid molecules. Each of the second test sequencing data sets is associated with the same nucleic acid molecule as the first test sequencing data set. That is, the nucleic acid molecule is associated with both the first sequencing data set and the second sequencing data set. The second sequencing data set is determined by sequencing one or more test nucleic acid molecules using non-terminal nucleotides provided in a separate nucleotide flow according to a second flow cycle sequence that is different from the first flow cycle sequence. The resulting one or more second test sequencing data sets each include a flow signal corresponding to a flow position at a nucleotide flow. In step 608, for each first sequencing data set and the second sequencing data set, a matching score is determined. The matching score indicates that the first test sequencing data set, the sequencing data set, or both match a candidate sequence from one or more candidate sequences (which may include, for example, a reference sequence). In step 610, the determined matching score is used to determine the presence or absence of a short genetic variant in the test sample.
[0154] In some embodiments, a method for detecting the presence or absence of a short genetic variant in a test sample includes: (a) obtaining one or more first test sequencing data sets, each first test sequencing data set being associated with a different test nucleic acid molecule derived from the test sample, wherein the first test sequencing data set is determined by sequencing the one or more test nucleic acid molecules according to a first flow cycle order using non-termination nucleotides provided in separate nucleotide flows, and wherein the one or more first test sequencing data sets include flow signals at flow positions corresponding to the nucleotide flows; (b) obtaining one or more second test sequencing data sets, each second test sequencing data set being associated with the same test nucleic acid molecule as the first test sequencing data set, wherein the second test sequencing data set is determined by sequencing the one or more test nucleic acid molecules according to a second flow cycle order using non-termination nucleotides provided in separate nucleotide flows, wherein the first flow cycle order and the second flow cycle order are different, and wherein the test sequencing data set includes flow signals at flow positions corresponding to the nucleotide flows; (c) for each of the first sequencing data set and the second sequencing data set, determining a match score for one or more candidate sequences, wherein the match score indicates a likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from the one or more candidate sequences; and (d) using the determined match score to determine the presence or absence of the short genetic variant in the test sample.
[0155] In some embodiments, a method for detecting the presence or absence of a short genetic variant in a test sample comprises: (a) sequencing one or more test nucleic acid molecules derived from the test sample using non-terminating nucleotides provided in separate nucleotide flows according to a first flow cycle order to obtain one or more first test sequencing data sets, the first test sequencing data sets comprising flow signals at flow positions corresponding to the nucleotide flows, each first test sequencing data set being associated with a different test nucleic acid molecule; (b) sequencing the same one or more test nucleic acid molecules derived from the test sample using non-terminating nucleotides provided in separate nucleotide flows according to a second flow cycle order, wherein the second flow cycle order is different from the first flow cycle order to obtain one or more second test sequencing data sets, the second test sequencing data sets comprising flow signals at flow positions corresponding to the nucleotide flows, each second test sequencing data set being associated with the same test nucleic acid molecule in the first test sequencing data set; (c) for each first sequencing data set and second sequencing data set, determining a match score for one or more candidate sequences, wherein the match score indicates the likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from one or more candidate sequences; and (d) using the determined match score to determine the presence or absence of a short genetic variant in the test sample.
[0156] In some embodiments, a method for detecting the presence or absence of a short genetic variant in a test sample comprises: (a) obtaining one or more first test sequencing data sets, each first test sequencing data set being associated with a different test nucleic acid molecule derived from the test sample, wherein the first test sequencing data sets are determined by sequencing the one or more test nucleic acid molecules according to a first flow cycle order using a non-termination nucleotide provided in a separate nucleotide flow, and wherein the one or more first test sequencing data sets include flow signals at flow positions corresponding to the nucleotide flow; (b) obtaining one or more second test sequencing data sets, each second test sequencing data set being associated with the same test nucleic acid molecule as the first test sequencing data set, wherein the one or more test nucleic acid molecules are sequenced according to a second flow cycle order using a non-termination nucleotide provided in a separate nucleotide flow, and wherein the one or more first test sequencing data sets include flow signals at flow positions corresponding to the nucleotide flow; One or more test nucleic acid molecules are sequenced to determine a second test sequencing data set, wherein the first flow cycle order and the second flow cycle order are different, and wherein the test sequencing data set includes flow signals at flow positions corresponding to nucleotide flows; (c) for each first sequencing data set and the second sequencing data set, determine a match score for one or more candidate sequences, wherein the match score indicates the likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from one or more candidate sequences; (d) select a candidate sequence from two or more different candidate sequences, wherein the selected candidate sequence has the highest likelihood of matching the first test sequencing data set, the second test sequencing data set, or both; and (e) use the selected candidate sequence to determine the presence or absence of a short genetic variant in the test sample. In some embodiments, at least one unselected candidate sequence from two or more different candidate sequences is different from the selected candidate sequence at two or more (or three or more, or across one or more flow cycles) flow positions (which can be continuous or non-continuous) according to the first flow cycle order and / or the second flow cycle order.
[0157] In some embodiments, a method for detecting the presence or absence of a short genetic variant in a test sample comprises: (a) sequencing one or more test nucleic acid molecules derived from the test sample using non-termination nucleotides provided in a separate nucleotide flow according to a first flow cycle order to obtain one or more first test sequencing data sets, which include flow signals at flow positions corresponding to the nucleotide flows, each first test sequencing data set is associated with a different test nucleic acid molecule; (b) sequencing the same one or more test nucleic acid molecules derived from the test sample using non-termination nucleotides provided in a separate nucleotide flow according to a second flow cycle order, wherein the second flow cycle order is different from the first flow cycle order to obtain one or more second test sequencing data sets. (c) for each first and second sequencing data sets, determining a match score for one or more candidate sequences, wherein the match score indicates the likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from one or more candidate sequences; (d) selecting a candidate sequence from two or more different candidate sequences, wherein the selected candidate sequence has the highest likelihood of matching the first test sequencing data set, the second test sequencing data set, or both; and (e) using the selected candidate sequence to determine the presence or absence of a short genetic variant in the test sample. In some embodiments, at least one unselected candidate sequence from two or more different candidate sequences is different from the selected candidate sequence at two or more (or three or more, or across one or more flow cycles) flow positions (which may be continuous or non-continuous) according to the first flow cycle order and / or the second flow cycle order.
[0158] In some embodiments, a method for detecting the presence or absence of a short genetic variant in a test sample includes: (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence according to a first flow cycle order or a second flow cycle order using a non-termination nucleotide provided in a separate nucleotide flow, the target sequencing dataset associated with the target sequence including the target short genetic variant is different from a reference sequencing dataset associated with the reference sequence at two or more flow positions, wherein the first flow cycle order is different from the second flow cycle order, and wherein the flow positions correspond to the nucleotide flows; (b) obtaining one or more first test sequences, each first test sequencing dataset being associated with a different test nucleic acid molecule derived from the test sample, wherein the first test sequencing dataset is determined by sequencing the one or more test nucleic acid molecules according to the first flow cycle order using a non-termination nucleotide provided in a separate nucleotide flow. (c) obtaining one or more second test sequencing data sets, each second test sequencing data set being associated with the same test nucleic acid molecule as the first test sequencing data set, wherein the second test sequencing data set is determined by sequencing the one or more test nucleic acid molecules sequentially using a non-termination nucleotide provided in a separate nucleotide flow, wherein the test sequencing data sets include flow signals at flow positions corresponding to the nucleotide flow; (d) for each of the first sequencing data set and the second sequencing data set, determining a match score for one or more candidate sequences, wherein the match score indicates a likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from the one or more candidate sequences; and (e) using the determined match score to determine the presence or absence of a short genetic variant in the test sample.
[0159] In some embodiments, a method for detecting the presence or absence of a short genetic variant in a test sample comprises: (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using a non-terminating nucleotide provided in a separate nucleotide flow according to a first flow cycle order or a second flow cycle order, the target sequencing dataset associated with the target sequence including the target short genetic variant is different from a reference sequencing dataset associated with the reference sequence at two or more flow positions, wherein the first flow cycle order is different from the second flow cycle order, and wherein the flow positions correspond to the nucleotide flows; (b) sequencing one or more test nucleic acid molecules derived from the test sample using a non-terminating nucleotide provided in a separate nucleotide flow according to the first flow cycle order to obtain one or more first test sequencing datasets, which include the non-terminating nucleotides at the flow positions corresponding to the nucleotide flows. (c) sequencing the same one or more test nucleic acid molecules derived from the test sample using non-termination nucleotides provided in separate nucleotide flows according to a second flow cycle order to obtain one or more second test sequencing data sets, which include flow signals at flow positions corresponding to the nucleotide flows, each second test sequencing data set is associated with an identical test nucleic acid molecule in the first test sequencing data set; (d) for each of the first sequencing data set and the second sequencing data set, determining a match score for one or more candidate sequences, wherein the match score indicates a likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from the one or more candidate sequences; and (e) using the determined match score to determine the presence or absence of a short genetic variant in the test sample.
[0160] In some embodiments, a method for detecting the presence or absence of a short genetic variant in a test sample includes: (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence according to a first flow cycle order or a second flow cycle order using a non-termination nucleotide provided in a separate nucleotide flow, the target sequencing dataset associated with the target sequence including the target short genetic variant is different from a reference sequencing dataset associated with the reference sequence at two or more flow positions, wherein the first flow cycle order is different from the second flow cycle order, and wherein the flow positions correspond to the nucleotide flows; (b) obtaining one or more first test sequencing datasets, each first test sequencing dataset being associated with a different test nucleic acid molecule derived from the test sample, wherein the first test sequencing datasets are determined by sequencing the one or more test nucleic acid molecules according to the first flow cycle order using a non-termination nucleotide provided in a separate nucleotide flow, and wherein the one or more first test sequencing datasets include a non-termination nucleotide at a flow position corresponding to the nucleotide flow. (c) obtaining one or more second test sequencing data sets, each second test sequencing data set being associated with the same test nucleic acid molecule as the first test sequencing data set, wherein the second sequencing data sets are determined by sequencing the one or more test nucleic acid molecules according to a second flow cycle order using non-termination nucleotides provided in separate nucleotide flows, wherein the test sequencing data sets include flow signals at flow positions corresponding to the nucleotide flows; (d) for each of the first sequencing data set and the second sequencing data set, determining a match score for one or more candidate sequences (which may include a reference sequence), wherein the match score indicates a likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from the one or more candidate sequences; (e) selecting a candidate sequence from two or more different candidate sequences, wherein the selected candidate sequence has the highest likelihood of matching the first test sequencing data set, the second test sequencing data set, or both; and (f) determining the presence or absence of a short genetic variant in the test sample using the selected candidate sequence. In some embodiments, at least one unselected candidate sequence from two or more different candidate sequences differs from the selected candidate sequence at two or more (or three or more, or across one or more flow cycles) flow positions (which may be continuous or non-contiguous) according to the first flow cycle order and / or the second flow cycle order.
[0161] In some embodiments, a method for detecting the presence or absence of a short genetic variant in a test sample comprises: (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using a non-terminating nucleotide provided in a separate nucleotide flow according to a first flow cycle order or a second flow cycle order, the target sequencing dataset associated with the target sequence including the target short genetic variant is different from a reference sequencing dataset associated with the reference sequence at two or more flow positions, wherein the first flow cycle order is different from the second flow cycle order, and wherein the flow position corresponds to the nucleotide flow; (b) sequencing one or more test nucleic acid molecules derived from the test sample using a non-terminating nucleotide provided in a separate nucleotide flow according to the first flow cycle order to obtain one or more first test sequencing datasets, which include flow signals at flow positions corresponding to the nucleotide flow, each first test sequencing dataset being associated with a different test nucleic acid molecule; and (c) sequencing the target sequence according to the second flow cycle order. The two flow cycle sequences use non-terminal nucleotides provided in separate nucleotide flows to sequence the same one or more test nucleic acid molecules derived from the test sample to obtain one or more second test sequencing data sets, which include flow signals at flow positions corresponding to the nucleotide flows, each second test sequencing data set is associated with the same test nucleic acid molecule in the first test sequencing data set; (d) for each first sequencing data set and the second sequencing data set, determine the match score of one or more candidate sequences, wherein the match score indicates the possibility that the first test sequencing data set, the second test sequencing data set, or both match the candidate sequence from the one or more candidate sequences; (e) select a candidate sequence from two or more different candidate sequences (which may include a reference sequence), wherein the selected candidate sequence has the highest possibility of matching the first test sequencing data set, the second test sequencing data set, or both; and (f) use the selected candidate sequence to determine the presence or absence of short genetic variants in the test sample. In some embodiments, at least one unselected candidate sequence from two or more different candidate sequences is different from the selected candidate sequence at two or more (or three or more, or across one or more flow cycles) flow positions (which may be continuous or non-continuous) according to the first flow cycle sequence and / or the second flow cycle sequence.
[0162] Systems, Equipment and Reports
[0163] The above operations (including those described with reference to the accompanying drawings) are optionally performed by Figure 7 A person skilled in the art will clearly know how to implement the Figure 7The components depicted in the figure can be used to implement other processes, for example, a combination or sub-combination of all or part of the operations described above. It will also be clear to those skilled in the art how the methods, techniques, systems and devices described herein can be combined with each other in whole or in part, and whether those methods, techniques, systems and / or devices are composed of Figure 7 The component depicted implements and / or provides.
[0164] Figure 7 An example of a computing device according to one embodiment is illustrated. Device 700 may be a host computer connected to a network. Device 700 may be a client computer or a server. Figure 7 As shown, device 700 can be any suitable type of microprocessor-based device, such as a personal computer, workstation, server, or handheld computing device (portable electronic device), such as a phone or tablet computer. The device can include, for example, one or more of a processor 710, an input device 720, an output device 730, a storage device 740, and a communication device 760. The input device 720 and the output device 730 can generally correspond to those described above and can be connected or integrated with the computer.
[0165] Input device 720 may be any suitable device that provides input, such as a touch screen, a keyboard or keypad, a mouse, or a voice recognition device. Output device 730 may be any suitable device that provides output, such as a touch screen, a tactile device, or a speaker.
[0166] Memory 740 may be any suitable device that provides storage, such as electrical, magnetic or optical memory, including RAM, cache, hard drive or removable storage disk. Communication device 760 may include any suitable device capable of sending and receiving signals over a network, such as a network interface chip or device. The components of the computer may be connected in any suitable manner, such as via a physical bus or wireless connection.
[0167] Software 750 , which may be stored in memory 740 and executed by processor 710 , may include, for example, programming that embodies the functionality of the present disclosure (eg, as embodied in the devices described above).
[0168] The software 750 may also be stored and / or transmitted in any non-transitory computer-readable storage medium for use by or in conjunction with an instruction execution system, apparatus, or device, such as those described above, which may retrieve instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a computer-readable storage medium may be any medium, such as storage device 740, which may include or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0169] The software 750 may also be transmitted in any transmission medium for use by or in conjunction with an instruction execution system, device, or apparatus, such as those described above, which may obtain instructions associated with the software from the instruction execution system, device, or apparatus and execute the instructions. In the context of this disclosure, a transmission medium may be any medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, device, or apparatus. Transmission-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless communication media.
[0170] Device 700 can be connected to a network, which can be any suitable type of interconnected communication system. The network can implement any suitable communication protocol and can be protected by any suitable security protocol. The network can include any suitable arrangement of network links that can implement transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL or telephone lines.
[0171] Device 700 may implement any operating system suitable for operating on a network. Software 1850 may be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying the functionality of the present disclosure may be deployed in different configurations, such as in a client / server arrangement or as a web-based application or web service, for example, through a web browser.
[0172] Methods described herein optionally further include reporting the information determined using analytical methods and / or generating a report including the information determined using analytical methods. For example, in some embodiments, the method also includes reporting or generating a report including the identification of a variant in the polynucleotide derived from a subject (for example, in the subject's genome). The information reported or the information in the report can be related to the validation statistics of the locus of the coupled sequencing read pair, detected variants (such as detected structural variants or detected SNPs), the consensus sequences of one or more assemblies, and / or the consensus sequences of one or more assemblies, for example, mapped to a reference sequence. The report can be distributed to a recipient, or the information can be reported to a recipient, such as a clinician, a subject, or a researcher.
[0173] In some embodiments, there is a system comprising: one or more processors; and a non-transitory computer-readable medium storing one or more programs, the programs comprising instructions for the following operations: (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide flows according to a flow cycle order, the target sequencing dataset associated with the target sequence including the target short genetic variant is different from the reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to nucleotide flows; (b) obtaining one or more test sequencing datasets, each test sequencing dataset being associated with a test nucleic acid molecule, each test nucleic acid molecule At least partially overlapping with a locus associated with a target short genetic variant and derived from a test sample, wherein one or more test sequencing data sets are determined by sequencing a test nucleic acid molecule using non-terminating nucleotides provided in a separate nucleotide stream according to a flow cycle order, and wherein the test sequencing data set includes flow signals at multiple flow positions; (c) for each test nucleic acid molecule associated with the test sequencing data set, determining a match score indicating the likelihood that the test sequencing data set associated with the nucleic acid molecule matches the target sequence, or a match score indicating the likelihood that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence; and (d) using one or more determined match scores to determine the presence or absence of the target short genetic variant in the test sample. In some embodiments, the method also includes generating a personalized biomarker group for a subject associated with the test sample, the biomarker group including the target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0174] In some embodiments, there is a system comprising: one or more processors; and a non-transitory computer-readable medium storing one or more programs, the programs comprising instructions for the following operations: (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order, the target sequencing dataset associated with the target sequence including the target short genetic variant is different from the reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to the nucleotide streams; (b) using the non-termination nucleotides provided in the separate nucleotide streams according to the flow cycle order to obtain a target sequencing dataset and a reference sequencing dataset. or multiple test nucleic acid molecules are sequenced to obtain one or more test sequencing data sets including flow signals at multiple flow positions, each test sequencing data set is associated with a test nucleic acid molecule, and each test nucleic acid molecule at least partially overlaps with a locus associated with a target short genetic variant and derived from a test sample; (c) for each test nucleic acid molecule associated with the test sequencing data set, a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the target sequence is determined, or a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence is determined; and (d) using one or more determined match scores to determine the presence or absence of the target short genetic variant in the test sample. In some embodiments, the method also includes generating a personalized biomarker group for a subject associated with the test sample, the biomarker group including the target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0175] In some embodiments, there is a system comprising: one or more processors; and a non-transitory computer-readable medium storing one or more programs, the programs comprising instructions for the following operations: (a) pre-selecting target short genetic variants, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-terminal nucleotides provided in separate nucleotide flows according to a flow cycle order, the target sequencing dataset associated with the target sequence including the pre-selected target short genetic variant is different from the reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to nucleotide flows; (b) obtaining one or more test sequencing datasets, each test sequencing dataset being associated with a test nucleic acid molecule, each test nucleic acid molecule being at least A small portion overlaps with a locus associated with a pre-selected target short genetic variant and derived from a test sample, wherein one or more test sequencing data sets are determined by sequencing the test nucleic acid molecule using non-terminal nucleotides provided in separate nucleotide flows according to a flow cycle order, and wherein the test sequencing data set includes flow signals at multiple flow positions; (c) for each test nucleic acid molecule associated with the test sequencing data set, a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the target sequence is determined, or a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence is determined; and (d) using one or more determined match scores to determine the presence or absence of a pre-selected target short genetic variant in the test sample. In some embodiments, the method also includes generating a personalized biomarker group for a subject associated with the test sample, the biomarker group including the target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0176] In some embodiments, there is a system comprising one or more processors; and a non-transitory computer-readable medium storing one or more programs, the programs comprising instructions for the following operations: (a) pre-selecting target short genetic variants, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order, the target sequencing dataset associated with the target sequence including the pre-selected target short genetic variant is different from the reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to the nucleotide streams; (b) using non-termination nucleotides provided in separate nucleotide streams according to the flow cycle order to obtain a target sequencing dataset and a reference sequencing dataset. One or more test nucleic acid molecules are sequenced to obtain one or more test sequencing data sets including flow signals at multiple flow positions, each test sequencing data set is associated with the test nucleic acid molecule, and each test nucleic acid molecule at least partially overlaps with a locus associated with a target short genetic variant and derived from the test sample; (c) for each test nucleic acid molecule associated with the test sequencing data set, a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the target sequence is determined, or a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence is determined; and (d) using one or more determined match scores to determine the presence or absence of a pre-selected target short genetic variant in the test sample. In some embodiments, the method also includes generating a personalized biomarker group for a subject associated with the test sample, the biomarker group including the target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0177] In some embodiments, a system is provided, comprising: one or more processors; and a non-transitory computer-readable medium storing one or more programs, the programs comprising instructions for the following operations: (a) pre-selecting target short genetic variants and flow cycle sequences, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide flows according to the flow cycle sequence, the target sequencing dataset associated with the target sequence including the pre-selected target short genetic variant is different from the reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to nucleotide flows; (b) obtaining one or more test sequencing datasets, each test sequencing dataset being associated with a test nucleic acid molecule, each test nucleic acid molecule At least partially overlapping with a locus associated with a pre-selected target short genetic variant and derived from a test sample, wherein one or more test sequencing data sets are determined by sequencing the test nucleic acid molecule using non-terminating nucleotides provided in a separate nucleotide flow according to a pre-selected flow cycle order, and wherein the test sequencing data set includes flow signals at multiple flow positions; (c) for each test nucleic acid molecule associated with the test sequencing data set, determining a match score indicating the likelihood that the test sequencing data set associated with the nucleic acid molecule matches the target sequence, or indicating the likelihood that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence; and (d) using one or more determined match scores to determine the presence or absence of a pre-selected target short genetic variant in the test sample. In some embodiments, the method also includes generating a personalized biomarker group for a subject associated with the test sample, the biomarker group including the target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0178] In some embodiments, there is a system comprising: one or more processors; and a non-transitory computer-readable medium storing one or more programs, the programs comprising instructions for the following operations: (a) pre-selecting target short genetic variants and flow cycle orders, wherein when a target sequencing data set and a reference sequencing data set are obtained by sequencing the target sequence using non-terminal nucleotides provided in separate nucleotide streams according to the flow cycle order, the target sequencing data set associated with the target sequence including the pre-selected target short genetic variant is different from the reference sequencing data set associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to the nucleotide streams; (b) using non-terminal nucleotides provided in separate nucleotide streams according to the pre-selected flow cycle order. (c) for each test nucleic acid molecule associated with the test sequencing data set, determining a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the target sequence, or a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence; and (d) using one or more determined match scores to determine the presence or absence of a pre-selected target short genetic variant in the test sample. In some embodiments, the method further includes generating a personalized biomarker panel for a subject associated with the test sample, the biomarker panel including the target short genetic variant. In some embodiments, the method further includes generating a personalized biomarker group for a subject associated with the test sample, the biomarker group including the target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing dataset differs from the reference sequencing dataset across one or more flow cycles.
[0179] In some embodiments, there is a system comprising one or more processors; and a non-transitory computer-readable medium storing one or more programs, wherein the one or more programs include instructions for the following operations: (a) obtaining one or more first test sequencing data sets, each first test sequencing data set is associated with a different test nucleic acid molecule derived from a test sample, wherein the first test sequencing data set is determined by sequencing the one or more test nucleic acid molecules using non-termination nucleotides provided in separate nucleotide flows according to a first flow cycle sequence, and wherein the one or more first test sequencing data sets include flow signals at flow positions corresponding to the nucleotide flows; (b) obtaining one or more second test sequencing data sets, each second test sequencing data set being associated with the first test sequencing data set; (c) for each of the first sequencing data set and the second sequencing data set, determining a match score for one or more candidate sequences, wherein the match score indicates a likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from the one or more candidate sequences; and (d) using the determined match score to determine the presence or absence of a short genetic variant in the test sample.
[0180] In some embodiments, there is a system comprising: one or more processors; and a non-transitory computer-readable medium storing one or more programs, the one or more programs comprising instructions for the following operations: (a) sequencing one or more test nucleic acid molecules derived from a test sample using non-termination nucleotides provided in separate nucleotide streams according to a first flow cycle order to obtain one or more first test sequencing data sets including flow signals at flow positions corresponding to the nucleotide streams, each first test sequencing data set being associated with a different test nucleic acid molecule; (b) sequencing the same one or more test nucleic acid molecules derived from the test sample using non-termination nucleotides provided in separate nucleotide streams according to a second flow cycle order to obtain one or more second test sequencing data sets including flow signals at flow positions corresponding to the nucleotide streams, wherein the second flow cycle order is different from the first flow cycle order, and each second test sequencing data set is associated with the same test nucleic acid molecule in the first test sequencing data set; (c) for each first sequencing data set and the second sequencing data set, determining a match score for one or more candidate sequences, wherein the match score indicates the likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from the one or more candidate sequences; and (d) using the determined match score to determine the presence or absence of a short genetic variant in the test sample.
[0181] In some embodiments, there is a system comprising: one or more processors; and a non-transitory computer-readable medium storing one or more programs, the one or more programs comprising instructions for the following operations: (a) obtaining one or more first test sequencing data sets, each first test sequencing data set being associated with a different test nucleic acid molecule derived from a test sample, wherein the first test sequencing data sets are determined by sequencing the one or more test nucleic acid molecules using non-termination nucleotides provided in separate nucleotide streams according to a first flow cycle sequence, and wherein the one or more first test sequencing data sets include flow signals corresponding to flow positions at the nucleotide streams; (b) obtaining one or more second test sequencing data sets, each second test sequencing data set being associated with the same test nucleic acid molecule as the first test sequencing data set, wherein the one or more first test sequencing data sets are determined by sequencing the one or more test nucleic acid molecules using non-termination nucleotides provided in separate nucleotide streams according to a second flow cycle sequence, and wherein the one or more first test sequencing data sets include flow signals corresponding to flow positions at the nucleotide streams; (c) for each of the first and second sequencing data sets, determining a match score for one or more candidate sequences, wherein the match score indicates the likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from one or more candidate sequences; (d) selecting a candidate sequence from two or more different candidate sequences, wherein the selected candidate sequence has the highest likelihood of matching the first test sequencing data set, the second test sequencing data set, or both; and (e) using the selected candidate sequence to determine the presence or absence of a short genetic variant in the test sample. In some embodiments, at least one unselected candidate sequence from the two or more different candidate sequences is different from the selected candidate sequence at two or more (or three or more, or across one or more flow cycles) flow positions (which may be continuous or non-continuous) according to the first flow cycle order and / or the second flow cycle order.
[0182] In some embodiments, there is a system comprising: one or more processors; and a non-transitory computer-readable medium storing one or more programs, wherein the one or more programs include instructions for the following operations: (a) sequencing one or more test nucleic acid molecules derived from a test sample using non-termination nucleotides provided in separate nucleotide streams according to a first flow cycle sequence to obtain one or more first test sequencing data sets including flow signals at flow positions corresponding to the nucleotide streams, each first test sequencing data set being associated with a different test nucleic acid molecule; (b) sequencing the same one or more test nucleic acid molecules derived from the test sample using non-termination nucleotides provided in separate nucleotide streams according to a second flow cycle sequence to obtain one or more flow signals including flow positions corresponding to the nucleotide streams. (c) for each first and second sequencing data sets, determining a match score for one or more candidate sequences, wherein the match score indicates the likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from one or more candidate sequences; (d) selecting a candidate sequence from two or more different candidate sequences, wherein the selected candidate sequence has the highest likelihood of matching the first test sequencing data set, the second test sequencing data set, or both; and (e) using the selected candidate sequence to determine the presence or absence of a short genetic variant in the test sample. In some embodiments, at least one unselected candidate sequence from the two or more different candidate sequences is different from the selected candidate sequence at two or more (or three or more, or across one or more flow cycles) flow positions (which may be continuous or non-continuous) according to the first flow cycle order and / or the second flow cycle order.
[0183] In some embodiments, there is a system comprising: one or more processors; and a non-transitory computer-readable medium storing one or more programs, the programs comprising instructions for the following operations: (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence according to a first flow cycle order or a second flow cycle order using non-termination nucleotides provided in separate nucleotide streams, the target sequencing dataset associated with the target sequence including the target short genetic variant is different from the reference sequencing dataset associated with the reference sequence at two or more flow positions, wherein the first flow cycle order is different from the second flow cycle order, and wherein the flow positions correspond to the nucleotide streams; (b) obtaining one or more first test sequencing datasets, each first test sequencing dataset being associated with a different test nucleic acid molecule derived from a test sample, wherein the one or more test nucleic acids are sequenced according to the first flow cycle order using non-termination nucleotides provided in separate nucleotide streams. (c) obtaining one or more second test sequencing data sets, each second test sequencing data set being associated with the same test nucleic acid molecule as the first test sequencing data set, wherein the second test sequencing data set is determined by sequencing the one or more test nucleic acid molecules according to a second flow cycle order using non-termination nucleotides provided in separate nucleotide flows, wherein the test sequencing data sets include flow signals at flow positions corresponding to the nucleotide flows; (d) for each of the first sequencing data set and the second sequencing data set, determining a match score for one or more candidate sequences, wherein the match score indicates a likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from the one or more candidate sequences; and (e) using the determined match score to determine the presence or absence of a short genetic variant in the test sample.
[0184] In some embodiments, there is a system comprising: one or more processors; and a non-transitory computer-readable medium storing one or more programs, the programs comprising instructions for the following operations: (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using a non-terminating nucleotide provided in a separate nucleotide stream according to a first flow cycle order or a second flow cycle order, the target sequencing dataset associated with the target sequence including the target short genetic variant is different from a reference sequencing dataset associated with the reference sequence at two or more flow positions, wherein the first flow cycle order is different from the second flow cycle order, and wherein the flow position corresponds to the nucleotide stream; (b) sequencing one or more test nucleic acid molecules derived from a test sample using a non-terminating nucleotide provided in a separate nucleotide stream according to the first flow cycle order to obtain one or more test nucleic acid molecules including at positions corresponding to the nucleotides. (c) sequencing the same one or more test nucleic acid molecules derived from the test sample using non-termination nucleotides provided in separate nucleotide flows according to a second flow cycle order to obtain one or more second test sequencing data sets including flow signals at flow positions corresponding to the nucleotide flows, each second test sequencing data set being associated with the same test nucleic acid molecule in the first test sequencing data set; (d) for each of the first sequencing data set and the second sequencing data set, determining a match score for one or more candidate sequences, wherein the match score indicates a likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from the one or more candidate sequences; and (e) using the determined match score to determine the presence or absence of a short genetic variant in the test sample.
[0185] In some embodiments, there is a system comprising: one or more processors; and a non-transitory computer-readable medium storing one or more programs, the programs comprising instructions for the following operations: (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence according to a first flow cycle order or a second flow cycle order using a non-termination nucleotide provided in a separate nucleotide stream, the target sequencing dataset associated with the target sequence including the target short genetic variant is different from the reference sequencing dataset associated with the reference sequence at two or more flow positions, wherein the first flow cycle order is different from the second flow cycle order, and wherein the flow positions correspond to the nucleotide streams; (b) obtaining one or more first test sequencing datasets, each of the first test sequencing datasets being associated with a different test nucleic acid molecule derived from a test sample, wherein the first test sequencing datasets are determined by sequencing the one or more test nucleic acid molecules according to the first flow cycle order using a non-termination nucleotide provided in a separate nucleotide stream, and wherein the one or more first test sequencing datasets comprise (c) obtaining one or more second test sequencing data sets, each second test sequencing data set being associated with the same test nucleic acid molecule as the first test sequencing data set, wherein the second test sequencing data sets are determined by sequencing the one or more test nucleic acid molecules according to a second flow cycle order using non-termination nucleotides provided in a separate nucleotide flow, wherein the test sequencing data sets include flow signals at flow positions corresponding to the nucleotide flow; (d) for each of the first sequencing data set and the second sequencing data set, determining a match score for one or more candidate sequences (which may include a reference sequence), wherein the match score indicates a likelihood that the first test sequencing data set, the second test sequencing data set, or both, matches a candidate sequence from the one or more candidate sequences; (e) selecting a candidate sequence from two or more different candidate sequences, wherein the selected candidate sequence has the highest likelihood of matching the first test sequencing data set, the second test sequencing data set, or both; and (f) determining the presence or absence of a short genetic variant in the test sample using the selected candidate sequence. In some embodiments, at least one unselected candidate sequence from two or more different candidate sequences differs from the selected candidate sequence at two or more (or three or more, or across one or more flow cycles) flow positions (which may be continuous or non-contiguous) according to the first flow cycle order and / or the second flow cycle order.
[0186] In some embodiments, there is a system comprising: one or more processors; and a non-transitory computer-readable medium storing one or more programs, the programs comprising instructions for the following operations: (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide streams according to a first flow cycle order or a second flow cycle order, the target sequencing dataset associated with the target sequence including the target short genetic variant is different from a reference sequencing dataset associated with the reference sequence at two or more flow positions, wherein the first flow cycle order is different from the second flow cycle order, and wherein the flow positions correspond to nucleotide streams; (b) sequencing one or more test nucleic acid molecules derived from a test sample using non-termination nucleotides provided in separate nucleotide streams according to the first flow cycle order to obtain one or more first test sequencing datasets including flow signals at flow positions corresponding to the nucleotide streams, each first test sequencing dataset being different from a different flow position; (c) sequencing the same one or more test nucleic acid molecules derived from the test sample using non-termination nucleotides provided in a separate nucleotide flow according to a second flow cycle order to obtain one or more second test sequencing data sets including flow signals at flow positions corresponding to the nucleotide flows, each second test sequencing data set being associated with the same test nucleic acid molecule in the first test sequencing data set; (d) for each first sequencing data set and second sequencing data set, determining a match score for one or more candidate sequences, wherein the match score indicates a likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from the one or more candidate sequences; (e) selecting a candidate sequence from two or more different candidate sequences (which may include a reference sequence), wherein the selected candidate sequence has the highest likelihood of matching the first test sequencing data set, the second test sequencing data set, or both; and (f) using the selected candidate sequence to determine the presence or absence of a short genetic variant in the test sample. In some embodiments, at least one unselected candidate sequence from two or more different candidate sequences differs from the selected candidate sequence at two or more (or three or more, or across one or more flow cycles) flow positions (which may be continuous or non-contiguous) according to the first flow cycle order and / or the second flow cycle order.
[0187] In some embodiments, the methods described herein are computer-implemented methods that can be used Figure 7For example, in some embodiments, a computer-implemented method for detecting short genetic variants in a test sample includes: (a) selecting a target short genetic variant using one or more processors, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order, the target sequencing dataset associated with the target sequence including the target short genetic variant is different from the reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to nucleotide streams; (b) receiving one or more test sequencing datasets at the one or more processors, each test sequencing dataset being associated with a test nucleic acid molecule, each test nucleic acid molecule at least partially overlapping with a target sequence associated with the target short genetic variant. And derived from the loci of the test sample, wherein one or more test sequencing data sets are determined by sequencing the test nucleic acid molecule using non-terminal nucleotides provided in a separate nucleotide flow according to the flow cycle order, and wherein the test sequencing data set includes flow signals at multiple flow positions; (c) for each test nucleic acid molecule associated with the test sequencing data set, one or more processors are used to determine a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the target sequence, or a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence; and (d) in one or more processors and using one or more determined match scores to determine the presence or absence of a target short genetic variant in the test sample. In some embodiments, the method also includes generating a personalized biomarker group for a subject associated with the test sample, the biomarker group including a target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0188] In some embodiments, a computer-implemented method for detecting short genetic variants in a test sample comprises: (a) preselecting a target short genetic variant using one or more processors, wherein when a target sequencing data set is obtained by sequencing a target sequence using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order and a reference sequencing data set associated with a reference sequence, the target sequencing data set associated with a target sequence including the preselected target short genetic variant differs from the reference sequencing data set associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to nucleotide streams; (b) receiving, at the one or more processors, one or more test sequencing data sets, each test sequencing data set associated with a test nucleic acid molecule, each test nucleic acid molecule at least partially overlapping with the preselected target sequence. (c) for each test nucleic acid molecule associated with the test sequencing data set, a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the target sequence is determined in one or more processors, or a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence is determined; and (d) in one or more processors, and using one or more determined match scores to determine the presence or absence of a pre-selected target short genetic variant in the test sample. In some embodiments, the method further includes generating a personalized biomarker group for a subject associated with the test sample, the biomarker group including a target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0189] In some embodiments, a computer-implemented method for detecting short genetic variants in a test sample comprises: (a) preselecting a target short genetic variant and a flow cycle sequence using one or more processors, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide streams according to the preselected flow cycle sequence, the target sequencing dataset associated with the target sequence including the preselected target short genetic variant differs from the reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to nucleotide streams; (b) receiving, at the one or more processors, one or more test sequencing datasets, each test sequencing dataset being associated with a test nucleic acid molecule, each test nucleic acid molecule at least partially overlapping a region corresponding to the preselected target sequence. A method for preparing a ...
[0190] In some embodiments, a computer-implemented method for detecting short genetic variants in a test sample includes (a) selecting a target short genetic variant at one or more processors, wherein when a target sequencing data set is obtained by sequencing a target sequence using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order, the target sequencing data set associated with the target sequence including the target short genetic variant is different from the reference sequencing data set associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to nucleotide streams; (b) receiving one or more test sequencing data sets at one or more processors, each test sequencing data set being associated with a test nucleic acid molecule, each test nucleic acid molecule at least partially overlapping with a target short genetic variant. (c) for each test nucleic acid molecule associated with the test sequencing data set, a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the target sequence is determined in one or more processors, or a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence is determined; and (d) in one or more processors, and using one or more determined match scores to determine the presence or absence of a target short genetic variant in the test sample. In some embodiments, the method also includes generating a personalized biomarker group for a subject associated with the test sample, the biomarker group including a target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0191] In some embodiments, a computer-implemented method for detecting short genetic variants in a test sample includes: (a) pre-selecting a target short genetic variant at one or more processors, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order, the target sequencing dataset associated with the target sequence including the pre-selected target short genetic variant is different from a reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to nucleotide streams; (b) receiving one or more test sequencing datasets at one or more processors, each test sequencing dataset being associated with a test nucleic acid molecule, each test nucleic acid molecule at least partially overlapping with a target sequence that is associated with the pre-selected target short genetic variant. (c) for each test nucleic acid molecule associated with the test sequencing data set, a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the target sequence is determined in one or more processors, or a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence is determined; and (d) at one or more processors and using one or more determined match scores to determine the presence or absence of a pre-selected target short genetic variant in the test sample. In some embodiments, the method also includes generating a personalized biomarker group for a subject associated with the test sample, the biomarker group including a target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0192] In some embodiments, a computer-implemented method for detecting short genetic variants in a test sample comprises: (a) preselecting a target short genetic variant and a flow cycle sequence at one or more processors, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide streams according to the preselected flow cycle sequence, the target sequencing dataset associated with the target sequence including the preselected target short genetic variant differs from the reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to nucleotide streams; (b) receiving at one or more processors one or more test sequencing datasets, each test sequencing dataset being associated with a test nucleic acid molecule, each test nucleic acid molecule at least partially overlapping a region corresponding to the preselected target sequence. (c) for each test nucleic acid molecule associated with the test sequencing data set, a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the target sequence is determined in one or more processors, or a match score indicating the possibility that the test sequencing data set associated with the nucleic acid molecule matches the reference sequence is determined; and (d) at one or more processors and using one or more determined match scores to determine the presence or absence of a pre-selected target short genetic variant in the test sample. In some embodiments, the method further includes generating a personalized biomarker group for a subject associated with the test sample, the biomarker group including a target short genetic variant. In some embodiments, the target sequencing data set is different from the reference sequencing data set at more than two flow positions (e.g., more than two continuous flow positions or more than two non-continuous flow positions). In some embodiments, the target sequencing data set is different from the reference sequencing data set across one or more flow cycles.
[0193] Exemplary embodiments
[0194] The following embodiments are exemplary and are not intended to limit the scope of the claimed invention.
[0195] Embodiment 1. A method for detecting short genetic variants in a test sample, comprising:
[0196] (a) selecting a target short genetic variant, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-terminal nucleotides provided in separate nucleotide flows according to a flow cycle order, the target sequencing dataset associated with the target sequence including the target short genetic variant differs from a reference sequencing dataset associated with the reference sequence at more than two flow positions, wherein the flow positions correspond to the nucleotide flows;
[0197] (b) obtaining one or more test sequencing data sets, each test sequencing data set being associated with a test nucleic acid molecule, each test nucleic acid molecule at least partially overlapping a locus associated with a target short genetic variant and derived from a test sample, wherein the one or more test sequencing data sets are determined by sequencing the test nucleic acid molecules using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order, and wherein the test sequencing data sets include flow signals at a plurality of flow positions;
[0198] (c) for each test nucleic acid molecule associated with the test sequencing dataset, determining a match score indicating a likelihood that the test sequencing dataset associated with the nucleic acid molecule matches the target sequence, or a match score indicating a likelihood that the test sequencing dataset associated with the nucleic acid molecule matches the reference sequence; and
[0199] (d) using the one or more determined match scores to determine the presence or absence of the target short genetic variant in the test sample.
[0200] Embodiment 2. The method of embodiment 1, wherein obtaining comprises sequencing the test nucleic acid molecule using non-terminating nucleotides provided in separate nucleotide streams according to a flow cycle order.
[0201] Embodiment 3. The method of embodiment 1 or embodiment 2, wherein the target short genetic variant is pre-selected before determining the presence or absence of the target short genetic variant in the test sample.
[0202] Embodiment 4. The method of embodiment 1 or embodiment 2, wherein the target short genetic variant is selected after determining the presence or absence of the target short genetic variant in the test sample based on the confidence of the determination.
[0203] Embodiment 5. The method of any one of embodiments 1-4, comprising generating a personalized biomarker panel for a subject associated with the test sample, the biomarker panel comprising the target short genetic variant.
[0204] Embodiment 6. The method of any one of embodiments 1-5, comprising selecting a flow cycle sequence.
[0205] Embodiment 7. The method of any one of embodiments 1-6, wherein the target sequencing dataset is an expected target sequencing dataset, or the reference sequencing dataset is an expected reference sequencing dataset.
[0206] Embodiment 8. The method of embodiment 7, wherein the target sequence and the reference sequence are sequenced by computer to obtain the expected target sequencing data set and the expected reference sequencing data set.
[0207] Embodiment 9. The method of any one of embodiments 1-8, wherein the target sequencing data set differs from the reference sequencing data at more than two non-contiguous flow positions.
[0208] Embodiment 10. The method of any one of embodiments 1-9, wherein the target sequencing data set differs from the reference sequencing data at more than two consecutive flow positions.
[0209] Embodiment 11. The method of any of embodiments 1-10, wherein the target sequence differs from the reference sequence at X base positions, and wherein the target sequencing data set differs from the reference sequencing data at (X+2) or more continuous flow positions.
[0210] Embodiment 12. The method of Embodiment 11, wherein the (X+2) flow position difference comprises a difference between a value substantially equal to zero and a value substantially greater than zero.
[0211] Embodiment 13. The method of any one of embodiments 1-12, wherein the target sequencing dataset differs from the reference sequencing dataset across one or more flow cycles.
[0212] Embodiment 14. The method of any of embodiments 1-13, wherein the flow signal comprises a base count indicating the number of bases of the test nucleic acid molecule sequenced at each flow position.
[0213] Embodiment 15. The method of any of embodiments 1-14, wherein the flow signal includes a statistical parameter indicating the likelihood of at least one base count at each flow position, wherein the base count indicates the number of bases of the test nucleic acid molecule sequenced at the flow position.
[0214] Embodiment 16. The method of any of embodiments 1-15, wherein the flow signal includes a statistical parameter indicating the likelihood of multiple base counts at each flow position, wherein each base count indicates the number of bases of the test nucleic acid molecule sequenced at the flow position.
[0215] Embodiment 17. The method of Embodiment 16, wherein step (c) comprises:
[0216] selecting a statistical parameter at each flow position in the test sequencing data set that corresponds to the base count of the target sequence at the flow position, and determining a match score that indicates a likelihood that the test sequencing data set matches the target sequence;
[0217] or
[0218] A statistical parameter at each flow position in the test sequencing data set that corresponds to the base count of the reference sequence at the flow position is selected, and a match score is determined that indicates the likelihood that the test sequencing data set matches the reference sequence.
[0219] Embodiment 18. The method of embodiment 17, wherein the match score determined in step (c) is a combined value of statistical parameters selected across flow positions in the test sequencing data set.
[0220] Embodiment 19. The method of any one of embodiments 1-18, wherein step (c) comprises determining a match score indicating the likelihood that the test sequencing dataset matches the target sequence.
[0221] Embodiment 20. The method of any one of embodiments 1-19, wherein step (c) comprises determining a match score indicating the likelihood that the test sequencing data set matches the reference sequence.
[0222] Embodiment 21. The method of any one of Embodiments 1-20, wherein the one or more test sequencing data sets comprises a plurality of test sequencing data sets.
[0223] Embodiment 22. The method of embodiment 21, wherein the presence or absence of the target short genetic variant is determined separately for each of the one or more test sequencing data sets.
[0224] Embodiment 23. The method of embodiment 21 or 22, wherein at least a portion of the plurality of test sequencing data sets are associated with different test nucleic acid molecules having different sequencing start positions.
[0225] Embodiment 24. The method of any of Embodiments 1-23, wherein the flow cycle sequence comprises four separate flows repeated in the same order.
[0226] Embodiment 25. The method of any of Embodiments 1-24, wherein the flow cycle sequence comprises 5 or more separate flows.
[0227] Embodiment 26. The method of any one of embodiments 1-25, wherein the method is a computer-implemented method comprising:
[0228] selecting a target short genetic variant using one or more processors;
[0229] Obtaining one or more test sequencing data sets by receiving one or more test sequencing data sets at one or more processors;
[0230] determining one or more matching scores using one or more processors; and
[0231] One or more processors are used to determine the presence or absence of a target short genetic variant in a test sample.
[0232] Embodiment 27. A system comprising:
[0233] one or more processors; and
[0234] A non-transitory computer-readable medium storing one or more programs including instructions for implementing any one of the methods of embodiments 1-26.
[0235] Embodiment 28. A method for detecting short genetic variants in a test sample, comprising:
[0236] (a) obtaining one or more first test sequencing data sets, each first test sequencing data set being associated with a different test nucleic acid molecule derived from a test sample, wherein the first test sequencing data sets are determined by sequencing the one or more test nucleic acid molecules according to a first flow cycle sequence using non-termination nucleotides provided in separate nucleotide flows, and wherein the one or more first test sequencing data sets include flow signals at flow positions corresponding to the nucleotide flows;
[0237] (b) obtaining one or more second test sequencing data sets, each second test sequencing data set being associated with the same test nucleic acid molecule as the first test sequencing data set, wherein the second test sequencing data set is determined by sequencing the one or more test nucleic acid molecules according to a second flow cycle order using a non-termination nucleotide provided in a separate nucleotide flow, wherein the first flow cycle order and the second flow cycle order are different, and wherein the test sequencing data set comprises flow signals at flow positions corresponding to the nucleotide flows.
[0238] (c) for each of the first test sequencing data set and the second test sequencing data set, determining a match score for the one or more candidate sequences, wherein the match score indicates a likelihood that the first test sequencing data set, the second test sequencing data set, or both, matches a candidate sequence from the one or more candidate sequences; and
[0239] (d) using the determined match score to determine the presence or absence of the short genetic variant in the test sample.
[0240] Embodiment 29. The method of embodiment 28 comprises sequencing the test nucleic acid molecule according to a first flow cycle sequence using non-termination nucleotides provided in a separate nucleotide flow, and sequencing the test nucleic acid molecule according to a second flow cycle sequence using non-termination nucleotides provided in a separate nucleotide flow.
[0241] Embodiment 30. The method of embodiment 28 or 29, wherein the match score indicates the likelihood that the first test sequencing data set matches the candidate sequence, or the likelihood that the second test sequencing data set matches the candidate sequence.
[0242] Embodiment 31. The method of embodiment 28 or 29, wherein the match score indicates a likelihood that both the first test sequencing data set and the second sequencing data set match the candidate sequence.
[0243] Embodiment 32. The method of any one of embodiments 28-31, wherein the one or more candidate sequences include two or more different candidate sequences, and for each nucleic acid molecule associated with the first sequencing data set and the second sequencing data set, the method comprises:
[0244] selecting a candidate sequence from two or more different candidate sequences, wherein the selected candidate sequence has a highest likelihood of matching the first test sequencing dataset, the second test sequencing dataset, or both; and
[0245] The selected candidate sequences are used to call for the presence or absence of short genetic variants in the test sample.
[0246] Embodiment 33. The method of embodiment 32, wherein at least one unselected candidate sequence from two or more different candidate sequences is different from the selected candidate sequence at two or more flow positions according to the first flow cycle order or the second flow cycle order.
[0247] Embodiment 34. The method of embodiment 32, wherein at least one unselected candidate sequence from the two or more different candidate sequences is different from the selected candidate sequence at two or more flow positions according to both the first flow cycle order and the second flow cycle order.
[0248] Embodiment 35. The method of embodiment 32, wherein at least one unselected candidate sequence from the two or more different candidate sequences is different from the selected candidate sequence at two or more non-contiguous flow positions according to the first flow cycle sequence or the second flow cycle sequence.
[0249] Embodiment 36. The method of embodiment 32, wherein at least one unselected candidate sequence from the two or more different candidate sequences is different from the selected candidate sequence at two or more non-contiguous flow positions according to both the first flow cycle order and the second flow cycle order.
[0250] Embodiment 37. The method of embodiment 32, wherein at least one unselected candidate sequence from the two or more different candidate sequences is different from the selected candidate sequence at two or more consecutive flow positions according to the first flow cycle sequence or the second flow cycle sequence.
[0251] Embodiment 38. The method of embodiment 32, wherein at least one unselected candidate sequence from the two or more different candidate sequences is different from the selected candidate sequence at two or more consecutive flow positions according to both the first flow cycle order and the second flow cycle order.
[0252] Embodiment 39. The method of embodiment 32, wherein at least one unselected candidate sequence from two or more different candidate sequences is different from the selected candidate sequence at 3 or more flow positions according to the first flow cycle order or the second flow cycle order.
[0253] Embodiment 40. The method of embodiment 32, wherein at least one unselected candidate sequence from two or more different candidate sequences differs from the selected candidate sequence at 3 or more flow positions according to both the first flow cycle order and the second flow cycle order.
[0254] Embodiment 41. A method according to embodiment 32, wherein at least one unselected candidate sequence from two or more different candidate sequences differs from the selected candidate sequence at X base positions, and wherein the test sequencing data set associated with the test nucleic acid molecule differs from at least one unselected candidate sequence from the two or more different candidate sequences at (X+2) or more flow positions according to the first flow cycle order or the second flow cycle order.
[0255] Embodiment 42. The method of embodiment 32, wherein at least one unselected candidate sequence from two or more different candidate sequences differs from the selected candidate sequence at X base positions, and wherein according to both the first flow cycle order and the second flow cycle order, the test sequencing data set associated with the test nucleic acid molecule differs from at least one unselected candidate sequence from the two or more different candidate sequences at (X+2) or more flow positions.
[0256] Embodiment 43. The method of Embodiment 41 or 42, wherein the (X+2) flow position difference comprises the difference between a value substantially equal to zero and a value substantially greater than zero.
[0257] Embodiment 44. The method of embodiment 32, wherein at least one unselected candidate sequence from the two or more different candidate sequences is different from the selected candidate sequence across one or more flow cycles according to the first flow cycle order or the second flow cycle order.
[0258] Embodiment 45. The method of embodiment 32, wherein at least one unselected candidate sequence from at least one of the two or more different candidate sequences is different from the selected candidate sequence across one or more flow cycles according to both the first flow cycle order and the second flow cycle order.
[0259] Embodiment 46. The method of any one of Embodiments 28-45, wherein the flow signal comprises a base count indicating the number of bases of the test nucleic acid molecule sequenced at each flow position.
[0260] Embodiment 47. The method of any one of Embodiments 28-46, wherein the flow signal includes a statistical parameter indicating the likelihood of at least one base count at each flow position, wherein the base count indicates the number of bases of the test nucleic acid molecule sequenced at the flow position.
[0261] Embodiment 48. The method of any one of Embodiments 28-47, wherein the flow signal includes a statistical parameter indicating the likelihood of multiple base counts at each flow position, wherein each base count indicates the number of bases of the test nucleic acid molecule sequenced at the flow position.
[0262] Embodiment 49. A method according to embodiment 48, wherein for each of one or more different candidate sequences, determining a match score comprises selecting a statistical parameter corresponding to a base count of the candidate sequence at each flow position in the first test sequencing data set and the second test sequencing data set.
[0263] Embodiment 50. The method of embodiment 49, for one or more different candidate sequences, comprises: generating a candidate sequencing data set including base counts of the candidate sequence at each flow position.
[0264] Embodiment 51. The method of embodiment 50, wherein the candidate sequencing data set is generated in a computer.
[0265] Embodiment 52. The method of any one of Embodiments 49-51, wherein the match score is a combined value of the selected statistical parameter across flow positions in the first test sequencing data set and the second test sequencing data set.
[0266] Embodiment 53. The method of any one of embodiments 28-52, wherein at least a portion of the test nucleic acid molecules have different sequencing start positions.
[0267] Embodiment 54. The method of any one of Embodiments 28-52, comprising:
[0268] selecting a target short genetic variant, wherein the target sequencing dataset associated with the target sequence including the target short genetic variant differs from the reference sequencing dataset associated with the reference sequence at two or more flow positions when the target sequencing dataset is obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide flows according to a first flow cycle order or a second flow cycle order, wherein the first flow cycle order is different from the second flow cycle order, and wherein the flow positions correspond to the nucleotide flows;
[0269] The one or more candidate sequences include a target sequence and a reference sequence.
[0270] Embodiment 55. The method of embodiment 54, wherein the target short genetic variant is pre-selected before determining the presence or absence of the target short genetic variant in the test sample.
[0271] Embodiment 56. The method of embodiment 54, wherein the target short genetic variant is selected after determining the presence or absence of the target short genetic variant in the test sample based on the confidence of the determination.
[0272] Embodiment 57. The method of embodiment 56, comprising: generating a personalized biomarker panel for a subject associated with the test sample, the biomarker panel comprising a target short genetic variant present in the test sample.
[0273] Embodiment 58. The method of any one of Embodiments 54-57, wherein the reference sequencing data set is obtained by determining an expected reference sequencing data set if the reference sequence is sequenced according to the first flow cycle order or the second flow cycle order using non-termination nucleotides provided in a separate flow.
[0274] Embodiment 59. The method of any one of Embodiments 54-57, wherein the reference sequencing data set is obtained by determining an expected reference sequencing data set if the reference sequence is sequenced according to the first flow cycle sequence and the second flow cycle sequence using non-termination nucleotides provided in separate flows.
[0275] Embodiment 60. The method of any one of Embodiments 54-57, wherein the target sequence differs from the reference sequence at two or more flow positions according to both the first flow cycle sequence and the second flow cycle sequence.
[0276] Embodiment 61. The method of any one of Embodiments 54-57, wherein the target sequence differs from the reference sequence at two or more non-contiguous flow positions according to the first flow cycle sequence or the second flow cycle sequence.
[0277] Embodiment 62. The method of any one of Embodiments 54-57, wherein the target sequence differs from the reference sequence at two or more non-contiguous flow positions according to both the first flow cycle sequence and the second flow cycle sequence.
[0278] Embodiment 63. The method of any one of Embodiments 54-57, wherein the target sequence differs from the reference sequence at two or more consecutive flow positions according to the first flow cycle sequence or the second flow cycle sequence.
[0279] Embodiment 64. The method of any one of Embodiments 54-57, wherein the target sequence differs from the reference sequence at two or more consecutive flow positions according to both the first flow cycle sequence and the second flow cycle sequence.
[0280] Embodiment 65. The method of any one of Embodiments 54-57, wherein the target sequence differs from the reference sequence at three or more flow positions according to the first flow cycle sequence or the second flow cycle sequence.
[0281] Embodiment 66. The method of any one of Embodiments 54-57, wherein the target sequence differs from the reference sequence at three or more flow positions according to both the first flow cycle sequence and the second flow cycle sequence.
[0282] Embodiment 67. The method of any one of Embodiments 54-57, wherein the target sequence differs from the reference sequence across one or more flow cycles according to the first flow cycle sequence or the second flow cycle sequence.
[0283] Embodiment 68. The method of any one of Embodiments 54-57, wherein the target sequence differs from the reference sequence across one or more flow cycles according to the first flow cycle sequence and the second flow cycle sequence.
[0284] Embodiment 69. The method of any one of Embodiments 28-68, wherein the first flow cycle sequence or the second flow cycle sequence comprises four separate flows repeated in the same order.
[0285] Embodiment 70. The method of any of Embodiments 28-68, wherein the first flow cycle sequence or the second flow cycle sequence comprises 5 or more separate flows repeated in the same sequence.
[0286] Embodiment 71. The method of any one of Embodiments 28-70, comprising:
[0287] sequencing the test nucleic acid molecule, including providing non-terminating nucleotides in separate nucleotide streams according to a first flow cycle sequence, extending a sequencing primer, and detecting the presence or absence of a nucleotide incorporated into the sequencing primer after each nucleotide stream to generate a first test sequencing data set;
[0288] removing the extended sequencing primer; and
[0289] The same test nucleic acid molecule is sequenced, including providing non-terminating nucleotides in separate nucleotide streams according to a second flow cycle order, extending a sequencing primer, and detecting the presence or absence of a nucleotide incorporated into the sequencing primer after each nucleotide stream to generate a second test sequencing data set.
[0290] Embodiment 72. The method of any one of Embodiments 28-71, wherein the method is a computer-implemented method comprising:
[0291] receiving, at one or more processors, one or more first sequencing data sets;
[0292] receiving, at one or more processors, one or more second sequencing data sets;
[0293] determining a match score using one or more processors; and
[0294] One or more processors are used to determine the presence or absence of a target short genetic variant in a test sample.
[0295] Embodiment 73. A system comprising:
[0296] one or more processors; and
[0297] A non-transitory computer-readable medium storing one or more programs comprising instructions for implementing the method of any of embodiments 28-72.
[0298] Embodiment 74. The method or system of any of Embodiments 1-73, wherein the separate flows comprise a single base type.
[0299] Embodiment 75. The method or system of any of Embodiments 1-74, wherein at least one separate flow comprises 2 or 3 different base types.
[0300] Embodiment 76. The method or system of any one of embodiments 1-75, comprising generating or updating a variant call file indicating the presence, identity, or absence of short genetic variants in the test sample.
[0301] Embodiment 77. The method or system of any of Embodiments 1-76, comprising generating a report indicating the presence, identity, or absence of a short genetic variant in the test sample.
[0302] Embodiment 78. The method or system of embodiment 77, wherein the report includes text, probability, numerical or graphical output indicating the presence, identity or absence of the short genetic variant in the test sample.
[0303] Embodiment 79. The method or system of Embodiment 77 or 78, comprising providing the report to the patient or the patient's healthcare representative.
[0304] Embodiment 78. The method or system of any one of Embodiments 1-77, wherein the short genetic variants comprise single nucleotide polymorphisms.
[0305] Embodiment 79. The method or system of any one of Embodiments 1-77, wherein the short genetic variant comprises an indel.
[0306] Embodiment 80. The method or system of any one of Embodiments 1-79, wherein the test sample comprises fragmented DNA.
[0307] Embodiment 81. The method or system of any of Embodiments 1-80, wherein the test sample comprises cell-free DNA.
[0308] Embodiment 82. The method or system of Embodiment 81, wherein the cell-free DNA comprises circulating tumor DNA (ctDNA).
[0309] Embodiment 83. A method for sequencing a nucleic acid molecule, comprising:
[0310] hybridizing the nucleic acid molecule to the primer to form a hybridization template;
[0311] extending the primer using a labeled non-termination nucleotide provided in the separate nucleotide streams according to a repeated flow cycle sequence comprising five or more separate nucleotide streams; and
[0312] As the primer is extended through the nucleotide stream, the signal or absence of signal from the incorporated labeled nucleotide is detected.
[0313] Embodiment 84. The method of embodiment 83, comprising detecting a signal or the absence of a signal after each nucleotide flow.
[0314] Embodiment 85. The method of embodiment 83 or 84, comprising sequencing multiple nucleic acid molecules.
[0315] Embodiment 86. The method of embodiment 85, wherein the nucleic acid molecules in the plurality of nucleic acid molecules have different sequencing start positions relative to the locus.
[0316] Embodiment 87. The method of any one of Embodiments 83-86, wherein the test sample is cell-free DNA.
[0317] Embodiment 88. The method of any one of Embodiments 83-86, wherein the cell-free DNA comprises circulating tumor DNA (ctDNA).
[0318] Embodiment 89. The method of any one of embodiments 83-86, wherein for 50% or more of the possible SNP arrangements at 5% or more of the random sequencing starting positions, the flow cycle sequence induces a signal change at more than two flow positions.
[0319] Embodiment 90. The method of any one of Embodiments 83-86, wherein the flow cycling sequence has an efficiency of 0.6 or more bases per flow integration. Example
[0320] By reference to the following non-limiting examples provided as exemplary embodiments of the application, the application can be better understood. The following examples are provided to more fully illustrate the embodiments, but should never be construed as limiting the broad scope of the application. Although some embodiments of the application have been shown and described herein, it is apparent that these embodiments are provided only by way of example. Without departing from the spirit and scope of the present invention, it will be appreciated by those skilled in the art that many variations, changes and replacements can be made. It should be understood that the various alternatives of the embodiments described herein can be used to practice the methods described herein.
[0321] Example 1 - SNP Detection
[0322] The hypothetical nucleic acid molecule is sequenced according to the flow cycle order ATGC using the non-terminating nucleotides provided in the separate nucleotide streams to obtain Figure 1A The test sequencing data set shown in . Each value in the sequencing data set indicates the correct probability of the base count indicated at each flow position. Based on the sequencing data set, the preliminary sequence is determined to be TATGGTCGTCGA (SEQ ID NO: 1), which is mapped to the locus of the reference genome. The locus of the reference genome is related to potential haplotype sequences TATGGTCGTCGA (SEQ ID NO: 1) (H1) and TATGGTCATCGA (SEQ ID NO: 2) (H2). For each haplotype, the probability value related to the base count of the haplotype sequence of each flow position is selected. The probability of the sequencing data set of each given haplotype is determined by multiplying the probability values related to the base count of the haplotype sequence of each flow position. If H1 is the correct sequence, the logarithm probability of the sequencing data set is -0.015, and if H2 is the correct sequence, the logarithm probability of the sequencing data set is -27.008. Therefore, the sequence of H1 is selected for this nucleic acid molecule.
[0323] Example 2 - Indel Detection
[0324] The hypothetical nucleic acid molecule is sequenced according to the flow cycle order ATGC using the non-terminating nucleotides provided in the separate nucleotide streams to obtain Figure 8The test sequencing data set shown in . Each value in the sequencing data set indicates the correct possibility of the base count indicated at each flow position. Based on the sequencing data set (that is, by selecting the most likely base count at each flow position), the preliminary sequence is determined to be TATGGTCGATCG (SEQ ID NO:8), which is mapped to the locus of the reference genome. The locus of the reference genome is related to the potential haplotype sequence TATGGTCG-TCGA (SEQ ID NO:7) (H1) and TATGGTCGATCG (SEQ ID NO:8) (H2). For each haplotype, the possibility value related to the base count of the haplotype sequence of each flow position is selected. The possibility of the sequencing data set of each given haplotype is determined by multiplying the possibility value related to the base count of the haplotype sequence of each flow position. If H1 is the correct sequence, the logarithm possibility of the sequencing data set is -24.009, and if H2 is the correct sequence, the logarithm possibility of the sequencing data set is -0.015. Therefore, the sequence of H2 is selected for this nucleic acid molecule.
[0325] Example 3 - Expanded Sequencing Flow Order
[0326] More than one million extended sequencing flow sequences were tested in silico to determine their likelihood of inducing a signal change at more than two flow positions over a set of all possible SNPs (XYZ→XQZ, where Q≠Y (and Q, X, Y, and Z are each any of A, C, G, and T)). The extended flow sequences were designed to have a minimum of 12 base sequences, with all valid 2-base flow permutations, and flow sequences with sequential base repeats were removed. All possible starting positions for the flow sequence were tested to assess the sensitivity of the extended flow sequence to induce a signal change at more than two flow positions. Fig. 9 Table 2 shows exemplary results of this analysis. Fig. 9 In the , the x-axis indicates the fraction of flow phases (or fragmentation starting positions), and the y-axis indicates the fraction of SNP arrangements that induce signal changes at more than two flow positions. For approximately 10% of the reads (or flow starting positions), several flow orders induce two or more signal differences at all possible (87.5%) SNP arrangements. The four-base periodic flow only induces cyclic shifts in 42% of possible SNPs, but it does so for all reads or flow phases. A final assessment of efficiency was performed on a million-base subset of the human reference genome to establish feasibility. This is a practical measure of how effectively the flow order can extend the sequence given the pattern and bias in real tissue.
[0327] Example 4 - SNP Detection Accuracy
[0328] The genome of DNA sample NA12878 (sample available from Coriell Institute for Medical Research) was sequenced using non-terminated fluorescently labeled nucleotides according to four flow cycles (TACG). The sequencing run generated 415,900,002 reads with an average length of 176 bases. 399,804,925 reads were aligned to the hg38 reference genome (using BWA, version 0.7.17-r1188).
[0329] After alignment, reads that were perfectly aligned to the reference genome (178,634,625 reads) or reads that contained a single mismatch with the reference genome and were aligned with a mapping quality score of 20 or higher were selected (27,265,661 reads). That is, 193,904,639, for example, were excluded for further analysis due to insertion / deletion mutations (indels), multiple mismatches, or potential incorrect (artificial) alignments with the reference genome. Therefore, it was assumed that 27,265,661 reads included true positive NA12878 SNPs, as well as any false positive SNPs caused by sequencing errors. From this pool of 27,265,661 reads, sequencing reads that spanned the mismatched loci more than once were removed to reduce the impact of true positive NA12878 SNP variants, obtaining a total of 3,413,700 reads containing one mismatch at depth 1).
[0330] Each of the remaining 3,413,700 reads included mismatches that: (1) would be expected to induce a cyclic shift if the flow map flow signal was offset by one full cycle (e.g., 4 flow positions) relative to the reference based on the flow cycle order, (2) would potentially induce a cyclic shift if a different flow cycle was used (e.g., it produced a new zero or a new non-zero signal in the flow map), or (3) would not induce a cyclic shift regardless of the flow cycle order. Of the 3,413,700 mismatches, 1,184,954 (34%) caused cyclic shifts, while 1,546,588 (43%) could cause cyclic shifts with a different flow order (i.e., “potential cyclic shifts”). In comparison, theoretical expectations for random mismatches would nominally indicate 42% cyclic shift mismatches and 46% potential cyclic shift mismatches. Overall, the mismatch rate for induced cyclic shifts was 3.7×10 -5 events / base, and the mismatch rate that induced potential circular shifts was 4.8×10 -5 Table 3 shows the 10 most common single mismatches that induce circular shifts and their relative incidence percentages.
[0331] Table 3
[0332] refer to Read segment %example TTT TCT 7.18 AAA AGA 7.18 GAG GGG 4.63 CTC CCC 4.62 CAG CGG 4.12 CTG CCG 4.09 AAC AGC 3.86 GTT GCT 3.83 CAT CGT 3.63 GAT GGT 3.62
[0333] Then evaluate the performance of variant determination based on mispairing (i.e., inducing circular shift, potential inducing circular shift or not inducing and not inducing circular shift) in each of the three different categories. Use BWA to compare reads with reference genome, and use the HaplotypeCaller tool of GATK (version 4) to determine variants. Filter obtained mispairing determination by discarding variants in homopolymers longer than 10 bases or in 10 bases adjacent to homopolymers of 10 bases or longer in length.
[0334] The mismatch calls were compared with the calls generated by the Genome in a Bottle (GIAB) project for the same NA12878 to determine the accuracy #TP / (#FP+#FN+#TP) for each category. The sequencing data were randomly downsampled to the specified average genome depth. Mismatches that induce circular shifts and mismatches that potentially induce circular shifts have higher accuracy than mismatches that do not induce circular shifts, as shown in Table 4.
[0335] Table 4
[0336] Mismatch type 30x 22x 15x 8x Cyclic shift 0.9834 0.981 0.981 0.9772 No circular shift 0.9799 0.9759 0.9775 0.9696 Potential cyclic shift 0.9826 0.9808 0.9795 0.9767
Claims
1. A computer-implemented method for detecting short genetic variants in a test sample, comprising: (a) selecting a target short genetic variant using one or more processors, wherein when a target sequencing dataset and a reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order, the target sequencing dataset associated with the target sequence comprising the target short genetic variant differs from a reference sequencing dataset associated with the reference sequence at more than two flow positions; (b) receiving, at one or more processors, one or more test sequencing data sets, each test sequencing data set associated with a test nucleic acid molecule, each test nucleic acid molecule at least partially overlapping a locus associated with a target short genetic variant and derived from a test sample, wherein the one or more test sequencing data sets are determined by sequencing the test nucleic acid molecules using non-termination nucleotides provided in separate nucleotide streams according to a flow cycle order, and wherein the test sequencing data sets include flow signals at a plurality of flow positions, wherein the flow signals include a statistical parameter indicating a likelihood of at least one base count at each flow position, wherein the base count indicates a number of bases of the test nucleic acid molecule sequenced at the flow position, and wherein the statistical parameter is determined from simulated signals detected during sequencing using a machine learning algorithm; (c) determining, for each test nucleic acid molecule associated with the test sequencing dataset, using one or more processors, a match score indicating a likelihood that the test sequencing dataset associated with the test nucleic acid molecule matches the target sequence as determined by a product of the likelihood of the base counts associated with the target sequence at each flow position, or a match score indicating a likelihood that the test sequencing dataset associated with the nucleic acid molecule matches the reference sequence as determined by a product of the likelihood of the base counts associated with the reference sequence at each flow position; and (d) using one or more processors to determine the presence or absence of the target short genetic variant in the test sample using the one or more determined match scores; wherein the flow cycle sequence comprises 4 or more separate nucleotide streams, each separate nucleotide stream being a group of one non-terminal nucleotide base type, the group of one non-terminal nucleotide base type being labeled or a portion of which is labeled, and wherein the flow cycle sequence is repeated in the same order.
2. The method of claim 1, comprising sequencing the test nucleic acid molecule using non-terminating nucleotides provided in separate nucleotide streams according to a flow cycle order.
3. The method of claim 1, wherein the target short genetic variant is pre-selected before determining the presence or absence of the target short genetic variant in the test sample.
4. The method of claim 1, comprising generating, using one or more processors, a personalized biomarker panel for a subject associated with the test sample, the biomarker panel comprising target short genetic variants.
5. The method of claim 1, comprising selecting a flow cycle sequence using one or more processors.
6. The method of claim 1, wherein the target sequencing dataset and the reference sequencing dataset are obtained by sequencing the target sequence and the reference sequence in a computer.
7. The method of claim 1, wherein the target sequencing data set differs from the reference sequencing data at more than two non-contiguous flow positions.
8. The method of claim 1, wherein the target sequencing data set differs from the reference sequencing data at more than two consecutive flow positions.
9. The method of claim 1, wherein the target sequence differs from the reference sequence at X base positions, and wherein the target sequencing data set differs from the reference sequencing data at (X+2) or more consecutive flow positions.
10. The method of claim 1, wherein the target sequencing dataset differs from the reference sequencing dataset across one or more flow cycles.
11. The method of claim 1, wherein the flow signal comprises a base count indicating the number of bases of the test nucleic acid molecule sequenced at each flow position.
12. The method of claim 1, wherein the flow signal includes a statistical parameter indicating the likelihood of multiple base counts at each flow position, wherein each base count indicates the number of bases of the nucleic acid molecule sequenced at the flow position.
13. The method of claim 12, wherein step (c) comprises: selecting, using one or more processors, a statistical parameter of the base counts of the target sequence at each flow position in the test sequencing data set corresponding to the flow position, and determining a match score indicating a likelihood that the test sequencing data set matches the target sequence; or selecting, using one or more processors, statistical parameters of base counts of a reference sequence at each flow position in a test sequencing data set corresponding to the flow position, and determining, using one or more processors, a match score indicating a likelihood that the test sequencing data set matches the reference sequence.
14. The method of claim 13, wherein the match score determined in step (c) is a combined value of a selected statistical parameter across flow positions in the test sequencing data set.
15. The method of claim 1, wherein step (c) comprises determining, using one or more processors, a match score that indicates a likelihood that the test sequencing data set matches the target sequence.
16. The method of claim 1, wherein step (c) comprises determining, using one or more processors, a match score indicating a likelihood that the test sequencing data set matches the reference sequence.
17. The method of claim 1, wherein the one or more test sequencing data sets comprises a plurality of test sequencing data sets.
18. The method of claim 17, wherein the presence or absence of the target short genetic variant is determined individually for each of the one or more test sequencing data sets.
19. The method of claim 17, wherein at least a portion of the plurality of test sequencing data sets are associated with different test nucleic acid molecules having different sequencing start positions.
20. A computer-implemented method for detecting short genetic variants in a test sample, comprising: (a) receiving, at one or more processors, one or more first test sequencing data sets, each first test sequencing data set associated with a different test nucleic acid molecule derived from a test sample, wherein the first test sequencing data sets are determined by sequencing the one or more test nucleic acid molecules using non-termination nucleotides provided in separate nucleotide streams according to a first flow cycle sequence, and wherein the one or more first test sequencing data sets include flow signals at flow positions corresponding to the nucleotide streams; (b) receiving, at one or more processors, one or more second test sequencing data sets, each second test sequencing data set associated with the same test nucleic acid molecule as the first test sequencing data set, wherein the second test sequencing data set is determined by sequencing the one or more test nucleic acid molecules using non-termination nucleotides provided in separate nucleotide streams according to a second flow cycle order, wherein the first flow cycle order and the second flow cycle order are different, wherein the test sequencing data set comprises flow signals at flow positions corresponding to the nucleotide streams, wherein the flow signals comprise a statistical parameter indicating a likelihood of at least one base count at each flow position, wherein the base count indicates a number of bases of the test nucleic acid molecule sequenced at the flow position, and wherein the statistical parameter is determined from simulated signals detected during sequencing using a machine learning algorithm; (c) for each of the first test sequencing data set and the second test sequencing data set, determining, using one or more processors, a match score for the one or more candidate sequences, wherein the match score indicates a likelihood that the first test sequencing data set, the second test sequencing data set, or both match a candidate sequence from the one or more candidate sequences as determined by a product of likelihoods of base counts associated with the one or more candidate sequences at each flow position; and (d) using the one or more processors to determine the presence or absence of the short genetic variant in the test sample using the determined match score; wherein the flow cycle sequence comprises 4 or more separate nucleotide streams, each separate nucleotide stream being a group of one non-terminal nucleotide base type, the group of one non-terminal nucleotide base type being labeled or a portion of which is labeled, and wherein the flow cycle sequence is repeated in the same order.
21. The method of claim 20, comprising sequencing the test nucleic acid molecule using non-termination nucleotides provided in separate nucleotide streams according to a first flow cycle order, and sequencing the test nucleic acid molecule using non-termination nucleotides provided in separate nucleotide streams according to a second flow cycle order.
22. The method of claim 20, wherein the match score indicates a likelihood that the first test sequencing data set matches the candidate sequence, or a likelihood that the second test sequencing data set matches the candidate sequence.
23. The method of claim 20, wherein the match score indicates a likelihood that both the first test sequencing data set and the second sequencing data set match the candidate sequence.
24. The method of claim 20, wherein the one or more candidate sequences include two or more different candidate sequences, the method comprising, for each nucleic acid molecule associated with the first sequencing data set and the second sequencing data set: selecting, using one or more processors, a candidate sequence from two or more different candidate sequences, wherein the selected candidate sequence has a highest likelihood of matching with the first test sequencing data set, the second test sequencing data set, or both; and The selected candidate sequences are used to determine the presence or absence of the short genetic variant in the test sample using one or more processors.
25. The method of claim 24, wherein at least one unselected candidate sequence from the two or more different candidate sequences differs from the selected candidate sequence at two or more flow positions according to the first flow cycle order or the second flow cycle order.
26. The method of claim 24, wherein at least one unselected candidate sequence from the two or more different candidate sequences differs from the selected candidate sequence at two or more non-contiguous flow positions according to the first flow cycle sequence or the second flow cycle sequence.
27. The method of claim 24, wherein at least one unselected candidate sequence from the two or more different candidate sequences differs from the selected candidate sequence at three or more flow positions according to the first flow cycle sequence or the second flow cycle sequence.
28. The method of claim 24, wherein at least one unselected candidate sequence from the two or more different candidate sequences differs from the selected candidate sequence at X base positions, and wherein the test sequencing data set associated with the test nucleic acid molecule differs from at least one unselected candidate sequence from the two or more different candidate sequences at (X+2) or more flow positions according to the first flow cycle order or the second flow cycle order.
29. The method of claim 24, wherein at least one unselected candidate sequence from two or more different candidate sequences differs from the selected candidate sequence across one or more flow cycles according to the first flow cycle order or the second flow cycle order.
30. The method of claim 20, wherein the flow signal comprises a base count indicating the number of bases of the test nucleic acid molecule sequenced at each flow position.
31. The method of claim 20, wherein the flow signal includes a statistical parameter indicating the likelihood of a plurality of base counts at each flow position, wherein each base count indicates the number of bases of a nucleic acid molecule sequenced at that flow position.
32. The method of claim 31, wherein determining the match score comprises, for each of one or more different candidate sequences, selecting, using one or more processors, statistical parameters of base counts corresponding to the candidate sequence at each flow position in the first test sequencing data set and the second test sequencing data set.
33. The method of claim 31, comprising: For one or more different candidate sequences, a candidate sequencing data set is generated that includes base counts of the candidate sequence at each flow position.
34. The method of claim 33, wherein the candidate sequencing data set is generated on a computer.
35. The method of claim 31, wherein the match score is a combined value of the selected statistical parameter across flow positions in the first test sequencing data set and the second test sequencing data set.
36. The method of claim 20, wherein at least a portion of the test nucleic acid molecules have different sequencing start positions.
37. The method of claim 20, comprising: selecting a target short genetic variant using one or more processors, wherein the target sequencing dataset associated with the target sequence including the target short genetic variant differs from the reference sequencing dataset associated with the reference sequence at two or more flow positions when the target sequencing dataset and the reference sequencing dataset are obtained by sequencing the target sequence using non-termination nucleotides provided in separate nucleotide streams according to a first flow cycle order or a second flow cycle order, wherein the first flow cycle order is different from the second flow cycle order, and wherein the flow positions correspond to the nucleotide streams; The one or more candidate sequences include a target sequence and a reference sequence.
38. The method of claim 37, wherein the target short genetic variant is pre-selected before determining the presence or absence of the target short genetic variant in the test sample.
39. The method of claim 37, comprising generating a personalized biomarker panel for a subject associated with the test sample, the biomarker panel comprising target short genetic variants present in the test sample.
40. The method of claim 37, wherein the reference sequencing data set is obtained by determining an expected reference sequencing data set if the reference sequence is sequenced using non-terminating nucleotides provided in a separate flow according to the first flow cycle order or the second flow cycle order.
41. The method of claim 37, wherein the target sequence differs from the reference sequence at two or more flow positions according to both the first flow cycle sequence and the second flow cycle sequence.
42. The method of claim 37, wherein according to the first flow, the target sequence differs from the reference sequence at two or more non-contiguous flow positions.
43. The method of claim 37, wherein the target sequence differs from the reference sequence at three or more flow positions according to the first flow cycle sequence or the second flow cycle sequence.
44. The method of claim 37, wherein the target sequence differs from the reference sequence across one or more flow cycles according to the first flow cycle sequence or the second flow cycle sequence.
45. The method of claim 20, comprising: sequencing the test nucleic acid molecule, including providing non-terminating nucleotides in separate nucleotide streams according to a first flow cycle sequence, extending a sequencing primer, and detecting the presence or absence of incorporation of a nucleotide into the sequencing primer after each nucleotide stream to generate a first test sequencing data set; Removal of extended sequencing primers; and The same test nucleic acid molecule is sequenced, including providing non-terminating nucleotides in separate nucleotide streams according to a second flow cycle order, extending a sequencing primer, and detecting the presence or absence of incorporation of a nucleotide into the sequencing primer after each nucleotide stream to generate a second test sequencing data set.
46. The method of any one of claims 1-45, wherein the individually separated flows comprise a single base type.
47. The method of any one of claims 1-45, wherein at least one of the individually separated flows comprises 2 or 3 different base types.
48. The method of any one of claims 1-45, comprising using one or more processors to generate or update a variant call file indicating the presence, identity, or absence of short genetic variants in a test sample.
49. The method of any one of claims 1-45, comprising generating a report indicating the presence, identity, or absence of a short genetic variant in the test sample.
50. The method of claim 49, wherein the report comprises a textual, probabilistic, numerical, or graphical output indicating the presence, identity, or absence of the short genetic variant in the test sample.
51. The method of claim 49, comprising providing the report to the patient or a health care representative of the patient.
52. The method of any one of claims 1-45, wherein the short genetic variants comprise single nucleotide polymorphisms or indels.
53. A system for detecting short genetic variants in a test sample, comprising: one or more processors; and A non-transitory computer-readable medium storing one or more programs including instructions for implementing the computer-implemented method of any of claims 1-52.
Citation Information
Patent Citations
Methods for biological sample processing and analysis
US10344328B2
Mostly natural DNA sequencing by synthesis
US8772473B2
Methods and systems for sequence calling
WO2019084158A1
Mostly Natural DNA Sequencing by Synthesis
US20120046177A1
Systems and methods for identifying sequence variation
US20130073214A1