Methods and systems for determining barcode sequences

The method and system for determining barcode sequences in sequencing machines address the issue of data loss by calculating and refining matching probabilities between observed and expected barcode sequences, ensuring high assignment specificity and minimizing data loss.

WO2025245518A1PCT designated stage Publication Date: 2025-11-27RGT UNIV OF CALIFORNIA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/030911
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-24
Filing Date
2025-05-26
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Sequencing machines face a failure mode where a large fraction of reads are assigned to the 'unknown' bin due to errors in index or barcode cycles, leading to significant data loss, with up to 90% of NovaSeq X reads becoming unusable.

Method used

A method and system for determining barcode sequences using a processor to calculate matching probabilities between observed and expected barcode sequences based on event rates, iteratively refining these probabilities to accurately assign barcode sequences to their respective samples.

Benefits of technology

This approach effectively demultiplexes unassignable reads, minimizing data loss and ensuring high assignment specificity, thereby maximizing the usable data output from sequencing runs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025030911_27112025_PF_FP_ABST
    Figure US2025030911_27112025_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are methods and systems for determining barcode sequences generated from sequencing multiplexed samples. In some embodiments, the methods comprises determining a probability that an observed barcode sequence of a sequence read originates from an expected barcode sequence of a plurality of expected barcode sequences using a probability matrix and assigning the observed barcode sequence to the expected barcode sequence based on the calculated probability. The methods and systems disclosed herein can demultiplex error-prone sequences for previously unassignable reads with high yield and minimal sample crosstalk.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR DETERMINING BARCODE SEQUENCESRELATED APPLICATIONS

[0001] The present application claims priority to U.S. Provisional Patent Application Ser. No. 63 / 651,688, filed May 24, 2024, the entire content of which is hereby expressly incorporated by reference in its entirety.STATEMENT REGARDING FEDERALLY SPONSORED R&D

[0002] This invention was made with government support under Grant No. DE- AC02-05CH11231 awarded by the U.S. Department of Energy. The government has certain rights in the invention.BACKGROUNDField

[0003] The present application generally relates to the field of nucleic acid sequencing.Description of the Related Art

[0004] Sequencing machines, such as DNA sequencing machines, can have a failure mode of pooled runs, such as errors in index or barcode cycles, leading to a large fraction of reads assigned to the “unknown” bin rather than a specific library. This can lead to the loss of 40% to 90% of sequencing reads, such as NovaSeq X reads, which can make the data unusable. There is a need for methods and systems that can demultiplex unassignable reads without loss of assignment specificity.SUMMARY

[0005] Disclosed herein include methods and systems for determining barcode sequences. In some embodiments, a method for determining a barcode sequence (or one or more actions thereof) can be under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: (a) obtaining a sequence read comprising a sequence being sequenced (or a sequence, a sample sequence, a sequence of interest, or a target sequence) and a first observed barcode sequence that does not have an identical match in a plurality of expected barcode sequences. The method can comprise: (b) determining a first matching probability between the first observed barcode sequence and each expected barcode sequence of the plurality of expected barcode sequences based on, for each expected barcode sequence, Nevent rates one for each base position in a plurality of N base positions of the first observed barcode sequence (for example, by, for each expected barcode sequence, multiplying N event rates one for each base position in a plurality of N base positions of the first observed barcode sequence) An event rate can be the probability (e.g., a calculated probability) of observing a first base at a base position when the actual base is a second base. For example, the first base and the second base can be identical or different. The method can comprise: (c) identifying, from the plurality of expected barcode sequences, a first expected barcode sequence having a highest first matching probability with the first observed barcode sequence. The method can comprise: (d) assigning the first observed barcode sequence to the first expected barcode sequence having the highest first matching probability.

[0006] In some embodiments, the sequence read further comprises a second observed barcode sequence, and the method further comprises (e) determining a second matching probability between each expected barcode sequence of the plurality of expected barcode sequences and the second observed barcode sequence by, for each expected barcode sequence, multiplying N event rates one for each base position in a plurality of N base positions of the second observed barcode sequence, (f) identifying, from the plurality of expected barcode sequences, a second expected barcode sequence having a highest second matching probability with the second observed barcode sequence, and (g) assigning the second observed barcode sequence to the second expected barcode sequence having the highest second matching probability.

[0007] In some embodiments, the method further comprises retrieving the plurality of expected barcode sequences from a barcode data set. An event rate can be, for example, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.2, 0.21, 0.22, 0.23, 0.24, 0.25, 0.26, 0.27, 0.28, 0.29, 0.3, 0.31, 0.32, 0.33, 0.34, 0.35, 0.36, 0.37, 0.38, 0.39, 0.4, 0.41, 0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49, 0.5, 0.51, 0.52, 0.53, 0.54, 0.55, 0.56, 0.57, 0.58, 0.59, 0.6, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.7, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.8, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, or 0.99. In some embodiments, the method further comprises: for each base position in N, randomly assigning each event rate with an initial value such that the initial value for an event rate of the observed first base being identical to the actual second base is the highest. The initial value for such event rate can be greater than 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, or 0.9. The remaining initial values for event rates of the observed first base being different from the actual second base can be the same or different. A remaining initial value can be, for example, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, or 0.1.

[0008] In some embodiments, the event rate for each base position in N e.g., for an iteration, is pre-determined, e.g., in a prior iteration (such as the immediate prior iteration). For example, the event rate for each base position in N can be pre-determined by: (i) assigning an observed barcode sequence of each sequence read of a plurality of sequence reads to one of the plurality of expected barcode sequences. The event rate for each base position in N can be predetermined by: (ii) for each base position in N, calculating the probability of observing the first base that is expected to be the second base by determining the number of times the observed first base corresponds to the expected second base based on the assignment from step (i). In some embodiments, the method further comprises iterating steps (i) and (ii) for a number of iterations (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10 times), or until the event rates remain essentially the same (e.g., within 0.05, 0.04, 0.03, 0.02, 0.01, 0.009, 0.008, 0.007, 0.006, 0.005, 0.004, 0.003, 0.002, 0.001, or smaller) as the event rates from the previous iteration or the difference between the event rates from two consecutive iterations is below a threshold (e.g., 0.05, 0.04, 0.03, 0.02, 0.01, 0.009, 0.008, 0.007, 0.006, 0.005, 0.004, 0.003, 0.002, 0.001, or smaller). In some embodiments, the step (i) of assigning an observed barcode to an expected barcode sequence can be performed for an iteration (e.g., the first iteration) based on a Hamming distance between an observed barcode sequence and an expected barcode sequence. The maximum Hamming distance can be between 0 and N, inclusive The maximum Hamming distance can be equal to or less than 2, 3, 4, 5, 6, 7, 8, 9, or 10. The number of iterations (or the maximum number of iterations) can be pre-determined. In some embodiments, the step (ii) of calculating the probability comprises generating at least one event probability matrix (also referred to herein as a Position-Call-Ref (PCR) matrix or a table of error probabilities). An event probability matrix can comprise N submatrices each for a position in N. Each sub-matrix) can have rows and columns representing different observed bases and expected bases and event rates each corresponding to a pair of observed base and expected base.

[0009] In some embodiments, the method further comprises: for each base position in N, updating the event rate based on the assignment, and repeating steps (b)-(d). Updating the event rate can comprise calculating the probability of observing the first base that is expected to be the second base. Updating the event rate can comprise determining the number of times the observed first base corresponds to the expected second base based on the assignment.

[0010] In some embodiments, the method further comprises multiplying the highest first matching probability by a relative abundance value of the identified first expected barcode sequence. Alternatively or additionally, the method further comprises multiplying the highest second matching probability by a relative abundance value of the identified second expected barcode sequence. In some embodiments, the method further comprises comparing the highestfirst matching probability of the identified first expected barcode sequence to the sum of other matching probabilities of other expected barcode sequences.

[0011] In some embodiments, the first observed barcode sequence can be assigned to the identified first expected barcode sequence if the highest first matching probability is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 million times greater than the sum of other matching probabilities of other expected barcode sequences (also referred to herein as minratio). In some embodiments, the first observed barcode sequence is assigned to the identified first expected barcode sequence if the highest first matching probability is above a first threshold. In some embodiments, the method further comprises comparing the highest second matching probability to the sum of other matching probabilities of other expected barcode sequences. In some embodiments, the second observed barcode sequence is assigned to the identified second expected barcode sequence if the highest second matching probability is above a second threshold. In some embodiments, the first threshold and / or the second threshold (also referred to herein as minprob) is at least 10’1, 10'2, 10'3, 10'4, 10'5, IO'5 5, 10'6, 10'7, 10'8, 10'9, 10'10, or a number or a range between any two of these values. In some embodiments, the first observed barcode sequence is assigned to the identified first expected barcode sequence if the Hamming distance between the first observed barcode sequence and the identified first expected barcode sequence (also referred to herein as maxhdist) is below a certain value, optionally equal to or less than, for example, 3, 4, 5, 6, 7, 8, 9, or 10. In some embodiments, the maximum Hamming distance between the observed barcode sequence and the identified expected barcode sequence is below a certain value. Such maximum Hamming distance can be equal to or less than, for example, 3, 4, 5, 6, 7, 8, 9, or 10.

[0012] In some embodiments, the method further comprises discarding the first and second observed barcode sequences if the combination of the first observed barcode sequence and the second observed barcode sequence (also referred to herein as a chimeric pair) does not exist in the barcode data set (but exists in a ritual pool of cross products of each first expected barcode sequence and second expected barcode sequence).

[0013] In some embodiments, the method further comprises assigning the sequence read to a respective sample in a plurality of samples (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, or 50 samples) using the assigned expected barcode sequence. The first observed barcode sequence and / or the second observed barcode sequence can be sample-specific barcode sequences.

[0014] In some embodiments, the method can further comprise providing a plurality of sequence reads each comprising a sequence being sequenced (or a sequence, a sample sequence, a sequence of interest, or a target sequence) and at least one observed barcodesequence. The method can comprise comparing each observed barcode sequence of the plurality sequence reads to each of the plurality of expected barcode sequences. The method can further comprise for a sequence read in the plurality of sequence reads comprising an observed barcode sequence having an identical match in the plurality of expected barcode sequences, assigning the observed barcode sequence to the sequence read as a true barcode sequence. The method can further comprise assigning the sequence read to a respective sample in a plurality of samples using the observed barcode sequence.

[0015] In some embodiments, the method can further comprise sequencing a plurality of barcoded nucleotide molecules with a sequencing system, thereby obtaining the plurality of sequence reads. The plurality of barcoded nucleotide molecules can be pooled from a plurality of samples. Each barcode nucleotide molecule can comprise a sample-specific barcode. The plurality of samples can be different.

[0016] In some embodiments, the number of expected barcode sequences can be from about 3 to about 50. The number of expected barcode sequences can be, be about, be at least, be at least about, be at most, or be at most about, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or a range between any two of these values. The first observed barcode sequence and / or the second observed barcode sequence can be from about 3 to about 12 bases in length. A barcode sequence can be, be about, be at least, be at least about, be at most, or be at most about, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, or a number or a range between any two of these values, bases in length. In some embodiments, the first base (a base of an observed barcode sequence) is A, C, T, G, or N, and the second base (a base of an expected barcode sequence) is A, C, T, or G. Alternatively or additionally, (b) the first base is the same as or different from the second base.

[0017] Disclosed herein also include a system for determining a barcode sequence. In some embodiments, the system can comprise non-transitory memory configured to store executable instructions. The system can comprise a processor (e.g., a hardware processor or a virtual processor) in communication with the non-transitory memory. The processor can be programmed by the executable instructions to perform: receiving a sequence read comprising a sequence being sequenced and an observed barcode sequence that does not have an identical match in a plurality of expected barcode sequences. The processor can be programmed by the executable instructions to perform: determining a matching probability between the observed barcode sequence and each expected barcode sequence of the plurality of expected barcode sequences based on, for each expected barcode sequence, N event rates one for each base position in a plurality of N base positions of the observed barcode sequence (for example, by,for each expected barcode sequence, multiplying N event rates one for each base position in a plurality of N base positions of the observed barcode sequence). An event rate can be the probability of observing a first base at a base position when the actual base is a second base. The processor can be programmed by the executable instructions to perform: identifying, from the plurality of expected barcode sequences, an expected barcode sequence having a highest matching probability with the observed barcode sequence. The processor can be programmed by the executable instructions to perform: assigning the observed barcode sequence to the identified expected barcode sequence.

[0018] Disclosed herein include embodiments of a method for determining a pattern. In some embodiments, a method for determining a pattern (or one or more actions thereof) can be under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: (a) obtaining a first observed pattern that does not have an identical match in a plurality of expected patterns. The method can comprise: (b) determining a first matching probability between the first observed pattern and each expected pattern of the plurality of expected patterns, based on N event rates one for each pattern component position in a plurality of N pattern component positions of the first observed pattern. An event rate can be the probability of observing a first pattern component at a pattern component position when the actual pattern component is a second pattern component. In some embodiments, the first pattern component is the same as the second base component. In some embodiments, the first pattern component is different from the second base component. The method can comprise: (c) identifying, from the plurality of expected patterns, a first expected pattern having a highest first matching probability with the first observed pattern. The method can comprise: (d) assigning the first observed pattern to the first expected pattern having the highest first matching probability.

[0019] An event rate can be, for example, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.2, 0.21, 0.22, 0.23, 0.24,0.25, 0.26, 0.27, 0.28, 0.29, 0.3, 0.31, 0.32, 0.33, 0.34, 0.35, 0.36, 0.37, 0.38, 0.39, 0.4, 0.41,0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49, 0.5, 0.51, 0.52, 0.53, 0.54, 0.55, 0.56, 0.57, 0.58, 0.59, 0.6, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.7, 0.71, 0.72, 0.73, 0.74, 0.75,0.76, 0.77, 0.78, 0.79, 0.8, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92,0.93, 0.94, 0.95, 0.96, 0.97, 0.98, or 0.99. In some embodiments, the method further comprises: for each pattern component position in N, randomly assigning each event rate with an initial value such that the initial value for an event rate of the observed first pattern component being identical to the actual second pattern component is the highest. The initial value for such event rate can be greater than 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, or 0.9. The remaining initial values for event rates of the observed first pattern component being different from the actualsecond pattern component are the same or different. A remaining initial value can be, for example, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, or 0.1.

[0020] In some embodiments, the event rate for each pattern component position inN, e.g., for an iteration, is pre-determined (e.g., in a prior iteration (such as the immediate prior iteration). For example, the event rate for each pattern component position in N for a iteration is determined in the prior iteration by: (i) assigning each of a plurality of observed patterns to one of the plurality of expected patterns. The event rate for each pattern component position in N can be determined by: (ii) for each pattern component position in N, calculating the probability of observing the first pattern component that is expected to be the second pattern component by determining the number of times the observed first pattern component corresponds to the expected second pattern component based on the assignment from step (i). In some embodiments, the method further comprises: iterating steps (i) and (ii) for a number of iterations (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10), or until the event rates remain essentially the same (e.g., withinO.05, 0.04, 0.03, 0.02, 0.01, 0.009, 0.008, 0.007, 0.006, 0.005, 0.004, 0.003, 0.002, 0.001, or smaller) as the event rates from the previous iteration or the difference between the event rates from two consecutive iterations is below a threshold (e.g., 0.05, 0.04, 0.03, 0.02, 0.01, 0.009, 0.008, 0.007, 0.006, 0.005, 0.004, 0.003, 0.002, 0.001, or smaller) The number of iterations (or the maximum number of iterations) can be pre-determined.

[0021] In some embodiments, step (i) for an iteration (e.g., the first iteration) is performed based on a Hamming distance between an observed pattern and an expected pattern. The maximum Hamming distance can be between 0 and N, inclusive. The maximum Hamming distance can be equal to or less than 2, 3, 4, 5, 6, 7, 8, 9, or 10. The number of iterations (or the maximum number of iterations) can be pre-determined. In some embodiments, step (ii) comprises generating at least one event probability matrix (also referred to herein as a Position- Call-Ref (PCR) matrix or a table of error probabilities). An event probability matrix can comprise N sub-matrices each for a position in N. Each sub-matrix can have rows and columns representing different observed pattern components and expected pattern components and event rates each corresponding to a pair of observed pattern component and expected pattern component.

[0022] In some embodiments, the method further comprises: for each pattern component position in N, updating the event rate based on the assignment, and repeating steps (b)-(d). Updating the event rate can comprising calculating the probability of observing the first pattern component that is expected to be the second pattern component. Updating the event rate can comprise: determining the number of times the observed first pattern component corresponds to the expected second pattern component based on the assignment.

[0023] In some embodiments, obtaining the first observed pattern comprises receiving or retrieving the first observed pattern. Obtaining the first observed pattern can comprise capturing the first observed pattern using an imaging sensor. The imaging sensor can be a CCD sensor or a CMOS sensor. The imaging sensor can be comprised in a camera or a phone.

[0024] In some embodiments, determining the first matching probability between the first observed pattern and each expected pattern of the plurality of expected patterns is based on a relative abundance value of the identified first expected pattern and / or a relative abundance value of the identified second expected pattern. Determining the first matching probability between the first observed pattern and each expected pattern of the plurality of expected patterns can comprise: for each expected pattern, multiplying N event rates one for each pattern component position in a plurality of N pattern component positions of the first observed pattern.

[0025] In some embodiments, the method further comprising multiplying the highest first matching probability by a relative abundance value of the identified first expected pattern. Alternatively or additionally, the method further comprises: multiplying the highest second matching probability by a relative abundance value of the identified second expected pattern. In some embodiments, the method further comprises comparing the highest first matching probability of the identified first expected pattern to the sum of other matching probabilities of other expected patterns.

[0026] In some embodiments, the first observed pattern can be assigned to the identified first expected pattern if the highest first matching probability is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 million times greater than the sum of other matching probabilities of other expected patterns (also referred to herein as minratio). In some embodiments, the first observed pattern is assigned to the identified first expected pattern if the highest first matching probability is above a first threshold. In some embodiments, the method further comprises comparing the highest second matching probability to the sum of other matching probabilities of other expected patterns. In some embodiments, the second observed pattern is assigned to the identified second expected pattern if the highest second matching probability is above a second threshold. In some embodiments, the first threshold and / or the second threshold (also referred to herein as minprob) is at least 10’1, 10'2, 10'3, 10'4, 10'5, 10'5 5, 10'6, 10'7, 10'8, 10'9, 10'10, or a number or a range between any two of these values. In some embodiments, the first observed pattern is assigned to the identified first expected pattern if the Hamming distance between the first observed pattern and the identified first expected pattern (also referred to herein as maxhdist) is below a certain value, optionally equal to or less than, for example, 3, 4, 5, 6, 7, 8, 9, or 10. In someembodiments, the maximum Hamming distance between the observed pattern and the identified expected pattern is below a certain value. Such maximum Hamming distance can be equal to or less than, for example, 3, 4, 5, 6, 7, 8, 9, or 10.

[0027] In some embodiments, the method further comprises assigning the first observed pattern, or a sample associated with the first observed pattern, to a respective sample in a plurality of samples using the assigned expected pattern. The first observed pattern can be a sample-specific pattern.

[0028] In some embodiments, the method further comprises: providing a plurality of observed patterns. The method can comprise: comparing each observed pattern of the plurality observed patterns to each of the plurality of expected patterns. In some embodiments, the method further comprises: for an observed pattern in the plurality of observed patterns having an identical match in the plurality of expected patterns, assigning the observed pattern as a true pattern. In some embodiments, the method further comprises assigning the observed pattern, or a sample associated with the observed pattern, to a respective sample in a plurality of samples using the observed pattern. The plurality of samples can be different.

[0029] In some embodiments, a pattern (or each pattern), the pattern being an observed pattern or an expected pattern, comprises a barcode sequence. A barcode sequences can be associated with a sequence being sequenced (or a sequence, a sample sequence, a sequence of interest, or a target sequence). A barcode sequence can be associated with (e.g., comprised in) a sequence read. A sequence read can comprise a barcode sequence and a sequence being sequenced (or a sequence, a sample sequence, a sequence of interest, or a target sequence). The sequence read can comprise one or both sequence reads of paired-end sequence reads. The sequence read can comprise a single-end sequence read. The first pattern component can be A, C, T, or G (or A, C, T, G or N). The second pattern component can be A, C, T, or G.

[0030] In some embodiments, a pattern (or each pattern), the pattern being an observed pattern or an expected pattern, comprises a barcode, a ID barcode, a Universal Product Code (UPC), a 2D barcode, a matrix barcode, a two-dimensional matrix barcode, a Quick Response code (QR code), a static QR code, a dynamic QR code, a High Capacity Colored two- Dimensional code (HCC2D code), a digital code embedded in a presentation or image, a digital code represented by a presentation or image, a predefined custom pattern, or an encoding using an image (or an encoding in an image). The encoding can be decoded after the image is captured.

[0031] In some embodiments, the number of expected patterns can be from about 3 to about 50. The number of expected patterns can be, be about, be at least, be at least about, be at most, or be at most about, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40,50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or a range between any two of these values. A pattern can comprise a number of pattern components (corresponding to a number of pattern component positions), such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or a range between any two of these values, pattern components. A pattern component can be selected from one of a number of possibilities. The number of possibilities a pattern component can be selected from can be, be about, be at least, be at least about, be at most, or be at most about, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or a range between any two of these values.

[0032] Disclosed herein include embodiments of a system for determining a pattern. In some embodiments, a system for determining a pattern comprises: non-transitory memory configured to store executable instructions. The system can comprise: a processor (e.g., a hardware processor or a virtual processor) in communication with the non-transitory memory. A processor can be programmed by the executable instructions to perform: receiving a sequence read comprising a sequence being sequenced and an observed pattern that does not have an identical match in a plurality of expected patterns. The processor can be programmed by the executable instructions to perform: determining a matching probability between the observed pattern and each expected pattern of the plurality of expected patterns, by, for each expected pattern, multiplying N event rates one for each pattern component position in a plurality of N pattern component positions of the observed pattern wherein an event rate is the probability of observing a first pattern component at a pattern component position when the actual pattern component is a second pattern component. The processor can be programmed by the executable instructions to perform: identifying, from the plurality of expected patterns, an expected pattern having a highest matching probability with the observed pattern. The processor can be programmed by the executable instructions to perform: assigning the observed pattern to the identified expected pattern.

[0033] Disclosed herein include embodiments of a method of determining barcode sequences. In some embodiments, a method of determining barcode sequences (or one or more actions thereof) can be under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: (a) receiving a plurality of sequence reads each comprising a sequence (e.g., a sequence being sequenced, a sample sequence, a sequence of interest, or a target sequence) and an observed barcode sequence. A barcode sequence can be anindex sequence. An observed barcode sequence can yield (or originate) from an expected barcode sequence of a plurality of expected barcode sequences. The observed barcode sequence of a sequence read of the plurality sequences can be different from each expected barcode sequence (or all expected barcode sequences) of the plurality of expected barcode sequences. An (or each) observed barcode sequence and an (or each) expected barcode sequence can have the same length. Two (or all) observed barcode sequences can have the same length. Two (or all) expected barcode sequences can have the same length. The method can comprise (b) for each of one or more sequence reads of the plurality of sequence reads: determining a probability that the observed barcode sequence of the sequence read yields (or originates) from each expected barcode sequence of the plurality of expected barcode sequences using a probability matrix (also referred to herein as a Position-Call-Ref (PCR) matrix, event probabilities matrix, or table of error probabilities). The probability matrix can comprises a probability of an observed base (e.g., A, C, G, T, or N) yields (or originates) from each of a plurality of possible (or expected) bases (e.g., A, C, G, or T) for each base position of a plurality of base positions. The method can comprise ((b) for each of one or more sequence reads of the plurality of sequence reads): assigning the observed barcode sequence of the sequence read to an expected barcode sequence of the plurality of expected barcode sequences based on the probabilities of the observed barcode sequence of the sequence read yielding (originating) from the expected barcode sequences.

[0034] Disclosed herein include embodiments of a method of determining barcode sequences. In some embodiments, a method of determining barcode sequences (or one or more actions thereof) can be under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: (a) receiving a plurality of sequence reads each comprising a sequence (e.g., a sequence being sequenced, a sample sequence, a sequence of interest, or a target sequence) and an observed barcode sequence. A barcode sequence can be an index sequence. An observed barcode sequence can yield (or originates) from an expected barcode sequence of a plurality of expected barcode sequences. The observed barcode sequence of a sequence read of the plurality sequences can be different from each expected barcode sequence (or all expected barcode sequences) of the plurality of expected barcode sequences. An (or each) observed barcode sequence and an (or each) expected barcode sequence can have the same length. Two (or all) observed barcode sequences can have the same length. Two (or all) expected barcode sequences can have the same length. The method can comprise (b) for each of one or more sequence reads of the plurality of sequence reads: if the observed barcode sequence of the sequence read is identical to (or the same as) an expected barcode sequence: assigning the observed barcode sequence to the identical expected barcode sequence. The method cancomprise ((b) for each of one or more sequence reads of the plurality of sequence reads): if the observed barcode sequence of the sequence read is not identical to (or is different from) any expected barcode sequence of the plurality of expected barcode sequences: determining a probability that the observed barcode sequence of the sequence read yields (or originates) from each expected barcode sequence of the plurality of expected barcode sequences using a probability matrix (also referred to herein as a Position-Call-Ref (PCR) matrix, event probabilities matrix, or table of error probabilities). The probably matrix can comprise a probability of an observed base (e.g., A, C, G, T, or N) yields (or originates) from each of a plurality of possible (or expected) bases (e.g., A, C, G, or T) for each base position of a plurality of base positions. The method can comprise ((b) for each of one or more sequence reads of the plurality of sequence reads): if the observed barcode sequence of the sequence read is not identical to (or different from) any expected barcode sequence of the plurality of expected barcode sequences: assigning the observed barcode sequence of the sequence read to an expected barcode sequence of the plurality of expected barcode sequences based on the probabilities of the observed barcode sequence of the sequence read yielding (or originating) from the expected barcode sequences.

[0035] Disclosed herein include embodiments of a method of determining barcode sequences. In some embodiments, a method of determining barcode sequences (or one or more actions thereof) can be under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: (a) receiving a plurality of sequence reads each comprising a sequence (e.g., a sequence being sequenced, a sample sequence, a sequence of interest, or a target sequence) and an observed barcode sequence. A barcode sequence can be an index sequence. An observed barcode sequence can yield (or originate) from an expected barcode sequence of a plurality of expected barcode sequences. The observed barcode sequence of a sequence read of the plurality sequences can be different from each expected barcode sequence of the plurality of expected barcode sequences. An (or each) observed barcode sequence and an (or each) expected barcode sequence can have the same length. Two (or all) observed barcode sequences can have the same length. Two (or all) expected barcode sequences can have the same length. The method can comprise: iteratively, (b) for each of one or more sequence reads of the plurality of sequence reads: determining a probability that the observed barcode sequence of the sequence read yields (or originates) from each expected barcode sequence of the plurality of expected barcode sequences using a probability matrix (also referred to herein as a Position-Call-Ref (PCR) matrix, event probabilities matrix, or table of error probabilities). The probably matrix can comprise a probability of an observed base (e.g., A, C, G, T, or N) yields (or originates) from each of a plurality of possible (or expected) bases (e.g.,A, C, G, or T) for each base position of a plurality of base positions. The method can comprise ((b) for each of one or more sequence reads of the plurality of sequence reads): assigning the observed barcode sequence of the sequence read to an expected barcode sequence of the plurality of expected barcode sequences based on the probabilities of the observed barcode sequence of the sequence read yielding (or originating) from the expected barcode sequences. The method can comprise (after step (b) and part of iteratively): (c) determining an updated probability matrix based on assignments of the observed barcode sequences of the sequence reads to the expected barcode sequences if the current iteration is not the last iteration, wherein the updated probability matrix is the probability matrix of the subsequent iteration (e.g., immediate subsequent iteration).

[0036] In some embodiments, assigning the observed barcode sequence of the sequence read comprises: assigning the observed barcode sequence of the sequence read to the assigned expected barcode sequence if the probability that the observed barcode sequence of the sequence read yields from the assigned expected barcode is higher than the probabilities that the observed barcode sequence yield from the remaining expected barcode sequences of the plurality of expected barcode sequences (or the remaining expected barcode sequence(s) of the plurality of expected barcode sequences, or the remaining one or more expected barcode sequences of the plurality of expected barcode sequences, or one or more of the remaining expected barcode sequences of the plurality of expected barcode sequences).

[0037] In some embodiments, step (b) is performed for a plurality of iterations. The number of iterations can be, for example, 2, 3, 4, 5, 6, 7, 8, 9, or 10).

[0038] In some embodiments, the method further comprises: determining an updated probability matrix. In some embodiments, the method, further comprises: determining an updated probability matrix based on assignments of the observed barcode sequences of the sequence reads to the expected barcode sequences. In some embodiments, the method further comprises: determining an updated probability matrix based on the observed barcode sequences of the sequence reads assigned to each expected barcode sequence of the plurality of expected barcode sequences. In some embodiments, the method further comprises: determining an updated probability matrix based on (x) the observed barcode sequences of the sequence reads and (y) the expected barcode sequences of the plurality of expected barcode sequences to which the observed barcode sequences are assigned.

[0039] In some embodiments, the method further comprises: determining an updated probability matrix based on (x) bases (or identities of bases) of the observed barcode sequences of the sequence reads at each base position of the plurality of base positions and (y) bases (or identities of bases) of the expected barcode sequences, to which the observed barcode sequencesare assigned, at each base position of the plurality of base positions. The method can comprise: determining an updated probability matrix based on (x) bases (or identities of bases) of the observed barcode sequences of the sequence reads at each base position of the plurality of base positions and (y) one or more bases (or identities of bases) of the expected barcode sequences of the expected barcode sequences, to which the observed barcode sequences are assigned, at each base position of the plurality of base positions. The method can comprise: determining an updated probability matrix based on, for each base position of the plurality of base positions, (x) bases of the observed barcode sequences of the sequence reads at the base position and (y) bases of the expected barcode sequences, to which the observed barcode sequences are assigned, at the base position.

[0040] In some embodiments, the method further comprises: determining an updated probability matrix based on (x) a number of each base, of the observed barcode sequences of the sequence reads, at each base position of the plurality of base positions and (y) bases of the expected barcode sequences of the expected barcode sequences, to which the observed barcode sequences are assigned, at each base position of the plurality of base positions. In some embodiments,, the method further comprises: determining an updated probability matrix based on, for each base position of the plurality of base positions, (x) a number of each base and the identity of the base, of the observed barcode sequences of the sequence reads, at the base position and (y) identities of bases of the expected barcode sequences of the expected barcode sequences, to which the observed barcode sequences are assigned, at the base position.

[0041] In some embodiments, the updated probability matrix is determined in one iteration of the plurality of iterations, and wherein the updated probability matrix is used to determine probabilities of the observed barcode sequences in a subsequent iteration (e.g., immediate subsequent iteration) of the plurality of iterations. In some embodiments, the plurality of sequence reads comprise paired-end sequence reads and / or single-end sequence reads.

[0042] Disclosed herein include embodiments of a method of determining patterns. In some embodiments, a method of determining patterns (or one or more actions thereof) can be under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: (a) receiving a plurality of observed patterns. An observed pattern can yield (or originate) from an expected pattern of a plurality of expected patterns. A pattern can comprises a plurality of pattern components. Each of the plurality of pattern components is at one (or unique or a different) pattern component position of the plurality of pattern component positions (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 pattern component positions). At least one observed pattern of the plurality observedpatterns can be different from each expected pattern (or all expected patterns) of the plurality of expected patterns. For example, a number of observed patterns (e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000 observed patterns; or 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, or 50%, of the plurality of observed patterns) of the plurality observed patterns (e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000 observed patterns) can be different from each expected pattern (or all expected patterns) of the plurality of expected patterns. An (or each) observed pattern and an (or each) expected pattern can have the same number of pattern components. Two (or all) observed patterns can have the same number of pattern components. Two (or all) expected patterns can have the same number of pattern components. The method can comprise: (b) for each of one or more observed patterns of the plurality of observed patterns: determining a probability that the observed pattern yields (or originates) from each expected pattern of the plurality of expected patterns using a probability matrix (also referred to herein as a Position-Call-Ref (PCR) matrix, event probabilities matrix, or table of error probabilities). The probability matrix can comprise a probability of an observed pattern component yields (or originates) from each of a plurality of possible (or expected) pattern components for each pattern component position of a plurality of pattern component positions. The method can comprise ((b) for each of one or more observed patterns of the plurality of observed patterns): assigning the observed pattern to an expected pattern of the plurality of expected patterns based on the probabilities of the observed pattern yielding (or originating) from the expected patterns.

[0043] Disclosed herein include embodiments of a method of determining patterns. In some embodiments, a method of determining patterns (or one or more actions thereof) can be under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: (a) receiving a plurality of observed patterns. An observed pattern can yield (or originate) from an expected pattern of a plurality of expected patterns (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 pattern component positions). A pattern can comprises a plurality of pattern components. Each of the plurality of pattern components is at one (or unique or a different) pattern component position of the plurality of pattern component positions (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16,17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 pattern component positions). At least one observed pattern of the plurality observed patterns can be different from each expected pattern (or all expected patterns) of the plurality of expected patterns. For example, a number of observed patterns (e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000 observed patterns; or 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, or 50%, of the plurality of observed patterns) of the plurality observed patterns (e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000 observed patterns) can be different from each expected pattern (or all expected patterns) of the plurality of expected patterns. An (or each) observed pattern and an (or each) expected pattern can have the same number of pattern components. Two (or all) observed patterns can have the same number of pattern components. Two (or all) expected patterns can have the same number of pattern components. The method can comprise: (b) for each of one or more observed patterns of the plurality of observed patterns: if the observed pattern is identical to (or the same as) an expected pattern: assigning the observed pattern to the identical expected pattern. The method can comprise ((b) for each of one or more observed patterns of the plurality of observed patterns): if the observed pattern is not identical (or is different from) to any expected pattern of the plurality of expected patterns: determining a probability that the observed pattern yields (or originates) from each expected pattern of the plurality of expected patterns using a probability matrix, (also referred to herein as a Position-Call-Ref (PCR) matrix, event probabilities matrix, or table of error probabilities). The probably matrix can comprise a probability of an observed pattern component yields (or originates) from each of a plurality of possible pattern components for each pattern component position of a plurality of pattern component positions. The method can comprise: assigning the observed pattern to an expected pattern of the plurality of expected patterns based on the probabilities of the observed pattern yielding (or originating) from the expected patterns.

[0044] Disclosed herein include embodiments of a method of determining patterns. In some embodiments, a method of determining patterns (or one or more actions thereof) can be under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: (a) receiving a plurality of observed patterns. An observed pattern can yield (or originates) from an expected pattern of a plurality of expected patterns. A pattern can comprisesa plurality of pattern components. Each of the plurality of pattern components is at one (or unique or a different) pattern component position of the plurality of pattern component positions (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 pattern component positions). At least one observed pattern of the plurality of observed patterns can be different from each expected pattern of the plurality of expected patterns. For example, a number of observed patterns (e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000 observed patterns; or 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, or 50%, of the plurality of observed patterns) of the plurality observed patterns (e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000 observed patterns) can be different from each expected pattern (or all expected patterns) of the plurality of expected patterns. An (or each) observed pattern and an (or each) expected pattern can have the same number of pattern components. Two (or all) observed patterns can have the same number of pattern components. Two (or all) expected patterns can have the same number of pattern components. The method can comprise: iteratively, (b) for each of one or more observed patterns of the plurality of observed patterns: determining a probability that the observed pattern yields (or originates) from each expected pattern of the plurality of expected patterns using a probability matrix (also referred to herein as a Position-Call-Ref (PCR) matrix, event probabilities matrix, or table of error probabilities). The probably matrix can comprise a probability of an observed pattern component yields (or originates) from each of a plurality of possible pattern components for each pattern component position of a plurality of pattern component positions. The method can comprise ((b) for each of one or more observed patterns of the plurality of observed patterns): assigning the observed pattern of the sequence read to an expected pattern of the plurality of expected patterns based on the probabilities of the observed pattern yielding (or originating) from the expected patterns. The method can comprise (after step (c) and part of iteratively): (c) determining an updated probability matrix based on assignments of the observed patterns of the sequence reads to the expected patterns if the current iteration is not the last iteration, wherein the updated probability matrix is the probability matrix of the subsequent iteration (e.g., immediate subsequent iteration).

[0045] In some embodiments, assigning the observed pattern comprises: assigning the observed pattern to the assigned expected pattern if the probability that the observed pattern originates from the assigned expected barcode is higher than the probabilities that the observed pattern originate from the remaining expected patterns of the plurality of expected patterns (or the remaining expected pattern(s) of the plurality of expected patterns, or the remaining one or more expected patterns of the plurality of expected patterns, or one or more of the remaining expected patterns of the plurality of expected patterns).

[0046] In some embodiments, step (b) is performed for a plurality of iterations. The number of iterations can be, for example, 2, 3, 4, 5, 6, 7, 8, 9, or 10).

[0047] In some embodiments, the method further comprising: determining an updated probability matrix. In some embodiments, the method further comprises: determining an updated probability matrix based on assignments of the observed patterns to the expected patterns. In some embodiments, the method further comprises: determining an updated probability matrix based on the observed patterns assigned to each expected pattern of the plurality of expected patterns. In some embodiments, the method further comprises: determining an updated probability matrix based on (x) the observed patterns and (y) the expected patterns of the plurality of expected patterns to which the observed patterns are assigned.

[0048] In some embodiments, the method further comprises: determining an updated probability matrix based on (x) pattern components (or identifies of pattern components) of the observed patterns at each pattern component position of the plurality of pattern component positions and (y) pattern components (or identities of pattern components) of the expected patterns, to which the observed patterns are assigned, at each pattern component position of the plurality of pattern component positions. The method can comprise: determining an updated probability matrix based on (x) pattern components (or identities of pattern components) of the observed patterns at each pattern component position of the plurality of pattern component positions and (y) one or more pattern components (or identities of pattern components) of the expected patterns of the expected patterns, to which the observed patterns are assigned, at each pattern component position of the plurality of pattern component positions. The method can comprise: determining an updated probability matrix based on, for each pattern component position of the plurality of pattern component positions, (x) pattern components of the observed patterns at the pattern component position and (y) pattern components of the expected patterns, to which the observed patterns are assigned, at the pattern component position.

[0049] In some embodiments, the method further comprises: determining an updated probability matrix based on (x) a number of each pattern component, of the observed patterns, at each pattern component position of the plurality of pattern component positions and (y) patterncomponents of the expected patterns of the expected patterns, to which the observed patterns are assigned, at each pattern component position of the plurality of pattern component positions. In some embodiments, the method further comprises: determining an updated probability matrix based on, for each pattern component position of the plurality of pattern component positions, (x) a number of each pattern component and the identity of the pattern component, of the observed patterns, at the pattern component position and (y) identities of pattern components of the expected patterns of the expected patterns, to which the observed patterns are assigned, at the pattern component position.

[0050] In some embodiments, the updated probability matrix is determined in one iteration of the plurality of iterations, and wherein the updated probability matrix is used to determine probabilities of the observed patterns in a subsequent iteration (e.g., immediate subsequent iteration) of the plurality of iterations.

[0051] In some embodiments, a pattern (or each pattern), the pattern being an observed pattern or an expected pattern, comprises a barcode sequence. A barcode sequences can be associated with a sequence being sequenced (or a sequence, a sample sequence, a sequence of interest, or a target sequence). A barcode sequence can be associated with (e.g., comprised in) a sequence read. A sequence read can comprise a barcode sequence and a sequence being sequenced (or a sequence, a sample sequence, a sequence of interest, or a target sequence). The sequence read can comprise one or both sequence reads of paired-end sequence reads. The sequence read can comprise a single-end sequence read. The first pattern component can be A, C, T, or G (or A, C, T, G or N). The second pattern component can be A, C, T, or G.

[0052] In some embodiments, a pattern (or each pattern), the pattern being an observed pattern or an expected pattern, comprises a barcode, a ID barcode, a Universal Product Code (UPC), a 2D barcode, a matrix barcode, a two-dimensional matrix barcode, a Quick Response code (QR code), a static QR code, a dynamic QR code, a High Capacity Colored two- Dimensional code (HCC2D code), a digital code embedded in a presentation or image, a digital code represented by a presentation or image, a predefined custom pattern, or an encoding using an image (or an encoding in an image). The encoding can be decoded after the image is captured.

[0053] In some embodiments, obtaining the first observed pattern comprises receiving or retrieving the first observed pattern. Obtaining the first observed pattern can comprise capturing the first observed pattern using an imaging sensor. The imaging sensor can be a CCD sensor or a CMOS sensor. The imaging sensor can be comprised in a camera or a phone.

[0054] Disclosed herein include devices for performing any method of the present disclosure. Disclosed herein include a computer readable medium comprising executable instructions that when executed by a processor (e.g., a hardware processor or a virtual processor) programs the processor to perform any method of the present disclosure. Disclosed herein include a computer readable medium comprising codes representing a neural network constructed using any method of the present disclosure.

[0055] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Neither this summary nor the following detailed description purports to define or limit the scope of the inventive subject matter.BRIEF DESCRIPTION OF THE DRAWINGS

[0056] FIG. 1 illustrates a non-limiting schematic of a sequence read used in some embodiments.

[0057] FIG. 2 depicts an exemplary scenario showing weakness of Hamming distance approach.

[0058] FIG. 3 depicts a non-limiting exemplary embodiment of calculating matching probability between an observed barcode sequence and an expected barcode sequence using a probability matrix.

[0059] FIG. 4 depicts a non-limiting exemplary embodiment of updating a probability matrix (i.e., PCR (Position-Call-Ref) matrix iteration).

[0060] FIG. 5 depicts a plot showing lane yield at various demultiplexing conditions using the disclosed probabilistic method in comparison to Hamming distance approach (HDistO, HDistl, and HDist2). These Hamming distances are per barcode, not per pair. Yield is the fraction of reads assignable to a library. HDist2 is shown dashed because it incurs the highest crosstalk.

[0061] FIG. 6 depicts a plot showing average crosstalk after demultiplexing at various demultiplexing conditions. The average crosstalk is similar for HDistl and the disclosed probabilistic method, and slightly lower when requiring perfect matches only. Presumably it is mainly driven by factors other than barcode misidentification.

[0062] FIG. 7 depicts a plot showing a pooled lane crosstalk correlation of two demultiplexing methods. Lane 5 data from FIG. 6 is used for the correlation analysis. Each library is a blue dot; the red dots are linear markers. Blue dots below the red linear markers indicate lower crosstalk (better performance) for the probabilistic demultiplexing.

[0063] FIG. 8 depicts a plot showing a pooled lane crosstalk correlation of two demultiplexing methods. Lane 6 data from FIG. 6 is used for the correlation analysis. Values above 4000 PPM are ignored as artifacts of physical contamination (of libraries or reagents) or high (>99%) identity between two organisms.

[0064] FIG. 9 depicts a plot showing the cumulative sum of per-library contamination in PPM after sorting the data in an ascending order.

[0065] FIG. 10 depicts a non-limiting exemplary embodiment of a probability matrix showing Hamming Distance being a poor proxy for probability. An observed base “A” at position 13 or 17 is more likely to originate from a reference “C” than from a reference “A”.

[0066] FIG. 11 is a block diagram of an illustrative computing system that can be used in some embodiments to execute the processes and implement the features described herein, for example, pattern (e.g., barcode sequence) determination.

[0067] Throughout the drawings, reference numbers may be re-used to indicate correspondence between referenced elements. The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the disclosure.DETAILED DESCRIPTION

[0068] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the Figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein and made part of the disclosure herein.

[0069] All patents, published patent applications, other publications, and sequences from GenBank, and other databases referred to herein are incorporated by reference in their entirety with respect to the related technology.

[0070] Sequencing machines, such as DNA sequencing machines, can have a failure mode of pooled runs, such as errors in index (or barcode) cycles, leading to a large fraction of reads assigned to the ‘unknown’ bin rather than a specific library. This can lead to the loss of 40% to 90% of sequencing reads, such as NovaSeq X reads, which can make the data unusable.

[0071] Provided herein include methods and systems for determining barcode sequences. The methods and systems disclosed herein can demultiplex barcode sequences for atleast a portion of previously unassignable reads without loss of assignment specificity. In some embodiments, the present methods and systems can demultiplex about 70% of previously unassignable reads without loss of assignment specificity. The methods and systems disclosed herein can be applied to any sequencing machine, such as ILLUMINA™, Inc. sequencing machines.

[0072] The methods and systems disclosed herein can be used to correct sequencing data generated from the current generation of sequencing machines that contain one or more errors. In some embodiments, the methods and systems disclosed herein can recover data from well-functioning machines post-sequencing, thus saving users time and energy, and preventing loss of irreplaceable data in case samples are destroyed during sequencing.

[0073] The methods and systems disclosed herein can greatly increase yield and decrease crosstalk for sequencing data. In some embodiments, the present invention can also be used in other fields, such as barcode scanning, image processing, text scanning, and the like.

[0074] Disclosed herein include a method for determining a barcode sequence. The method can comprise (a) obtaining a sequence read comprising a sequence being sequenced and an observed barcode sequence that does not have an identical match in a plurality of expected barcode sequences, (b) determining a matching probability between the observed barcode sequence and each expected barcode sequence of the plurality of expected barcode sequences, by, for each expected barcode sequence, multiplying N event rates one for each base position in a plurality of N base positions of the observed barcode sequence, wherein an event rate is the probability of observing a first base at a base position when the actual base is a second base, (c) identifying, from the plurality of expected barcode sequences, an expected barcode sequence having a highest matching probability with the observed barcode sequence, and (d) assigning the observed barcode sequence to the expected barcode sequence having the highest matching probability.

[0075] Disclosed herein also include a system for determining a barcode sequence. The system can comprise non-transitory memory configured to store executable instructions and a processor in communication with the non-transitory memory, the processor programmed by the executable instructions to perform: receiving a sequence read comprising a sequence being sequenced and an observed barcode sequence that does not have an identical match in a plurality of expected barcode sequences, determining a matching probability between the observed barcode sequence and each expected barcode sequence of the plurality of expected barcode sequences, by, for each expected barcode sequence, multiplying N event rates one for each base position in a plurality of N base positions of the observed barcode sequence, wherein an event rate is the probability of observing a first base at a base position when the actual base is a secondbase, identifying, from the plurality of expected barcode sequences, an expected barcode sequence having a highest matching probability with the observed barcode sequence, and assigning the observed barcode sequence to the expected barcode sequence having the highest matching probability.

[0076] Disclosed herein also include a method of determining barcode sequences. The method can be under control of a processor and comprise: (a) receiving a plurality of sequence reads each comprising a sequence being sequenced and an observed barcode sequence, wherein an observed barcode sequence yields from an expected barcode sequence of a plurality of expected barcode sequences, and wherein the observed barcode sequence of a sequence read of the plurality sequences is different from each expected barcode sequence of the plurality of expected barcode sequences; (b) for each of one or more sequence reads of the plurality of sequence reads: determining a probability that the observed barcode sequence of the sequence read originates from each expected barcode sequence of the plurality of expected barcode sequences using a probability matrix, wherein the probability matrix comprises a probability of an observed base yielding from each of a plurality of possible bases for each base position of a plurality of base positions; and assigning the observed barcode sequence of the sequence read to an expected barcode sequence of the plurality of barcode sequences based on the probabilities of the observed barcode sequence of the sequence read yielding from the expected barcode sequences.

[0077] Disclosed herein also include a method of determining barcode sequences. The method can be under control of a processor and comprise: (a) receiving a plurality of sequence reads each comprising a sequence being sequenced and an observed barcode sequence, wherein an observed barcode sequence yields from an expected barcode sequence of a plurality of expected barcode sequences, and (b) for each of one or more sequence reads of the plurality of sequence reads: if the observed barcode sequence of the sequence read is identical to an expected barcode sequence: assigning the observed barcode sequence to the identical expected barcode sequence, if the observed barcode sequence of the sequence read is not identical to any expected barcode sequence of the plurality of expected barcode sequences: determining a probability that the observed barcode sequence of the sequence read originates from each expected barcode sequence of the plurality of expected barcode sequences using a probability matrix, wherein the probably matrix comprises a probability of an observed base yielding from each of a plurality of possible bases for each base position of a plurality of base positions; and assigning the observed barcode sequence of the sequence read to an expected barcode sequence of the plurality of barcode sequences based on the probabilities of the observed barcode sequence of the sequence read yielding from the expected barcode sequences.

[0078] Disclosed herein also include a method of determining barcode sequences. The method can be under control of a processor and comprise: (a) receiving a plurality of sequence reads each comprising a sequence being sequenced and an observed barcode sequence, wherein an observed barcode sequence yields from an expected barcode sequence of a plurality of expected barcode sequences, and wherein the observed barcode sequence of a sequence read of the plurality sequences is different from each expected barcode sequence of the plurality of expected barcode sequences; iteratively, (b) for each of one or more sequence reads of the plurality of sequence reads: determining a probability that the observed barcode sequence of the sequence read originates from each expected barcode sequence of the plurality of expected barcode sequences using a probability matrix, wherein the probably matrix comprises a probability of an observed base yields from each of a plurality of possible bases for each base position of a plurality of base positions; assigning the observed barcode sequence of the sequence read to an expected barcode sequence of the plurality of barcode sequences based on the probabilities of the observed barcode sequence of the sequence read yielding from the expected barcode sequences; (c) determining an updated probability matrix based on assignments of the observed barcodes sequences of the sequence reads to the expected barcode sequences if the current iteration is not the last iteration, wherein the updated probability matrix is the probability matrix of the subsequent iteration.

[0079] Disclosed herein include a method for determining a pattern, comprising: (a) obtaining a first observed pattern that does not have an identical match in a plurality of expected patterns; (b) determining a first matching probability between the first observed pattern and each expected pattern of the plurality of expected patterns, based on N event rates one for each pattern component position in a plurality of N pattern component positions of the first observed pattern, wherein an event rate is the probability of observing a first pattern component at a pattern component position when the actual pattern component is a second pattern component; (c) identifying, from the plurality of expected patterns, a first expected pattern having a highest first matching probability with the first observed pattern ; and (d) assigning the first observed pattern to the first expected pattern having the highest first matching probability.

[0080] Disclosed herein include method of determining patterns comprising: under control of a processor: (a) receiving a plurality of observed patterns, wherein an observed pattern yields from an expected pattern of a plurality of expected patterns, and wherein at least one observed pattern of the plurality observed patterns is different from each expected pattern of the plurality of expected patterns; (b) for each of one or more observed patterns of the plurality of observed patterns: determining a probability that the observed pattern originates from each expected pattern of the plurality of expected patterns using a probability matrix, wherein theprobability matrix comprises a probability of an observed pattern component yields from each of a plurality of possible pattern components for each pattern component position of a plurality of pattern component positions; assigning the observed pattern to an expected pattern of the plurality of expected patterns based on the probabilities of the observed pattern yielding from the expected patterns.

[0081] Disclosed herein include a method of determining patterns comprising: under control of a processor: (a) receiving a plurality of observed patterns, wherein an observed pattern yields from an expected pattern of a plurality of expected patterns, and wherein at least one observed pattern of the plurality observed patterns is different from each expected pattern of the plurality of expected patterns; and (b) for each of one or more observed patterns of the plurality of observed patterns: if the observed pattern is identical to an expected pattern: assigning the observed pattern to the identical expected pattern, if the observed pattern is not identical to any expected pattern of the plurality of expected patterns: determining a probability that the observed pattern yields from each expected pattern of the plurality of expected patterns using a probability matrix, wherein the probably matrix comprises a probability of an observed pattern component yields from each of a plurality of possible pattern components for each pattern component position of a plurality of pattern component positions; and assigning the observed pattern to an expected pattern of the plurality of expected patterns based on the probabilities of the observed pattern yielding from the expected patterns.

[0082] Disclosed herein include a method of determining patterns comprising: under control of a processor: (a) receiving a plurality of observed patterns, wherein an observed pattern yields from an expected pattern of a plurality of expected patterns, and wherein an observed pattern of the plurality of observed patterns is different from each expected pattern of the plurality of expected patterns; iteratively, (b) for each of one or more observed patterns of the plurality of observed patterns: determining a probability that the observed pattern yields from each expected pattern of the plurality of expected patterns using a probability matrix, wherein the probably matrix comprises a probability of an observed pattern component yields from each of a plurality of possible pattern components for each pattern component position of a plurality of pattern component positions; assigning the observed pattern of the sequence read to an expected pattern of the plurality of expected patterns based on the probabilities of the observed pattern yielding from the expected patterns; (c) determining an updated probability matrix based on assignments of the observed patterns of the sequence reads to the expected patterns if the current iteration is not the last iteration, wherein the updated probability matrix is the probability matrix of the subsequent iteration.Definitions

[0083] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs. See, e.g. Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For purposes of the present disclosure, the following terms are defined below.

[0084] As used herein, the term “about” means a range of values including the specified value, which a person of ordinary skill in the art would consider reasonably similar to the specified value. In embodiments, the term “about” means within a standard deviation using measurements generally acceptable in the art. In embodiments, about means a range extending to + / - 10% of the specified value. In embodiments, about means the specified value.

[0085] As may be used herein, the terms “nucleic acid,” “nucleic acid molecule,” “nucleic acid sequence,” “nucleic acid fragment” and “polynucleotide” are used interchangeably and are intended to include, but are not limited to, a polymeric form of nucleotides covalently linked together that may have various lengths, either deoxyribonucleotides or ribonucleotides, or analogs, derivatives or modifications thereof. Different polynucleotides may have different three-dimensional structures, and may perform various functions, known or unknown. Nonlimiting examples of polynucleotides include a gene, a gene fragment, an exon, an intron, intergenic DNA (including, without limitation, heterochromatic DNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, a ribozyme, cDNA, a recombinant polynucleotide, a branched polynucleotide, a plasmid, a vector, isolated DNA of a sequence, isolated RNA of a sequence, a nucleic acid probe, and a primer. Polynucleotides useful in the methods of the disclosure may comprise natural nucleic acid sequences and variants thereof, artificial nucleic acid sequences, or a combination of such sequences. As may be used herein, the terms “nucleic acid oligomer” and “oligonucleotide” are used interchangeably and are intended to include, but are not limited to, nucleic acids having a length of 200 nucleotides or less. In some embodiments, an oligonucleotide is a nucleic acid having a length of 2 to 200 nucleotides, 2 to 150 nucleotides, 5 to 150 nucleotides or 5 to 100 nucleotides.

[0086] As used herein, the term “barcode” or “index” or “unique molecular identifier (UMI)” refers to a known nucleic acid sequence that allows some feature with which the barcode is associated to be identified. Typically, a barcode is unique to a particular feature in a pool of barcodes that differ from one another in sequence, and each of which is associated with a different feature. In embodiments, barcodes are about or at least about 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75 or more nucleotides in length. In embodiments, barcodes are shorter than20, 15, 10, 9, 8, 7, 6, or 5 nucleotides in length. In embodiments, barcodes are 10-50 nucleotides in length, such as 15-40 or 20-30 nucleotides in length. In a pool of different barcodes, barcodes may have the same or different lengths. In general, barcodes are of sufficient length and comprise sequences that are sufficiently different to allow the identification of associated features based on barcodes with which they are associated. In some embodiments, a barcode can be identified accurately after the mutation, insertion, or deletion of one or more nucleotides in the barcode sequence, such as the mutation, insertion, or deletion of 1, 2, 3, 4, 5, or more nucleotides. In some embodiments, each barcode in a plurality of barcodes differs from every other barcode in the plurality by at least three nucleotide positions, such as at least 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotide positions.

[0087] As used herein, the terms “sequencing”, “sequence determination”, “determining a nucleotide sequence”, and the like include determination of a partial or complete sequence information (e.g., a sequence) of a polynucleotide being sequenced, and particularly physical processes for generating such sequence information. That is, the term includes sequence comparisons, consensus sequence determination, contig assembly, fingerprinting, and like levels of information about a target polynucleotide, as well as the express identification and ordering of nucleotides in a target polynucleotide. The term also includes the determination of the identification, ordering, and locations of one, two, or three of the four types of nucleotides within a target polynucleotide. In some embodiments, a sequencing process described herein comprises contacting a template and an annealed primer with a suitable polymerase under conditions suitable for polymerase extension and / or sequencing. The sequencing methods are preferably carried out with the target polynucleotide arrayed on a solid substrate. Multiple target polynucleotides can be immobilized on the solid support through linker molecules, or can be attached to particles, e.g., microspheres, which can also be attached to a solid substrate. In embodiments, the solid substrate is in the form of a chip, a bead, a well, a capillary tube, a slide, a wafer, a filter, a fiber, a porous media, or a column. In some embodiments, the solid substrate is gold, quartz, silica, plastic, glass, diamond, silver, metal, or polypropylene. In some embodiments, the solid substrate is porous.

[0088] As used herein, the term “sequencing read” is used in accordance with its plain and ordinary meaning and refers to an inferred sequence of nucleotide bases (or nucleotide base probabilities) corresponding to all or part of a single polynucleotide fragment. A sequencing read may include 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, or more nucleotide bases. In some embodiments, a sequencing read includes a barcode and a template nucleotide sequence. In some embodiments, a sequencing read includes reading a barcode and not a template nucleotide sequence.Barcode Multiplexing and Demultiplexing

[0089] In order to reduce the time and cost associated with amplification and sequencing of nucleic acid sequences collected from a sample, multiple samples can be pooled (multiplexed) into the same run of a sequencing instrument. These multiple samples can be collected from different sources or different subjects. After sequencing, the sequenced reads are separated (demultiplexed) into sample-specific data files by sorting each read according to a sample-specific barcode sequence (index sequence) embedded into the amplified strands (see, for example, FIG. 1).

[0090] Barcodes are sequences incorporated into nucleic acid sequences and can be used to identify a sample from which the nucleic acids are taken. In general, every barcode in a set of barcodes is unique, that is, any two barcodes chosen out of a given set differ in at least one nucleotide position. Each barcode set includes at least one unique barcode for each sample desired to be processed in parallel. As would be understood by a person skilled in the art, DNA barcodes in a library can be designed to satisfy certain constraints related to, for example, GC content, homopolymer length, and Hamming distance using any suitable barcode designing process and applications.

[0091] Generally, if there were no errors in reading the barcode sequences, the demultiplexing process would be error-free, and an observed barcode sequence can be assigned to an actual barcode sequence (expected barcode sequence) in a barcode set based on sequence identity. However, in practice, errors can arise in various parts of the process, including barcode oligonucleotide synthesis, barcode oligonucleotide purification, sample preparation, and the sequencing process, causing the barcodes to be read incorrectly.

[0092] For example, the actual barcode AAAA may be read incorrectly as AAAG. In a simple scenario with 4 expected barcodes - AAAA, CCCC, GGGG, and TTTT, one could measure the observed barcode AAAG as being one mismatch away from AAAA, four mismatches from CCCC, three mismatches from GGGG, and four from TTTT. In this case, it may be safe to assume the observed barcode AAAG had one read error (i.e., AAAA instead of AAAG) and assign the observed barcode AAAG to the actual barcode AAAA.

[0093] These errors can complicate the demultiplexing process and cause demultiplexing errors, such as a sequence read belonging to sample A is incorrectly assigned to sample B because the error caused its barcode to appear identical to that of sample B or more like the barcode for sample B than the barcode for sample A. In some scenarios, a barcode sequence that does not correspond to any barcode in the set of barcodes is either discarded or assigned to an “unknown” bin. Demultiplexing errors can be considered as a source of samplecross-contamination that need to be avoided in applications that require high accuracy such as clinical diagnostics. Demultiplexing errors can also cause a large fraction of reads discarded or assigned to an “unknown” bin rather than a specific library, leading to the loss of sequence reads.

[0094] Existing barcode demultiplexing techniques consider that barcodes are distinct from one another by a specified minimum edit distance, also referred to a Hamming distance, which is the number of mismatches between two barcodes or the number of base substitutions required to transform one barcode sequence to another. For example, the Hamming distance of the sequences AACC and AACT is 1, the Hamming distance of the sequences AACC and GTCC is 2, and the Hamming distance of the sequences AACC and GTGC is 3.

[0095] In some cases, a given observed barcode sequence that differs from a barcode in the set by a Hamming distance of no more than 1, no more than 2, or no more than 3, or no more than 4 can be assigned to that barcode, as long as there is no other barcode in the set to which the observed barcode has a closer Hamming distance. In some cases, observed barcode sequences are assigned to actual barcode sequence with the fewest barcode mismatches, i.e., minimum Hamming distance. In some scenarios, to prevent cross-contamination barcodes with 3 or more mismatches away from any barcode in the set will be discarded as unrecognizable. For example, Illumina software demultiplexes data allowing a maximum of 0, 1, or 2 barcode mismatches (N-Mismatch Demultiplexing), and discards barcode sequences with 3 or more mismatches.

[0096] Barcodes are typically designed so that any pair has a Hamming distance (number of mismatches) of at least H between them, wherein H can be 1, 2, 3, or higher. If H=l, then a single sequence error can change one barcode into another valid one. At H=2, any error will yield an invalid barcode, but it may still be one Hamming distance (HD1) away from two different possible barcodes that cannot be distinguished. A Hamming distance of 3 or greater is needed to guarantee that any single error will still yield a barcode closest to its original. Then one can demultiplex allowing one mismatch and safely assign reads to libraries under the assumption that no barcode has more than one error in the sequence. This assumption is often not true in some sequencing systems such as NovaSeq X platform. Additionally, the Hamming distance based barcode approach also incorrectly assumes that the errors in the measured barcode sequences are uniform in any position and base composition.

[0097] As shown in FIG. 2, a pooled sequencing library is multiplexed with two barcodes {AACC, AATT} having a Hamming distance of 2. Considering the mutants shown in FIG. 2, an error can be recoverable depending on the position and type of the error. Based on the Hamming distance rule, a given observed barcode of AACT with one error at position 4 can beassigned to either AACC or AATT, therefore unrecoverable. An observed barcode of GACC with one error at position 1 can be correctly assigned to AACC. An observed barcode of AATT with two errors (AACC~ AATT) may be misassigned to AATT.

[0098] It is also notable that while an observed AAAG may arise from a real AAAA due to one error, it could also arise from a real GGGG via three errors. These events are both possible, and which is more likely is assumed to be the one involving one error, but the probability is not quantified in N-Mismatch demultiplexing. As a result, the maximum number of mismatches allowed must be chosen conservatively, and depending on the error profile, there is a possibility that a 3-mismtach event is more likely than a 1-mismatch event, ruining the assumptions behind N-Mismatch demultiplexing, and leading to low yield and / or high contamination.

[0099] The methods and systems disclosed reduce demultiplexing errors by calculating the frequency of different error types for different positions to determine the probability that an expected barcode yields an observed barcode. With high error rates in sequencing, this approach can dramatically increase yield of DNA sequencing machines. For example, for a given barcode position, a true base A may yield an observed A at a 80% rate, an observed C at a 1% rate, an observed G at a 19% rate, and an observed T at a 0% rate. The exact probability of a specific expected barcode yielding a given observed barcode can thus be determined by multiplying all positional event probabilities for the pair, and multiplying by the relative abundance of the expected barcode. By comparing the observed barcode to all possible expected barcodes, the methods can calculate the probability of each expected barcode emitting the given observed barcode. Once all the probabilities are determined, the observed barcode can be provisionally assigned to the expected barcode having the highest probability. In some embodiments, the highest probability event may not correspond to the event with the lowest Hamming distance. In contrast to the Hamming distance based demultiplexing approach which assigns or discards a barcode based on a minimum Hamming distance, the instant methods makes the match based on two factors: the absolute event probability and the ratio of the highest probability to the sum of all other probabilities. An event having a ratio above a certain threshold (e.g., 106) will be considered as a matched event between the observed barcode and the expected barcode. An event having a ratio below the certain threshold will be discarded.

[0100] The methods and systems described herein allow the actual quantification of tradeoffs between yield and contamination, and with fixed settings, can automatically adjust the decisions to keep or discard observed barcodes based on error rates and the specific set of barcodes pooled together in a given run, to maximize the yield without exceeding a specified threshold of crosstalk from barcode misassignment.

[0101] The methods and systems described herein can significantly improve the data loss issue from a sequencing system which typically loses 40% to 90% data to the “UNKNOWN” bin (unassigned reads) due to mismatch between expected barcode sequences and observed barcode sequences. Increasing the lowest Hamming distance (e.g., from 1 to 2) only recovers an additional 20-30% of the data, but substantially increases crosstalk especially for crosstalk-sensitive applications.

[0102] In some embodiments, the methods and systems disclosed herein can recover at least 50% (e.g., at least 50%, 60%, 70%, 80%, 90%, or 95%) unassigned reads. In some embodiments, the methods and systems disclosed herein can reduce the number of unassigned reads by at least or at least about 50%, 60%, 70%, 80%, 90%, 95%, or a number or a range between any two of these values. In some embodiments, the methods and systems disclosed herein can increase yield of the sequence reads by at least about 2-fold (e.g., 2-fold, 3-fold, 4- fold, 5-fold, 6-fold, 7-fold, 8-fold, or a number or a range between any of these values.) In an exemplary embodiment, the methods and systems disclosed herein increases yield from one pool from about 13% usable to about 70% usable. In another exemplary embodiment, the methods and systems disclosed herein increases yield from one pool from about 40% usable to about 90% usable.

[0103] The methods and systems disclosed herein can reduce the demultiplexing error rate by at least 40%. For example, the demultiplexing error rate can be reduced by at least or at least about 40%, 50%, 60%, 70%, 80%, 90%, 95%, or a number or a ranged between any two of these values. Error rate as used herein refers to the probability that a barcode having a first nucleotide sequence is assigned to another, different barcode sequence in the set.

[0104] Although this section is described with respect to barcode sequences, the disclosure is applicable to determining patterns. For example, a pattern can be (or can be analogous to) a barcode sequence. A pattern component can be (or can be analogous to) a base. The number of pattern components in a pattern can be (or can be analogous to) the length of a barcode sequence (or the number of bases in a barcode sequence).Methods and Systems for Determining Barcodes

[0105] Disclosed herein include methods and systems for determining barcode sequences, demultiplexing sequence reads, and / or reducing error in identifying nucleic acid sequence or barcode data. In some embodiments, a method for determining a barcode sequence is disclosed. The method can comprise (a) obtaining a sequence read comprising a sequence being sequenced and a first observed barcode sequence that does not have an identical match in a plurality of expected barcode sequences, (b) determining a first matching probability betweenthe first observed barcode sequence and each expected barcode sequence of the plurality of expected barcode sequences, by, for each expected barcode sequence, multiplying N event rates one for each base position in a plurality of N base positions of the first observed barcode sequence, wherein an event rate is the probability of observing a first base at a base position when the actual base (expected base) is a second base, (c) identifying, from the plurality of expected barcode sequences, an expected barcode sequence having a highest first matching probability with the first observed barcode sequence, and (d) assigning the first observed barcode sequence to the expected barcode sequence having the highest first matching probability. A barcoded sequence read used herein can include a genomic sequence of interest (e.g., a target sequence or a sequence being sequenced). In some embodiments, the barcoded sequence read can include one or more genomic sequences of interest.

[0106] In some embodiments, the first base is the same as or different from the second base, i.e., the observed base is the same or different from the actual base (expected base). The observed can be A, C, T, G or N. The actual base can be A, C, T, or G.

[0107] The plurality of expected barcode sequences can be retrieved from a barcode data set. The number of expected barcode sequences in a barcode data set can vary in different embodiments. In some embodiments, the number of expected barcode sequences is about, at least, at least about, at most, or at most about 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, or a number or a range between any two of these values. In some embodiments, the number of expected barcode sequences is from about 3 to about 15, for example, about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the expected barcode sequences in the plurality of expected barcode sequences are unique with respect to one another.

[0108] In some embodiments, the sequence read can comprise two barcodes. For example, the sequence read can be obtained from paired-end sequencing technique which sequences both ends of a fragment. The two barcodes can flank an insert sequence (the sequence being sequenced) on each end. Accordingly, the method can further comprise (e) determining a second matching probability between each expected barcode sequence of the plurality of expected barcode sequences and the second observed barcode sequence by, for each expected barcode sequence, multiplying N event rates one for each base position in a plurality of N base positions of the second observed barcode sequence, (f) identifying, from the plurality of expected barcode sequences, an expected barcode sequence having a highest second matching probability with the second observed barcode sequence, and (g) assigning the second observed barcode sequence to the expected barcode sequence having the highest second matching probability.

[0109] The length of a barcode sequence (e.g., the first / second observed barcode sequence) can be different in different embod ments. In some embodiments, a barcode sequence is, is about, is at least, is at least about, is at most, or is at most about, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or a number or a range between any two of these values, nucleotides in length. The plurality of barcode sequences are the same or different in length. In some embodiments, the barcode sequence (e.g., the first / second observed barcode sequence) is from about 3 to about 12 bases in length. In some embodiment, the observed barcode sequences and the expected barcode sequences are the same in length.

[0110] In some embodiments, an observed barcode sequence is identical to one of the plurality of expected barcode sequences. In some embodiments, an observed barcode sequence can comprise at least one nucleotide mismatch (e.g., one, two, three, four, or more) relative to every sequence of the plurality of expected barcode sequences, i.e., the observed barcode sequence does not have an identical match with any of the expected barcode sequences. A matching probability between the observed barcode sequence and each expected barcode sequence can be determined using a probability matrix. The matching probability is the probability that an observed barcode sequence of a sequence read yields or originates from an expected barcode sequence of the plurality of expected barcode sequences. The probability matrix contains a probability of an observed base yielding or originating from each of a plurality of possible bases for each position of a plurality of base positions in a barcode sequence. An expected barcode sequence having a highest matching probability with the observed barcode sequence is then identified. The observed barcode sequence can be assigned to the identified expected barcode sequence based on the matching probability.[OHl] In some embodiments, the matching probability between an observed barcode sequence and an expected barcode sequence is determined by multiplying N event rates, one for each base position, in a plurality of N base positions of the observed barcode sequence. N is equal to the length of the observed barcode sequence and / or the expected barcode sequence.

[0112] An even rate is the probability of observing a first base at a base position when the actual base (expected base) is a second base. An even rate can also be referred to as a probability of an observed base yielding or originating from each of a plurality of possible bases. An even rate is determined for each position of a plurality of base positions in the barcode sequence.

[0113] In some embodiments, the event rates can be presented in an event probability matrix (e.g., Position-Call-Ref (PCR) matrix or Table of Error matrix). The event probability matrix can comprise N sub-matrices each for a position in N. Each sub-matrix can contain rowsand columns representing different observed bases and actual bases. Each value in a sub-matrix is an event rate value corresponding to a pair of observed base and actual base.

[0114] For example, FIG. 3 depicts a non-limiting exemplary event probability matrix for a barcode sequence having three bases in length. The event probability matrix of FIG.3 can be thought of as 3 5x4 sub-matrices, wherein 3 corresponds to the number of bases in the barcode sequence, 5 is the number of possibilities for an observed base (base call), and 4 is the number of possibilities for an actual base. Each sub-matrix corresponds to one position, and each sub-matrix has rows and columns labeled with nucleotides. The observed base can be A, C, G, T, or N. The actual base (expected based) can be A, C, G, or T. Different rows represent different base calls (observed bases) in the output of a sequencer, while different columns represent different actual bases. The values in the matrix correspond to the event rates. As shown in FIG. 3, the first row corresponds to the scenario where the actual A is observed as A at the first position “0” at an event rate of 0.715, the actual C is observed as A at an event rate of 0.278, the actual G is observed as A at an event rate of 0.007, and the actual T is observed as A at an event rate of 0.001. Similarly, the second row corresponds to the scenario where the actual A, C, G, and T is observed as C at the first position “0” at an event rate of 0.10, 0.976, 0.000, and 0.013, respectively.

[0115] The matching probability between an observed barcode sequence and an expected barcode sequence can then be determined based on the event probability matrix, by multiplying N event rates, one for each base position, in the N base positions of the observed barcode sequence. For example, in an embodiment where an observed barcode is ACT and the expected barcode sequences are AAA, CCT, and AGG, one can calculate the matching probability between ACT and each of AAA, CCT, and AGG using the event probability matrix of FIG. 3. The matching probability between ACT and AAA can be calculated as 0.715*0.008*0.000, i.e., 0. The matching probability between ACT and CCT can be calculated as 0.278*0.984*0.988, i.e., 0.270. The matching probability between ACT and AGG can be calculated as 0.715*0.001*0.005, i.e., 0.0000036. These are the probabilities of each expected barcode (AAA, CCT, or AGG) yielding ACT as an observed barcode. CCT is the expected barcode sequence having the highest matching probability (0.270), therefore the observed ACT can be assigned to the actual barcode sequence CCT.

[0116] In some embodiments, for each base position in N base positions, an event rate is randomly assigned with an initial value such that the initial value for an event when the observed first base is the same as the actual second base is the highest, for example, greater than 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, or 0.9 and less than 1. The remaining values for events when the observed first base is different from the actual second base can be the same or different. Forexample, an event rate for an event when the observed first base A corresponds an actual base A can be 0.76. An event rate for an event when the observed first base A corresponds to an actual base C, G, or T can be 0.08.

[0117] In some embodiments, an event probability matrix (e.g., the matrix containing randomly assigned values) can be updated. In some embodiments, the method further comprises determining an updated probability matrix based on assignments of the observed barcodes sequences of the sequence reads to the expected barcode sequences. In some embodiments, the method further comprises determining an updated probability matrix based on the observed barcodes sequences of the sequence reads assigned to each expected barcode sequence of the plurality of expected barcode sequences.

[0118] For example, FIG. 4, left panel depicts an initial event probability matrix containing randomly assigned values. These initial values can be uniform across all positions and all bases with the observed base always favoring the same real base. After using this initial matrix for the initial assignments, new event rates can be generated (e.g., FIG. 4, right panel), and these new event rates can be used for the next round of assignment. This iteration can be performed for 2, 3, 4, 5 or more time. In some embodiments, the iteration is performed until the event rates remain essentially the same (e.g., differ by at most 15%, 10%, 5%, or 1%) as the event rates from the previous iteration or the difference between the event rates obtained from two consecutive iterations is below a threshold (e.g., less than 15%, 10%, 5%, or 1%).

[0119] In some embodiments, the event probability matrix is pre-determined. In some embodiments, the event rate for each base position in N is pre-determined by assigning an observed barcode sequence of the plurality of sequence reads to one of the plurality of expected barcode sequences, and generating an event probability matrix based on the assignment. In some embodiments, the method comprises assigning each observed barcode sequence of the plurality of sequence reads to one of the plurality of expected barcode sequences. To generate the matrix, an even rate can be calculated by counting the number of times an observed first base corresponds to an expected second base based on the assignment. In some embodiments, the initial assignment of an observed barcode sequence of the plurality of sequence reads to one of the plurality of expected barcode sequence can be performed based on a Hamming distance between the observed barcode sequence and an expected barcode sequence. For example, a Hamming distance between the observed barcode sequence and an expected barcode sequence cannot be greater than a criterion value. In some embodiments, the maximum Hamming distance the observed barcode sequence and an expected barcode sequence is less than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12. In some embodiments, the maximum Hamming distance between the observed barcode sequence and an expected barcode sequence is between 0 and N, inclusive,wherein N is the length of the barcode sequence. In some embodiments, the maximum Hamming distance is equal to or less than 10. Once an event probability matrix is generated based on the initial assignment of the observed barcode sequences to the expected barcode sequences, the observed barcode sequences can be reassigned based on the even probability matrix (see, for example, FIG. 3). Next, the event probability matrix can be updated based on the reassignment. This process of updating the event probability matrix and reassignment of observed barcode sequences to expected barcode sequences can be repeated for as many times as needed (e.g., 2, 3, 4, 5, or more) or until the event rates remain essentially the same (e.g., differ by at most 15%, 10%, 5%, or 1%) as the event rates from the previous iteration or the difference between the event rates from two consecutive iterations is below a threshold (e.g., less than 15%, 10%, 5%, or 1%). The matching probability between an observed barcode sequence and an expected barcode sequence can then be calculated based on the event rates by multiplying N event rates one for each base position in the plurality of N base position of the barcode sequence. The expected barcode sequence having a highest first matching probability with the observed barcode sequence can be identified based on the calculated matching probability.

[0120] In some embodiments, the method further comprises multiplying the highest matching probability of the observed barcode sequence by a relative abundance value of the identified expected barcode sequence. A relative abundance refers to the proportion of a specific barcode (or sequence) compared to the total number of barcodes (or sequences) in a sample (or samples). A relative abundance can be calculated by dividing the number of times a specific barcode sequence appears in a sample (or samples) by the total number of barcode sequences in the sample. A relative abundance can be any value in a range between 0 and 1. In some embodiments, the barcode sequences are equally frequent, i.e., having the same relative abundance values. In some embodiments, the barcode sequences have different relative abundance values. For example, as shown in FIG. 3, if the true barcode CCT is only 50% as common as AAA and AGG, a relative matching probability can be obtained by multiplying the matching probability of 0.270 by 0.5, i.e., 0.185.

[0121] In some embodiments, the observed barcode sequence of the sequence read is assigned to the identified expected barcode sequence if the highest matching probability is above a threshold. In some embodiments, the threshold is at least 10'5 5. In some embodiments, the observed barcode sequence of the sequence read is assigned to the identified expected barcode sequence if the probability that the observed barcode sequence of the sequence read originates from the identified expected barcode is higher than the probabilities that the observed barcode sequence originate from the remaining expected barcode sequences of the plurality of barcodesequences. In some embodiments, the method comprises comparing the highest matching probability of the observed barcode sequence to the sum of other matching probabilities of other expected barcode sequences. In some embodiments, the observed barcode sequence is assigned to the identified expected barcode sequence if the highest matching probability associated with the identified expected barcode sequence is at least 20 million times greater than the sum of other matching probabilities of other expected barcode sequences. In the non-limiting exemplary embodiment of FIG. 3, the relative maximal probability (0.185) is compared to the sum of all other probabilities: 0.185 / (0+0.0000036)=51, 389. The observed barcode sequence ACT is 51 thousand times more likely to originate from CCT than other possibilities (AAA or AGG) combined. Given that 51 thousand is below the cutoff value of 20 million, the observed barcode sequence ACT and its associated sequence read is assigned into an UNKNOWN bin because it is not sufficiently confident to assign ACT to CCT. In some embodiments, the method further comprises assigning the assigned expected barcode sequence (i.e., the expected barcode sequence having the highest matching probability with the observed barcode sequence) to the sequence read as a true barcode sequence of the sequence read.

[0122] In some embodiments, the observed barcode sequence is assigned to the identified expected barcode sequence if the Hamming distance between the observed barcode sequence and the identified expected barcode sequence is below a certain value, optionally equal to or less than 10. In some embodiments, the maximum Hamming distance between the observed barcode sequence and the assigned expected barcode sequence is below a certain value. For example, the maximum Hamming distance between the observed barcode sequence and the assigned expected barcode sequence can be less than 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12. In some embodiments, the maximum Hamming distance between the observed barcode sequence and the assigned expected barcode sequence is equal to or less than 10.

[0123] In some embodiments, the sequence read comprises a first observed barcode sequence and a second observed barcode sequence, and the method comprises assigning the first observed barcode sequence to the first identified expected barcode sequence and assigning the second observed barcode sequence to the second identified expected barcode sequences. The number of barcode pairs in a barcode data set can vary in different embodiments. In some embodiments, a barcode data set can comprise about, at least, at least about, at most or at most about 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or a number or a range between any two of these values, barcode pairs. In some embodiments, the method can further comprise discarding the sequence read comprising the first and second observed barcode sequences if the combination of the first and second observed barcode sequences does not exist in the barcode data set. For example, an intended barcode pool contains two pairs ofbarcodes: {AAA+AAA, CCC+CCC}. For an observed barcode pair of AAA + ACC, the probability is calculated from PCR matrix and observed frequency of each of virtual pool pair. The occurrence frequency of barcode pairs (expected or chimeric) is initially unknown but stabilizes after iterations. As a result, AAA+ACC would likely be assigned to AAA+CCC (a chimeric pair) and discarded, rather than to AAA+AAA or CCC+CCC.

[0124] In some embodiments, the method comprises providing a plurality of sequence reads each comprising a sequence being sequenced (e.g., a target sequence) and at least one observed barcode sequence. The method can further comprise comparing an observed barcode sequence of the plurality sequence reads to an expected barcode sequence of the plurality of expected barcode sequences. For example, the method can comprise comparing each observed barcode sequence of the plurality sequence reads to each of the plurality of expected barcode sequences. If the observed barcode sequence of a sequence read has an identical match in the expected barcode sequences, the observed barcode sequence (i.e., the matching expected barcode sequence) can be assigned to the sequence read as the true barcode sequence.

[0125] In some embodiments, the method further comprises assigning a sequence read to a respective sample in a plurality of samples using the barcode sequence assigned to the sequence read. The observed barcode sequence can comprise a sample-specific barcode sequence. In some embodiments, the observed barcode sequence is a sample-specific barcode sequence.

[0126] In some embodiments, the method further comprises sequencing a plurality of barcoded nucleotide molecules with a sequencing system, thereby obtaining the plurality of sequence reads. The plurality of barcoded nucleotide molecules can be pooled from a plurality of samples, wherein each barcode nucleotide molecule comprises a sample-specific barcode. The samples can be different, e.g., from different biological sources, from different organisms, from different subjects, from different developmental stages of the same or different individuals, from individuals subjected to different treatments for a disease, from individuals subjected to different environmental factors, and the like.

[0127] Disclosed herein also include a system for determining a barcode sequence. In some embodiments, the system can comprise non-transitory memory configured to store executable instructions, and a hardware processor in communication with the non-transitory memory, the hardware processor programmed by the executable instructions to perform: receiving a sequence read comprising a sequence being sequenced and an observed barcode sequence that does not have an identical match in a plurality of expected barcode sequences; determining a matching probability between the observed barcode sequence and each expected barcode sequence of the plurality of expected barcode sequences, by, for each expected barcodesequence, multiplying N event rates one for each base position in a plurality of N base positions of the observed barcode sequence, wherein an event rate is the probability of observing a first base at a base position when the actual base is a second base; identifying, from the plurality of expected barcode sequences, an expected barcode sequence having a highest matching probability with the observed barcode sequence, and assigning the observed barcode sequence to the expected barcode sequence having the highest matching probability.

[0128] Disclosed herein also include a method for determining barcode sequences. In some embodiments, the method can comprise under control of a hardware processor: (a) receiving a plurality of sequence reads each comprising a sequence being sequenced and an observed barcode sequence (e.g., index sequence), wherein an observed barcode sequence yields or originates from an expected barcode sequence of a plurality of expected barcode sequences, and wherein the observed barcode sequence of a sequence read of the plurality sequences is different from each expected barcode sequence of the plurality of expected barcode sequences. The plurality of sequence reads can comprise paired-end sequence reads and / or single-end sequence reads. In some embodiments, a sequence read can comprise one observed barcode sequence. In some embodiments, a sequence read can comprise two observed barcode sequence, one at 5’ end and the other at 3’ end. The sequence being sequence can comprise one or more sequence of interest, one or more target sequence, or one or more genomic sequence. The expected barcode sequence can be in the same length as the observed barcode sequence.

[0129] In some embodiments, the method further comprises (b) for each of one or more sequence reads of the plurality of sequence reads: determining a probability that the observed barcode sequence of the sequence read originates from each expected barcode sequence of the plurality of expected barcode sequences using a probability matrix (Position- Call-Ref matrix or Table of error probabilities), wherein the probability matrix comprises a probability (event probability or even rate) of an observed base yielding from each of a plurality of possible bases (expected bases) for each base position of a plurality of base positions; and assigning the observed barcode sequence of the sequence read to an expected barcode sequence of the plurality of barcode sequences based on the probabilities of the observed barcode sequence of the sequence read yielding from the expected barcode sequences. The observed base can be A, C, G, T, or N, while the expected based can be A, C, G, or T.

[0130] In some embodiments, assigning the observed barcode sequence to the sequence read comprises assigning the observed barcode sequence of the sequence read to the assigned expected barcode sequence if the probability that the observed barcode sequence of the sequence read originates from the assigned expected barcode is higher than the probabilities that the observed barcode sequence originate from the remaining expected barcode sequences of theplurality of barcode sequences or one or more of the remaining expected barcode sequences of the plurality of barcode sequences.

[0131] In some embodiments, the step (b) of determining a probability that the observed sequence originates from each expected barcode sequence and assigning the observed barcode sequence to an expected barcode sequence is performed for a plurality of iterations (e.g., 2, 3, 4, or 5).

[0132] In some embodiments, the method further comprises determining an updated probability matrix. The updated probability matrix can be determined based on assignments of the observed barcodes sequences of the sequence reads to the expected barcode sequences. The updated probability matrix can be determined based on the observed barcodes sequences of the sequence reads assigned to each expected barcode sequence of the plurality of expected barcode sequences. The updated probability matrix can be determined based on (x) the observed barcode sequences of the sequence reads and (y) the expected barcode sequences of the expected barcode sequences to which the observed barcode sequences are assigned. In some embodiments, the updated probability matrix is determined based on (x) one or more bases (or identity of bases) of the observed barcode sequences of the sequence reads at each base position of the plurality of base positions and (y) one or more bases (or identify of bases) of the expected barcode sequences of the expected barcode sequences, to which the observed barcode sequences are assigned, at each base position of the plurality of base positions. In some embodiments, the updated probability matrix is determined based on, for each base position of the plurality of base positions, (x) bases of the observed barcode sequences of the sequence reads at the base position and (y) bases of the expected barcode sequences of the expected barcode sequences, to which the observed barcode sequences are assigned, at the base position. In some embodiments, the updated probability matrix is determined based on (x) a number of each base, of the observed barcode sequences of the sequence reads, at each base position of the plurality of base positions and (y) bases of the expected barcode sequences of the expected barcode sequences, to which the observed barcode sequences are assigned, at each base position of the plurality of base positions. In some embodiments, the updated probability matrix is determined based on, for each base position of the plurality of base positions, (x) a number of each base and the identity of the base, of the observed barcode sequences of the sequence reads, at the base position and (y) identities of bases of the expected barcode sequences of the expected barcode sequences, to which the observed barcode sequences are assigned, at the base position. In some embodiments, the updated probability matrix is determined in one iteration of the plurality of iterations, and wherein the updated probability matrix is used to determine probabilities of the observedbarcode sequences in a subsequent iteration (e.g., an immediate subsequent to the iteration where the updated probability is generated) of the plurality of iterations.

[0133] Disclosed herein also includes a method of probabilistic demultiplexing barcode sequences. In some embodiments, the method can comprise under control of a hardware processor: (a) receiving a plurality of sequence reads each comprising a sequence being sequenced (e.g., a target sequence or a sample sequence) and an observed barcode sequence (index sequence), wherein an observed barcode sequence yields or originates from an expected barcode sequence of a plurality of expected barcode sequences. The expected barcode sequence can be in the same length as the observed barcode sequence. The method can further comprise (b) for each of one or more sequence reads of the plurality of sequence reads: if the observed barcode sequence of the sequence read is identical to an expected barcode sequence: assigning the observed barcode sequence to the identical expected barcode sequence. If the observed barcode sequence of the sequence read is not identical to any expected barcode sequence of the plurality of expected barcode sequences: determining a probability that the observed barcode sequence of the sequence read originates from each expected barcode sequence of the plurality of expected barcode sequences using a probability matrix (Position-Call-Ref Matrix or Table of error probabilities), wherein the probably matrix comprises a probability (even probability / rate) of an observed base (A, C, G, T or N) yields from each of a plurality of possible bases (A, C, G, or T) for each base position of a plurality of base positions; and assigning the observed barcode sequence of the sequence read to an expected barcode sequence of the plurality of barcode sequences based on the probabilities of the observed barcode sequence of the sequence read yielding from the expected barcode sequences.

[0134] Disclosed herein also includes a method of probabilistic demultiplexing barcode sequences. In some embodiments, the method can comprise under control of a hardware processor: (a) receiving a plurality of sequence reads each comprising a sequence being sequenced (e.g., a target sequence, a sample sequence, and / or a genomic sequence) and an observed barcode sequence (index sequence), wherein an observed barcode sequence yields or originates from an expected barcode sequence of a plurality of expected barcode sequences, and wherein the observed barcode sequence of a sequence read of the plurality sequences is different from each expected barcode sequence of the plurality of expected barcode sequences. The expected barcode sequence can be in the same length as the observed barcode sequence. The method further comprises: iteratively, (b) for each of one or more sequence reads of the plurality of sequence reads: determining a probability that the observed barcode sequence of the sequence read originates from each expected barcode sequence of the plurality of expected barcode sequences using a probability matrix (Position-Call-Ref Matrix or Table of error probabilities),wherein the probably matrix comprises a probability of an observed base (A, C, G, or T) yields from each of a plurality of possible bases (A, C, G or T) for each base position of a plurality of base positions; assigning the observed barcode sequence of the sequence read to an expected barcode sequence of the plurality of barcode sequences based on the probabilities of the observed barcode sequence of the sequence read yielding from the expected barcode sequences; and (c) determining an updated probability matrix based on assignments of the observed barcodes sequences of the sequence reads to the expected barcode sequences if the current iteration is not the last iteration, wherein the updated probability matrix is the probability matrix of the subsequent iteration (e.g., immediate subsequent iteration).

[0135] Although this section is described with respect to barcode sequences, the disclosure is applicable to determining patterns. For example, a pattern can be (or can be analogous to) a barcode sequence. A pattern component can be (or can be analogous to) a base. The number of pattern components in a pattern can be (or can be analogous to) the length of a barcode sequence (or the number of bases in a barcode sequence).NovaDemux Software

[0136] In some embodiments, the NovaDemux software, such as NovaDemux version 39.07, is suitable for use in the method of the present disclosure. This program is a sequence demultiplexer intended primarily for, but not limited to, sequencing machines, such as ILLUMINA™, Inc. sequencing machines. Typically, multiple experiments ("libraries") are pooled together and sequenced at once, with genetic molecules of these libraries tagged with a synthetic DNA "barcode". After sequencing, the data is demultiplexed into one file per library based on the barcode. However, errors in barcode reading cause misassignment and decrease yield. The NovaDemux software uses advanced statistical methods to maximize yield while minimizing misassignment compared to existing software.

[0137] The NovaDemux software is a module which extends the open-source software BBTools (released through IPO, and also authored by me, current public version 39.06). It makes use of some of the BBTools code and is designed such that BBTools can be downloaded and used freely as usual, while all code for probabilistic demultiplexing is contained within two files that can be present or absent (and do not share any code with earlier software).

[0138] BBTools can be found in the webpage at: “jgi.doe.gov / data-and- tools / software-tools / bbtools / ”. BBTools is described in Brian Bushnell, Jonathan Rood, Esther Singer, “BBMerge - Accurate paired shotgun read merging via overlap,” PLoS ONE 12(10):eO185O56; “doi.org / 10.1371 / joumal.pone.0185056” (published: October 26, 2017), the content of which is incorporated herein by reference in its entirety.

[0139] BBTools is a suite of fast, multithreaded bioinformatics tools designed for analysis of DNA and RNA sequence data. BBTools can handle common sequencing file formats such as fastq, fasta, sam, scarf, fasta+qual, compressed or raw, with autodetection of quality encoding and interleaving. It is written in Java and works on any platform supporting Java, including Linux, MacOS, and Microsoft Windows and Linux; there are no dependencies other than Java (version 7 or higher). Program descriptions and options are shown when running the shell scripts with no parameters.Sequencing

[0140] In some embodiments, the methods and systems described herein can further comprise, prior to demulipluxing the sequence reads, generating a plurality of sequence reads from sequencing barcoded nucleic acids. In some embodiments, the methods described herein comprise utilizing next generation sequencing technologies (NGS) that allow multiple, pooled samples to be sequenced on a single sequencing run. The sequencing technologies of NGS include, but are not limited to, pyrosequencing, sequencing-by-synthesis with reversible dye terminators, sequencing by oligonucleotide probe ligation, ion torrent sequencing, and others identifiable to a person skilled in the art. DNA multiple samples can be pooled and sequenced as indexed genomic molecules (i.e., multiplex sequencing) on a single sequencing run, to generate a plurality of barcoded sequence reads. The plurality of sequence reads can comprise paired-end sequence reads and / or single-end sequence reads.

[0141] In some embodiments, the methods and systems disclosed herein can be used in sequencing-by-synthesis (SBS) methods. In SBS, extension of a nucleic acid primer along a nucleic acid template is monitored to determine the sequence of nucleotides in the template. The underlying chemical process can be catalyzed by a polymerase, wherein fluorescently labeled nucleotides are added to a primer (thereby extending the primer) in a template dependent fashion such that detection of the order and type of nucleotides added to the primer can be used to determine the sequence of the template. Briefly, SBS can be initiated by contacting target nucleic acids, attached to sites in a flow cell, with one or more labeled nucleotides, DNA polymerase, etc. Those sites where a primer is extended using the target nucleic acid as template will incorporate a labeled nucleotide that can be detected. Detection can include scanning using an apparatus or method set forth herein. Optionally, the labeled nucleotides can further include a reversible termination property that terminates further primer extension once a nucleotide has been added to a primer. For example, a nucleotide analog having a reversible terminator moietycan be added to a primer such that subsequent extension cannot occur until a deblocking agent is delivered to remove the moiety. Thus, for embodiments that use reversible termination, a deblocking reagent can be delivered to the vessel (before or after detection occurs). Washes can be carried out between the various delivery steps. The cycle can be performed n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary SBS procedures, reagents and detection components that can be readily adapted for use with a method, system or apparatus of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; WO 91 / 06678; WO 07 / 123744; U.S. Pat. Nos. 7,057,026; 7,329,492; 7,211,414; 7,315,019 or 7,405,281, and US Pat. App. Pub. No. 2008 / 0108082 Al, each of which is incorporated herein by reference.

[0142] In some embodiments, the methods disclosed herein can further comprise a sequencing step. The sequencing steps can include annealing and extending a sequencing primer to incorporate a detectable label that indicates the identity of a nucleotide in a nucleic acid (e.g., a barcoded nucleic acid described herein), detecting the detectable label, and optionally repeating the extending and detecting steps.

[0143] In some embodiments, the methods and systems described herein can further comprise preparing a sequencing library prior to sequencing. The preparation typically involves fragmenting the DNA (sonication, nebulization or shearing), followed by DNA repair and end polishing (blunt end or A overhang), and platform-specific adapter ligation. Methods for preparing library preparation include, for example, ligation-base library preparation, tagmentation-based library preparation, amplicon library preparation, and others identifiable to a person skilled in the art. The specific protocol is selected depending on the sequencing platform and downstream analysis. In some embodiments, preparing a sequencing library can comprise fragmentation of the nucleic acids, end repair of fragmented nucleic acids, A-tailing of fragmented nucleic acids (to add a few nucleotides with adenosine (A) bases) that have been end-repaired, addition of adaptors, and optionally PCR amplification. Adapters can be covalently attached to the ends of the DNA fragments. An adapter may include a platform binding sequence, such as the P5 and P7 sequences, which are platform-specific and used to bind the flow cell, a sequencing primer binding sequence allowing the binding of sequencing primers, and one or two barcode / indexes (see, for example, FIG. 1).

[0144] In some embodiments, the method further comprises pooling the plurality of barcoded nucleic acids and amplifying the pooled barcoded nucleic acids before sequencing the pooled barcoded nucleic acids.Biological Samples

[0145] The barcoded nucleotide molecules described herein are pooled from a plurality of samples, wherein each barcode nucleotide molecule comprises a sample-specific barcode. Samples used herein can include samples taken from any cell, fluid, tissue, organ, any source including nucleic acids in which sequences of interest are to be determined. In some embodiments, the nucleic acids to be sequenced are purified or isolated by any of a number of well-known methods.

[0146] A “sample” or a “biological sample” used herein encompasses essentially any sample type obtained from a source containing nucleic acids in which sequences of interest are io be determined. A sample (e.g., a sample comprising nucleic acid) can be obtained from a suitable subject. A sample can be isolated or obtained directly from a subject or part thereof. In some embodiments, a sample is obtained indirectly from an individual or medical professional. In some embodiments, the sample used herein includes or consists essentially of purified or isolated polynucleotides. The biological sample can be any bodily fluid, tissue or any other suitable sample. The definition encompasses blood and other liquid samples of biological origin, solid tissue samples such as a biopsy specimen or tissue cultures or cells derived therefrom and the progeny thereof. The definition also includes samples that have been manipulated in any way after their procurement, such as by treatment with reagents, solubilization, or enrichment for certain components, such as cells, polypeptides, or proteins. The term “biological sample” encompasses a clinical sample, but also, in some instances, includes cells in culture, cell supernatants, cell lysates, blood, plasma, serum, sweat, tears, sputum, urine, sputum, ear flow, lymph, saliva, cerebrospinal fluid, ravages, bone marrow suspension, vaginal flow, transcervical lavage, brain fluid, ascites, milk, secretions of the respiratory, intestinal and genitourinary tracts, amniotic fluid, milk, and leukophoresis samples. In some embodiments, the sample is a sample that is easily obtainable by non-invasive procedures, e.g., blood, plasma, serum, sweat, tears, sputum, urine, stool, sputum, ear flow, saliva or feces. In certain embodiments the sample is a peripheral blood sample, or the plasma and / or serum fractions of a peripheral blood sample. In other embodiments, the biological sample is a swab or smear, a biopsy specimen, or a cell culture. In another embodiment, the sample is a mixture of two or more biological samples, e.g., a biological sample can include two or more of a biological fluid sample, a tissue sample, and a cell culture sample. As used herein, the terms “blood,” “plasma” and “serum” expressly encompass fractions or processed portions thereof. Similarly, where a sample is taken from a biopsy, swab, smear, etc., the “sample” expressly encompasses a processed fraction or portion derived from the biopsy, swab, smear, etc.

[0147] A sample can be any specimen that is isolated or obtained from a subject or part thereof. A sample can be any specimen that is isolated or obtained from multiple subjects. Non-limiting examples of specimens include fluid or tissue from a subject, including, without limitation, blood or a blood product (e.g., serum, plasma, platelets, huffy coats, or the like), umbilical cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, spinal fluid, lavage fluid (e.g., lung, gastric, peritoneal, ductal, ear, arthroscopic), a biopsy sample, celocentesis sample, cells (blood cells, lymphocytes, placental cells, stem cells, bone marrow derived cells, embryo or fetal cells) or parts thereof (e.g., mitochondrial, nucleus, extracts, or the like), urine, feces, sputum, saliva, nasal mucous, prostate fluid, lavage, semen, lymphatic fluid, bile, tears, sweat, breast milk, breast fluid, the like or combinations thereof. A fluid or tissue sample from which nucleic acid is extracted may be acellular (e.g., cell-free). Non-limiting examples of tissues include organ tissues (e.g., liver, kidney, lung, thymus, adrenals, skin, bladder, reproductive organs, intestine, colon, spleen, brain, the like or parts thereof), epithelial tissue, hair, hair follicles, ducts, canals, bone, eye, nose, mouth, throat, ear, nails, the like, parts thereof or combinations thereof. A sample may comprise cells or tissues that are normal, healthy, diseased (e.g., infected), and / or cancerous (e.g., cancer cells). A sample obtained from a subject may comprise cells or cellular material (e.g., nucleic acids) of multiple organisms (e.g., virus nucleic acid, fetal nucleic acid, bacterial nucleic acid, parasite nucleic acid).

[0148] In some embodiments, the plurality of samples are different samples. For example, the samples can be obtained from different individuals, samples from different developmental stages of the same or different individuals, samples from different diseased individuals (e.g., individuals suspected of having a genetic disorder), normal individuals, samples obtained at different stages of a disease in an individual, samples obtained from an individual subjected to different treatments for a disease, samples from individuals subjected to different environmental factors, samples from individuals with predisposition to a pathology, samples individuals with exposure to an infectious disease agent, and the like. In some embodiments, the plurality of samples comprise samples from different biological sources and / or different organisms.

[0149] In some embodiments, samples can also be obtained from in vitro cultured tissues, cells, or other polynucleotide-containing sources. The cultured samples can be taken from sources including, but not limited to, cultures (e.g., tissue or cells) maintained in different media and conditions (e.g., pH, pressure, or temperature), cultures (e.g., tissue or cells) maintained for different periods of length, cultures (e.g., tissue or cells) treated with different factors or reagents (e.g., a drug candidate, or a modulator), or cultures of different types of tissue and / or cells.Determining Patterns

[0150] The disclosure herein with respect to sample demultiplexing and determining barcode sequences can be applied to determining patterns. For example, a pattern can be (or can be analogous to) a barcode sequence. A pattern component can be (or can be analogous to) a base. The number of pattern components in a pattern can be (or can be analogous to) the length of a barcode sequence (or the number of bases in a barcode sequence).

[0151] Disclosed herein include embodiments of a method for determining a pattern. In some embodiments, a method for determining a pattern (or one or more actions thereof) can be under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: (a) obtaining a first observed pattern that does not have an identical match in a plurality of expected patterns. The method can comprise: (b) determining a first matching probability between the first observed pattern and each expected pattern of the plurality of expected patterns, based on N event rates one for each pattern component position in a plurality of N pattern component positions of the first observed pattern. An event rate can be the probability of observing a first pattern component at a pattern component position when the actual pattern component is a second pattern component. The method can comprise: (c) identifying, from the plurality of expected patterns, a first expected pattern having a highest first matching probability with the first observed pattern. The method can comprise: (d) assigning the first observed pattern to the first expected pattern having the highest first matching probability.

[0152] An event rate can be, for example, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.2, 0.21, 0.22, 0.23, 0.24,0.25, 0.26, 0.27, 0.28, 0.29, 0.3, 0.31, 0.32, 0.33, 0.34, 0.35, 0.36, 0.37, 0.38, 0.39, 0.4, 0.41,0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49, 0.5, 0.51, 0.52, 0.53, 0.54, 0.55, 0.56, 0.57, 0.58, 0.59, 0.6, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.7, 0.71, 0.72, 0.73, 0.74, 0.75,0.76, 0.77, 0.78, 0.79, 0.8, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92,0.93, 0.94, 0.95, 0.96, 0.97, 0.98, or 0.99. In some embodiments, the method further comprises: for each pattern component position in N, randomly assigning each event rate with an initial value such that the initial value for an event rate of the observed first pattern component being identical to the actual second pattern component is the highest. The initial value for such event rate can be greater than 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, or 0.9. The remaining initial values for event rates of the observed first pattern component being different from the actual second pattern component are the same or different. A remaining initial value can be, for example, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, or 0.1.

[0153] In some embodiments, the event rate for each pattern component position inN, e.g., for an iteration, is pre-determined (e.g., in a prior iteration (such as the immediate prior iteration). For example, the event rate for each pattern component position in N for a iteration is determined in the prior iteration by: (i) assigning each of a plurality of observed patterns to one of the plurality of expected patterns. The event rate for each pattern component position in N can be determined by: (ii) for each pattern component position in N, calculating the probability of observing the first pattern component that is expected to be the second pattern component by determining the number of times the observed first pattern component corresponds to the expected second pattern component based on the assignment from step (i). In some embodiments, the method further comprises: iterating steps (i) and (ii) for a number of iterations (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10), or until the event rates remain essentially the same (e.g., withinO.05, 0.04, 0.03, 0.02, 0.01, 0.009, 0.008, 0.007, 0.006, 0.005, 0.004, 0.003, 0.002, 0.001, or smaller) as the event rates from the previous iteration or the difference between the event rates from two consecutive iterations is below a threshold (e.g., 0.05, 0.04, 0.03, 0.02, 0.01, 0.009, 0.008, 0.007, 0.006, 0.005, 0.004, 0.003, 0.002, 0.001, or smaller) The number of iterations (or the maximum number of iterations) can be pre-determined.

[0154] In some embodiments, step (i) for an iteration (e.g., the first iteration) is performed based on a Hamming distance between an observed pattern and an expected pattern. The maximum Hamming distance can be between 0 and N, inclusive. The maximum Hamming distance can be equal to or less than 2, 3, 4, 5, 6, 7, 8, 9, or 10. The number of iterations (or the maximum number of iterations) can be pre-determined. In some embodiments, step (ii) comprises generating at least one event probability matrix (also referred to herein as a Position- Call-Ref (PCR) matrix or a table of error probabilities). An event probability matrix can comprise N sub-matrices each for a position in N. Each sub-matrix can have rows and columns representing different observed pattern components and expected pattern components and event rates each corresponding to a pair of observed pattern component and expected pattern component.

[0155] In some embodiments, the method further comprises: for each pattern component position in N, updating the event rate based on the assignment, and repeating steps (b)-(d). Updating the event rate can comprising calculating the probability of observing the first pattern component that is expected to be the second pattern component. Updating the event rate can comprise: determining the number of times the observed first pattern component corresponds to the expected second pattern component based on the assignment.

[0156] In some embodiments, determining the first matching probability between the first observed pattern and each expected pattern of the plurality of expected patterns is based ona relative abundance value of the identified first expected pattern and / or a relative abundance value of the identified second expected pattern. Determining the first matching probability between the first observed pattern and each expected pattern of the plurality of expected patterns can comprise: for each expected pattern, multiplying N event rates one for each pattern component position in a plurality of N pattern component positions of the first observed pattern.

[0157] In some embodiments, the method further comprising multiplying the highest first matching probability by a relative abundance value of the identified first expected pattern. Alternatively or additionally, the method further comprises: multiplying the highest second matching probability by a relative abundance value of the identified second expected pattern. In some embodiments, the method further comprises comparing the highest first matching probability of the identified first expected pattern to the sum of other matching probabilities of other expected patterns.

[0158] In some embodiments, the first observed pattern can be assigned to the identified first expected pattern if the highest first matching probability is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 million times greater than the sum of other matching probabilities of other expected patterns (also referred to herein as minratio). In some embodiments, the first observed pattern is assigned to the identified first expected pattern if the highest first matching probability is above a first threshold. In some embodiments, the method further comprises comparing the highest second matching probability to the sum of other matching probabilities of other expected patterns. In some embodiments, the second observed pattern is assigned to the identified second expected pattern if the highest second matching probability is above a second threshold. In some embodiments, the first threshold and / or the second threshold (also referred to herein as minprob) is at least 10’1, 10'2, 10'3, 10'4, 10'5, 10'5 5, 10'6, 10'7, 10'8, 10'9, 10'10, or a number or a range between any two of these values. In some embodiments, the first observed pattern is assigned to the identified first expected pattern if the Hamming distance between the first observed pattern and the identified first expected pattern (also referred to herein as maxhdist) is below a certain value, optionally equal to or less than, for example, 3, 4, 5, 6, 7, 8, 9, or 10. In some embodiments, the maximum Hamming distance between the observed pattern and the identified expected pattern is below a certain value. Such maximum Hamming distance can be equal to or less than, for example, 3, 4, 5, 6, 7, 8, 9, or 10.

[0159] In some embodiments, the method further comprises assigning the first observed pattern, or a sample associated with the first observed pattern, to a respective sample in a plurality of samples using the assigned expected pattern. The first observed pattern can be a sample-specific pattern.

[0160] In some embodiments, the method further comprises: providing a plurality of observed patterns. The method can comprise: comparing each observed pattern of the plurality observed patterns to each of the plurality of expected patterns. In some embodiments, the method further comprises: for an observed pattern in the plurality of observed patterns having an identical match in the plurality of expected patterns, assigning the observed pattern as a true pattern. In some embodiments, the method further comprises assigning the observed pattern, or a sample associated with the observed pattern, to a respective sample in a plurality of samples using the observed pattern. The plurality of samples can be different.

[0161] In some embodiments, the number of expected patterns can be from about 3 to about 50. The number of expected patterns can be, be about, be at least, be at least about, be at most, or be at most about, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or a range between any two of these values. A pattern can comprise a number of pattern components (corresponding to a number of pattern component positions), such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or a range between any two of these values, pattern components. A pattern component can be selected from one of a number of possibilities. The number of possibilities a pattern component can be selected from can be, be about, be at least, be at least about, be at most, or be at most about, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or a range between any two of these values.

[0162] Disclosed herein include embodiments of a method of determining patterns. In some embodiments, a method of determining patterns (or one or more actions thereof) can be under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: (a) receiving a plurality of observed patterns. An observed pattern can yield (or originate) from an expected pattern of a plurality of expected patterns. At least one observed pattern of the plurality observed patterns can be different from each expected pattern (or all expected patterns) of the plurality of expected patterns. The method can comprise: (b) for each of one or more observed patterns of the plurality of observed patterns: determining a probability that the observed pattern yields (or originates) from each expected pattern of the plurality of expected patterns using a probability matrix (also referred to herein as a Position-Call-Ref (PCR) matrix, event probabilities matrix, or table of error probabilities). The probability matrix can comprise a probability of an observed pattern component yields (or originates) from each ofa plurality of possible (or expected) pattern components for each pattern component position of a plurality of pattern component positions. The method can comprise ((b) for each of one or more observed patterns of the plurality of observed patterns): assigning the observed pattern to an expected pattern of the plurality of expected patterns based on the probabilities of the observed pattern yielding (or originating) from the expected patterns.

[0163] Disclosed herein include embodiments of a method of determining patterns. In some embodiments, a method of determining patterns (or one or more actions thereof) can be under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: (a) receiving a plurality of observed patterns. An observed pattern can yield (or originate) from an expected pattern of a plurality of expected patterns (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 pattern component positions). At least one observed pattern of the plurality observed patterns can be different from each expected pattern (or all expected patterns) of the plurality of expected patterns. The method can comprise: (b) for each of one or more observed patterns of the plurality of observed patterns: if the observed pattern is identical to (or the same as) an expected pattern: assigning the observed pattern to the identical expected pattern. The method can comprise ((b) for each of one or more observed patterns of the plurality of observed patterns): if the observed pattern is not identical (or is different from) to any expected pattern of the plurality of expected patterns: determining a probability that the observed pattern yields (or originates) from each expected pattern of the plurality of expected patterns using a probability matrix, (also referred to herein as a Position-Call-Ref (PCR) matrix, event probabilities matrix, or table of error probabilities). The probably matrix can comprise a probability of an observed pattern component yields (or originates) from each of a plurality of possible pattern components for each pattern component position of a plurality of pattern component positions. The method can comprise: assigning the observed pattern to an expected pattern of the plurality of expected patterns based on the probabilities of the observed pattern yielding (or originating) from the expected patterns.

[0164] Disclosed herein include embodiments of a method of determining patterns. In some embodiments, a method of determining patterns (or one or more actions thereof) can be under control of a processor (e.g., a hardware processor or a virtual processor). The method can comprise: (a) receiving a plurality of observed patterns. An observed pattern can yield (or originates) from an expected pattern of a plurality of expected patterns. At least one observed pattern of the plurality of observed patterns can be different from each expected pattern of the plurality of expected patterns. The method can comprise: iteratively, (b) for each of one or moreobserved patterns of the plurality of observed patterns: determining a probability that the observed pattern yields (or originates) from each expected pattern of the plurality of expected patterns using a probability matrix (also referred to herein as a Position-Call-Ref (PCR) matrix, event probabilities matrix, or table of error probabilities). The probably matrix can comprise a probability of an observed pattern component yields (or originates) from each of a plurality of possible pattern components for each pattern component position of a plurality of pattern component positions. The method can comprise ((b) for each of one or more observed patterns of the plurality of observed patterns): assigning the observed pattern of the sequence read to an expected pattern of the plurality of expected patterns based on the probabilities of the observed pattern yielding (or originating) from the expected patterns. The method can comprise (after step (c) and part of iteratively): (c) determining an updated probability matrix based on assignments of the observed patterns of the sequence reads to the expected patterns if the current iteration is not the last iteration, wherein the updated probability matrix is the probability matrix of the subsequent iteration (e.g., immediate subsequent iteration).

[0165] A pattern can comprises a plurality of pattern components. Each of the plurality of pattern components is at one (or unique or a different) pattern component position of the plurality of pattern component positions (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 pattern component positions). A number of observed patterns (e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000 observed patterns; or 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%,11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%,27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%,43%, 44%, 45%, 46%, 47%, 48%, 49%, or 50%, of the plurality of observed patterns) of the plurality observed patterns (e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000 observed patterns) can be different from each expected pattern (or all expected patterns) of the plurality of expected patterns. An (or each) observed pattern and an (or each) expected pattern can have the same number of pattern components. Two (or all) observed patterns can have the same number of pattern components. Two (or all) expected patterns can have the same number of pattern components.

[0166] In some embodiments, assigning the observed pattern comprises: assigning the observed pattern to the assigned expected pattern if the probability that the observed pattern originates from the assigned expected barcode is higher than the probabilities that the observedpattern originate from the remaining expected patterns of the plurality of expected patterns (or the remaining expected pattern(s) of the plurality of expected patterns, or the remaining one or more expected patterns of the plurality of expected patterns, or one or more of the remaining expected patterns of the plurality of expected patterns).

[0167] In some embodiments, step (b) is performed for a plurality of iterations. The number of iterations can be, for example, 2, 3, 4, 5, 6, 7, 8, 9, or 10).

[0168] In some embodiments, the method further comprising: determining an updated probability matrix. In some embodiments, the method further comprises: determining an updated probability matrix based on assignments of the observed patterns to the expected patterns. In some embodiments, the method further comprises: determining an updated probability matrix based on the observed patterns assigned to each expected pattern of the plurality of expected patterns. In some embodiments, the method further comprises: determining an updated probability matrix based on (x) the observed patterns and (y) the expected patterns of the plurality of expected patterns to which the observed patterns are assigned.

[0169] In some embodiments, the method further comprises: determining an updated probability matrix based on (x) pattern components (or identifies of pattern components) of the observed patterns at each pattern component position of the plurality of pattern component positions and (y) pattern components (or identities of pattern components) of the expected patterns, to which the observed patterns are assigned, at each pattern component position of the plurality of pattern component positions. The method can comprise: determining an updated probability matrix based on (x) pattern components (or identities of pattern components) of the observed patterns at each pattern component position of the plurality of pattern component positions and (y) one or more pattern components (or identities of pattern components) of the expected patterns of the expected patterns, to which the observed patterns are assigned, at each pattern component position of the plurality of pattern component positions. The method can comprise: determining an updated probability matrix based on, for each pattern component position of the plurality of pattern component positions, (x) pattern components of the observed patterns at the pattern component position and (y) pattern components of the expected patterns, to which the observed patterns are assigned, at the pattern component position.

[0170] In some embodiments, the method further comprises: determining an updated probability matrix based on (x) a number of each pattern component, of the observed patterns, at each pattern component position of the plurality of pattern component positions and (y) pattern components of the expected patterns of the expected patterns, to which the observed patterns are assigned, at each pattern component position of the plurality of pattern component positions. In some embodiments, the method further comprises: determining an updated probability matrixbased on, for each pattern component position of the plurality of pattern component positions, (x) a number of each pattern component and the identity of the pattern component, of the observed patterns, at the pattern component position and (y) identities of pattern components of the expected patterns of the expected patterns, to which the observed patterns are assigned, at the pattern component position.

[0171] In some embodiments, the updated probability matrix is determined in one iteration of the plurality of iterations, and wherein the updated probability matrix is used to determine probabilities of the observed patterns in a subsequent iteration (e.g., immediate subsequent iteration) of the plurality of iterations.

[0172] In some embodiments, obtaining the first observed pattern comprises receiving or retrieving the first observed pattern. Obtaining the first observed pattern can comprise capturing the first observed pattern using an imaging sensor. The imaging sensor can be a CCD sensor or a CMOS sensor. The imaging sensor can be comprised in a camera or a phone.

[0173] In some embodiments, a pattern (or each pattern), the pattern being an observed pattern or an expected pattern, comprises a barcode sequence. A barcode sequences can be associated with a sequence being sequenced (or a sequence, a sample sequence, a sequence of interest, or a target sequence). A barcode sequence can be associated with (e.g., comprised in) a sequence read. A sequence read can comprise a barcode sequence and a sequence being sequenced (or a sequence, a sample sequence, a sequence of interest, or a target sequence). The sequence read can comprise one or both sequence reads of paired-end sequence reads. The sequence read can comprise a single-end sequence read. The first pattern component can be A, C, T, or G (or A, C, T, G or N). The second pattern component can be A, C, T, or G.

[0174] In some embodiments, a pattern (or each pattern), the pattern being an observed pattern or an expected pattern, comprises a barcode, a ID barcode, a Universal Product Code (UPC), a 2D barcode, a matrix barcode, a two-dimensional matrix barcode, a Quick Response code (QR code), a static QR code, a dynamic QR code, a High Capacity Colored two- Dimensional code (HCC2D code), a digital code embedded in a presentation or image, a digital code represented by a presentation or image, a predefined custom pattern, or an encoding using an image (or an encoding in an image). The encoding can be decoded after the image is captured.Execution Environment

[0175] FIG. 11 depicts a general architecture of an example computing device 1100 that can be used in some embodiments to execute the processes and implement the featuresdescribed herein. The general architecture of the computing device 1100 depicted in FIG. 11 includes an arrangement of computer hardware and software components. The computing device 1100 may include many more (or fewer) elements than those shown in FIG. 11. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the computing device 1100 includes a processing unit 1110, a network interface 1120, a computer readable medium drive 1130, an input / output device interface 1140, a display 1150, and an input device 1160, all of which may communicate with one another by way of a communication bus. The network interface 1120 may provide connectivity to one or more networks or computing systems. The processing unit 1110 may thus receive information and instructions from other computing systems or services via a network. The processing unit 1110 may also communicate to and from memory 1170 and further provide output information for an optional display 1150 via the input / output device interface 1140. The input / output device interface 1140 may also accept input from the optional input device 1160, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.

[0176] The memory 1170 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 1110 executes in order to implement one or more embodiments. The memory 1170 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer-readable media. The memory 1170 may store an operating system 1172 that provides computer program instructions for use by the processing unit 1110 in the general administration and operation of the computing device 1100. The memory 1170 may further include computer program instructions and other information for implementing aspects of the present disclosure.

[0177] For example, in one embodiment, the memory 1170 includes a module 1174 for determining patterns, such as determining barcode sequences and / or probabilistic demultiplexing barcode sequences. In addition, memory 1170 may include or communicate with the data store 1190 and / or one or more other data stores that store a plurality of patterns (e.g., a plurality of sequence reads), observed patterns (e.g., observed barcode sequences), expected patterns (e.g., expected barcode sequences), probability matrix, updated probability matrix (during one or more iterations), assignments between observed patterns and expected patterns (e.g., assignments between observed barcode sequences and expected barcode sequences), and the like.Additional Considerations

[0178] In at least some of the previously described embodiments, one or moreelements used in an embodiment can interchangeably be used in another embodiment unless such a replacement is not technically feasible. It will be appreciated by those skilled in the art that various other omissions, additions and modifications may be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and changes are intended to fall within the scope of the subject matter, as defined by the appended claims.

[0179] One skilled in the art will appreciate that, for this and other processes and methods disclosed herein, the functions performed in the processes and methods can be implemented in differing order. Furthermore, the outlined steps and operations are only provided as examples, and some of the steps and operations can be optional, combined into fewer steps and operations, or expanded into additional steps and operations without detracting from the essence of the disclosed embodiments.

[0180] With respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Accordingly, phrases such as “a device configured to” are intended to include one or more recited devices. Such one or more recited devices can also be collectively configured to carry out the stated recitations. For example, “a processor configured to carry out recitations A, B and C can include a first processor configured to carry out recitation A and working in conjunction with a second processor configured to carry out recitations B and C. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.

[0181] It will be understood by those within the art that, in general, terms used herein, and especially in the appended claims (e.g., bodies of the appended claims) are generally intended as “open” terms (e.g., the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “includes but is not limited to,” etc.). It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles “a” or “an” limits any particular claimcontaining such introduced claim recitation to embodiments containing only one such recitation, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more”); the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., the bare recitation of “two recitations,” without other modifiers, means at least two recitations, or two or more recitations). Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “ a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In those instances where a convention analogous to “at least one of A, B, or C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “ a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B.”

[0182] In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.

[0183] As will be understood by one skilled in the art, for any and all purposes, such as in terms of providing a written description, all ranges disclosed herein also encompass any and all possible sub-ranges and combinations of sub-ranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like include the number recited and refer to ranges which can be subsequently broken down into sub-ranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a grouphaving 1-3 articles refers to groups having 1, 2, or 3 articles. Similarly, a group having 1-5 articles refers to groups having 1, 2, 3, 4, or 5 articles, and so forth.

[0184] It will be appreciated that various embodiments of the present disclosure have been described herein for purposes of illustration, and that various modifications may be made without departing from the scope and spirit of the present disclosure. Accordingly, the various embodiments disclosed herein are not intended to be limiting, with the true scope and spirit being indicated by the following claims.

[0185] It is to be understood that not necessarily all objects or advantages may be achieved in accordance with any particular embodiment described herein. Thus, for example, those skilled in the art will recognize that certain embodiments may be configured to operate in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other objects or advantages as may be taught or suggested herein.

[0186] All of the processes described herein may be embodied in, and fully automated via, software code modules executed by a computing system that includes one or more computers or processors. The code modules may be stored in any type of non-transitory computer-readable medium or other computer storage device. Some or all the methods may be embodied in specialized computer hardware.

[0187] Many other variations than those described herein will be apparent from this disclosure. For example, depending on the embodiment, certain acts, events, or functions of any of the algorithms described herein can be performed in a different sequence, can be added, merged, or left out altogether (for example, not all described acts or events are necessary for the practice of the algorithms). Moreover, in certain embodiments, acts or events can be performed concurrently, for example through multi -threaded processing, interrupt processing, or multiple processors or processor cores or on other parallel architectures, rather than sequentially. In addition, different tasks or processes can be performed by different machines and / or computing systems that can function together.

[0188] The various illustrative logical blocks and modules described in connection with the embodiments disclosed herein can be implemented or performed by a machine, such as a processing unit or processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processor can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor can include electrical circuitry configured to process computer-executable instructions. In another embodiment, a processorincludes an FPGA or other programmable device that performs logic operations without processing computer-executable instructions. A processor can also be implemented as a combination of computing devices, for example a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Although described herein primarily with respect to digital technology, a processor may also include primarily analog components. For example, some or all of the signal processing algorithms described herein may be implemented in analog circuitry or mixed analog and digital circuitry. A computing environment can include any type of computer system, including, but not limited to, a computer system based on a microprocessor, a mainframe computer, a digital signal processor, a portable computing device, a device controller, or a computational engine within an appliance, to name a few.

[0189] Any process descriptions, elements or blocks in the flow diagrams described herein and / or depicted in the attached figures should be understood as potentially representing modules, segments, or portions of code which include one or more executable instructions for implementing specific logical functions or elements in the process. Alternate implementations are included within the scope of the embodiments described herein in which elements or functions may be deleted, executed out of order from that shown, or discussed, including substantially concurrently or in reverse order, depending on the functionality involved as would be understood by those skilled in the art.

[0190] It should be emphasized that many variations and modifications may be made to the above-described embodiments, the elements of which are to be understood as being among other acceptable examples. All such modifications and variations are intended to be included herein within the scope of this disclosure and protected by the following claims.

[0191] In at least some of the previously described embodiments, one or more elements used in an embodiment can interchangeably be used in another embodiment unless such a replacement is not technically feasible. It will be appreciated by those skilled in the art that various other omissions, additions and modifications may be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and changes are intended to fall within the scope of the subject matter, as defined by the appended claims.

[0192] With respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include pluralreferences unless the context clearly dictates otherwise. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.

[0193] It will be understood by those within the art that, in general, terms used herein, and especially in the appended claims (e.g., bodies of the appended claims) are generally intended as “open” terms (e.g., the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “includes but is not limited to,” etc.). It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles “a” or “an” limits any particular claim containing such introduced claim recitation to embodiments containing only one such recitation, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more”); the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., the bare recitation of “two recitations,” without other modifiers, means at least two recitations, or two or more recitations). Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “ a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In those instances where a convention analogous to “at least one of A, B, or C, etc.” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “ a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms.

[0194] In addition, where features or aspects of the disclosure are described in termsof Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.

[0195] As will be understood by one skilled in the art, for any and all purposes, such as in terms of providing a written description, all ranges disclosed herein also encompass any and all possible sub-ranges and combinations of sub-ranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like include the number recited and refer to ranges which can be subsequently broken down into sub-ranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 articles refers to groups having 1, 2, or 3 articles. Similarly, a group having 1-5 articles refers to groups having 1, 2, 3, 4, or 5 articles, and so forth.

[0196] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.

Claims

WHAT IS CLAIMED IS:

1. A method for determining a barcode sequence, comprising:(a) obtaining a sequence read comprising a sequence being sequenced and a first observed barcode sequence that does not have an identical match in a plurality of expected barcode sequences;(b) determining a first matching probability between the first observed barcode sequence and each expected barcode sequence of the plurality of expected barcode sequences, by, for each expected barcode sequence, multiplying N event rates one for each base position in a plurality of N base positions of the first observed barcode sequence, wherein an event rate is the probability of observing a first base at a base position when the actual base is a second base;(c) identifying, from the plurality of expected barcode sequences, a first expected barcode sequence having a highest first matching probability with the first observed barcode sequence; and(d) assigning the first observed barcode sequence to the first expected barcode sequence having the highest first matching probability.

2. The method of claim 1, wherein the sequence read further comprises a second observed barcode sequence, and the method further comprises(e) determining a second matching probability between each expected barcode sequence of the plurality of expected barcode sequences and the second observed barcode sequence by, for each expected barcode sequence, multiplying N event rates one for each base position in a plurality of N base positions of the second observed barcode sequence;(f) identifying, from the plurality of expected barcode sequences, a second expected barcode sequence having a highest second matching probability with the second observed barcode sequence; and(g) assigning the second observed barcode sequence to the second expected barcode sequence having the highest second matching probability.

3. The method of claim 1 or 2, further comprising retrieving the plurality of expected barcode sequences from a barcode data set.

4. The method of any one of claims 1-3, further comprising, for each base position in N, randomly assigning each event rate with an initial value such that the initial value for an event rate of the observed first base being identical to the actual second base is the highest, optionally greater than 0.5, 0.6, 0.7, 0.8 or 0.9.

5. The method of claim 4, wherein the remaining initial values for event rates of theobserved first base being different from the actual second base are the same.

6. The method of any one of claims 1-3, wherein the event rate for each base position in N is pre-determined.

7. The method of claim 6, wherein the event rate for each base position in N for an iteration is pre-determined in a prior iteration by:(i) assigning an observed barcode sequence of each sequence read of a plurality of sequence reads to one of the plurality of expected barcode sequences; and(ii) for each base position in N, calculating the probability of observing the first base that is expected to be the second base by determining the number of times the observed first base corresponds to the expected second base based on the assignment from step (i).

8. The method of claim 7, further comprising iterating steps (i) and (ii) for 2, 3, 4, or 5 times, or until the event rates remain essentially the same as the event rates from the previous iteration or the difference between the event rates from two consecutive iterations is below a threshold.

9. The method of claim 7 or 8, wherein step (i) for an iteration is performed based on a Hamming distance between an observed barcode sequence and an expected barcode sequence, optionally, the maximum Hamming distance is between 0 and N, inclusive, further optionally, the maximum Hamming distance is equal to or less than 10.

10. The method of any one of claims 1-9, further comprising: for each base position in N, updating the event rate based on the assignment, and repeating steps (b)-(d), optionally updating the event rate comprises calculating the probability of observing the first base that is expected to be the second base, optionally updating the even rate comprises determining the number of times the observed first base corresponds to the expected second base based on the assignment.

11. The method of any one of claims 1-10, wherein (a) the first base is A, C, T, G, or N, and the second base is A, C, T, or G; and / or (b) the first base is the same as or different from the second base.

12. The method of any one of claims 7-11, wherein step (ii) comprises generating at least one event probability matrix comprising N sub-matrices each for a position in N, wherein each sub-matrix has rows and columns representing different observed bases and expected bases and event rates each corresponding to a pair of observed base and expected base.

13. The method of any one of claims 1-12, further comprising multiplying the highest first matching probability by a relative abundance value of the identified first expected barcode sequence, and / or multiplying the highest second matching probability by a relative abundance value of the identified second expected barcode sequence.

14. The method of any one of claims 1-13, further comprising comparing the highest first matching probability of the identified first expected barcode sequence to the sum of other matching probabilities of other expected barcode sequences.

15. The method of claim 14, wherein the first observed barcode sequence is assigned to the identified first expected barcode sequence if the highest first matching probability is at least 20 million times greater than the sum of other matching probabilities of other expected barcode sequences.

16. The method of any one of claims 1-15, further comprising comparing the highest second matching probability to the sum of other matching probabilities of other expected barcode sequences.

17. The method of any one of claims 1-16, wherein the first observed barcode sequence is assigned to the identified first expected barcode sequence if the highest first matching probability is above a first threshold.

18. The method of any one of claims 1-17, wherein the second observed barcode sequence is assigned to the identified second expected barcode sequence if the highest second matching probability is above a second threshold.

19. The method of claim 17 or 18, wherein the first threshold and / or the second threshold is at least 10'5 5.

20. The method of any one of claims 1-19, wherein the first observed barcode sequence is assigned to the identified first expected barcode sequence if the Hamming distance between the first observed barcode sequence and the identified first expected barcode sequence is below a certain value, optionally equal to or less than 10.

21. The method of any one of claims 1-20, wherein the maximum Hamming distance between the observed barcode sequence and the identified expected barcode sequence is below a certain value, optionally equal to or less than 10.

22. The method of any one of claims 1-21, further comprising discarding the first and second observed barcode sequences if the combination of the first observed barcode sequence and the second observed barcode sequence does not exist in the barcode data set.

23. The method of any one of claims 1-22, further comprising assigning the sequence read to a respective sample in a plurality of samples using the assigned expected barcode sequence.

24. The method of any one of claims 1-23, wherein the first observed barcode sequence and / or the second observed barcode sequence are sample-specific barcode sequences.

25. The method of any one of claims 1-24, further comprisingproviding a plurality of sequence reads each comprising a sequence being sequenced and at least one observed barcode sequence; and comparing each observed barcode sequence of the plurality sequence reads to each of the plurality of expected barcode sequences.

26. The method of claim 25, further comprising for a sequence read in the plurality of sequence reads comprising an observed barcode sequence having an identical match in the plurality of expected barcode sequences, assigning the observed barcode sequence to the sequence read as a true barcode sequence.

27. The method of claim 26, further comprising assigning the sequence read to a respective sample in a plurality of samples using the observed barcode sequence.

28. The method of any one of claims 1-27, further comprising sequencing a plurality of barcoded nucleotide molecules with a sequencing system, thereby obtaining the plurality of sequence reads.

29. The method of claim 28, wherein the plurality of barcoded nucleotide molecules are pooled from a plurality of samples, wherein each barcode nucleotide molecule comprises a sample-specific barcode.

30. The method of claim 29, wherein the plurality of samples are different.

31. The method of any one of claims 1-30, wherein the number of expected barcode sequences is from about 3 to about 50.

32. The method of any one of claims 1-31, wherein the first observed barcode sequence and / or the second observed barcode sequence is from about 3 to about 12 bases in length.

33. A system for determining a barcode sequence, comprising: non-transitory memory configured to store executable instructions; and hardware processor in communication with the non-transitory memory, the hardware processor programmed by the executable instructions to perform: any one of methods 1-32.

34. A system for determining a barcode sequence, comprising: non-transitory memory configured to store executable instructions; and a hardware processor in communication with the non-transitory memory, the hardware processor programmed by the executable instructions to perform: receiving a sequence read comprising a sequence being sequenced and an observed barcode sequence that does not have an identical match in a plurality of expected barcode sequences; determining a matching probability between the observed barcode sequence and each expected barcode sequence of the plurality of expected barcode sequences, by, foreach expected barcode sequence, multiplying N event rates one for each base position in a plurality of N base positions of the observed barcode sequence, wherein an event rate is the probability of observing a first base at a base position when the actual base is a second base; identifying, from the plurality of expected barcode sequences, an expected barcode sequence having a highest matching probability with the observed barcode sequence; and assigning the observed barcode sequence to the identified expected barcode sequence.

35. A method for determining a pattern, comprising:(a) obtaining a first observed pattern that does not have an identical match in a plurality of expected patterns;(b) determining a first matching probability between the first observed pattern and each expected pattern of the plurality of expected patterns, based on N event rates one for each pattern component position in a plurality of N pattern component positions of the first observed pattern, wherein an event rate is the probability of observing a first pattern component at a pattern component position when the actual pattern component is a second pattern component;(c) identifying, from the plurality of expected patterns, a first expected pattern having a highest first matching probability with the first observed pattern; and(d) assigning the first observed pattern to the first expected pattern having the highest first matching probability.

36. The method of claim 35, further comprising, for each pattern component position in N, randomly assigning each event rate with an initial value such that the initial value for an event rate of the observed first pattern component being identical to the actual second pattern component is the highest, optionally greater than 0.5, 0.6, 0.7, 0.8 or 0.9.

37. The method of claim 36, wherein the remaining initial values for event rates of the observed first pattern component being different from the actual second pattern component are the same.

38. The method of claim 35, wherein the event rate for each pattern component position in N is pre-determined.

39. The method of claim 38, wherein the event rate for each pattern component position in N for a iteration is determined in the prior iteration by:(i) assigning each of a plurality of observed patterns to one of the plurality of expected patterns; and(ii) for each pattern component position in N, calculating the probability of observing the first pattern component that is expected to be the second pattern component by determining the number of times the observed first pattern component corresponds to the expected second pattern component based on the assignment from step (i).

40. The method of claim 39, further comprising iterating steps (i) and (ii) for a number of iterations, or until the event rates remain essentially the same as the event rates from the previous iteration or the difference between the event rates from two consecutive iterations is below a threshold, optionally wherein the number of iterations is 2, 3, 4, 5, 6, 7, 8, 9, or 10.

41. The method of claim 39 or 40, wherein step (i) for a iteration is performed based on a Hamming distance between an observed pattern and an expected pattern, optionally, the maximum Hamming distance is between 0 and N, inclusive, further optionally, the maximum Hamming distance is equal to or less than 10.

42. The method of any one of claims 35-41, further comprising: for each pattern component position in N, updating the event rate based on the assignment, and repeating steps (b)-(d), optionally updating the event rate comprises calculating the probability of observing the first pattern component that is expected to be the second pattern component, updating the event rate comprises: determining the number of times the observed first pattern component corresponds to the expected second pattern component based on the assignment.

43. The method of any one of claims 35-42, wherein a pattern, the pattern being an observed pattern or an expected pattern, comprises a barcode sequence, optionally wherein a barcode sequences is associated with a sequence being sequenced, optionally wherein a barcode is associated with a sequence read, optionally a sequence read comprises a barcode sequence and a sequence being sequenced, and optionally wherein sequence read comprises paired-end sequence reads and / or a single-end sequence read.

44. The method of claim 43, wherein (a) the first pattern component is A, C, T, G, or N, and the second pattern component is A, C, T, or G.

45. The method of any one of claims 35-42, wherein a pattern, the pattern being an observed pattern or an expected pattern, comprises a barcode, a ID barcode, a Universal Product Code (UPC), a 2D barcode, a matrix barcode, a two-dimensional matrix barcode, a Quick Response code (QR code), a static QR code, a dynamic QR code, a High Capacity Colored two- Dimensional code (HCC2D code), a digital code embedded in a presentation or image, a digital code represented by a presentation or image, a predefined custom pattern, or an encoding using an image.

46. The method of any one of claims 35-45, the first pattern component is the same as or different from the second base component.

47. The method of any one of claims 35-46, wherein obtaining the first observed pattern comprises receiving or retrieving the first observed pattern, or wherein obtaining the first observed pattern comprises capturing the first observed pattern using an imaging sensor, optionally wherein the imaging sensor is comprised in a camera or a phone.

48. The method of any one of claims 39-47, wherein step (ii) comprises generating at least one event probability matrix comprising N sub-matrices each for a position in N, wherein each sub-matrix has rows and columns representing different observed pattern components and expected pattern components and event rates each corresponding to a pair of observed pattern component and expected pattern component.

49. The method of any one of claims 35-48, wherein determining the first matching probability between the first observed pattern and each expected pattern of the plurality of expected patterns is based on a relative abundance value of the identified first expected pattern, and / or a relative abundance value of the identified second expected pattern.

50. The method of any one of claims 35-49, wherein determining the first matching probability between the first observed pattern and each expected pattern of the plurality of expected patterns comprises: for each expected pattern, multiplying N event rates one for each pattern component position in a plurality of N pattern component positions of the first observed pattern.

51. The method of claim 50, further comprising multiplying the highest first matching probability by a relative abundance value of the identified first expected pattern, and / or multiplying the highest second matching probability by a relative abundance value of the identified second expected pattern.

52. The method of any one of claims 35-51, further comprising comparing the highest first matching probability of the identified first expected pattern to the sum of other matching probabilities of other expected patterns.

53. The method of claim 52, wherein the first observed pattern is assigned to the identified first expected pattern if the highest first matching probability is at least 20 million times greater than the sum of other matching probabilities of other expected patterns.

54. The method of any one of claims 35-53, further comprising comparing the highest second matching probability to the sum of other matching probabilities of other expected patterns.

55. The method of any one of claims 35-54, wherein the first observed pattern is assigned to the identified first expected pattern if the highest first matching probability is above a first threshold.

56. The method of any one of claims 35-55, wherein the second observed pattern isassigned to the identified second expected pattern if the highest second matching probability is above a second threshold.

57. The method of claim 55 or 56, wherein the first threshold and / or the second threshold is at least IO'5 5.

58. The method of any one of claims 35-57, wherein the first observed pattern is assigned to the identified first expected pattern if the Hamming distance between the first observed pattern and the identified first expected pattern is below a certain value, optionally equal to or less than 10.

59. The method of any one of claims 35-58, wherein the maximum Hamming distance between the observed pattern and the identified expected pattern is below a certain value, optionally equal to or less than 10.

60. The method of any one of claims 35-59, further comprising assigning the first observed pattern, or a sample associated with the first observed pattern, to a respective sample in a plurality of samples using the assigned expected pattern.

61. The method of any one of claims 35-60, wherein the first observed pattern is a sample-specific pattern.

62. The method of any one of claims 35-61, further comprising providing a plurality of s observed patterns; and comparing each observed pattern of the plurality observed patterns to each of the plurality of expected patterns.

63. The method of claim 62, further comprising for an observed pattern in the plurality of observed patterns having an identical match in the plurality of expected patterns, assigning the observed pattern as a true pattern.

64. The method of claim 63, further comprising assigning the observed pattern, or a sample associated with the observed pattern, to a respective sample in a plurality of samples using the observed pattern.

65. The method of claim 64, wherein the plurality of samples are different.

66. The method of any one of claims 35-65, wherein the number of expected patterns is from about 3 to about 50.

67. The method of any one of claims 35-66, wherein the first observed pattern and / or the second observed pattern is from about 3 to about 12 pattern components in length.

68. A system for determining a pattern, comprising: non-transitory memory configured to store executable instructions; and hardware processor in communication with the non-transitory memory, the hardware processor programmed by the executable instructions to perform: any one of methods 35-67.

69. A system for determining a pattern, comprising: non-transitory memory configured to store executable instructions; and a hardware processor in communication with the non-transitory memory, the hardware processor programmed by the executable instructions to perform: receiving a sequence read comprising a sequence being sequenced and an observed pattern that does not have an identical match in a plurality of expected patterns; determining a matching probability between the observed pattern and each expected pattern of the plurality of expected patterns, by, for each expected pattern, multiplying N event rates one for each pattern component position in a plurality of N pattern component positions of the observed pattern, wherein an event rate is the probability of observing a first pattern component at a pattern component position when the actual pattern component is a second pattern component; identifying, from the plurality of expected patterns, an expected pattern having a highest matching probability with the observed pattern; and assigning the observed pattern to the identified expected pattern.

70. A method of determining barcode sequences comprising: under control of a hardware processor:(a) receiving a plurality of sequence reads each comprising a sequence being sequenced and an observed barcode sequence, wherein an observed barcode sequence yields from an expected barcode sequence of a plurality of expected barcode sequences, and wherein the observed barcode sequence of a sequence read of the plurality sequences is different from each expected barcode sequence of the plurality of expected barcode sequences;(b) for each of one or more sequence reads of the plurality of sequence reads: determining a probability that the observed barcode sequence of the sequence read yields from each expected barcode sequence of the plurality of expected barcode sequences using a probability matrix, wherein the probability matrix comprises a probability of an observed base yields from each of a plurality of possible bases for each base position of a plurality of base positions; assigning the observed barcode sequence of the sequence read to an expected barcode sequence of the plurality of expected barcode sequences based on the probabilities of the observed barcode sequence of the sequence read yielding from the expected barcode sequences.

71. A method of determining barcode sequences comprising: under control of a hardware processor:(a) receiving a plurality of sequence reads each comprising a sequence being sequenced and an observed barcode sequence, wherein an observed barcode sequence yields from an expected barcode sequence of a plurality of expected barcode sequences, and wherein the observed barcode sequence of a sequence read of the plurality sequences is different from each expected barcode sequence of the plurality of expected barcode sequences; and(b) for each of one or more sequence reads of the plurality of sequence reads: if the observed barcode sequence of the sequence read is identical to an expected barcode sequence: assigning the observed barcode sequence to the identical expected barcode sequence, if the observed barcode sequence of the sequence read is not identical to any expected barcode sequence of the plurality of expected barcode sequences: determining a probability that the observed barcode sequence of the sequence read yields from each expected barcode sequence of the plurality of expected barcode sequences using a probability matrix, wherein the probably matrix comprises a probability of an observed base yields from each of a plurality of possible bases for each base position of a plurality of base positions; and assigning the observed barcode sequence of the sequence read to an expected barcode sequence of the plurality of expected barcode sequences based on the probabilities of the observed barcode sequence of the sequence read yielding from the expected barcode sequences.

72. A method of determining barcode sequences comprising: under control of a hardware processor:(a) receiving a plurality of sequence reads each comprising a sequence being sequenced and an observed barcode sequence, wherein an observed barcode sequence yields from an expected barcode sequence of a plurality of expected barcode sequences, and wherein the observed barcode sequence of a sequence read of the plurality sequences is different from each expected barcode sequence of the plurality of expected barcode sequences; iteratively,(b) for each of one or more sequence reads of the plurality of sequence reads: determining a probability that the observed barcode sequence of the sequence read yields from each expected barcode sequence of the plurality of expected barcode sequences using a probability matrix, wherein the probablymatrix comprises a probability of an observed base yields from each of a plurality of possible bases for each base position of a plurality of base positions; assigning the observed barcode sequence of the sequence read to an expected barcode sequence of the plurality of expected barcode sequences based on the probabilities of the observed barcode sequence of the sequence read yielding from the expected barcode sequences;(c) determining an updated probability matrix based on assignments of the observed barcode sequences of the sequence reads to the expected barcode sequences if the current iteration is not the last iteration, wherein the updated probability matrix is the probability matrix of the subsequent iteration.

73. The method of any one of claims 70-72, wherein assigning the observed barcode sequence of the sequence read comprises: assigning the observed barcode sequence of the sequence read to the assigned expected barcode sequence if the probability that the observed barcode sequence of the sequence read yields from the assigned expected barcode is higher than the probabilities that the observed barcode sequence yield from the remaining expected barcode sequences of the plurality of expected barcode sequences.

74. The method of any one of claims 70-73, wherein step (b) is performed for a plurality of iterations.

75. The method of any one of claims 70-74, further comprising: determining an updated probability matrix.

76. The method of any one of claims 70-75, further comprising: determining an updated probability matrix based on assignments of the observed barcode sequences of the sequence reads to the expected barcode sequences.

77. The method of any one of claims 70-75, further comprising: determining an updated probability matrix based on the observed barcode sequences of the sequence reads assigned to each expected barcode sequence of the plurality of expected barcode sequences.

78. The method of any one of claims 70-75, further comprising: determining an updated probability matrix based on (x) the observed barcode sequences of the sequence reads and (y) the expected barcode sequences of the plurality of expected barcode sequences to which the observed barcode sequences are assigned.

79. The method of any one of claims 70-75, further comprising: determining an updated probability matrix based on (x) bases of the observed barcode sequences of the sequence reads at each base position of the plurality of base positions and (y) bases of the expected barcode sequences, to which the observed barcode sequences are assigned, at each base position of the plurality of base positions.

80. The method of any one of claims 70-75, further comprising: determining an updated probability matrix based on (x) a number of each base, of the observed barcode sequences of the sequence reads, at each base position of the plurality of base positions and (y) bases of the expected barcode sequences of the expected barcode sequences, to which the observed barcode sequences are assigned, at each base position of the plurality of base positions.

81. The method of any one of claims 70-75, further comprising: determining an updated probability matrix based on, for each base position of the plurality of base positions, (x) a number of each base and the identity of the base, of the observed barcode sequences of the sequence reads, at the base position and (y) identities of bases of the expected barcode sequences of the expected barcode sequences, to which the observed barcode sequences are assigned, at the base position.

82. The method of any one of claims 73-81, wherein the updated probability matrix is determined in one iteration of the plurality of iterations, and wherein the updated probability matrix is used to determine probabilities of the observed barcode sequences in a subsequent iteration of the plurality of iterations.

83. The method of any one of claims 70-82, wherein the plurality of sequence reads comprise paired-end sequence reads and / or single-end sequence reads.

84. A system for determining barcode sequences, comprising: non-transitory memory configured to store executable instructions; and hardware processor in communication with the non-transitory memory, the hardware processor programmed by the executable instructions to perform: any one of methods 70-83.

85. A method of determining patterns comprising: under control of a hardware processor:(a) receiving a plurality of observed patterns, wherein an observed pattern yields from an expected pattern of a plurality of expected patterns, and wherein at least one observed pattern of the plurality observed patterns is different from each expected pattern of the plurality of expected patterns;(b) for each of one or more observed patterns of the plurality of observed patterns: determining a probability that the observed pattern originates from each expected pattern of the plurality of expected patterns using a probability matrix, wherein the probability matrix comprises a probability of an observed pattern component yields from each of a plurality of possible pattern components for each pattern component position of a plurality of pattern component positions;assigning the observed pattern to an expected pattern of the plurality of expected patterns based on the probabilities of the observed pattern yielding from the expected patterns.

86. A method of determining patterns comprising: under control of a hardware processor:(a) receiving a plurality of observed patterns, wherein an observed pattern yields from an expected pattern of a plurality of expected patterns, and wherein at least one observed pattern of the plurality observed patterns is different from each expected pattern of the plurality of expected patterns; and(b) for each of one or more observed patterns of the plurality of observed patterns: if the observed pattern is identical to an expected pattern: assigning the observed pattern to the identical expected pattern, if the observed pattern is not identical to any expected pattern of the plurality of expected patterns: determining a probability that the observed pattern yields from each expected pattern of the plurality of expected patterns using a probability matrix, wherein the probably matrix comprises a probability of an observed pattern component yields from each of a plurality of possible pattern components for each pattern component position of a plurality of pattern component positions; and assigning the observed pattern to an expected pattern of the plurality of expected patterns based on the probabilities of the observed pattern yielding from the expected patterns.

87. A method of determining patterns comprising: under control of a hardware processor:(a) receiving a plurality of observed patterns, wherein an observed pattern yields from an expected pattern of a plurality of expected patterns, and wherein an observed pattern of the plurality of observed patterns is different from each expected pattern of the plurality of expected patterns; iteratively,(b) for each of one or more observed patterns of the plurality of observed patterns: determining a probability that the observed pattern yields from each expected pattern of the plurality of expected patterns using a probability matrix,wherein the probably matrix comprises a probability of an observed pattern component yields from each of a plurality of possible pattern components for each pattern component position of a plurality of pattern component positions; assigning the observed pattern of the sequence read to an expected pattern of the plurality of expected patterns based on the probabilities of the observed pattern yielding from the expected patterns;(c) determining an updated probability matrix based on assignments of the observed patterns of the sequence reads to the expected patterns if the current iteration is not the last iteration, wherein the updated probability matrix is the probability matrix of the subsequent iteration.

88. The method of any one of claims 85-87, wherein assigning the observed pattern comprises: assigning the observed pattern to the assigned expected pattern if the probability that the observed pattern originates from the assigned expected barcode is higher than the probabilities that the observed pattern originate from the remaining expected patterns of the plurality of expected patterns.

89. The method of any one of claims 85-88, wherein step (b) is performed for a plurality of iterations.

90. The method of any one of claims 85-89, further comprising: determining an updated probability matrix.

91. The method of any one of claims 85-90, further comprising: determining an updated probability matrix based on assignments of the observed patterns to the expected patterns.

92. The method of any one of claims 85-90, further comprising: determining an updated probability matrix based on the observed patterns assigned to each expected pattern of the plurality of expected patterns.

93. The method of any one of claims 85-90, further comprising: determining an updated probability matrix based on (x) the observed patterns and (y) the expected patterns of the plurality of expected patterns to which the observed patterns are assigned.

94. The method of any one of claims 85-90, further comprising: determining an updated probability matrix based on (x) pattern components of the observed patterns at each pattern component position of the plurality of pattern component positions and (y) pattern components of the expected patterns, to which the observed patterns are assigned, at each pattern component position of the plurality of pattern component positions.

95. The method of any one of claims 85-90, further comprising: determining an updated probability matrix based on (x) a number of each pattern component, of the observedpatterns, at each pattern component position of the plurality of pattern component positions and (y) pattern components of the expected patterns of the expected patterns, to which the observed patterns are assigned, at each pattern component position of the plurality of pattern component positions.

96. The method of any one of claims 85-90, further comprising: determining an updated probability matrix based on, for each pattern component position of the plurality of pattern component positions, (x) a number of each pattern component and the identity of the pattern component, of the observed patterns, at the pattern component position and (y) identities of pattern components of the expected patterns of the expected patterns, to which the observed patterns are assigned, at the pattern component position.

97. The method of any one of claims 85-96, wherein the updated probability matrix is determined in one iteration of the plurality of iterations, and wherein the updated probability matrix is used to determine probabilities of the observed patterns in a subsequent iteration of the plurality of iterations.

98. The method of any one of claims 85-97, wherein a pattern, the pattern being an observed pattern or an expected pattern, comprises a barcode sequence, optionally wherein a barcode sequences is associated with a sequence being sequenced, optionally wherein a barcode is associated with a sequence read, optionally a sequence read comprises a barcode sequence and a sequence being sequenced, and optionally wherein sequence read comprises paired-end sequence reads and / or a single-end sequence read.

99. The method of any one of claims 85-97, wherein a pattern, the pattern being an observed pattern or an expected pattern, comprises a barcode, a ID barcode, a Universal Product Code (UPC), a 2D barcode, a matrix barcode, a two-dimensional matrix barcode, a Quick Response code (QR code), a static QR code, a dynamic QR code, a High Capacity Colored two- Dimensional code (HCC2D code), a digital code embedded in a presentation or image, a digital code represented by a presentation or image, a predefined custom pattern, or an encoding using an image.

100. The method of any one of claims 85-99, wherein receiving the plurality of observed patterns comprises capturing the plurality of patterns, optionally using an imaging sensor, optionally wherein the imaging sensor is comprised in a camera or a phone.

101. A system for determining patterns, comprising: non-transitory memory configured to store executable instructions; and hardware processor in communication with the non-transitory memory, the hardware processor programmed by the executable instructions to perform: any one of methods 85-100.

Citation Information

Patent Citations

  • Systems and methods for sequencing in emulsion based microfluidics

    US20180073074A1

  • Methods for non-invasive prenatal ploidy calling

    US20190309358A1

  • Systems and methods for identification of nucleic acids in a sample

    US20200131506A1

  • Enzyme screening methods

    US20200283842A1

  • Methods for non-invasive prenatal testing

    WO2023034090A1