Systems and methods for generating nucleotide sequence synthesis related metrics - Patents.com

JP2024534353A5Pending Publication Date: 2025-09-26NUTCRACKER THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024515624
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-17
Filing Date
2022-09-16
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Errors occur during the manufacturing of nucleotide sequences, leading to mismatches between the produced sequence and the target sequence, which affects the accuracy and quality of nucleic acid synthesis and sequencing processes.

Method used

A method and system for generating nucleotide sequence synthesis metrics using processors to detect unique molecular indices, adjust metric scores to reduce errors, and modify the synthesis process to align with target sequences, including the use of microfluidic pathway devices and sequencing systems to enhance accuracy.

Benefits of technology

The method and system improve the accuracy of nucleic acid synthesis and sequencing by reducing errors and ensuring that the produced sequences closely resemble the target sequences, facilitating the production of high-quality nucleic acids for therapeutic applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In one example, a system is provided that includes one or more processors to receive an array data structure, apply at least one first metric of a plurality of metrics to the array data structure to generate at least one first metric score, determine that the at least one first metric score satisfies a first condition, apply at least one second metric of the plurality of metrics to the array data structure to generate at least one second metric score in response to the at least one first metric score satisfying the first condition, and output an indication of the at least one first metric score and the at least one second metric score.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 245,528, filed September 17, 2021, the disclosure of which is incorporated herein by reference in its entirety. [Background technology]

[0002] Nucleotide sequences, such as mRNA sequences, can be manufactured to have a variety of characteristics. Errors can occur in the manufacturing process, which can result in the manufactured nucleotide sequence not matching the target sequence. Summary of the Invention

[0003] Some examples provided herein relate generally to the field of nucleotide sequences. More specifically, the present disclosure relates to systems and methods for generating nucleotide sequence synthesis-related metrics.

[0004] At least one embodiment relates to a method of generating nucleic acid sequencing metrics, the method may include receiving, using one or more processors, a sequence data structure, applying, using one or more processors, at least one first metric of the plurality of metrics to the sequence data structure to generate at least one first metric score, determining, using one or more processors, that the at least one first metric score satisfies a first condition, applying, using one or more processors, in response to determining that the at least one first metric score satisfies the first condition, applying, using one or more processors, at least one second metric of the plurality of metrics to the sequence data structure to generate at least one second metric score, and outputting, using one or more processors, an indication of the at least one first metric score and the at least one second metric score.

[0005] The method may include detecting, by one or more processors, at least one unique molecular index in the sequence data structure; and generating, by the one or more processors, at least one second metric score using the at least one unique molecular index.

[0006] Generating the at least one second metric score may include adjusting the at least one second metric score to reduce an error rate associated with the sequence data structure.

[0007] The at least one unique molecular index can be adjacent to the i7 index of the sequence data structure.

[0008] The method can include identifying, by the one or more processors, at least one synthetic adaptor sequence in the sequence data structure; removing, by the one or more processors, the at least one synthetic adaptor sequence from the sequence data structure; merging, by the one or more processors, a plurality of unique molecular indexes in the sequence data structure from which the at least one synthetic adaptor sequence has been removed in response to removing, by the one or more processors; and generating, by the one or more processors, at least one second metric score to include a mismatch rate in response to merging the plurality of unique molecular indexes.

[0009] The subset of the plurality of metrics can include at least one metric of the flow cell used to generate the sequence data structure.

[0010] The method may include outputting, by the one or more processors, an indication of the failure condition in response to the at least one first metric score failing the first condition.

[0011] Outputting the instructions can include generating, by the one or more processors, the instructions such that the sequence data structure includes instructions for altering the synthesis of the target sequence that is detected.

[0012] The method can include using the instructions to modify the synthesis of the target sequence.

[0013] At least one second metric score can include a target mismatch rate.

[0014] The sequence data structure can be from parallel sequencing of nucleic acids.

[0015] The method can include producing the mRNA therapeutic using the instructions.

[0016] Producing an mRNA therapeutic using the instructions can include using one or more processors to control operation of the processor chip to reduce differences between the mRNA of the mRNA therapeutic and a target sequence of the mRNA using the instructions.

[0017] The method can include encapsulating at least one mRNA with at least one delivery vehicle composition to form an mRNA therapeutic.

[0018] The method may include generating a report that includes the instructions.

[0019] The method can include using the instructions to modify a given mRNA production process.

[0020] At least one aspect relates to a system that can include one or more processors to receive an array data structure, apply at least one first metric of a plurality of metrics to the array data structure to generate at least one first metric score, determine that the at least one first metric score satisfies a first condition, apply at least one second metric of the plurality of metrics to the array data structure to generate at least one second metric score in response to the at least one first metric score satisfying the first condition, and output an indication of the at least one first metric score and the at least one second metric score.

[0021] The one or more processors can be for detecting at least one unique molecular index in the sequence data structure and generating at least one second metric score using the at least one unique molecular index.

[0022] The one or more processors can be for generating the at least one second metric score by adjusting the at least one second metric score to reduce an error rate associated with the sequence data structure.

[0023] At least one molecule index can be adjacent to an i7 index of the sequence data structure.

[0024] The one or more processors can be for identifying at least one synthetic adaptor sequence in the sequence data structure, removing the at least one synthetic adaptor sequence from the sequence data structure, merging a plurality of unique molecular indexes in the sequence data structure from which the at least one synthetic adaptor sequence has been removed in response to removing the at least one synthetic adaptor sequence, and generating at least one second metric score to include a mismatch rate in response to merging the plurality of unique molecular indexes.

[0025] The subset of metrics can include at least one metric of the flow cell used to generate the sequence data structure.

[0026] The one or more processors can be for outputting an indication of the failure condition in response to the at least one first metric score failing the first condition.

[0027] The one or more processors can be for generating instructions such that the sequence data structure includes instructions for altering the synthesis of a target sequence to be detected.

[0028] The one or more processors can be for controlling the operation of the processor chip to synthesize the target sequence using the instructions.

[0029] At least one second metric score can include a target mismatch rate.

[0030] The sequence data structure can be from parallel sequencing of nucleic acids.

[0031] The one or more processors can be for controlling the operation of the processor chip to use the instructions to generate the mRNA therapeutic.

[0032] The one or more processors can be for performing a process of encapsulating at least one mRNA with at least one delivery vehicle composition to form an mRNA therapeutic.

[0033] The one or more processors can be for generating a report that includes the instructions.

[0034] The one or more processors can be for using the instructions to modify a given mRNA production process.

[0035] At least one embodiment relates to a sequencing device, which may include a sequencer for generating a sequence data structure based on a flow cell containing a target sequence, and one or more processors for receiving the sequence data structure, applying at least one first metric of the plurality of metrics to the sequence data structure to generate at least one first metric score, determining that the at least one first metric score satisfies a first condition, applying at least one second metric of the plurality of metrics to the sequence data structure to generate at least one second metric score in response to the at least one first metric score satisfying the first condition, and outputting an indication of the at least one first metric score and the at least one second metric score.

[0036] The one or more processors can be for detecting at least one unique molecular index in the sequence data structure and generating at least one second metric score using the at least one unique molecular index.

[0037] The one or more processors can be for generating the at least one second metric score by adjusting the at least one second metric score to reduce an error rate associated with the sequence data structure.

[0038] At least one molecule index can be adjacent to an i7 index of the sequence data structure.

[0039] The one or more processors can be for identifying at least one synthetic adaptor sequence in the sequence data structure, removing the at least one synthetic adaptor sequence from the sequence data structure, merging a plurality of unique molecular indexes in the sequence data structure from which the at least one synthetic adaptor sequence has been removed in response to removing the at least one synthetic adaptor sequence, and generating at least one second metric score to include a mismatch rate in response to merging the plurality of unique molecular indexes.

[0040] The subset of metrics can include at least one metric of the flow cell used to generate the sequence data structure.

[0041] The one or more processors can be for outputting an indication of the failure condition in response to the at least one first metric score failing the first condition.

[0042] The one or more processors can be for generating instructions such that the sequence data structure includes instructions for altering the synthesis of a target sequence to be detected.

[0043] The one or more processors can be for controlling the operation of the processor chip to synthesize the target sequence using the instructions.

[0044] At least one second metric score can include a target mismatch rate.

[0045] The sequence data structure can be from parallel sequencing of nucleic acids.

[0046] The one or more processors can be for controlling the operation of the processor chip to use the instructions to generate the mRNA therapeutic.

[0047] The one or more processors can be for performing a process of encapsulating at least one mRNA with at least one delivery vehicle composition to form an mRNA therapeutic.

[0048] The one or more processors can be for generating a report that includes the instructions.

[0049] The one or more processors can be for using the instructions to modify a given mRNA production process.

[0050] At least one aspect relates to a non-transitory processor-readable medium. The computer-readable medium can include computer-readable instructions that, when executed by one or more processors, can cause the one or more processors to receive an array data structure, apply at least one first metric of the plurality of metrics to the array data structure to generate at least one first metric score, determine that the at least one first metric score satisfies a first condition, apply at least one second metric of the plurality of metrics to the array data structure to generate at least one second metric score in response to the at least one first metric score satisfying the first condition, and output an indication of the at least one first metric score and the at least one second metric score.

[0051] The processor-readable medium can include instructions that cause one or more processors to detect at least one unique molecular index in a sequence data structure and generate at least one second metric score using the at least one unique molecular index.

[0052] The processor-readable medium may include instructions that cause the one or more processors to generate at least one second metric score by adjusting the at least one second metric score to reduce an error rate associated with the sequence data structure.

[0053] At least one molecule index can be adjacent to an i7 index of the sequence data structure.

[0054] The processor-readable medium can include instructions to cause the one or more processors to identify at least one synthetic adaptor sequence from a sequence data structure; to remove the at least one synthetic adaptor sequence from the sequence data structure; in response to removing the at least one synthetic adaptor sequence, to merge a plurality of unique molecular indexes from the sequence data structure from which the at least one synthetic adaptor sequence is removed; and in response to merging the plurality of unique molecular indexes, to generate at least one second metric score to include a mismatch rate.

[0055] The processor-readable medium can include instructions that cause one or more processors to detect at least one unique molecular index in a sequence data structure and generate at least one second metric score using the at least one unique molecular index.

[0056] The processor-readable medium may include instructions that cause the one or more processors to generate the at least one second metric score by adjusting the at least one second metric score to reduce an error rate associated with the sequence data structure.

[0057] At least one molecule index can be adjacent to an i7 index of the sequence data structure.

[0058] The processor-readable medium can include instructions to cause the one or more processors to identify at least one synthetic adaptor sequence from a sequence data structure; to remove the at least one synthetic adaptor sequence from the sequence data structure; in response to removing the at least one synthetic adaptor sequence, to merge a plurality of unique molecular indexes from the sequence data structure from which the at least one synthetic adaptor sequence is removed; and in response to merging the plurality of unique molecular indexes, to generate at least one second metric score to include a mismatch rate.

[0059] The subset of metrics can include at least one metric of the flow cell used to generate the sequence data structure.

[0060] The processor-readable medium may include instructions that cause the one or more processors to output an indication of a failure condition in response to at least one first metric score failing a first condition.

[0061] The processor-readable medium can include instructions that cause one or more processors to generate instructions that indicate the sequence data structure to modify the synthesis of a target sequence that is detected.

[0062] The processor-readable medium can include instructions that cause one or more processors to control the operation of a processor chip to synthesize a target sequence using the instructions.

[0063] At least one second metric score can include a target mismatch rate.

[0064] The sequence data structure can be from parallel sequencing of nucleic acids.

[0065] The processor-readable medium can include instructions that cause one or more processors to control operation of a processor chip to use the instructions to produce an mRNA therapeutic.

[0066] The processor-readable medium can include instructions that cause one or more processors to control operation of a processor chip to use the instructions to reduce the difference between the mRNA of an mRNA therapeutic and a target sequence of the mRNA.

[0067] The processor-readable medium can include instructions that cause one or more processors to perform a process of encapsulating at least one mRNA with at least one delivery vehicle composition to form an mRNA therapeutic.

[0068] The processor-readable medium may include instructions for causing one or more processors to generate a report that includes the instructions.

[0069] The processor-readable medium can include instructions that cause one or more processors to use the instructions to alter a given mRNA production process.

[0070] At least one embodiment relates to a system that can include a microfluidic pathway device for generating nucleic acids and a controller including one or more processors for receiving an indication of an error of a target sequence associated with the nucleic acid relative to a reference sequence and using the indication to control operation of the microfluidic pathway device.

[0071] The controller can be for controlling the operation of the processor chip to use the instructions to reduce the difference between the sequence data structure and the target sequence.

[0072] The controller can be for using the instructions to identify an association between the error and a particular reagent among the plurality of reagents used to generate the nucleic acid and for controlling the operation of the processor chip by causing the processor chip to modify the amount of the particular reagent used to generate the nucleic acid.

[0073] The controller can be for controlling operation of the processor chip by using the instructions to determine from the instructions that the error is associated with a particular contaminant and causing the processor chip to provide the nucleic acid to a purification pathway of the processor chip targeted to remove the particular contaminant from the nucleic acid.

[0074] The nucleic acid can be a first nucleic acid, the processor chip can be for generating a product comprising the first nucleic acid and at least one second nucleic acid, and the controller can be for controlling operation of the processor chip using the instructions to determine a difference between a ratio of the first nucleic acid to the at least one second nucleic acid from the instructions and a target ratio of the first nucleic acid to the at least one second nucleic acid, and modifying operation of the processor chip to reduce the difference.

[0075] These and other aspects and implementations are described in detail below. The preceding information and the following detailed description, including illustrative examples of the various aspects and implementations, provide an overview or framework for understanding the nature and characteristics of the claimed aspects and implementations. The drawings provide illustration and a further understanding of the various aspects and implementations, and are incorporated in and constitute a part of this specification.

[0076] It should be understood that all combinations of the foregoing concepts, and additional concepts discussed in more detail below, are contemplated as being part of the inventive subject matter disclosed herein and may be used in any combination to achieve the benefits described herein. [Brief description of the drawings]

[0077] [Figure 1] FIG. 1 is a block diagram of an example of a nucleic acid production system. [Diagram 2] FIG. 1 is a block diagram of an example of a nucleic acid sequencer. [Diagram 3] FIG. 1 is a block diagram of an example nucleic acid metric generator. [Figure 4] FIG. 1 is a flow diagram of an example of a method for generating nucleic acid sequencing metrics. [Diagram 5] FIG. 1 is a block diagram of an example of a nucleic acid production system. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0078] The systems, devices, and methods described herein can be used to accurately detect errors in analytical processes, polymerase chain reaction (PCR), sequencing, or sequencing data of nucleic acids, including errors resulting from underlying biological differences between the nucleic acid (including mRNA) and the target sequence to which the nucleic acid is intended to be synthesized. For example, various metrics can be determined from the sequencing data to identify and remove errors. The metrics can be determined in a particular order to make the metric determination process more efficient, including by reducing the calculations performed to evaluate failed samples or sequencing runs, for example to identify sequencing errors more quickly and trigger actions to address the errors. In response to the metrics (and the underlying errors detected using the metrics), the metrics can be used to trigger actions such as modifying how the nucleic acid is synthesized. Certain metrics can be determined based on unique molecular indexes (UMIs) contained in the nucleic acid, which can facilitate distinguishing analytical differences between the nucleic acid and the target sequence from biological differences between the nucleic acid and the biological differences. For example, mRNA manufacturing quality control can be implemented and improved by using a UMI to identify sequence data corresponding to nucleic acids with the same UMI and determining that errors detected from the sequence data structure correspond to analytical errors rather than underlying biological errors.

[0079] A. Preparation of the Nucleotide Sequence Nucleotide sequences, such as nucleic acids, including mRNA, can be manufactured or synthesized for use in a variety of applications, including therapeutic delivery to a subject. For example, therapeutics, such as mRNA therapeutics, can be used in multiple treatment modalities, including vaccination, immunotherapy, protein replacement therapy, tissue remodeling / regeneration, and treatment of genetic diseases by gene editing. mRNA can be manufactured to have a target sequence that corresponds to a target protein, such as a target protein for performing a particular treatment. A double-stranded DNA sequence can be used as a template for transcription of mRNA (e.g., by in vitro transcription (IVT)).

[0080] FIG. 1 illustrates an example of a nucleic acid production system 100. The nucleic acid production system 100 can be used to produce a target mRNA sequence, including for therapeutic delivery to a subject. The nucleic acid production system 100 can be implemented using one or more components of the system 500 described with reference to FIG.

[0081] The nucleic acid manufacturing system 100 can include at least one processor chip 104, such as a microfluidic pathway device or biochip. The processor chip 104 can include multiple reactors that receive reagents (e.g., via fluidic channels) and react using the received reagents to generate a target product 106. For example, various processor chips 104 can operate as template devices (e.g., to generate DNA templates corresponding to target mRNA sequences), IVT devices (e.g., to generate target mRNA sequences), formulation devices (e.g., to generate drugs for target mRNA sequences), or various combinations thereof. The processor chip 104 or other components of the nucleic acid manufacturing system 100 can generate an mRNA therapeutic, for example, by encapsulating at least one mRNA with at least one delivery vehicle composition to form an mRNA therapeutic.

[0082] The nucleic acid manufacturing system 100 can include at least one controller 108. The controller 108 can be configured to control the operation of the processor chip 104. For example, the controller 108 can control various components included in or coupled to the processor chip 104, such as flow control devices (e.g., pumps, valves), heating or cooling elements or fluid flows, optical elements, or other components, to control reactions performed by the processor chip 104. The controller 108 can include one or more processors and memories configured to execute computer-readable instructions (e.g., stored in memory) to perform various operations described herein, including receiving sensor data from the sensor 112 and using the sensor data to control the operation of the processor chip 104 or its components.

[0083] The nucleic acid manufacturing system 100 can include at least one sensor 112. The at least one sensor 112 can detect one or more parameters of the processor chip 104 or materials within the processor chip 104, such as fluids, reagents, or products 106.

[0084] The controller 108 may perform various operations to control the generation of the product 106, including based on an indication or metric of the error determined by the metric generator 300 as described further herein. For example, the controller 108 may control the operation of the processor chip 104 (or components thereof) using at least one of the indications or parameters detected by the at least one sensor 112. The controller 108 may control the operation of the processor chip 104 to control the temperature, pressure, flow rate of a reagent, time of introduction of a reagent, duration of reaction or mixing of a reagent, or various other process steps performed in one or more paths of the processor chip 104.

[0085] The controller 108 can control the operation of the processor chip 104 to reduce errors (including, for example, based on identifying a contaminant associated with the error) identified from the instructions (generated by the metric generator 300). For example, the controller 108 can identify or receive instructions to identify a particular reagent or nucleotide of the product 106 associated with the error, such as based on one or more UMIs associated with a material in the product 106 associated with the error (which may be traced back to the pathway in which the UMI was introduced to generate the reagent or product 106). The controller 108 can identify or receive instructions to identify a particular process step or pathway of the processor chip 104 associated with the error and modify the operation of the processor chip 104 to avoid using the particular process step or pathway or reduce the flow rate of a reagent used in the particular process step or pathway. The controller 108 can identify or receive instructions to identify a particular contaminant in the product 106 associated with the error and cause the processor chip 104 to provide the product 106 to a purification pathway to remove the contaminant. For example, the controller 108 can identify or receive instructions that identify a particular type of contaminant, identify or receive instructions (e.g., from a look-up table or other logical or heuristic data structure maintained in the memory of the controller 108) that identify a type of purification to be performed to remove the contaminant, and cause the processor chip 104 to flow the product 106 through the identified purification path to remove the contaminant. The controller 108 can identify or receive instructions that identify a target ratio of nucleic acids in the product 106, determine (from an error or a metric related to the error) a difference between the target ratio and the actual ratio of the product 106, and modify the operation of the processor chip 104 to reduce the difference. The controller 108 can control the production of the product 106 in various control flows and feedback loops using instructions determined using data received from or output by the metric generator 300, for example, to iteratively update the production of the product 106 to achieve a target metric or specification for the product 106.

[0086] B. Nucleic Acid Sequencing To assess the characteristics of the product 106, which may be, for example, a DNA sample, an mRNA sample, or a drug product having nucleic acids, the product 106 may be subjected to sequencing.

[0087] 2 illustrates an example of a sequencer 200. The sequencer 200 can be a next generation sequencing (NGS) system, such as a sequencer that performs parallel sequencing (e.g., by sequencing multiple nucleic acids in parallel). The sequencer 200 can be implemented as a device separate from the nucleic acid manufacturing system 100 and metric generator 300 described herein. The sequencer 200 can receive the product 106 as an input and output at least one sequence data structure 204 that represents the product 106. The sequence data structure 204 can include a plurality of nucleotide data elements 208 that can represent the nucleotides and order of the nucleotides of the nucleic acid of the product 106.

[0088] The sequencer 200 can perform various operations to detect and assign nucleotides of the products 106 to a sequence data structure 204, such as using a flow cell, amplification, staining (e.g., fluorescent dyes), or various combinations thereof. The sequencer 200 can include one or more processors and memories configured to identify each nucleotide of the products 106 and assign the nucleotides to respective data elements 208 of the sequence data structure 204.

[0089] Unique molecular indexes (UMIs) (or unique molecular identifiers) can be applied (e.g., ligated) to the products 106 to uniquely tag the starting molecules of the library preparation. UMIs can be random strands of nucleotides that can be ligated to the molecules of the products 106 prior to other operations that sequence the products 106, such as before PCR is applied to the molecules. UMIs can be applied to label one or more nucleic acid strands of the molecules, for example, UMIs can be used to facilitate double-stranded sequencing and separate UMIs can be used to facilitate tagging of each nucleic acid strand (so that errors can be detected for a particular nucleic acid strand).

[0090] C. Nucleic Acid Sequence Metric Generation As mentioned above, there may be various sources of error of the sequence data structure 204 relative to the actual sequence of the product 106 (e.g., analytical differences such as computational errors, PCR errors, and / or sequencing errors) as well as errors of the product 106 relative to the target sequence intended to be synthesized in manufacturing the product 106 (e.g., biological errors). For example, the product 106 may have misincorporated DNA / RNA bases or contaminating DNA or other possible contaminants. The sequence data structure 204 may be susceptible to factors such as amplification bias or sequencing errors, such as bases included in the sequence data structure 204 but not representing true or actual bases in the underlying product 106 as detected by the sequencing system 200, which may result, for example, from the sequencing by synthesis process performed by the sequencing system 200 becoming out of sync within the clone copies within a cluster, which may result in some molecules within the cluster transmitting inaccurate signals, increasing noise and reducing accuracy in assigning fluorescent signals to base calls.

[0091] Systems and methods according to the present disclosure can address various such computational and biological errors in the sequence data structure 204 and the product 106 by determining one or more metrics of the sequence data structure 204 that correspond to one or more such errors, allowing for more accurate generation of the sequence data structure 204. The metrics can be selectively determined in a particular order, reducing computational resources expended to generate the metrics. The metrics can be used to trigger actions such as modifying how the product 106 is synthesized, modifying how the sequence data structure 204 is generated, or various combinations thereof, allowing for the product 106 to be generated so that it more closely resembles the target sequence (e.g., does not have errors such as bases different from those intended in the target sequence). For example, the metrics can be used for various quality control actions.

[0092] FIG. 3 illustrates an example of a metric generator 300. The metric generator 300 can be used to perform various processes described herein for accurately and efficiently generating metrics for nucleic acids, such as generating metrics based on the sequence data structure 204. The metric generator 300 can be used to evaluate features of one or more nucleic acid samples (e.g., of the product 106 from which the sequence data structure 204 is detected), such as the degree to which the synthesized nucleic acids in the sample differ from their respective target sequences, the relative ratio of target sequences in the sample if multiple target sequences are intended, and the occurrence and classification composition of exogenous nucleic acid contaminants in the sample. The sequence metric generating system 300 can trigger or implement actions to reduce errors represented in the sequence data structure 204, including, for example, ligating UMI tags to library molecules prior to amplification to reduce analytical noise, and merging duplicate read pairs to find and resolve sequencing errors. The metric generator 300 can include one or more processors and memories configured to execute computer-readable instructions for performing various operations described herein.

[0093] The metric generator 300 can receive the array data structure 204. For example, the metric generator 300 can be transmitted the array data structure 204 or can access one or more databases in which the array data structure 204 is stored or maintained.

[0094] The metric generator 300 may apply one or more of the multiple metrics 304 to the array data structure 204 to generate a respective metric score 308 for the metrics 304 and may output an indication 312 of the metric score 308.

[0095] The metric 304 may be one or more functions, calculations, equations, algorithms, filters, rules, heuristics, policies, logic, or other operations that may be implemented by the processor and memory of the metric generator 300 to receive the sequence data structure 204 (or a portion of the data of the sequence data structure 204) and generate an output in response to receiving the sequence data structure 204. As described further herein, the metric generator 300 may selectively apply the metrics 304 to the sequence data structure 204, such as in a particular order, and in response to the applied metrics 304 not satisfying a respective condition (e.g., a threshold), discontinue metric application or otherwise output an error, which may enable the metric generator 300 to process the sequence data structure 204 more efficiently (e.g., using fewer computational resources) to identify and correct at least one of analytical errors of the sequencing process used to generate the sequence data structure 204 or biological errors of the product 106.

[0096] The metrics 304 can include at least one pre-processing metric 320. The pre-processing metric 320 can include various filters, metadata extractors, or other metrics that can be used to prepare the data in the array data structure 204 for further evaluation. The pre-processing metric 320 can include a file integrity check metric 320 that can analyze the array data structure 204 and compare the array data structure 204 to an expected file template (e.g., a list of expected files, such as an XML metadata file). In response to a comparison that indicates that the array data structure 204 does not match the expected file template (e.g., expected files or metadata are missing), the metric generator 300 can output an error.

[0097] The metrics 304 can include at least one sequencer metric 324. The sequencer metric 324 can evaluate features of the sequence data structure 204 related to the sequencing process performed on the product 106 to generate the sequence data structure 204. For example, the sequencer metric 324 can include metrics related to the flow cell used to sequence the product, such as whether the flow cell is expired, whether the flow cell has been rehybridized, the completion status of the flow cell, whether all planned cycles have been completed for all reads, or various combinations thereof. The metric generator 300 can use one or more sequencer metrics 324 as gating metrics to determine whether to evaluate additional metrics. For example, in response to determining whether one or more sequencer metrics 324 meet a target value (e.g., a threshold), the metric generator 300 can determine whether to evaluate additional metrics 304.

[0098] The sequencer metrics 324 can include a flow cell yield metric indicating the yield (e.g., the total yield at base level in the flow cell), which can be compared by the metric generator 300 to an expected value or specification for the type of flow cell (e.g., a value in the order of gigabases (Gb)).

[0099] The sequencer metric 324 can include a base quality metric, such as a metric indicating the percentage of bases (e.g., from one or more of the reads of the sequence data structure 204) that have a quality score that meets or exceeds a threshold. The sequencer metric 324 can indicate a probability of error relative to a threshold. For example, the base quality metric can be a Q30 metric that indicates the fraction of bases in the sequence data structure 204 that have a quality score of at least 30. The quality scores can be included in the sequence data structure 204. The metric generator 300 can evaluate a condition on the base quality metric by comparing the base quality metric to a threshold, such as a minimum threshold (e.g., 80 percent).

[0100] The metric generator 300 can output an error in response to the base quality metric being less than a minimum threshold. The metric generator 300 can determine conditions to be met in response to the base quality metric meeting or exceeding a threshold, for example, and continue to generate various metrics 304 in response to the base quality metric meeting or exceeding a threshold.

[0101] The sequencer metrics 324 can include filter evaluation metrics. For example, the sequence data structure 204 can indicate the number of clusters of the product 106 that passed one or more filters used to generate the sequence data structure 204, such as a filter related to the purity of the intensity of the signal detected by the sequencer 200 to identify bases. The metric generator 300 can evaluate the condition for the filter evaluation metric by comparing the number of clusters that passed the filter to a threshold value (e.g., a threshold value corresponding to the type of flow cell, which can be on the order of millions, such as 7 million for a medium output flow cell or 22 million for a high output flow cell).

[0102] The sequencer metric 324 can include a control library metric. The control library metric can correspond to a control library (e.g., a nucleic acid library, such as a non-indexed library, such as a PhiX control library manufactured by Illumina, Inc. of San Diego, CA) included with the product 106 during sequencing of the product 106 by the sequencer 200. The sequencer 200 can determine the control library metric by comparing sequence reads detected while generating the sequence data structure 204 to the control library to identify a mismatch rate (e.g., the control library metric can be an error rate, such as a mismatch rate). The metric generator 300 can evaluate a condition on the control library metric by comparing the control library metric to a threshold (e.g., a maximum threshold) and determining the control library metric to satisfy the condition in response to the control library metric being equal to or less than the threshold. The metric generator 300 can determine that the condition is not satisfied in response to the control library metric exceeding the threshold.

[0103] The metric 304 can include a read size metric 328. The read size metric 328 can correspond to the number of read clusters of the library represented by the sequence data structure 204. A read cluster can represent one or more nucleic acids of the library represented by the sequence data structure 204 that are identified as duplicates by the sequencer 200 or the metric generator 300. The read size metric 328 can be related to the number of UMI clusters (e.g., read clusters corresponding to unique UMIs) of sufficient depth (e.g., 2 or more depth) because factors such as the molecular complexity (e.g., number of starting DNA molecules) and sequencing depth of the library can affect the number of UMI clusters of sufficient depth, and the number of read clusters can be indicative of such factors.

[0104] Thus, the metric generator 300 may determine a read size metric 328, compare the read size metric 328 to a threshold, such as a minimum threshold (e.g., 1 million read clusters), and determine that the read size metric 328 satisfies the condition in response to the read size metric 328 meeting or exceeding the threshold. The metric generator 300 may determine that the condition is not satisfied in response to the read size metric 328 being less than the threshold, and output, for example, an error.

[0105] As described above, UMIs can be applied to molecules of the product 106 prior to sequencing, and thus the UMIs can be represented in the sequence data structure 204. The metric generator 300 can use the sequence data structure 204 to assign read clusters to corresponding UMI families (each family having the same UMI and position of the UMI in the target sequence). For example, each UMI family can include one or more read clusters having the same UMI at the same position in the target sequence. Thus, differences in the sequence represented by the sequence data structure 204 may correspond to sources of error, such as PCR errors or sequencing errors, which can be attempted to be removed to leave primarily biological differences in the product 106 relative to the target sequence.

[0106] For example, the metric generator 300 can apply at least one UMI difference metric 332 to the sequence data structures 204 to identify base differences between nucleic acids (e.g., UMI clusters) represented by different sequence data structures 204. The metric generator 300 can perform actions such as discarding clusters of the sequence data structures 204 based on the UMI difference metric 332 (e.g., to remove PCR or sequencing errors) prior to further evaluation, or outputting an error in response to the UMI difference metric 332 not meeting or exceeding a threshold (e.g., a sample having less than a threshold number of UMI clusters covered by at least two reads, such as 2000 UMI clusters, results in an error).

[0107] The metric generator 300 can determine at least one target metric 336 (e.g., using the output of any of the metrics described herein or a combination thereof as input), e.g., how closely the sequence data structure 204 matches the target sequence. The metric generator 300 can perform various trimming or merging operations on the sequence data structure 204 before determining the at least one target metric 336. For example, the metric generator 300 can identify and remove bases of adaptors that were ligated during library preparation. The metric generator 300 can merge R1 / R2 read pairs, which can improve the quality of the data in the sequence data structure 204 for generation of the target metric 336 (e.g., for performing on-target mismatch analysis), and can discard unmerged read pairs. For example, merging overlapping read pairs can result in a lower sequencing error rate by addressing portions of the sequence that are most likely to have sequencing errors (e.g., from the 3' end of the sequence).

[0108] The metric generator 300 can align the sequence data structure 204 to a reference sequence (e.g., in response to performing a merge). The reference sequence can be a predetermined sequence that represents a sequence intended to be synthesized (e.g., a target sequence). The metric generator 300 can align the sequence data structure 204 to multiple reference sequences simultaneously (which can improve specificity).

[0109] The metric generator 300 can determine at least one target metric 336 by comparing the sequence data structure 204 to the reference sequence (e.g., in response to aligning the sequence data structure with the reference sequence). For example, the metric generator 300 can compare each base (e.g., A, T, C, G, N, insertion, deletion) of each nucleotide data element 208 of the sequence data structure 204 to a corresponding base (e.g., allele) of the reference sequence and determine a difference count based on the comparison to determine at least one target metric 336 as a mismatch rate. The metric generator 300 can determine a consensus sequence from the multiple sequences of the sequence data structure 204 by determining the most frequently observed base at each position, and determine at least one target metric 336 to include a percentage of positions that differ between the consensus sequence and the reference sequence.

[0110] The metric generator 300 can determine at least one target metric 336 to include a mismatch percentage for a particular position (e.g., for each position) as a percentage of the different alleles at a particular position relative to the sum of the different alleles and the reference allele. The metric generator 300 can determine the overall target mismatch rate of the at least one target metric 336 to include at least one of an average mismatch rate (e.g., an average of the mismatch percentages for a particular position), an average mismatch rate that discounts or does not include mismatch percentages with UMI coverage depths below a threshold depth (e.g., 200), or a weighted mismatch rate by dividing the number of errors for the entire reference sequence by the total number of bases in the reference sequence (e.g., weighting the contribution of the error calculation by the depth of each position).

[0111] The metric generator 300 can determine at least one target metric 336 to include a target fraction metric 340. The target fraction metric 340 can indicate the relative fraction of the product 106 that is composed of a particular transcript (e.g., versus another transcript, contaminating DNA, or various combinations thereof). The metric generator 300 can determine the target fraction metric 340 by counting the number of particular UMIs (e.g., unique, distinct, or discrete UMIs) associated with a particular transcript and comparing the count of particular UMIs identified from the sequence data structure 204 that are associated (e.g., mapped) with a particular transcript to the count of particular UMIs not associated with a particular transcript.

[0112] The metric generator 300 can determine at least one off-target metric 344 based on the sequence data structure 204, for example, to identify and characterize off-target sequences. An off-target sequence can be a non-aligned sequence that does not align with a reference sequence (e.g., with any intended target sequence), which can be due to factors such as contamination with foreign cells or nucleic acids, too many mismatches to the reference sequence, or representing a chimeric library preparation artifact. The metric generator 300 can identify the non-aligned sequence and compare the non-aligned sequence to one or more off-target reference sequences (which can be retrieved, for example, from various databases or other sequencing runs performed by the sequencer 200).

[0113] The metric generator 300 can include a report generator 348. The report generator 348 can generate an output providing an indication 312 of the metric score 308 to indicate, for example, the value of the metric 304 or an error condition resulting from an evaluation of the metric 304. The indication 312 can be used to determine an instruction, action, or other operation to modify the synthesis of the nucleic acid, such as modifying the synthesis of the mRNA for which the metric 304 is determined, to cause, for example, at least one of: re-preparation of the product 106 or re-sequencing of the product 106 (e.g., on a new flow cell).

[0114] FIG. 4 illustrates an example of a method 400 for generating nucleic acid sequencing metrics. Method 400 may be performed using various systems and devices described herein, including but not limited to sequence metric generating system 300. Method 400 or operations thereof may be performed following or in response to the generation of an output by a nucleic acid sequencing system (e.g., sequencer 200). Method 400 may be performed during operation of nucleic acid manufacturing system 100, sequencer 200, or various combinations thereof, and may provide an output of a metric that may be used by a control scheme for manufacturing or sequencing nucleic acids to reduce analytical, sequencing, or biological errors, such as differences between sequence data representing a nucleic acid and a target sequence of the nucleic acid. Method 400 may include the generation of various metrics described herein (e.g., metrics described with reference to FIG. 3 and metric generator 300).

[0115] At 405, a sequence data structure is received. The sequence data structure may be received in response to a request for the sequence data structure from a sequencer or a database that stores or maintains the sequence data structure. The sequence data structure may be received in one or more batches or streams of data, such as from different sequencers or databases.

[0116] At least one first metric is evaluated at 410. The first metric can be a metric related to the sequencing process used to determine the sequence data structure. For example, the first metric can be determined based on data from a sequencing component such as a flow cell, such as flow cell expiration, rehybridization, or completion status. The first metric can be a flow cell yield metric. The first metric can be a base quality metric. The first metric can be evaluated to determine whether the data in the sequence data structure is of sufficient quality to determine further metrics from the data in the sequence data structure.

[0117] At 415, it is determined whether at least one first metric satisfies a first condition. The first condition can be at least one of a quantitative (e.g., threshold) or qualitative (e.g., categorical) condition. For example, in response to the flow cell yield metric score and the base quality score each meeting or exceeding a respective threshold, the first condition can be determined to be met. In response to the first condition not being met, at least one of the errors can be output (which can indicate the metric that did not meet the first condition), or an alteration of the process (e.g., nucleic acid synthesis, nucleic acid sequencing) that resulted in the sequence data structure can be made.

[0118] At least one second metric is evaluated at 420. The at least one second metric can be a metric corresponding to one or more UMIs associated with the sequence data structure, e.g., UMIs ligated to one or more nucleic acids (or strands or portions thereof) represented by the sequence data structure. For example, the at least one second metric can be a UMI difference metric indicating a base difference between UMI portions of the nucleic acids represented by the sequence data structure.

[0119] At 425, it is determined whether at least one second metric satisfies a second condition. For example, in response to the UMI difference metric score being below a threshold, it can be determined that there is insufficient data to effectively evaluate further metrics related to the sequence data structure, and an error can be output or other action can be taken to modify the synthesis or sequencing of the nucleic acid.

[0120] At least one target metric may be evaluated at 430. For example, the target metric may be evaluated in response to the first and second conditions being met (e.g., these conditions may indicate that a sufficient amount and quality of sequencing data is available to effectively evaluate the target and off-target metrics). The target metric may indicate a match (or mismatch) rate between the sequence data structure and the target sequence. The target metric may be evaluated in response to at least one of trimming adapters or merging read pairs of the sequence data structure to reduce errors that may be present in the sequence data structure prior to determining the target metric (which may make the target metric and any actions taken based on the target metric more accurate).

[0121] At least one off-target metric is evaluated in 435. The off-target metric can identify off-target sequences. For example, the off-target metric can be evaluated by identifying sequences that are not aligned with reference sequences, comparing the non-aligned sequences with non-reference sequences, and detecting the match between the non-aligned sequences and the non-reference sequences. The count of the non-aligned sequences can be compared with the count of the sequences in the sequence data structure to determine the off-target metric (e.g., determine the proportion of off-target sequences).

[0122] D. Systems Comprising Microfluidic Processor Chips FIG. 5 illustrates an example of a system 500, at least some features of which may be used to implement the nuclear manufacturing system 100 described with reference to FIG. 1. The system 500 may include a housing 503 enclosing a seating mount 515 that may removably receive one or more processor chips 511 (e.g., microfluidic processing chips). The system 500 may include chip-receiving components configured to removably house the processor chip 511, which itself defines one or more microfluidic channels or fluid paths. Components of the system 500 (e.g., within the housing 503) that fluidically interact with the processor chip 511 may include fluid channels or paths that may not be microfluidic (e.g., such fluid channels or paths are larger than the microfluidic channels or fluid paths within the processor chip 511). The processor chip 511 may be provided and utilized as a single-use device, while the remainder of the system 500 may be reusable. The housing 503 may be in the form of a chamber, enclosure, etc., with an opening that may be closed (e.g., via a lid or door, etc.) to thereby seal the interior. The housing 503 may enclose a temperature regulator and / or may be configured to be enclosed within a temperature-regulated environment (e.g., a refrigeration unit, etc.). The housing 503 may form a sterility barrier. The housing 503 may form a humidified or humidity-controlled environment. The system 500 may be disposed within a cabinet (not shown). Such a cabinet may provide a temperature-regulated (e.g., refrigerated) environment. Such a cabinet may also provide air filtration and airflow management to facilitate reagents being kept at a desired temperature throughout the manufacturing process. In addition, such a cabinet may include UV lamps for sterilization of the processor chip 511 and other components of the system 500. Other suitable features may be incorporated into the cabinet housing the system 500.

[0123] The assembly formed by the housing 503 and the components of the system 500 within the housing 503 without the processor chip 511 can be considered to be an appliance. Although the controller 521 and the user interface 523 are shown in FIG. 5 as being outside the housing 503, the controller 521 and the user interface 523 can be provided in or on the housing 503 and thus form part of the appliance. As will be explained in more detail below, the appliance can removably receive the processor chip 511 via a seating mount 515. When the processor chip 511 is seated in the seating mount 515, the appliance and the processor chip 511 can cooperate together to form the system 500. When the processor chip 511 is removed from the seating mount 515, the remaining portion of the system 500 can be considered an appliance. The appliance, the system 500, and the processor chip 511 can each be considered an apparatus.

[0124] The seating mount 515 may be configured to secure the processor chip 511 using one or more pins or other components configured to hold the processor chip 511 in a fixed and predefined orientation. Thus, the seating mount 515 may facilitate the processor chip 511 being held in a proper position and orientation relative to other components of the system 500. The seating mount 515 may hold the processor chip 511 in a horizontal orientation such that the processor chip 511 is parallel to the ground.

[0125] The thermal controller 513 is located adjacent to the seating mount 515 and can modulate the temperature of any processor chip 511 mounted within the seating mount 515. The thermal controller 513 can include one or more heat sinks and / or thermoelectric components (e.g., Peltier elements, etc.) to control the temperature of all or a portion of any processor chip 511 mounted within the seating mount 515. Two or more thermal controllers 513 can be included, such as to separately regulate the temperature of different ones of one or more regions of the processor chip 511. The thermal controller 513 can include one or more thermal sensors (e.g., thermocouples), etc., that can be used for feedback control of the processor chip 511 and / or the thermal controller 513.

[0126] As shown in FIG. 5, the fluid interface assembly 509 couples the processor chip 511 to a pressure source 517, thereby providing one or more pathways for a fluid (e.g., gas) at positive or negative pressure to be communicated from the pressure source 517 to one or more interior regions of the processor chip 511, as described in more detail below. Although only one pressure source 517 is shown, the system 500 can include two or more pressure sources 517. Pressure can be generated by one or more sources other than the pressure source 517. For example, one or more vials or other fluid sources in the reagent storage frame 507 can be pressurized. Reactions and / or other processes performed on the processor chip 511 can generate additional fluid pressure. The fluid interface assembly 509 can also couple the processor chip 511 to the reagent storage frame 507, thereby providing one or more pathways for liquid reagents and the like to be communicated from the reagent storage frame 507 to one or more interior regions of the processor chip 511, as described in more detail below.

[0127] Pressurized fluid (e.g., gas) from at least one pressure source 517 can reach the fluid interface assembly 509 through the reagent storage frame 507, such that the reagent storage frame 507 includes one or more components interposed in a fluid path between the pressure source 517 and the fluid interface assembly 509. The one or more pressure sources 517 can be directly coupled to the fluid interface assembly 509 such that positive pressure fluid (e.g., positive pressure gas) or negative pressure fluid (e.g., suction gas or other negative pressure gas) bypasses the reagent storage frame 507 to reach the fluid interface assembly 509. Regardless of whether the reagent storage frame 507 is interposed in a fluid path between the pressure source 517 and the fluid interface assembly 509, the fluid interface assembly 509 can be removably coupled to the remainder of the system 500 such that at least a portion of the fluid interface assembly 509 can be removed for sterilization between uses. As described in more detail below, the pressure source 517 can selectively pressurize one or more chamber regions on the processor chip 511. The pressure source 517 can also selectively pressurize one or more vials or other fluid storage containers held by the reagent storage frame 507 .

[0128] The reagent storage frame 507 can accommodate a plurality of fluid sample holders, each of which can hold a fluid vial configured to hold a reagent (e.g., nucleotides, solvents, water, etc.) for delivery to the processor chip 511. One or more fluid vials or other storage containers in the reagent storage frame 507 can receive products from the interior of the processor chip 511. The second processor chip 511 can receive products from the interior of the first processor chip 511 such that one or more fluids are transferred from one processor chip 511 to another processor chip 511. The first processor chip 511 can perform a first dedicated function (e.g., synthesis, etc.) and the second processor chip 511 can perform a second dedicated function (e.g., encapsulation, etc.). The reagent storage frame 507 can include a plurality of pressure lines and / or manifolds configured to split one or more pressure sources 517 into a plurality of pressure lines that can be applied to the processor chip 511. Such pressure lines may be controlled independently or collectively (in subcombinations).

[0129] The fluid interface assembly 509 may include multiple fluid and / or pressure lines, where each such line may include a biased (e.g., spring-loaded) holder or tip that individually and independently drives each fluid and / or pressure line to the processor chip 511 when the processor chip 511 is held in the seating fixture 515. Any associated tubing (e.g., fluid and / or pressure lines) may be part of and / or connected to the fluid interface assembly 509. Each fluid line may include flexible tubing that connects between the reagent storage frame 507 and the processor chip 511 via connectors that couple vials to the tubing in a locking engagement (e.g., ferrules). The ends of the fluid / pressure lines may be configured to seal against the processor chip 511 (e.g., at corresponding sealing ports formed in the processor chip 511), as described below. The connections between the pressure source 517 and the processor chip 511, and the connections between the vials in the reagent storage frame 507 and the processor chip 511 all form sealed, closed pathways that are isolated when the processor chip 511 is seated in the seating fixture 515. Such sealed, closed pathways can provide protection against contamination when processing therapeutic polynucleotides.

[0130] The vials in the reagent storage frame 507 may be pressurized (e.g., pressures greater than 1 atm, such as 2 atm, 3 atm, 5 atm, or more). In some versions, the vials may be pressurized by a pressure source 517. Thus, negative or positive pressures may be applied. For example, the fluid vials may be pressurized to about 1 to about 20 psig (e.g., 5 psig, 10 psig, etc.). At the end of the process, a vacuum (e.g., about -7 psig or about 7 psia) may be applied to draw the fluid back into the vial (e.g., the vial serving as a reservoir). The fluid vials may be driven at a lower pressure than the pneumatic valves, as described below, which may prevent or reduce leakage. The pressure differential between the fluid valves and the pneumatic valves may be about 1 psi to about 25 psi (e.g., about 3 psi, about 5 psi, 7 psi, 10 psi, 12 psi, 15 psi, 20 psi, etc.).

[0131] The system 500 may further comprise a magnetic field applicator 519 configured to generate a magnetic field in the region of the processor chip 511. The magnetic field applicator 519 may comprise a movable head operable to move the magnetic field, thereby selectively isolating products attached to magnetic capture beads in vials or other storage containers in the reagent storage frame 507.

[0132] The system 500 may include one or more sensors 505. The sensors 505 may include one or more cameras and / or other types of optical sensors. The sensors 505 may sense one or more of a barcode, a fluid level in a fluid vial held in a reagent storage frame 507, fluid movement in a processor chip 511 mounted in a seating fixture 515, and / or other optically detectable conditions. The sensors 505 may be used to sense barcodes included on vials of the reagent storage frame 507 such that the sensors 505 may be used to identify vials in the reagent storage frame 507. A single sensor 505 may be positioned and configured to simultaneously view such barcodes on vials in the reagent storage frame 507, a fluid level in a vial in the reagent storage frame 507, fluid movement in a processor chip 511 mounted in a seating fixture 515, and / or other optically detectable conditions. Two or more sensors 505 may be used to view such conditions. Different sensors 505 may be positioned and configured to view corresponding optically detectable conditions separately, such that a sensor 505 may be dedicated to a particular corresponding optically detectable condition.

[0133] In versions where the sensor 505 includes at least one optical sensor, visual / optical markers can be used to estimate yield. For example, fluorescence can be used to detect process yield or residual material by tagging with a fluorophore. Dynamic light scattering (DLS) can be used to measure particle size distribution within a portion of the processor chip 511 (e.g., a mixed portion of the processor chip 511, etc.). The sensor 505 can provide measurements using one or two optical fibers to transmit light (e.g., laser light) to the processor chip 511 and detect light signals emerging from the processor chip 511. For example, to optically detect process yield or residual material, etc., the sensor 505 can be configured to detect visible light, fluorescence, ultraviolet (UV) absorbance signals, infrared (IR) absorbance signals, and / or any other suitable type of optical feedback.

[0134] In versions in which the sensor 505 comprises at least one optical sensor configured to capture video images, the sensor 505 can record at least some activity on the processor chip 511. For example, an entire run for synthesizing and / or processing a material (e.g., a therapeutic RNA) can be recorded by one or more video sensors 505, including a video sensor 505 that can visualize the processor chip 511 (e.g., from above). The process on the processor chip 511 can be tracked visually, and this video recording can be retained for later quality control and / or processing. Thus, the video recording of the process can be saved, stored, and / or transmitted for subsequent review and / or analysis. In addition, as described in more detail below, the video can be used as a real-time feedback input that can use at least visually observable conditions captured in the video to affect the process.

[0135] The system 500 of the present embodiment may be controlled by a controller 521. The controller 521 may include one or more processors, one or more memories, and various other suitable electrical components. One or more components of the controller 521 (e.g., one or more processors, etc.) may be embedded within the system 500 (e.g., housed within the housing 503). One or more components of the controller 521 (e.g., one or more processors, etc.) may be removably attached to or removably connected to other components of the system 500. Thus, at least a portion of the controller 521 may be removable. Furthermore, at least a portion of the controller 521 may be separate from the housing 503.

[0136] Control by the controller 521 may comprise actuating the pressure source 517 to apply pressure through the processor chip 511 to drive fluid movement. The controller 521 may be completely or partially outside the housing 503 or may be completely or partially inside the housing 503. The controller 521 may be configured to receive user input via a user interface 523 of the system 500 and provide output to a user via the user interface 523. The controller 521 may be fully automated to the point that no user input is required, such that the user interface 523 provides only output to the user. The user interface 523 may comprise a monitor, a touch screen, a keyboard, and / or any other suitable features. The controller 521 may coordinate processing including moving one or more fluids onto the processor chip 511, mixing one or more fluids onto the processor chip 511, adding one or more components to the processor chip 511, metering fluids within the processor chip 511, adjusting the temperature of the processor chip 511, applying a magnetic field (e.g., when using magnetic beads), and the like. The controller 521 can receive real-time feedback from the sensor 505 and execute control algorithms according to such feedback from the sensor 505. Such feedback from the sensor 505 can include, but need not be limited to, identification of reagents in vials in the reagent storage frame 507, detected fluid levels in vials in the reagent storage frame 507, detected movement of fluid in the processor chip 511, fluorescence of fluorophores in fluid in the processor chip 511, and the like. The controller 521 can include software, firmware, and / or hardware. The controller 521 can also communicate with a remote server, for example, to track operation of the device, to sort materials (e.g., components such as nucleotides, the processor chip 511, etc.), and / or to download protocols, etc.

[0137] It should be understood that all combinations of the foregoing concepts, and additional concepts discussed in more detail below, are contemplated as being part of the inventive subject matter disclosed herein and may be used in any combination to achieve the benefits described herein.

[0138] The processes described herein and their various modifications (hereinafter referred to as "processes") may be implemented, at least in part, in whole or in part, via a computer program product, i.e., a computer program tangibly embodied in one or more tangible physical hardware storage devices, which are data processing apparatuses, e.g., programmable processors, computers, or computers and / or machine-readable storage devices for execution by or for controlling the operation of multiple computers. The computer programs may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, such as as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The computer programs may be deployed to be executed on one computer, or on multiple computers at one site, or on multiple computers distributed across multiple sites and interconnected by a network.

[0139] Processors suitable for executing a computer program include, by way of example, both general purpose and special purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from a read-only or random access memory area, or both. Elements of a computer (including a server) comprise one or more processors for executing instructions and one or more storage devices for storing instructions and data. Typically, a computer also comprises one or more machine-readable storage media, such as mass storage devices for storing data, e.g., magnetic, magneto-optical, or optical disks, or is operatively coupled to receive data from or transfer data to, or both.

[0140] The computer program product is stored in tangible form on non-transitory computer readable media and non-transitory physical hardware storage devices suitable for embodying computer program instructions and data, including, by way of example, all forms of non-volatile storage including semiconductor storage devices, such as EPROM, EEPROM, and flash storage devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks, as well as volatile computer memory, such as RAM, including static RAM and dynamic RAM, and erasable memory, such as flash memory and other non-transitory devices.

[0141] The configurations and arrangements of the systems and methods shown in the various embodiments are merely exemplary. Although only a few embodiments have been described in detail in this disclosure, many modifications are possible (e.g., variations in size, dimensions, structure, shape and proportions of various elements, parameter values, mounting arrangements, use of materials, colors, orientations, etc.). For example, the positions of elements can be reversed or otherwise varied, and the nature or number of separate elements or positions can be modified or varied. Accordingly, all such modifications are intended to be included within the scope of this disclosure. The order or arrangement of any process or method steps can be varied or reordered. Other substitutions, modifications, changes, and omissions can be made in the design, operating conditions, and arrangements of the embodiments without departing from the scope of this disclosure.

[0142] As used herein, the terms "approximately," "about," "substantially," and similar terms are intended to include any given range or number + / - 10%. These terms include minor or insignificant modifications or variations of the subject matter described and claimed that are believed to be within the scope of the present disclosure as set forth in the appended claims.

[0143] The term "coupled" and its variations as used herein means joining two members directly or indirectly to each other. Such joining can be static (e.g., permanent or fixed) or movable (e.g., removable or releasable). Such joining can be achieved by two members being directly joined to each other, by two members being joined to each other using a separate intervening member and any additional intermediate members joined to each other, or by two members being joined to each other using an intervening member integrally formed as a single unit with one of the two members. When "coupled" or its variations are modified by an additional term (e.g., directly coupled), the general definition of "coupled" provided above is modified by the plain language meaning of the additional term (e.g., "directly coupled" means joining two members without the use of a separate intervening member), resulting in a definition narrower than the general definition of "coupled" provided above. Such joining can be mechanical, electrical, or fluid.

[0144] As used herein, the term "or," when used to connect a list of elements, is used in its inclusive sense (and not its exclusive sense), such that the term "or" means one, some, or all of the elements in the list. Conjunctive language, such as the phrase "at least one of X, Y, and Z," is understood to convey that an element can be any of X, Y, Z; X and Y; X and Z; Y and Z; or X, Y and Z (i.e., any combination of X, Y, and Z), unless specifically stated otherwise. Thus, such conjunctive language is generally not intended to imply that a particular embodiment requires that at least one of X, at least one of Y, and at least one of Z each be present, unless otherwise indicated.

[0145] References to the location of elements herein (e.g., "top," "bottom," "upper," "lower") are merely used to describe the orientation of the various elements in the drawings. It should be noted that the orientation of the various elements may vary according to other exemplary embodiments, and such variations are intended to be encompassed by the present disclosure.

[0146] The present disclosure contemplates methods, systems, and program products on any machine-readable medium for accomplishing various operations. The embodiments of the present disclosure may be implemented using existing computer processors, or by dedicated computer processors for suitable systems incorporated for this or another purpose, or by hardwired systems. Embodiments within the scope of the present disclosure include program products comprising machine-readable media for carrying or storing machine-executable instructions or data structures. Such machine-readable media may be any available medium that can be accessed by a general purpose or special purpose computer or other machine having a processor. By way of example, such machine-readable media may comprise RAM, ROM, EPROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of machine-executable instructions or data structures and that can be accessed by a general purpose or special purpose computer or other machine having a processor. Combinations of the above are also included within the scope of machine-readable media. Machine-executable instructions include, for example, instructions and data that cause a general purpose computer, a special purpose computer, or a special purpose processing machine to perform a certain function or group of functions.

[0147] Although the figures show a particular order of method steps, the order of steps may differ from that shown. Also, two or more steps may be performed simultaneously or with partial concurrence. Such variations depend on the software and hardware systems selected and the designer's choices. All such variations are within the scope of this disclosure. Similarly, software implementations may be accomplished using standard programming techniques with rule-based logic and other logic to accomplish the various connecting, processing, comparing, and deciding steps.

Claims

1. 1. A method for generating nucleic acid sequencing metrics, comprising: receiving, by one or more processors, a sequence data structure that is from parallel processing of the nucleic acids; applying, by the one or more processors, at least one first metric of a plurality of metrics to the array data structure to generate at least one first metric score; determining, by the one or more processors, that the at least one first metric score satisfies a first condition; In response to determining, by the one or more processors, that the at least one first metric score satisfies the first condition, applying at least one second metric of the plurality of metrics to the array data structure to generate at least one second metric score; outputting, by the one or more processors, an indication of the at least one first metric score and the at least one second metric score; A method comprising:

2. detecting, by the one or more processors, at least one unique molecular index in the sequence data structure; generating, by the one or more processors, the at least one second metric score using the at least one unique molecular index; The method of claim 1 , comprising:

3. 3. The method of claim 2, wherein the nucleic acid comprises mRNA or DNA, and generating the at least one second metric score comprises adjusting the at least one second metric score to reduce an error rate associated with the sequence data structure to perform quality control for the mRNA.

4. The method of claim 2 , wherein the at least one unique molecular index is adjacent to an i7 index of the sequence data structure.

5. identifying, by the one or more processors, at least one synthetic adaptor sequence in the sequence data structure; removing, by the one or more processors, the at least one synthetic adaptor sequence from the sequence data structure; merging, by the one or more processors, a plurality of unique molecule indexes of the sequence data structure from which the at least one synthetic adaptor sequence has been removed in response to removing the at least one synthetic adaptor sequence; generating, by the one or more processors, the at least one second metric score to include a mismatch rate in response to merging the plurality of unique molecular indexes; The method of claim 1 , comprising:

6. 2. The method of claim 1, wherein the at least one first metric comprises at least one metric of a flow cell used to generate the sequence data structure.

7. The method of claim 1 , comprising outputting, by the one or more processors, an indication of a failure condition in response to the at least one first metric score not satisfying the first condition.

8. 8. The method of claim 7, wherein outputting the instructions comprises generating, by the one or more processors, the instructions such that the sequence data structure includes instructions for modifying the synthesis of the target sequence detected.

9. 2. The method of claim 1, wherein the at least one second metric score comprises a target mismatch rate, the target mismatch rate corresponding to the amount of mismatch between the sequence data structure and a data structure of a target sequence for synthesis of the nucleic acid.

10. 10. The method of claim 1, further comprising using the instructions to evaluate one or more characteristics of an mRNA therapeutic corresponding to the nucleic acid.

11. A method for generating a sequence data structure comprising: receiving a sequence data structure from parallel processing of nucleic acids; applying at least one first metric of a plurality of metrics to the array data structure to generate at least one first metric score; determining that the at least one first metric score satisfies a first condition; responsive to the at least one first metric score satisfying the first condition, applying at least one second metric of the plurality of metrics to the array data structure to generate at least one second metric score; A system comprising one or more processors for outputting an indication of the at least one first metric score and the at least one second metric score.

12. the one or more processors: detecting at least one unique molecular index in said sequence data structure; 12. The system of claim 11, wherein the system is for generating the at least one second metric score using the at least one unique molecular index.

13. the one or more processors: identifying at least one synthetic adaptor sequence in said sequence data structure; removing the at least one synthetic adaptor sequence from the sequence data structure; merging a plurality of unique molecule indexes of the sequence data structure from which the at least one synthetic adaptor sequence is removed in response to removing the at least one synthetic adaptor sequence; 12. The system of claim 11, wherein in response to merging the plurality of unique molecular indexes, the system is configured to generate the at least one second metric score to include a mismatch rate of the sequence data structure relative to a target data structure for synthesis of the nucleic acid.

14. 12. The system of claim 11, wherein the one or more processors are for controlling operation of a processor chip to produce an mRNA therapeutic using the instructions.

15. 15. The system of claim 14, wherein the one or more processors are for controlling operation of the processor chip to use the instructions to reduce the difference between the mRNA and a target sequence of the mRNA.

16. 15. The system of claim 14, wherein the one or more processors are for performing a process of encapsulating at least one mRNA with at least one delivery vehicle composition to form the mRNA therapeutic.

17. The system of claim 11 , wherein the one or more processors are for generating a report that includes the indication.

18. The method of claim 1, wherein the nucleic acid is a first nucleic acid, the sequence data structure is from parallel processing of the first nucleic acid and a second nucleic acid, and the method includes determining, by the one or more processors, a ratio of the first nucleic acid to the second nucleic acid in the sequence data structure.

19. The method of claim 1, wherein the at least one first metric includes at least one of a flow cell metric or a base quality metric, and the at least one second metric includes a target metric of a match rate between the sequence data structure and a target sequence of the nucleic acid.

20. Aligning, by the one or more processors, the sequence data structure with a target sequence; determining, by the one or more processors, the instructions based on comparing the sequence data structure to the target sequence in response to aligning the sequence data structure with the target sequence; The method of claim 1 , comprising: