Systems and methods for sequence identity analysis

The method and system automate and enhance biological sample sequencing analysis by comparing reference and sample data to output consensus metrics, improving accuracy and reducing manual intervention.

WO2026122349A1PCT designated stage Publication Date: 2026-06-11LIFE TECHNOLOGIES CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LIFE TECHNOLOGIES CORP
Filing Date
2025-11-24
Publication Date
2026-06-11

AI Technical Summary

Technical Problem

Biological sample sequencing analysis is time-consuming and prone to errors due to the need for manual adjustment of analysis settings for each sequence, lacking flexibility and automation, leading to inaccurate results.

Method used

A method and system for automated biological sample sequence analysis that includes receiving reference and sample nucleic acid sequence data, comparing them, and outputting overall and region-level consensus metrics, allowing for customizable analysis settings and reducing the need for manual intervention.

Benefits of technology

Enhances automation and accuracy in biological sample sequencing by providing flexible, region-based analysis settings and metrics, reducing human error and time consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025056914_11062026_PF_FP_ABST
    Figure US2025056914_11062026_PF_FP_ABST
Patent Text Reader

Abstract

A method for biological sample analysis includes receiving reference nucleic acid sequence information and biological sample nucleic acid sequence data, comparing the biological sample nucleic acid sequence data to the reference nucleic acid sequence information, outputting overall consensus metrics of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information based on the comparison and outputting one or more region level consensus metrics of one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information.
Need to check novelty before this filing date? Find Prior Art

Description

PROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1SYSTEMS AND METHODS FOR SEQUENCE IDENTITY ANALYSISCROSS-REFERENCE

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 727,065, filed on December 2, 2024, which is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0002] This disclosure relates to biological sample analysis, and more specifically, to genetic sequence identification, including instruments, systems, and methods pertaining to the same.INTRODUCTION

[0003] Biological sample analysis techniques and systems can involve various steps, such as adjustment of various settings in instrumentation and analysis platforms used for biological sample sequence analysis. For example, gene sequencing instruments and analysis platforms utilized in genetic sequence identification involves adjustment of assembly and alignment settings. Review of output results to provide certain information also is performed, such as for example, review of alignment and consensus metrics in the context of genetic sequence analysis. Such steps in the overall workflow can be relatively time consuming and often require significant human intervention, which can be prone to error leading to inaccurate results.

[0004] Even in applications where all the settings for the analysis are adjusted prior to commencement of the analysis, which can minimize some of the timeconsuming aspects of biological sample sequencing analysis, these often involve having to set all the analysis settings for the whole process, such as specifying the settings for a whole consensus process each time. For a consensus result, often the analysis settings need to be setup for every reference (specific sample for analysis) file used. For example, the settings for assembling multiple related sequences into a single consensus sequence by using a reference sequence may need to be set each time.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01

[0005] Once the analysis settings are set, they will often apply to the whole sample analysis process (i.e., over the entire sequence), with relatively little flexibility to apply different settings for different regions of interest of a sequence.

[0006] Enhancing automation of overall workflows related to biological sample sequencing analysis is desirable. Further, allowing for more robust and comprehensive analysis results in biological sample sequencing analysis also is desirable.SUMMARY

[0007] Various embodiments of the present disclosure may solve one or more of the above-mentioned problems and / or may demonstrate one or more of the above-mentioned desirable features. Other features and / or advantages may become apparent from the description that follows.

[0008] Various embodiments contemplate a method including receiving, at a computer processor, reference nucleic acid sequence information; receiving, at the computer processor, biological sample nucleic acid sequence data; comparing, using the computer processor, the biological sample nucleic acid sequence data to the reference nucleic acid sequence information; outputting, using the computer processor, overall consensus metrics of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information based on the comparison; and outputting, using the computer processor, one or more region level consensus metrics of one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information.

[0009] Various implementations of the method may include one or more of the following features. The can further include receiving input of a file location for each of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data. The biological sample nucleic acid sequence data is generated by a remotely located instrument configured to perform a biological sample sequencing assay. The instrument can include a capillary electrophoresis analysis instrument. The reference nucleic acid sequencePROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 information is received from a laboratory information management (LIM) system. Receiving the reference nucleic acid sequence information may include receiving the reference nucleic acid sequence information from a remote database. Receiving the reference nucleic acid sequence information may include accessing a shared file system. The method can further include identifying, by the computer processor, the one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information. The one or more regions are stored at a data store communicatively coupled to the computer processor. The one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information are inputted by a user at the computer processor. The biological sample nucleic acid sequence data may include data collected from performing capillary electrophoresis on the biological sample. The one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information may include user-annotated regions of interest. The one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information can be formatted in a GenBank format standard. The method can further include determining the region level consensus metrics as a function of one or more region level analysis settings. The one or more region level analysis settings are based on minimum coverage. The one or more region level analysis settings are based on one or both of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information. The outputting quality metrics of the biological sample nucleic acid sequence data can be outputting as a function of the one or more region level consensus metrics. The method can further include comparing the one or more region level consensus metrics to threshold scores. The method can include determining a match between the biological sample nucleic acid sequence data and the reference nucleic acid sequence information as a function of the comparison.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01

[0010] Yet other embodiments contemplate a system including at least one processor; and at least one memory, where the at least one memory contains instructions configuring the at least one processor to: receive reference nucleic acid sequence information; receive biological sample nucleic acid sequence data; compare the biological sample nucleic acid sequence data to the reference nucleic acid sequence information; output overall consensus metrics of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information based on the comparison; and output one or more region level consensus metrics of one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information.

[0011] Implementations of the system may include one or more of the following aspects. The at least one processor can be further configured to compare the biological sample nucleic acid sequence data to the reference nucleic acid sequence information using a consensus module. The system may include one or more remote instruments. The at least one processor can be further configured to receive the biological sample nucleic acid sequence data from the one or more remote instruments. The processor can be further configured to retrieve the reference nucleic acid sequence information and the biological sample nucleic acid sequence data from a remote database. The system can be further configured to retrieve the reference nucleic acid sequence information and the biological sample nucleic acid sequence data from a shared file system. The reference nucleic acid sequence information may include one or more reference regions of interest. The biological sample nucleic acid sequence data may include one or more biological sample regions of interest. 28. The system can be further configured to output the one or more region level consensus metrics as a function of region level analysis settings. The region level analysis settings can be a function of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data.

[0012] Various embodiments further contemplate a non-transitory computer- readable medium storing instructions, which when executed by at least onePROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 processor, configure at least one processor to: receive reference nucleic acid sequence information; receive biological sample nucleic acid sequence data; compare the biological sample nucleic acid sequence data to the reference nucleic acid sequence information; output overall consensus metrics of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information based on the comparison; and output one or more region level consensus metrics of one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information. The instructions can further configure the processor to carry out any of the actions of the above method.

[0013] Various additional embodiments contemplate a method for analysis of biological sample sequence data that includes: receiving input at a computer processor, the input may include: a file location to retrieve reference nucleic acid sequence information, a file location to retrieve biological sample nucleic acid sequence data, analysis settings for comparison of sequences corresponding to each of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and trigger settings for performing analysis of the biological sample nucleic acid sequence data; and on a condition of meeting a trigger specified in the trigger settings, using the computer processor to: retrieve from the respective file location the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and output metrics corresponding to a consensus of the biological sample nucleic acid sequence data with the reference nucleic acid sequence information.

[0014] Implementations of the method can include one or more of the following. The analysis settings may include overall sequence consensus and / or region level sequence consensus. The method may include outputting overall consensus metrics. The method can include outputting one or more region level consensus metrics. The trigger settings may include a scheduler. The scheduler can be time-based. The trigger setting is generated based on a new sample file being uploaded. The method can further include scanning, by the computerPROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 processor, remotely located file storage location for the new sample file. The method can include selecting, by the computer processor, the reference nucleic acid sequence information from a plurality of reference nucleic acid sequence information based on the biological sample nucleic acid sequence data. The file location to retrieve nucleic acid sequence information is at one or more remotely located databases. The file location to retrieve biological sample nucleic acid sequence data is at one or more remotely located databases. The method can include exporting the output metrics to the file location of the biological sample nucleic acid sequence data.

[0015] Various embodiments contemplate a system including at least one processor; and at least one memory, where the at least one memory contains instructions configuring the at least one processor to: receive input may include: a file location to retrieve reference nucleic acid sequence information, a file location to retrieve biological sample nucleic acid sequence data, analysis settings for comparison of sequences corresponding to each of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and trigger settings for performing analysis of the biological sample sequence data; and on a condition of meeting a trigger specified in the trigger settings: retrieve from the respective file location the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and output metrics corresponding to a consensus of the biological sample nucleic acid sequence data with the reference nucleic acid sequence information.

[0016] Implementations of the system may include one or more of the following. The trigger settings may include a scheduler. The system can be configured to retrieve at least one of the reference nucleic acid sequence information, biological sample nucleic acid sequence data, and the analysis settings from a remote data store.

[0017] Yet other embodiments contemplate a non-transitory computer-readable medium storing instructions, when executed by at least one processor, configuring the at least one processor to: receive input may include: a filePROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 location to retrieve reference nucleic acid sequence information, a file location to retrieve biological sample nucleic acid sequence data, analysis settings for comparison of sequences corresponding to each of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and trigger settings for performing analysis of the biological sample sequence data; and on a condition of meeting a trigger specified in the trigger settings: retrieve from the respective file location the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and output metrics corresponding to a consensus of the biological sample nucleic acid sequence data with the reference nucleic acid sequence information. The instructions can further configure the processor to carry out any of the actions of the above method.BRIEF DESCRIPTION OF DRAWINGS

[0018] Various objects, features, characteristics, and / or advantages of embodiments of the present disclosure will become apparent and more readily appreciated from the following description of the embodiments, taken in conjunction with the accompanying drawings and the appended claims, all of which form a part of this application. In the Drawings, like reference numerals may be utilized to designate corresponding or similar parts in the various Figures, and the various elements depicted are not necessarily drawn to scale.

[0019] FIG.1 is a block diagram of an embodiment of a system for biological sample sequence analysis.

[0020] FIG. 2 is a flow diagram of an embodiment of a method for biological sample sequence analysis.

[0021] FIG. 3 is a block diagram of an embodiment of a computing system used for biological sample analysis.

[0022] FIG. 4 depicts an exemplary representation of consensus results output in accordance with aspects of the disclosure.

[0023] FIG. 5 depicts an exemplary representation of an assembly information interface.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01

[0024] FIG. 6 depicts an exemplary representation of an interface showing a specimen sequence layout against a segment of a reference sequence.

[0025] FIG. 7 is a block diagram of an embodiment of an automated biological sample sequence analysis system.

[0026] FIG. 8 is an illustrative flow diagram of an embodiment of an automated biological sample sequence analysis method.

[0027] FIG. 9 is a block diagram of another embodiment of a computing system used for automated biological sample sequence analysis.

[0028] FIG. 10 depicts an exemplary representation of an interface for a project input in accordance with aspects of the disclosure.

[0029] FIG. 11 depicts an exemplary representation of a specimen group autocreation section of an interface for generating a template for specifying analysis settings in accordance with aspects of the disclosure.

[0030] FIG. 12 depicts an exemplary representation of an analysis settings section of an interface for generating a template in accordance with aspects of the disclosure.

[0031] FIG. 13 depicts an exemplary representation of an interface for inputting trigger settings comprising a schedule for automated running of biological sample sequencing analysis.DETAILED DESCRIPTION

[0032] In accordance with aspects of the disclosure, systems and techniques are provided to enhance automation of biological sample genetic sequence analysis, including, for example, to automate and render more efficient project creation, analysis settings, and job scheduling in instruments used for biological sample sequence analysis.

[0033] In accordance with various additional aspects, the present disclosure contemplates robust integration between instrumentation used for biological sample sequence analysis and lab information management systems (LIMS).

[0034] Moreover, aspects of the disclosure contemplate providing an ability to customize analysis settings on a region-based level of an overall sequence, asPROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 well as to output region-based metrics of an overall sequence so as to be able to improve accuracy in determining the identity of the sequence produced from the biological sample sequence analysis.

[0035] Implementations in accordance with this disclosure provide for outputting consensus results, which can include overall consensus metrics based on comparing reference information to sequence sample data and / or region-based consensus metrics based on comparison at a region level of the sequence data and reference.

[0036] Various implementations in accordance with the disclosure further contemplate the ability for analysis settings of the instrument and analysis software to be set manually by a user, or automatically by the system, such as based on characteristics of the sample being analyzed and / or specified regions of a sequence being analyzed.

[0037] Aspects of the present disclosure further improve automation by reducing or eliminating the need for a user to specify settings for every analysis run of an instrument. To help address this issue, implementations disclosed herein allow a user to set up a project that specifies saved data file locations for where to retrieve genetic sequence reference information, the sample data to be analyzed, and the analysis settings applicable to those files, while also allowing for setting trigger settings as to when analysis is performed and / or output metrics generated. Once a project is created, files may only need to be changed at their original location as the project, with its timing and analysis / output settings, is agnostic to the specific contents of those data files, which can allow for multiple runs being performed without human intervention.

[0038] Depending on the type of trigger set, consensus metrics may be generated at specific dates and times, at set intervals, when a file is changed at the original location, or based on many other triggers a user may set. In other words, a trigger can be time-based or event-based. In some aspects of present disclosure, assembly and alignment results may also be generated based on trigger sets.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1

[0039] These and other implementations and embodiments in accordance with the disclosure will be described in greater detail below in relation to FIGs. 1 -13.

[0040] Referring now to FIG. 1 , a block diagram of an embodiment of a system100 for performing biological sample sequencing analysis which includes both overall consensus and region-level analysis capabilities is illustrated. In embodiments, system 100 includes a computing device 101. Computing device101 includes any system capable of processing, storing, and manipulating data. The system 100, in embodiments, includes at least one processor 102. The at least one processor 102 may be communicatively connected to, or a component of, computing device 101. As used herein, a “processor” is a component configured for executing instructions, performing calculations and managing tasks. The at least one processor 102 includes a microprocessor, a microcontroller, one or more central processing unit (CPU) cores, an applicationspecific integrated circuit (ASIC), one or more graphical processing unit (GPU) cores, a field programmable gate array (FPGA), and / or any other hardware device suitable for retrieval and execution of instructions. In some embodiments, the at least one processor 102 includes electronic circuitry for performing instructions described in this disclosure.

[0041] In embodiments, system 100 includes at least one memory 103. In embodiments, the at least one memory 103 is communicatively coupled to the at least one processor 102. In embodiments, the at least one memory 103 may be communicatively connected to, or a component of, computing device 101. As used in this disclosure, a “memory” is a data storage component configured to store instructions for a computing component, such as processor 102. In examples, without limitations, memory 103 may be configured for temporary storage of data, such as a random-access memory (RAM), or permanent data storage, such as Solid-State drives (SSD).

[0042] In embodiments, system 100 includes a consensus module 104. As used herein, “consensus module” is an executable program, or a combination of programs, configured to generate one or more consensus sequences. In anPROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 example, consensus module 104 may be a program configured to accept reference genetic sequences in multiple formats, such as, but not limited to, for example, FASTA and GenBank, perform sequence alignment using one or more alignment algorithms, such as, but not limited to, for example, a Smith-Waterman algorithm, and determine a consensus sequence for each aligned position between the reference and a biological sample for which a sequencing analysis has been conducted and data therefrom gathered.

[0043] In embodiments, system 100 is configured to receive reference nucleic acid sequence (NAS) information 110. As used herein, “reference NAS information” and variations thereof refers to a nucleic acid sequence that is used as a standard (reference) for comparison with sample sequences. In some embodiments, reference NAS information 110 includes reference regions 111. For example, a file, or any sort of data, that contains the reference NAS information 110 includes an identification or annotation of regions of interest (subsequences) of an overall sequence, such as a reference sequence that includes region identifiers. For example, the reference regions 111 may be metadata annotations to the reference NAS information 110.

[0044] In some embodiments, the computing device 101 is configured to identify the reference regions 111 or alternatively, receive the reference regions 111 as user input. For example, the reference regions 111 includes user-annotated regions of interest. In some examples, computing device 101 may be configured to identify the reference regions 111 based on known regions, such as, but not limited to, Exons, Proms, UTRs located at the 5’ (five prime) and 3’ (three prime) ends, and the like. In some embodiments, system 100 is configured to receive the reference NAS information 110 from a laboratory information management (LIM) system. In embodiments, the system 100 is configured to receive the reference NAS information and reference regions 111 in a standard format. For example, reference regions 111 can be provided in GenBank format and reference NAS information 110 may be included in a FASTA format file, a text file format, and / or a Genbank format.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01

[0045] In various embodiments, computing device 101 also is configured to receive biological sample NAS data 120. As used herein, “biological sample NAS data” and variations thereof refers to the raw sequencing reads of biological samples generated from a sequencing analysis assay and information derived from those raw reads. In embodiments, biological sample NAS data 120 includes sample regions 121 corresponding to regions of interest (subsequences) of the overall sequence of the biological sample that is the target sequence of interest.

[0046] In some embodiments, computing device 101 is configured to identify the sample regions 121 and / or to receive the sample regions 121 as user input. For example, the sample regions 121 includes user-annotated regions of interest. In some embodiments, computing device 101 is configured to receive the sample NAS data 120 and / or sample regions 121 in a standard format. For example, biological sample NAS data 120 may be included in a AB1 file format and sample regions 121 may be included in a Genbank file format.

[0047] In some embodiments, system 100 can include one or more remote data stores 140 and the reference NAS information 110 and / or biological sample NAS data 120 can be retrieved from the one or more remote data stores 140 (with only one being illustrated in FIG. 1 for simplicity). Examples of remote data store 140 includes file systems, object storage, remote databases, version control systems, and the like. In some cases, system 100 is configured to receive an input of a file location for each of the reference NAS information 110 and / or the biological sample NAS data 120. In an example, the input includes file locations for files stored in the remote data store 140. In some embodiments, system 100 is configured to access reference NAS information 110 and / or biological sample NAS data 120 through a shared file system.

[0048] In embodiments, the biological sample NAS data 120 may be generated by a remote instrument 141. In an embodiment, system 100 is communicatively coupled to the remote instrument 141 configured to perform a sequencing assay and generate sequencing data from a biological sample sequencing assay. In embodiments, system 100 receives biological sample NAS data 120 from thePROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1 instrument 141. In some embodiments, the remote data store 140 is communicatively connected to the instrument 141. For example, the remote instrument 141 may generate the biological sample NAS data 120 and the transmit it to the remote data store 140. In this example, system 100 then retrieves the biological sample NAS data 120 from the remote data store 140. In some examples, system 100 receives the biological sample NAS data 120 directly from the remote instrument 141 .

[0049] In embodiments, the remote instrument 141 includes a capillary electrophoresis analysis instrument, a qPCR analysis instrument, or other type of instrument familiar to those having ordinary skill in the art for performing a biological sample sequencing assay. In an example, biological sample NAS data may be generated by performing capillary electrophoresis and detection of amplified nucleic acid fragments from a biological sample. Examples of remote instrument 141 may include SeqStudio™ Genetic Analyzer, SeqStudio™ Flex Genetic Analyzer, Applied Biosystems® 3500 series Genetic Analyzer and Applied Biosystems® 3730 series Genetic Analyzer, all commercialized by Thermo Fisher Scientific Inc. headquartered in Waltham, Massachusetts, USA.

[0050] In some embodiments, computing device 101 and / or instrument 141 may be configured to generate the biological sample NAS data 120 from raw sequencing data. In embodiments, computing device 101 and / or instrument 141 may be configured to perform base calling on the raw sequencing data. In embodiments, computing device 101 and / or instrument 141 may generate a base call output as a function of performing the base calling process.

[0051] In various embodiments, computing device 101 and / or instrument 141 may be configured to perform quality value assignment on the base call output. In embodiments, computing device 101 and / or instrument 141 may be further configured to trim low-quality bases using the quality value assignment.

[0052] In some embodiments, computing device 101 and / or instrument 141 may be further configured to perform mixed base identification on the base call output.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01

[0053] In some embodiments, computing device 101 may be configured to identify poor-quality samples and remove the poor-quality samples from further analysis. Identifying sample quality may be based on read quality assessments, contaminant identification (such as contaminants introduced during sample preparation), duplicate reads, base composition assessment, and the like. For example, computing device 101 may remove a poor-quality biological sample NAS data 120 from further analysis. A person with ordinary skill in the art would readily recognize the many methods and tools that can be used for identifying poor-quality samples.

[0054] In various embodiments, the computing device 101 is configured to compare the biological sample NAS data 120 to the reference NAS information 110. For example, the consensus module 104 is used to compare the biological sample NAS data 120 to the reference NAS information 110. The consensus module 104 may be configured to compare the reference regions 111 to the sample regions 121. For example, computing device 101 may be configured to compare the remaining biological sample NAS data 120 (i.e. after removal of poor-quality samples) to reference NAS information 110. The data used for the comparison can be the data generated as a result of the base call output, quality value assignment, mixed base identification on the base call output and / or data after trimming of low-quality bases using the quality value assignment.

[0055] In some embodiments, the computing device 101 is configured to output overall consensus metrics 130 based on the comparison of the biological sample NAS data 120 to the reference NAS information 110. In embodiments, computing device 101 may be configured to output assembly results based on the comparison of the biological sample NAS data 120 to the reference NAS information 110. In embodiments, computing device 101 may be configured to output alignment results based on the comparison of the biological sample NAS data 120 to the reference NAS information 110. In embodiments, the assembly and / or alignment results can be output as part of the overall consensus metricsPROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1130, and the overall consensus metrics can further include the overall consensus sequence.

[0056] In some embodiments, computing device 101 is configured to output region-level coverage metrics 131 based on the comparison of the biological sample NAS data 120 to the reference NAS information 110. In an embodiment, the region-level coverage metrics 131 are based on the comparison of reference regions 111 to sample regions 121. The overall consensus metrics 130 and / or the region level consensus metrics 131 can include quality metrics.

[0057] In embodiments, computing device 101 may be configured to identify mismatches by aligning biological sample NAS data 120 and reference NAS information 110 and comparing overall consensus sequences output as overall consensus metrics 130 to reference NAS information 110 sequences. In some embodiments, computing device 101 may be further configured to generate mismatch metrics 132 based on the comparison. Mismatch metrics 132 can include, for example, mismatch frequency, mismatch distribution, quality scores at mismatched positions, and the like.

[0058] In embodiments, the computing device 101 is configured to compare one or more region level consensus metrics 131 to threshold scores. In some embodiments, the computing device 101 is configured to determine a match between the biological sample NAS data 120 and the reference NAS information 110. For example, a match between the biological sample NAS data 120 and the reference NAS information 110 may be determined based on the comparison made by the consensus module 104 and various threshold settings.

[0059] The overall consensus metrics 130 and region level consensus metrics 131 can be included in the same output. Output of overall consensus metrics 130 and region level consensus metrics 131 can be in a plurality of formats, such as, but not limited to, for example, HTML (hypertext markup language), CSV (comma separated values), GUI (graphical user interface) visualizations, and a variety of other formats those having ordinary skill in the art would be familiar with. In some embodiments, the output format can be selected from a plurality of differentPROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 formats, such as formats used by a sample sequencing analysis software. For example, overall consensus metrics 130 and region level consensus metrics 131 outputs may be, but is not limited to, a Binary Alignment / Map (BAM) format.

[0060] In embodiments, the consensus module 104 is configured to determine sample level overall consensus 130 and / or region level consensus metrics 131 based on one or more analysis settings 105. Analysis settings are specifications for a sample overall or for each of the regions of interest for analysis. In embodiments, analysis settings 105 can include preset minimum coverage to be used in the comparison of the reference regions 111 and the sample regions 121. For example, analysis settings 105 includes specifications for regions of an alignment that have highest scores. In embodiments, analysis settings 105 can be based on the biological sample analyzed to generate the reference NAS data 120 and / or the reference associated with the reference NAS information 110, and / or may be based on the particular sample and reference region of interest. Examples of analysis settings 105 are discussed further in reference to FIG. 4.

[0061] Referring to FIG. 4, exemplary consensus results output 400 is illustrated and includes overall consensus metrics and region level consensus metrics (e.g., metrics outputs 130 and 131 ), and analysis settings 105. In this example, analysis settings 105 are set at the alignment and assembly section 401 . Also in this example is an alignment section 402 that includes overall consensus metrics 130 and region level consensus metrics output 131. In this example, each specimen (i.e. each biological sample) includes consensus metrics for one or more regions. In some examples, the alignment section 402 includes, for each specimen, only region consensus metrics 131 for regions with high alignment scores based on analysis settings 105 that includes specifications for only outputting consensus metrics for regions of interest with an alignment score above a set threshold.

[0062] As used herein, an “alignment score” is a numerical value used to quantify the quality and accuracy of the match between individual sample sequences and the aligned reference sequences. For example, trace files with a score below aPROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1 set threshold may be removed from further analysis. An example of filtering based on alignment scores is shown in reference to FIG. 12 described further below.

[0063] With reference now to FIG. 2, a method 200 for analysis of biological sample sequence data is presented. At action 205, the method 200 includes, in embodiments, receiving reference NAS information, such as, for example, receiving NAS information 110 at a computing device 101. Receiving NAS information can include receiving a file location of reference NAS information and thereafter retrieving the reference NAS information from the file location. In embodiments, method 200 includes retrieving the reference NAS information from a remote data store, such as, for example, remote data store 140. In some embodiments, method 200 includes retrieving the reference NAS information from a shared file system. In embodiments, method 200 includes receiving the reference NAS information from a laboratory information management (LIM) system.

[0064] In embodiments, method 200 further includes, at action 210, receiving biological sample NAS data, such as receiving sample NAS data 120 at a computing device 101. In some embodiments, method 200 includes receiving a file location of biological sample NAS data and thereafter retrieving the biological sample NAS data from the file location. In some embodiments, method 200 includes retrieving biological sample NAS data from a remote data store, such as a remote data store 140. In some embodiments, method 200 includes retrieving biological samples NAS data from a shared file system. In some embodiments, method 200 includes receiving biological samples NAS data from one or more instruments configured to perform a biological sample sequencing assay, such as one or more instruments 141. In some embodiments, the one or more instruments are remotely located, such as being remotely located from computing device 101.

[0065] In embodiments, the biological sample NAS sequence data may be derived from raw sequencing data generated from a biological sequencing assayPROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 performed at a biological sample sequencing instrument and method 200 can include generating the sample NAS data. In some embodiments, generating sample NAS data may include performing base calling on the raw sequencing data and outputting base calls based on the base calling process. Further, various embodiments can perform quality value assignment on the base call output and trimming low-quality bases using the quality value assignment to generated the sample NAS data. In embodiments, mixed base identification of the base call output may be utilized to generate the sample NAS data.

[0066] In an embodiment, method 200 may include identifying poor-quality samples and removing the poor-quality samples from further analysis when generating the sample NAS data.

[0067] The method 200, in embodiments, includes, at action 215, comparing the biological sample NAS data to the reference NAS information. In some embodiments, method 200 includes comparing the one or more region level consensus metrics to threshold scores. In some embodiments, method 200 includes determining a match between the biological NAS data and the reference NAS information as a function of the comparison of the region level consensus metrics.

[0068] In embodiments, method 200, at action 220, includes outputting overall consensus metrics of the biological sample NAS data and the reference NAS information. For example, the overall consensus metrics can be output as overall consensus metrics 130 from the computing device 101. In some embodiments, the method 200 includes outputting overall consensus metrics comprising sequence alignments in a Binary Alignment / Mapping (BAM) format. In some embodiments, the method 200 includes storing the overall consensus metrics, such as for example, in a shared file system. Shared file system may be accessed by multiple components of, or connected to, a system. For example, shared file system may be accessed by computing device 101 and other computing devices in the same network. In some embodiments, the shared file system may be locally hosted. For example, shared file system may be hosted byPROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 a computing system connected to computing device 101 , such as by a Network- Attached Storage device. In some embodiments, shared file system may be remotely hosted, such as by a remote server. In these examples, method 200 includes outputting overall consensus metrics of the biological sample NAS data and the reference NAS information to the shared file system. In embodiments, the method 200 includes storing the overall consensus metrics in a local database, for example, at or communicatively coupled with a computing device such as in the memory 103 of computing device 101. In some embodiments, the method 200 includes exporting the overall consensus metrics to a remote data store.

[0069] In some embodiments, the method 200 may include generating consensus metrics based on the quality value assignments. In some embodiments, method 200 may include modifying consensus metrics based on the quality value assignments.

[0070] At action 225, method 200, in embodiments, includes outputting one or more region-level consensus metrics based on a comparison of one or more sample regions (e.g., sample regions 121 of sample NAS data 120) and corresponding one or more reference regions (e.g., reference regions 111 of reference NAS information 110). In an embodiment, the one or more reference and / or sample regions include user-annotated regions of interest. As shown with reference to FIG. 4 for example, the output may be a Comma-Separated Value (CSV) output, or other format used for data storage and exchange, that includes an alignment section 402 with one or more region level consensus metrics 131 for each specimen. In some embodiments, method 200 further includes identifying one or more regions of interest based on the received biological sample NAS data and / or the reference NAS information. In embodiments, method 200 includes outputting one or more region level consensus metrics based on one or more region level analysis settings, such as the settings 105 described above.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01

[0071] In some embodiments, method 200, at action 230, may include identifying mismatches by aligning sequences of biological sample NAS data, such as biological sample NAS data 120, to reference sequences of reference NAS information, such as reference NAS information 110, and comparing overall consensus metrics 130, such as an overall consensus sequence. At action 235, method 200 may include outputting mismatch metrics for the mismatches, such as mismatch metrics 132.

[0072] In some embodiments, the one or more region level analysis settings include a minimum coverage. In embodiments, specimen level analysis settings include, but are not limited to, for example, gap penalty, extension penalty, alignment score and minimum coverage, which are discussed in more detail in reference to FIG. 12. Specimens as used herein are portions of a sample, and can include the entirety of the sample or multiple portions of it, the amount of which may depend on the particular analysis assays being performed and instrument being used to perform the assays. It should be understood that these are only some examples and that method 200 includes sample (specimen) and region level analysis settings for other parameters not described herein but that those of ordinary skill in the art would be familiar with. The overall and / or regionlevel consensus metrics can be output based on the analysis settings. For example, the consensus output includes coverage of a sample sequence at each reference base position where for each region of interest, computing device 101 calculates minimum coverages and compares those minimum coverages to minimum coverage requirements specified in the analysis settings 105 and determines whether or not a threshold is met. The consensus metrics output can flag a user regarding the same, such as by highlighting passing and / or failing results in differing colors. Reference is made to FIG. 4 which will be discussed further below.

[0073] Now referring to FIG. 3, an example computing device 300 is presented. In embodiments, computing device 300 may be same, include, or be a part of computing device 101. In this example, computing device 300 includes at leastPROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1 one processor 391 communicably coupled to a computer-readable storage medium 392. The at least one processor 391 includes a microprocessor, a microcontroller, one or more central processing unit (CPU) cores, an applicationspecific integrated circuit (ASIC), one or more graphical processing unit (GPU) cores, a field programmable gate array (FPGA), and / or any other hardware device suitable for retrieval and execution of instructions from computer-readable storage medium 392. In instances, the at least one processor 391 includes electronic circuitry for performing instructions described in this disclosure.

[0074] In embodiments, computer-readable storage medium 392 may be any medium suitable for storing executable instructions. In examples, without limitation, computer-readable storage medium 392 includes RAM, ROM, EEPROM, HHD, SSD, optical disc, and the like. Computer-readable medium storage 392 may be disposed within computing device 300. In embodiments, computer-readable storage medium 392 may external, and communicably connected, to computing device 300. The instruction stored on computer- readable storage medium may be used to implement method described in reference to FIG. 2.

[0075] Continuing to refer to FIG. 3, in this example computer-readable storage medium 392 is encoded with set of instructions 393-397. In an embodiment, executable instructions included in each block may be included in different blocks and blocks not shown. In some embodiments, computer-readable storage medium 392 may be encoded with sets of instructions not shown in the blocks.

[0076] In embodiments, instruction 393, when executed by at least one processor 391 , configures the at least one processor 391 to receive reference NAS information (e.g., reference NAS information 110).

[0077] In embodiments, when executed by at least one processor 391 , instruction 394 configures the at least one processor 391 to receive biological sample NAS data (e.g., biological sample NAS data 120). In some embodiments, computer- readable storage medium 392 may include instructions configuring the at least one processor 391 to perform base calling on raw sequence data, quality valuePROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 assignment, and / or mixed based identification. In embodiments, computer- readable storage medium 392 may further include instructions configuring the at least one processor 391 to trim low-quality samples and modify the biological sample NAS data (i.e. remove the low-quality samples from the analysis process).

[0078] In some embodiments, instruction 395, when executed by at least one processor 391 , configures the at least one processor 391 to compare the biological sample NAS data to the reference NAS information.

[0079] In embodiments, instruction 396 further configures the at least one processor 391 to output overall consensus results (e.g., overall consensus 130) of the biological sample NAS data and the reference NAS information based on the comparison. In some embodiments, computer-readable storage medium 392 may include instructions configuring the at least one processor 391 to output overall consensus metrics based on a comparison with the modified biological sample NAS data. In an embodiment, computer-readable storage medium 392 may include instructions configuring the at least one processor 391 to generate, modify, or confirm overall consensus metrics, included in overall consensus metrics, based on the base call quality value assignment.

[0080] In embodiments, instruction 397 further configures the at least one processor 391 to output one or more region level consensus metrics (e.g., region level metrics output 131 ) of one or more regions of each of the biological sample NAS 120 and the reference NAS 110.

[0081] In an embodiment, computer-readable storage medium 392 may include instructions configuring the at least one processor 391 to identify mismatches by aligning sequences of biological sample NAS data to reference NAS sequences from the reference NAS information, and comparing the overall consensus metrics including an overall consensus sample sequence to the reference sequence. In embodiments, computer-readable storage medium 392 may include instructions further configuring the at least one processor 391 to output as part of the overall consensus metrics mismatch metrics for the identified mismatches.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1

[0082] With reference again to FIG. 4, results output 400 also includes the summary section 401 which is used to show overall consensus metrics 130 and other overall results for all the specimens of the sample. As used herein, each specimen refers to portions of a biological sample, which can include plasmid DNA, cDNA, and the like, that are sequenced and later assembled to produce a consensus sequence. This section 401 can also include multiple types of overall consensus metrics, such as numbers of reverse and forward reads of a nucleic acid sequence, number of mismatches in the sample, and the like. As also mentioned above, the results output 400 also includes an alignment section 402 that includes the region level consensus metrics and other region-based metrics. In this example, the alignment section 402 includes region-based results such as specimen scores for each region (subsequence) of a specimen, percentage match, average coverage, total coverage, forward minimum coverage, reverse minimum coverage, and the like. It should be noted that this example is provided for ease of description only. As such, a person of ordinary skill in the art would appreciate, upon reading this disclosure, that other output types, and the many other types of sections that could be included in the output, that are not described herein. Moreover, various data can be presented as part of the consensus output depending on the type of overall sequence and / or region for which the sequencing analysis is being performed.

[0083] Referring to FIG. 5, an example interface 500 for visualizing assembly information as part of the overall consensus metrics is presented. In this example, assembly information is shown for a first specimen of a sample, shown by label 501 , and an assembly information for a second specimen of a sample, shown by label 502. It should be noted that this interface is provided only as an example, as such more specimens than shown could be included in the interface. In examples, the interface may allow a user to sort the specimens for assembly visualization based on custom filtering, such as ascending order. In this example, a mismatch is shown at label 503.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01

[0084] Referring now to FIG. 6, an example output interface 600 for visualizing the layout of specimens is illustrated. In this example, the interface provides a schematic of the location and orientation of the specimen sequences (shown at 601 ) with respect to a segment of reference sequence (shown at 602). In this example, the forward trace direction for a sample is shown by the label “F” and the reverse trace direction is shown by the label “R.”

[0085] Referring now to FIG.7, a block diagram of another embodiment of a system 700 for performing biological sample sequencing analysis and outputting consensus metrics is illustrated. Among other aspects, the system 700 is configured to perform analysis and output consensus metrics based on a trigger setting, as will be described in more detail below. System 700 includes, or can be a part of, system 100. System 700 includes a computing device 701. Computing device 701 includes, or can be a part of, computing device 101. System 700 includes at least one processor 702 and at least one memory 703. In embodiments, system 700 includes a consensus module 704. Consensus module 704 may be the same as consensus module 104. In embodiments, consensus module 704 includes, or can be a part of, consensus module 104. In some embodiments, the at least one processor 702 and / or the at least one memory 703 may be communicatively connected to, or a component of, the computing device 701. Elements of FIG. 7 and elements of FIG. 1 which are labeled with reference numerals that have the same last two digits as one another, such as 140 and 740, correspond to one another. For ease of description, and to avoid duplication, aspects and functions of the components of system 100 are not repeated here, but should be understood to also apply to components of system 700 as well.

[0086] In embodiments, computing device 701 is configured to receive an input 750, such as, for example, a user input, an input from a remote device, data transmitted from a database, data accessed on a shared file system, and the like. In an embodiment, input 750 includes a reference file location 751 , a sample file location 752, analysis settings 753, and one or more trigger settings 754. InPROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 some examples, input 750 includes a file location reference for the analysis settings 753. In other examples, input 750 includes the data of analysis settings 753. For example, input 750 may be a file that includes the information of the analysis settings 753 and trigger setting 754, and paths for reference file location751 and sample file location 752. In embodiments, analysis settings include specimen and / or region level analysis settings. In embodiments, analysis settings 753 may include settings for assembly and alignment of sequences.

[0087] In embodiments, reference file location 751 is a file location to retrieve reference nucleic acid sequence (NAS) information 710. In embodiments, reference NAS information 710 may be the same as reference NAS information 110. In some embodiments, reference file location 751 be a system path for accessing the file, such as in a shared file system. In some embodiments, reference file location 751 may be a remote database location.

[0088] In embodiments, sample file location 752 is a file location to retrieve biological sample NAS data 720. In some embodiments, sample file location 752 may also be a system path for accessing the file, such as in a shared file system. In some embodiments, sample file location 752 may be a remote database location. In some embodiments, sample file location 752 may be a location of a file that includes biological sample nucleic acid sequence (NAS) data 720. In embodiments, biological sample NAS data 720 may be the same as biological sample NAS data 120. In embodiments, reference NAS information 710 and biological sample NAS data 720 may be located within the same file. In some embodiments, sample file location 752 may be a compressed file version of a file containing biological sample NAS data 720. For example, sample file location752 may be an AB1 file compressed using standard compression algorithms. In embodiments, reference file location 751 and sample file location 752 may be the same location (i.e. a location to the same file). An example of an AB1 file with both reference NAS information 710 and biological sample NAS data 720 is shown further below in reference to FIG. 11 . A person with ordinary skill in the art would appreciate that reference file location 751 and sample file location 752PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 includes locations in any medium of storage that can be accessed by the computing device 701 , computing device 101 and / or any other computing device not described herein.

[0089] In embodiments, input 750 includes analysis settings 753. In some embodiments, analysis settings 753 includes region level settings. In further embodiments, the region level settings may be the same as the analysis settings 105.

[0090] In embodiments, input 750 includes a trigger setting 754. In some embodiments, the trigger setting 754 may be a scheduler. As shown in the example scheduler 1100 in reference to FIG. 11 , trigger setting 754 includes specific dates and times to trigger a generating the consensus metrics 731. In some embodiments, the trigger setting 754 may be a trigger based on new or updated file. For example, computing device 701 may monitor the reference NAS information 710 and the biological sample NAS data 720 for changes, and the consensus module 704 may generate new consensus metrics 730, including region level consensus metrics 731 , when that change is detected. In some embodiments, the trigger setting 754 may be based on a number of changes above a set threshold. For example, the trigger setting 754 may be set at a scheduler to run at set dates / times or based on a number of updates to files at reference file location 751 and / or sample file location 752. In some embodiments, trigger setting 754 may also be based on changes to analysis settings 753. For example, if an auto-project is created based on a location for analysis settings 753, if the settings are changed at that location, a new run is triggered.

[0091] As discussed above, various implementations in accordance with the present disclosure provide the ability to enhance automation of biological sample genetic sequencing analysis and outputting of consensus metrics by allowing the setup of a project that specifies saved data file locations for where to retrieve genetic sequence reference information, the sample data to be analyzed, and the analysis settings applicable to those files. System 700 represents a system configured for such automation.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1

[0092] Referring to FIG. 10, an example of an interface 1000 to generate such a project (“job”) is illustrated. In this example, the project is a process that creates the input 750 of system 700. The interface 1000 for project generation includes sections for specifying the reference file location 751 , the biological sample file location 752, and a template section for specifying location of the analysis settings 753. An example of an interface for a template for analysis settings 753 is shown in reference to FIG. 12, which analysis settings can be set automatically by the system depending, for example, on the type of sample or reference provided, input by a user, or combinations thereof. In some examples, not shown, the analysis settings 753 may also be set directly at the interface 1000 for project creation. Still referring to FIG. 10, once the file locations are specified, along with other details such as project name, descriptions of the project, and the like, the “schedule” tab may be accessed to set the trigger settings 754, as shown in the example 1300 in reference to FIG. 13. FIGs. 10, 12, and 13 are described in more detail further below.

[0093] In embodiments, system 700 is configured to retrieve the reference NAS information 710 based on the reference file location 751. As an example, referring to FIG. 10, reference file location 751 may be selected under a reference file section 1051 , such as a drop-down menu. In some other examples, reference file location 751 may be specified as a file path, such as a through a text field.

[0094] In embodiments, system 700 is configured to retrieve the biological sample NAS data 720 based on the sample file location 752. As shown in example 1000 in reference to FIG. 10, the sample file location 752 may also be selected through a drop-down menu of a samples file 1052 section. In other examples, sample file location 752 may be specified as a file path, such as a user typing the file path to the biological sample NAS data 720.

[0095] In embodiments, system 700 is configured to output overall consensus metrics 730, including an overall consensus sequence based on comparison of the biological sample NAS data 720 with the reference NAS information 710. InPROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 some embodiments, system 700 may be configured to export consensus metrics 730 to a remote data store 740. Consensus metrics 730 can include overall sequence consensus, as noted above, and consensus quality metrics. In some embodiments, in addition to consensus metrics 730, region level consensus metrics 731 can also be output from consensus module 704. In some embodiments, consensus metrics 730 and region level consensus metrics 731 may be the same as, respectively, the overall consensus 130 and the region level consensus metrics 131.

[0096] In some embodiments, system 700 may be configured to select the reference NAS information 710 from multiple files containing reference information based on the biological sample NAS data 720. In some examples, the biological sample NAS data 720 includes information used by the system 700 to select the applicable reference NAS information 710. In some embodiments, system 700 may utilize machine learning algorithms to identify and select reference NAS information 710. For example, system 700 may use correlations of previously generated consensus metrics 730, 731 for biological sample NAS data 720 to identify patterns of reference NAS information 710 used in association with that biological sample.

[0097] Now referring to FIG. 8, another method 800 for analysis of biological sample genetic sequence data is presented. In embodiments, the method 800 includes at action 805 receiving an input (e.g., input 750) that includes a file location to retrieve reference nucleic acid sequence (NAS) information (e.g., NAS information 710), a file location to retrieve biological sample NAS data (e.g., biological sample NAS data 720), analysis settings (e.g., analysis settings 753) for comparison of sequences corresponding to each of the reference NAS data and the biological sample NAS data, and trigger settings (e.g., trigger settings 754) for performing analysis of the biological sample NAS data.

[0098] Monitoring for the trigger(s) set by the trigger settings can occur at action 810 and on a condition of meeting a trigger setting, method 800 includes, at action 815, accessing respective file locations and retrieving the reference NASPROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 information and the biological sample NAS data. For example, a processor, such as processor 102 or 702, accesses the respective file locations of storage of the reference NAS information 710 and the biological sample NAS data 720 upon meeting trigger setting 754. At step 820, the method further includes using a processor to output consensus metrics for consensus of the biological sample NAS data with the reference NAS information. For example, processor 102 or 702 is used to output consensus metrics 730 for consensus of the biological sample NAS data 720 with the reference NAS information 710.

[0099] Now referring to FIG. 9, an example computing device 900 is presented. In embodiments, computing device 900 may be same, include, or be a part of computing device 701. In this example, computing device 900 includes at least one processor 991 communicably coupled to a computer-readable storage medium 992. The at least one processor 991 includes a microprocessor, a microcontroller, one or more central processing unit (CPU) cores, an applicationspecific integrated circuit (ASIC), one or more graphical processing unit (GPU) cores, a field programmable gate array (FPGA), and / or any other hardware device suitable for retrieval and execution of instructions from computer-readable storage medium 992. In instances, the at least one processor 991 includes electronic circuitry for performing instructions described in this disclosure.

[0100] In embodiments, computer-readable storage medium 992 may be any medium suitable for storing executable instructions. In examples, without limitation, computer-readable storage medium 992 includes RAM, ROM, EEPROM, HHD, SSD, optical disc, and the like. Computer-readable medium storage 992 may be disposed within computing device 900. In embodiments, computer-readable storage medium 992 may external, and communicably connected, to computing device 900. The instruction stored on computer- readable storage medium may be used to implement the method described in reference to FIG. 8.

[0101] In this example, still referring to FIG.9, computer-readable storage medium 992 is encoded with set of instructions 993-997. In an embodiment,PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 executable instructions included in each block may be included in different blocks and blocks not shown.

[0102] In embodiments, instruction 993, when executed by at least one processor 991 , configures the at least one processor 991 to receive an input 750. Input 750 includes the following data attributes 993a-993d. Data attribute 993a includes a file location to retrieve reference NAS information (e.g., reference NAS information 710). Data attribute 993b includes a file location to retrieve biological sample NAS data (e.g., biological NAS data 720). Data attribute 993c includes analysis settings (e.g., analysis settings 753) for comparison of sequences corresponding to each of the reference NAS information and the biological sample NAS data. Data attribute 993d includes trigger settings (e.g., trigger settings 754) for performing analysis of the biological sample NAS data.

[0103] Referring again to FIG. 10, the interface 1000 for setting up a project (job) to be utilized by the system 700 is described in more detail. In this example, creating the project at interface 1000 includes a name tab 1002. The name tab 1002 is shows the step of job creation a user is in. Such as in this example, the first step for creating the project is highlighted, while further processes of creating trigger settings, which an example scheduler 1300 is described in reference to FIG. 13 further below, and a final review section are grayed out.

[0104] The interface 1000 also includes a job name 1003 section. This section may be prefilled with an automated name for the project. In some cases, a user may be able to override the prefilled naming by filling out the text field in this section. In some examples, the auto-project section 1003 may be prefilled by a name supplied in the naming step 1002. In examples, this section is used for dynamic naming, where the auto-project name is used by a computing system, which is mapped to a user set name, which may be set at other sections. The example 1000 further includes a description section 1004 for providing some context of the goal or type of analysis of the automated project. The project name section 1005 also allows a user to specify a name for the project which sets a naming convention. In some cases, the project name may only be set at thisPROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 section, which if left blank, the naming of the project becomes what is prefilled at section 1003. In some other cases, section 1005 provides a dynamic naming, where the name provided at section 1005 is used by users of the same, which is mapped to an auto-project name, prefilled at section 1003, used by the system. A person of ordinary skill would understand how computing system use dynamic naming for readability reasons, since naming used by the system are often difficult to be used by a human user.

[0105] Under an organizing section 1006, a computing system may automatically extract delimiters to be used on the project 1000 from the human set project title of section 1005. For example, project 1000 may be set up where the name under section 1005 follows a set convention, such as in the example 1000, the naming convention is “ProjectName_Reference_Sample_DateStamp.”

[0106] In this example, file location section 1007 specifies file locations for reference, sample, and template files 751 , 752, 753, respectively. In some examples, this location may be a network drive. In examples, the file location may be remote data store, such as remote data store 140, 740.

[0107] As mentioned above, the project example includes a reference file section 1051 for providing the reference file location (e.g., reference file 751 ), a sample file location 1052 for providing sample file location (e.g., sample file 752) and a templates file section 1053 for providing a location for analysis settings (e.g., analysis settings 753). An example of a template is described below with reference to FIGs. 11 and 12. As mentioned above, once the project is set up using interface 1000, the trigger settings can then be set, an example of which is as a scheduler 1300 in reference to FIG. 13.

[0108] Now referring to FIG. 11 , an embodiment of an auto-create specimen group section 1100 of an interface for generating a template is presented. Specimens and references from projects may be automatically specified through the interface 1100. This example includes an organize tab 1101 , a file section 1102, and attribute selection section 1103. Attribute in this example refers to the specimen and reference. In this example, the attribute selection section 1103PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 allows for specifying delimiters for both the specimen and the reference names. Once the naming is completed, the user can select the file name for usage by clicking the “choose file” button of file section 1102.

[0109] Now referring to FIG. 12, an embodiment of an analysis setting section 1200 of an interface for generating a template is described. Analysis settings can include overall consensus analysis settings and region level analysis settings (such as described above with reference to analysis settings 105). In this example, template 1200 includes setting the analysis settings (e.g., for analysis settings 753). Similar to example interface 1100, example interface 1200 includes an analysis step tab 1201 . This example interface 1200 includes a trim section 1202 for setting up trim threshold values for filtering out low-quality bases.

[0110] In this example, a user can specify a number of bases in the 5’ (five prime) and 3’ (three prime) regions, which generally includes lower quality bases. This example interface 1200 also includes a trace filtering section 1203, the filtering section allows for setting up thresholds for trace filtering, such as noisereduction. For example, a user may set an alignment score for the trace file as 22, where trace files with an alignment score below the set threshold are removed from analysis. Interface 1200 also includes an alignment and assembly section 1204 that allows for providing specific settings for each specimen and threshold values for multiple parameters, such as gap penalty, which is the cost of gaps introduced during assembly, extension penalty, which is the cost of extending an existing gap during assembly, and alignment stringency. As described in reference to FIG. 4, analysis setting may also be edited at the consensus results output 400 interface. In an example, a project may run with analysis settings 1200 of template at specified times, but a user may re-run the project with edited settings after result is outputted.

[0111] With reference now to FIG. 13, an example of a scheduler interface 1300 for trigger settings is shown. As mentioned above, the scheduler 1300 is an example of trigger settings, but such trigger settings could be based on an eventPROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1 as well, rather than the time-based scheduler 1300. This example scheduler 1300 shows a date / time based trigger. It should be note that a scheduler is only one example of trigger settings. Many other types of trigger settings are contemplated by the present disclosure, such as, as a trigger based on one or more events, such as, the creation or update of a file, machine learning based trigger such as triggers that occur based on an identified pattern of sample, references, and / or results.

[0112] FIGs. 10, 11 and 13 are provided only as examples for ease of description. It should be understood that once the parameters of those examples are set, the end result will be the input 750 that includes reference file location 751 , sample file location 752, analysis settings 753, and trigger settings 754 which are specified using interfaces configured to interact with a computing system, such as the interfaces provided as examples in reference to FIGs. 10, 11 , and 13.

[0113] Examples

[0114] The following numbered examples are embodiments:1 . A method for analysis of biological sample sequence data, the method comprising: receiving, at a computer processor, reference nucleic acid sequence information; receiving, at the computer processor, biological sample nucleic acid sequence data; comparing, using the computer processor, the biological sample nucleic acid sequence data to the reference nucleic acid sequence information; outputting, using the computer processor, overall consensus metrics of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information based on the comparison; and outputting, using the computer processor, one or more region level consensus metrics of one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 The method of example 1 , wherein the method further comprises receiving input of a file location for each of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data. The method of example 1 or 2, wherein the biological sample nucleic acid sequence data is generated by a remotely located instrument configured to perform a biological sample sequencing assay. The method of example 3, wherein the instrument comprises a capillary electrophoresis analysis instrument. The method of any one of examples 1 -4, wherein the reference nucleic acid sequence information is received from a laboratory information management (LIM) system. The method of any one of examples 1 -5, wherein receiving the reference nucleic acid sequence information comprises receiving the reference nucleic acid sequence information from a remote database. The method of any one of examples 1 -6, wherein receiving the reference nucleic acid sequence information comprises accessing a shared file system. The method of any one of examples 1 -7, further comprising identifying, by the computer processor, the one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information. The method of example 8, wherein the one or more regions are stored at a data store communicatively coupled to the computer processor. The method of any one of examples 1 -9, wherein the one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information are inputted by a user at the computer processor. The method of any one of examples 1 -10, wherein the biological sample nucleic acid sequence data comprises data collected from performing capillary electrophoresis on the biological sample.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1 The method of any one of examples 1 -11 , wherein the one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information comprise user-annotated regions of interest. The method of any one of examples 1 -12, wherein the one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information are formatted in a GenBank format standard. The method of any one of examples 1 -13, further comprising determining the region level consensus metrics as a function of one or more region level analysis settings. The method of example 14, wherein the one or more region level analysis settings are based on minimum coverage. The method of example 14, wherein the one or more region level analysis settings are based on one or both of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information. The method of any one of examples 1 -16, further comprising outputting quality metrics of the biological sample nucleic acid sequence data as a function of the one or more region level consensus metrics. The method of any one of examples 1 -17, further comprising comparing the one or more region level consensus metrics to threshold scores. The method of example 18, further comprising determining a match between the biological sample nucleic acid sequence data and the reference nucleic acid sequence information as a function of the comparison. A system comprising: at least one processor; and at least one memory, wherein the at least one memory contains instructions configuring the at least one processor to: receive reference nucleic acid sequence information; receive biological sample nucleic acid sequence data; compare the biological sample nucleic acid sequence data to the reference nucleic acid sequence information;PROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1 output overall consensus metrics of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information based on the comparison; and output one or more region level consensus metrics of one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information. The system of example 20, further configured to compare the biological sample nucleic acid sequence data to the reference nucleic acid sequence information using a consensus module. The system of example 20 or 21 , wherein the system comprises one or more remote instruments. The system of example 22, wherein the system is further configured to receive the biological sample nucleic acid sequence data from the one or more remote instruments. The system of any one of examples 20-23, wherein the system is further configured to retrieve the reference nucleic acid sequence information and the biological sample nucleic acid sequence data from a remote database. The system of any one of examples 20-24, wherein the system is further configured to retrieve the reference nucleic acid sequence information and the biological sample nucleic acid sequence data from a shared file system. The system of any one of examples 20-25, wherein the reference nucleic acid sequence information comprises one or more reference regions of interest. The system of any one of examples 20-26, wherein the biological sample nucleic acid sequence data comprises one or more biological sample regions of interest. The system of any one of examples 20-27, further configured to output the one or more region level consensus metrics as a function of region level analysis settings. The system of example 28, wherein the system is further configured to generate the region level analysis settings as a function of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1 A non-transitory computer-readable medium storing instructions, when executed by at least one processor, configuring the at least one processor to: receive reference nucleic acid sequence information; receive biological sample nucleic acid sequence data; compare the biological sample nucleic acid sequence data to the reference nucleic acid sequence information; output overall consensus metrics of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information based on the comparison; and output one or more region level consensus metrics of one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information. A method for analysis of biological sample sequence data, the method comprising: receiving input at a computer processor, the input comprising: a file location to retrieve reference nucleic acid sequence information, a file location to retrieve biological sample nucleic acid sequence data, analysis settings for comparison of sequences corresponding to each of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and trigger settings for performing analysis of the biological sample nucleic acid sequence data; and on a condition of meeting a trigger specified in the trigger settings, using the computer processor to: retrieve from the respective file location the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and output metrics corresponding to a consensus of the biological sample nucleic acid sequence data with the reference nucleic acid sequence information. The method of example 31 , wherein the analysis settings comprise overall sequence consensus.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1 The method of example 31 or 32, wherein the analysis settings comprise region level sequence consensus. The method of any one of examples 31 -33, further comprising outputting overall consensus metrics. The method of any one of examples 31 -34, further comprising outputting one or more region level consensus metrics. The method of any one of examples 31 -35, wherein the trigger settings comprise a scheduler. The method of example 36, wherein the scheduler is time based. The method of any one of examples 31 -37, wherein the trigger setting is generated based on a new sample file being uploaded. The method of example 38, further comprising scanning, by the computer processor, remotely located file storage location for the new sample file. The method of any one of examples 31 -39, wherein further comprising selecting, by the computer processor, the reference nucleic acid sequence information from a plurality of reference nucleic acid sequence information based on the biological sample nucleic acid sequence data. The method of any one of examples 31 -40, wherein the file location to retrieve nucleic acid sequence information is at one or more remotely located databases. The method of any one of examples 31 -41 , wherein the file location to retrieve biological sample nucleic acid sequence data is at one or more remotely located databases. The method of any one of examples 31 -42, further comprising exporting the output metrics to the file location of the biological sample nucleic acid sequence data. A system comprising: at least one processor; and at least one memory, wherein the at least one memory contains instructions configuring the at least one processor to: receive input comprising:PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 a file location to retrieve reference nucleic acid sequence information, a file location to retrieve biological sample nucleic acid sequence data, analysis settings for comparison of sequences corresponding to each of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and trigger settings for performing analysis of the biological sample sequence data; and on a condition of meeting a trigger specified in the trigger settings: retrieve from the respective file location the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and output metrics corresponding to a consensus of the biological sample nucleic acid sequence data with the reference nucleic acid sequence information. The system of example 44, wherein the trigger settings comprise a scheduler. The system of example 44 or 45, wherein the system is configured to retrieve at least one of the reference nucleic acid sequence information, biological sample nucleic acid sequence data and the analysis settings from a remote data store. A non-transitory computer-readable medium storing instructions, when executed by at least one processor, configuring the at least one processor to: receive input comprising: a file location to retrieve reference nucleic acid sequence information, a file location to retrieve biological sample nucleic acid sequence data, analysis settings for comparison of sequences corresponding to each of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and trigger settings for performing analysis of the biological sample sequence data; and on a condition of meeting a trigger specified in the trigger settings: retrieve from the respective file location the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, andPROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1 output metrics corresponding to a consensus of the biological sample nucleic acid sequence data with the reference nucleic acid sequence information.

[0115] It is to be understood that both the general description and the detailed description provide examples that are explanatory in nature and are intended to provide an understanding of the present disclosure without limiting the scope of the present disclosure. Various mechanical, compositional, structural, electronic, and operational changes may be made without departing from the scope of this description and the claims. In some instances, well-known circuits, structures, and techniques have not been shown or described in detail in order not to obscure the examples. Like numbers in two or more figures represent the same or similar elements.

[0116] In addition, the singular forms “a”, “an”, and “the” are intended to include the plural forms as well, unless the context indicates otherwise. Moreover, the terms “comprises”, “comprising”, “includes”, and the like specify the presence of stated features, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups. Components described as coupled may be electronically or mechanically directly coupled, or they may be indirectly coupled via one or more intermediate components, unless specifically noted otherwise. Mathematical and geometric terms are not necessarily intended to be used in accordance with their strict definitions unless the context of the description indicates otherwise, because a person having ordinary skill in the art would understand that, for example, a substantially similar element that functions in a substantially similar way could easily fall within the scope of a descriptive term even though the term also has a strict definition.

[0117] The phrase “and / or” is used herein in conjunction with a list of items. This phrase means that any combination of items in the list — from a single item to all of the items and any permutation in between — may be included. Thus, for example, “A, B, and / or C” means “one of {A}, {B}, {C}, {A, B}, {A, C}, {C, B}, and {A, C, B}”.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01

[0118] Elements and their associated aspects that are described in detail with reference to one example may, whenever practical, be included in other examples in which they are not specifically shown or described. For example, if an element is described in detail with reference to one example and is not described with reference to a second example, the element may nevertheless be claimed as included in the second example.

[0119] Unless otherwise noted herein or implied by the context, when terms of approximation such as “substantially,” “approximately,” “about,” “around,” “roughly,” and the like, are used, this should be understood as meaning that mathematical exactitude is not required and that instead a range of variation is being referred to that includes but is not strictly limited to the stated value, property, or relationship. In particular, in addition to any ranges explicitly stated herein (if any), the range of variation implied by the usage of such a term of approximation includes at least any inconsequential variations and also those variations that are typical in the relevant art for the type of item in question due to manufacturing or other tolerances. In any case, the range of variation includes at least values that are within ±1 % of the stated value, property, or relationship unless indicated otherwise.

[0120] Further modifications and alternative examples will be apparent to those of ordinary skill in the art in view of the disclosure herein. For example, the devices and methods includes additional components or steps that were omitted from the diagrams and description for clarity of operation. Accordingly, this description is to be construed as illustrative only and is for the purpose of teaching those skilled in the art the general manner of carrying out the present teachings. It is to be understood that the various examples shown and described herein are to be taken as exemplary. Elements and materials, and arrangements of those elements and materials, may be substituted for those illustrated and described herein, parts and processes may be reversed, and certain features of the present teachings may be utilized independently, all as would be apparent to one skilled in the art after having the benefit of the description herein. ChangesPROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 may be made in the elements described herein without departing from the scope of the present teachings and following claims.

[0121] It is to be understood that the particular examples set forth herein are nonlimiting, and modifications to structure, dimensions, materials, and methodologies may be made without departing from the scope of the present teachings.

[0122] Other examples in accordance with the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the following claims being entitled to their fullest breadth, including equivalents, under the applicable law.

Claims

PROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO1CLAIMSWhat is Claimed is1 . A method for analysis of biological sample sequence data, the method comprising: receiving, at a computer processor, reference nucleic acid sequence information; receiving, at the computer processor, biological sample nucleic acid sequence data; comparing, using the computer processor, the biological sample nucleic acid sequence data to the reference nucleic acid sequence information; outputting, using the computer processor, overall consensus metrics of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information based on the comparison; and outputting, using the computer processor, one or more region level consensus metrics of one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information.

2. The method of claim 1 , wherein the method further comprises receiving input of a file location for each of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data.

3. The method of claim 1 or 2, wherein the biological sample nucleic acid sequence data is generated by a remotely located instrument configured to perform a biological sample sequencing assay.

4. The method of claim 3, wherein the instrument comprises a capillary electrophoresis analysis instrument.

5. The method of any one of claims 1 -4, wherein the reference nucleic acid sequence information is received from a laboratory information management (LIM) system.

6. The method of any one of claims 1 -5, wherein receiving the reference nucleic acid sequence information comprises receiving the reference nucleic acid sequence information from a remote database.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO17. The method of any one of claims 1 -6, wherein receiving the reference nucleic acid sequence information comprises accessing a shared file system.

8. The method of any one of claims 1 -7, further comprising identifying, by the computer processor, the one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information.

9. The method of claim 8, wherein the one or more regions are stored at a data store communicatively coupled to the computer processor.

10. The method of any one of claims 1 -9, wherein the one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information are inputted by a user at the computer processor.11 . The method of any one of claims 1 -10, wherein the biological sample nucleic acid sequence data comprises data collected from performing capillary electrophoresis on the biological sample.

12. The method of any one of claims 1 -11 , wherein the one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information comprise user-annotated regions of interest.

13. The method of any one of claims 1 -12, wherein the one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information are formatted in a GenBank format standard.

14. The method of any one of claims 1 -13, further comprising determining the region level consensus metrics as a function of one or more region level analysis settings.

15. The method of claim 14, wherein the one or more region level analysis settings are based on minimum coverage.

16. The method of claim 14, wherein the one or more region level analysis settings are based on one or both of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information.

17. The method of any one of claims 1 -16, further comprising outputting quality metrics of the biological sample nucleic acid sequence data as a function of the one or more region level consensus metrics.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO118. The method of any one of claims 1 -17, further comprising comparing the one or more region level consensus metrics to threshold scores.

19. The method of claim 18, further comprising determining a match between the biological sample nucleic acid sequence data and the reference nucleic acid sequence information as a function of the comparison.

20. A system comprising: at least one processor; and at least one memory, wherein the at least one memory contains instructions configuring the at least one processor to: receive reference nucleic acid sequence information; receive biological sample nucleic acid sequence data; compare the biological sample nucleic acid sequence data to the reference nucleic acid sequence information; output overall consensus metrics of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information based on the comparison; and output one or more region level consensus metrics of one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information.21 . The system of claim 20, further configured to compare the biological sample nucleic acid sequence data to the reference nucleic acid sequence information using a consensus module.

22. The system of claim 20 or 21 , wherein the system comprises one or more remote instruments.

23. The system of claim 22, wherein the system is further configured to receive the biological sample nucleic acid sequence data from the one or more remote instruments.

24. The system of any one of claims 20-23, wherein the system is further configured to retrieve the reference nucleic acid sequence information and the biological sample nucleic acid sequence data from a remote database.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO125. The system of any one of claims 20-24, wherein the system is further configured to retrieve the reference nucleic acid sequence information and the biological sample nucleic acid sequence data from a shared file system.

26. The system of any one of claims 20-25, wherein the reference nucleic acid sequence information comprises one or more reference regions of interest.

27. The system of any one of claims 20-26, wherein the biological sample nucleic acid sequence data comprises one or more biological sample regions of interest.

28. The system of any one of claims 20-27, further configured to output the one or more region level consensus metrics as a function of region level analysis settings.

29. The system of claim 28, wherein the system is further configured to generate the region level analysis settings as a function of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data.

30. A non-transitory computer-readable medium storing instructions, when executed by at least one processor, configuring the at least one processor to: receive reference nucleic acid sequence information; receive biological sample nucleic acid sequence data; compare the biological sample nucleic acid sequence data to the reference nucleic acid sequence information; output overall consensus metrics of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information based on the comparison; and output one or more region level consensus metrics of one or more regions of each of the biological sample nucleic acid sequence data and the reference nucleic acid sequence information.31 .A method for analysis of biological sample sequence data, the method comprising: receiving input at a computer processor, the input comprising: a file location to retrieve reference nucleic acid sequence information, a file location to retrieve biological sample nucleic acid sequence data,PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 analysis settings for comparison of sequences corresponding to each of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and trigger settings for performing analysis of the biological sample nucleic acid sequence data; and on a condition of meeting a trigger specified in the trigger settings, using the computer processor to: retrieve from the respective file location the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and output metrics corresponding to a consensus of the biological sample nucleic acid sequence data with the reference nucleic acid sequence information.

32. The method of claim 31 , wherein the analysis settings comprise overall sequence consensus.

33. The method of claim 31 or 32, wherein the analysis settings comprise region level sequence consensus.

34. The method of any one of claims 31 -33, further comprising outputting overall consensus metrics.

35. The method of any one of claims 31 -34, further comprising outputting one or more region level consensus metrics.

36. The method of any one of claims 31 -35, wherein the trigger settings comprise a scheduler.

37. The method of claim 36, wherein the scheduler is time based.

38. The method of any one of claims 31 -37, wherein the trigger setting is generated based on a new sample file being uploaded.

39. The method of claim 38, further comprising scanning, by the computer processor, remotely located file storage location for the new sample file.

40. The method of any one of claims 31 -39, wherein further comprising selecting, by the computer processor, the reference nucleic acid sequence information from a plurality of reference nucleic acid sequence information based on the biological sample nucleic acid sequence data.PROVISIONAL PATENT APPLICATIONTF REF.: TP387869WO141 . The method of any one of claims 31 -40, wherein the file location to retrieve nucleic acid sequence information is at one or more remotely located databases.

42. The method of any one of claims 31 -41 , wherein the file location to retrieve biological sample nucleic acid sequence data is at one or more remotely located databases.

43. The method of any one of claims 31 -42, further comprising exporting the output metrics to the file location of the biological sample nucleic acid sequence data.

44. A system comprising: at least one processor; and at least one memory, wherein the at least one memory contains instructions configuring the at least one processor to: receive input comprising: a file location to retrieve reference nucleic acid sequence information, a file location to retrieve biological sample nucleic acid sequence data, analysis settings for comparison of sequences corresponding to each of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and trigger settings for performing analysis of the biological sample sequence data; and on a condition of meeting a trigger specified in the trigger settings: retrieve from the respective file location the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and output metrics corresponding to a consensus of the biological sample nucleic acid sequence data with the reference nucleic acid sequence information.

45. The system of claim 44, wherein the trigger settings comprise a scheduler.

46. The system of claim 44 or 45, wherein the system is configured to retrieve at least one of the reference nucleic acid sequence information, biological sample nucleic acid sequence data and the analysis settings from a remote data store.

47. A non-transitory computer-readable medium storing instructions, when executed by at least one processor, configuring the at least one processor to:PROVISIONAL PATENT APPLICATIONTF REF.: TP387869W01 receive input comprising: a file location to retrieve reference nucleic acid sequence information, a file location to retrieve biological sample nucleic acid sequence data, analysis settings for comparison of sequences corresponding to each of the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and trigger settings for performing analysis of the biological sample sequence data; and on a condition of meeting a trigger specified in the trigger settings: retrieve from the respective file location the reference nucleic acid sequence information and the biological sample nucleic acid sequence data, and output metrics corresponding to a consensus of the biological sample nucleic acid sequence data with the reference nucleic acid sequence information.

Citation Information

Patent Citations

  • Inferring microorganism of origin for antimicrobial resistance markers in targeted metagenomics

    US20240254570A1

  • Systems and methods for iterative and scalable population-scale variant analysis

    WO2023114415A2

  • AU2021396452A1