Systems and methods for systematic design of blueprints for pathogen detection
A computational approach using AI and machine learning automates the design of PCR and CRISPR blueprints for pathogen detection, addressing inefficiencies in existing methods by enhancing assay design efficiency and accuracy.
Patent Information
- Application Number
- PCT/US2025/036160
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-01
- Filing Date
- 2025-07-01
- Publication Date
- 2026-01-08
AI Technical Summary
Existing methods for designing PCR and CRISPR-based pathogen detection assays are time-consuming, require significant expertise, and lack efficiency in identifying high-quality amplicons and protospacers for accurate and rapid pathogen detection.
A systematic computational approach using artificial intelligence and machine learning algorithms to identify and evaluate nucleotide sequences for amplicons and protospacers, automating the design of PCR and CRISPR blueprints for pathogen detection, incorporating tools like Primer3, BLAST, and ChopChop to enhance specificity and sensitivity.
Facilitates rapid and accurate pathogen detection by automating the design of high-quality assays, reducing human intervention, and improving detection accuracy and speed through intelligent blueprint identification.
Smart Images

Figure US2025036160_08012026_PF_FP_ABST
Abstract
Description
Systems and Methods for Systematic Design of Blueprints for Pathogen DetectionREFERENCES
[0001] This application claims priority from U.S. provisional patent application serial no. 63 / 666,660, entitled "Systems and Methods for Systematic Designing of Assays for Pathogen Detection," filed on July 1, 2024. The priority application is hereby incorporated by reference, as if it is set forth in full in this specification.
[0002] Each publication, patent, and / or patent application mentioned in this specification is herein incorporated by reference in its entirety to the same extent as if each individual publication and / or patent application was specifically and individually indicated to be incorporated by reference.BACKGROUNDFIELD OF TECHNOLOGY
[0003] The disclosed technology relates to artificial intelligence (Al) type computers and digital data processing systems and corresponding data processing methods and products for emulation of intelligence (i.e., knowledge-based systems, reasoning systems, and knowledge acquisition systems); and includes systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. In particular, the disclosed technology relates to identification and evaluation of subsequences of nucleotide bases for detection of a target pathogen.CONTEXT
[0004] PCR and CRISPR based test panels are increasingly replacing conventional bacterial culture-based tests. PCR and CRISPR offer accuracy, speed, and low-cost detection. They may also detect pathogens that don't grow on a petri dish and viruses that require living host cells.SUMMARY
[0005] In some aspects, the techniques described herein relate to a method of systematically determining a blueprint including three sequences of nucleotides for a subsequence of nucleotide bases representing a genome of a target pathogen, including: identifying one or more feasible subsequences of nucleotide bases in the subsequence ofnucleotide bases representing the genome of the target pathogen; identifying, for the one or more feasible subsequences, a plurality of amplicons wherein an amplicon is a portion of a feasible subsequence of the plurality of feasible subsequences; matching the respective amplicons with at least one non-target pathogen, and assigning a specificity score to the respective amplicons based on mismatches of nucleotides of the respective amplicons with nucleotides of the at least one non-target pathogen; selecting an amplicon with a respective specificity score above a minimum-specificity threshold; identifying the blueprint for the selected amplicon wherein the blueprint includes three sequences of nucleotide bases that are used for generating the selected amplicon; and storing the blueprint for the selected amplicon for use in detecting the target pathogen in a biological sample.
[0006] In some aspects, the techniques described herein relate to a system including one or more processors coupled to memory, the memory loaded with computer instructions to systematically determine a blueprint including three sequences of nucleotides for a subsequence of nucleotide bases representing a genome of a target pathogen, the instructions, when executed on the processors, implementing actions including: identifying one or more feasible subsequences of nucleotide bases in the subsequence of nucleotide bases representing the genome of the target pathogen; identifying, for the one or more feasible subsequences, a plurality of amplicons wherein an amplicon is a portion of a feasible subsequence of the plurality of feasible subsequences; matching the respective amplicons with at least one non-target pathogen, and assigning a specificity score to the respective amplicons based on mismatches of nucleotides of the respective amplicons with nucleotides of the at least one non-target pathogen; selecting an amplicon with a respective specificity score above a minimum-specificity threshold; identifying the blueprint for the selected amplicon wherein the blueprint includes three sequences of nucleotide bases that are used for generating the selected amplicon; and storing the blueprint for the selected amplicon for use in detecting the target pathogen in a biological sample.
[0007] In some aspects, the techniques described herein relate to a non-transitory computer-readable storage medium storing computer program instructions to systematically determine a blueprint including three sequences of nucleotides for a subsequence of nucleotide bases representing a genome of a target pathogen, wherein the computer program instructions, when executed on a processor, implement actions including: identifying one or more feasible subsequences of nucleotide bases in the subsequence of nucleotide bases representing the genome of the target pathogen; identifying, for the one or more feasible subsequences, a plurality of amplicons wherein an amplicon is a portion of a feasible subsequence of the plurality of feasible subsequences; matching the respective amplicons with at least one non-target pathogen, and assigning a specificity score to the respective amplicons based on mismatches of nucleotides of the respective amplicons with nucleotides of the at least one non-target pathogen; selecting an amplicon with a respective specificityscore above a minimum-specificity threshold; identifying the blueprint for the selected amplicon wherein the blueprint includes three sequences of nucleotide bases that are used for generating the selected amplicon; and storing the blueprint for the selected amplicon for use in detecting the target pathogen in a biological sample.
[0008] In some aspects, the techniques described herein relate to logic configured to execute functionality including: comparing a target genome with one or more non-target genomes to generate a heatmap of mismatches for each of the one or more non-target genomes; on the target genome, identifying one or more high-specificity areas, wherein a high-specificity area is a sequence of nucleotides with a density of mismatches higher than a minimum mismatch density for all non-target genomes; using a sliding window to identify, in one or more of the one or more high-specificity areas, multiple candidate blueprints; for each of the multiple candidate blueprints, designing an associated candidate assay; determining a quality score for each candidate blueprint and each associated candidate assay; and ranking the multiple candidate blueprints based on their quality scores.
[0009] In some aspects, the techniques described herein relate to a system for systematic design of blueprints for pathogen detection, including: blueprint and assay design logic; logic to remove low-quality molecular data from a molecular dataset to obtain curated molecular data; a database to store the curated molecular data; a molecular knowledge database to store and retrieve information about molecules processed by the blueprint and assay design logic; and an assays database to store candidate assays; wherein the blueprint and assay design logic is configured to: compare a target genome with one or more non-target genomes to generate a heatmap of mismatches for each of the one or more non-target genomes; on the target genome, identify one or more high-specificity areas, wherein a high-specificity area is a sequence of nucleotides with a density of mismatches higher than a minimum mismatch density for all non-target genomes; use a sliding window to identify, in one or more of the one or more high-specificity areas, multiple candidate blueprints; for each of the multiple candidate blueprints, design an associated candidate assay; determine a quality score for each candidate blueprint and each associated candidate assay; and rank the multiple candidate blueprints based on their quality scores.
[0010] In some aspects, the techniques described herein relate to a non-transitory computer-readable storage medium storing computer program instructions to design blueprints for pathogen detection, wherein the computer program instructions, when executed on a processor, implement actions including: comparing a target genome with one or more non-target genomes to generate a heatmap of mismatches for each of the one or more non-target genomes; on the target genome, identifying one or more high-specificity areas, wherein a high-specificity area is a sequence of nucleotides with a density of mismatches higher than a minimum mismatch density for all non-target genomes; using a sliding window to identify, in one or more of the one or more high-specificity areas, multiplecandidate blueprints; for each of the multiple candidate blueprints, designing an associated candidate assay; determining a quality score for each candidate blueprint and each associated candidate assay; and ranking the multiple candidate blueprints based on their quality scores.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, with an emphasis instead generally being placed upon illustrating the principles of the disclosed technology. In the following description, various implementations of the disclosed technology are described with reference to the following drawings.
[0012] FIG. 1 illustrates an overview of an example system in which a molecular assay development platform is used to identify at least one blueprint for a target pathogen.
[0013] FIG. 2 illustrates an example of the pathogen detection process using Polymerase Chain Reaction (PCR).
[0014] FIG. 3 illustrates an example of the pathogen detection process using clustered regularly interspaced short palindromic repeats based on the enzyme Casl2a (CRISP- Casl2a).
[0015] FIG. 4 illustrates an example method of amplification of an intended target using recombinase polymerase amplification (RPA) isothermal amplification at room temperature.
[0016] FIG. 5 presents an example architecture of a system to systematically design assays for pathogen detection using PCR.
[0017] FIG. 6 presents an example architecture of a system to systematically design assays for pathogen detection using CRISPR.
[0018] FIG. 7 presents an example method to systematically find and evaluate blueprints for pathogen detection.
[0019] FIG. 8 illustrates an example process based on the method in FIG. 7.
[0020] FIG. 9 shows an implementation of a split search query blueprint generation process that may be used for evaluating blueprints and assays for PCR.
[0021] FIG. 10 shows an implementation for the generation of mutated genomes for a target organism genome that includes sequences identifying a target pathogen.
[0022] FIG. 11 shows an implementation of the generation of blueprints that are robust to potential future genome mutations.
[0023] FIG. 12 illustrates example pseudocode for Stage 1.
[0024] FIG. 13 illustrates example pseudocode for Stage 2 for PCR.
[0025] FIG. 14 illustrates example pseudocode for Stage 3.
[0026] FIG. 15 shows an example computer system that can be used to implement the disclosed technology.DETAILED DESCRIPTION
[0027] Polymerase Chain Reaction (PCR) is an important methodology in molecular biology. It is used to make millions of exact copies of specific DNA sequences, thereby enriching a particular segment of DNA, which can reduce false negatives when detecting a pathogen. Similarly, CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) is another important technology, originally designed for gene editing but that can also be used in diagnostics. Different proteins associated with CRISPR are Cas9, Casl2, Cas 13, and Casl4. Casl2a is rapidly gaining attention for its high precision and low cost in pathogen detection.
[0028] Examples are the detection and diagnosis of pathogens causing infectious disease, cancer causing mutations in oncology, and mutations involved in inherited disease. To test a sample, it is placed in an environment with one or more reagents (assays) to detect the pathogen DNA. Pathogen DNA is amplified to increase sensitivity, and the test result can be read from a fluorescent marker on a fluorescence reader.
[0029] The design of effective assays for PCR and CRISPR technologies is critical to having the highest sensitivity, speed, specificity (i.e., accuracy of detecting specific pathogens), and a clear fluorescence reading.
[0030] PCR utilizes DNA sequences called primers and probes that anneal to their target DNA. The sequences of primers and probes are specific to their target. For example, if we want to know whether a sample includes a particular Escherichia coli (E. coli) bacterium, we may design primers and probes with a sequence that matches a sequence of the bacterium. The specific primers and probes can anneal to the DNA of the bacterium, which initiates amplification of the targeted region. A fluorescent dye of the probe confirms that the bacterium is present in the sample. If there is no fluorescence, the bacterium has not been detected.
[0031] Casl2a-based CRISPR detection, on the other hand, uses single-stranded guide RNA (crRNA - CRISPR RNA) that anneals to target DNA. Here we design the crRNA sequence that matches the unique sub-sequence of a pathogen's DNA. To allow pathogen detection, the Casl2a protein is attached to the crRNA sequence, to form a ribonucleoprotein (RNP). The RNP can bind to the double-stranded DNA (dsDNA) of the pathogen target. It includes a protospacer adjacent motif (PAM) and a protospacer. The PAM, for Casl2a, includes the nucleotides TTTV, where T is A, C, or G. The protospacer itself may be a sequence of 20-24 base pairs. The singlestranded guide RNA is designed to match the single-stranded DNA protospacer, including the same sequence of nucleotides, but with the T nucleotide replaced with a U nucleotide.
[0032] This binding reaction causes the RNP to go into a collateral cleaving state, where all RNAs and DNAs are cut indiscriminately, including the reporter which then releases its fluorescence. This indicates detection of the pathogen. If there is no fluorescence, then pathogen DNA is not present.
[0033] However, not all protospacers are equally good. Some protospacers provide better detection results than others, based on base pairs included and the order of the included base pairs. Thus, designing a good test requires finding pathogen DNA sequences that include the PAM followed by a strong protospacer. The eventual assay must include one or more RNPs that bind properly to protospacers, as well as an ssDNA reporter that provides high fluorescence at the time of collateral cleavage. Implementations of the technology disclosed herein provide identification of the strongest protospacers, and thus design of the strongest RNP, and a reporter design with the highest possible fluorescence reading.
[0034] PCR-based tests for infectious diseases like Urinary Tract Infection (UTI) or Gastrointestinal Infection (GI) are increasingly replacing conventional bacterial culture-based tests. PCR-based tests can be faster and less expensive. These tests can detect a variety of pathogens simultaneously. For each pathogen represented in the test, primers and probes must be designed. This can be done manually using publicly available software tools like Primer3 to design the primer and probe sequences, and BLAST to check for specificity to avoid detecting additional organisms that show high sequence similarity to the target pathogen.
[0035] Similarly, for Casl2a based tests, tools such as ChopChop, CRISPOR and others can be used manually to select targets. Again, BLAST can be used to check for specificity and avoid detection of additional organisms that show high sequence similarity to the target pathogen.
[0036] The disclosed technologies provide a systematic and efficient computational approach to the stated problem, resulting in high-quality assays. The technologies can be implemented as part of a molecular assay development platform such as provided by Medix Biotechnologies, Inc. The molecular assay development platform provides solutions quickly as new pathogens arise.
[0037] In PCR, the primers anneal (or bind) to particular locations of the target pathogen DNA, but not those of other possible pathogens. The sequence between primers is then copied multiple times, including the probe section. Fluorescence indicates detection of the targeted area of a genome pathogen.
[0038] Similarly, in CRISPR-Casl2a, crRNA anneals to the target pathogen DNA, which initiates collateral cleavage including the reporter. This releases fluorescence, indicating presence of targeted pathogen. The CRISPR system does not amplify the target. If amplification or duplication of target is needed, then a step prior to detection called RPA (Recombinase Polymerase Amplification) or LAMP (Loop-Mediated Isothermal Amplification)can be used to increase the total target count and accuracy of detection. These amplification methods require the design of forward and reverse primers, similar to PCR methodology. RPA and LAMP amplification methodologies are well known but may have different design criteria than PCR. For example, for RPA you can select 30-45 base pairs as a forward primer (left of pathogen target) and 30-45 base pairs as reverse primer and complimentary of the sequence to the right of the target. This end-to-end region defined by the primers, including the target region gets amplified with the help of certain reagents. Amplification is fast and carries a low cost. These amplified millions of copies of the region can easily be detected by Casl2a or any other Cas systems.
[0039] The disclosed technology may utilize publicly available molecular and genomic datasets, software tools such as Primer3 or ChopChop, BLAST, and laboratory wet-lab testing to create pathogen detection assays. The technology provides an intelligent search engine that can automatically produce good assay designs for higher detection accuracy of pathogens.TERMINOLOGY"Amplicon" - a target DNA region on a target pathogen genome."Assay" - a combination of reagents, primers, probes, buffer and water (in the case of PCR), or reagents, the Casl2a protein, crRNA, reporter and buffer, and water (Casl2a detection method)."Blueprint" - see "PCR blueprint" or "CRISP blueprint"."Cas", "Cas9", "Casl2", "Casl3", "Casl4" - CRISPR-associated proteins 9, 12, 13, and 14 respectively - enzymes that uses CRISPR sequences as a guide to recognize and open up specific strands of DNA that are complementary to the CRISPR sequence. Cas enzymes together with CRISPR sequences form the basis of technologies known as CRISPR-Cas9, CRISPR-Casl2, etc., that can be used to edit genes within living organisms. Casl2a is a type of endonuclease that cuts DNA in a programmable, sequence-specific way. It is guided by a single CRISPR RNA (crRNA) that recognizes a specific sequence called the PAM with format TTTV, where V can be A, C, or G."CRISPR" - clustered regularly interspaced short palindromic repeats - a natural defense system found in bacteria and archaea that allows them to recognize and destroy invading viruses. This system has been adapted into a set of powerful tools for gene editing and molecular diagnostics."CRISPR blueprint" - a unique nucleotide sequence including the PAM sequence and protospacer within a pathogen genome. crRNA" - guide RNA."Database" - as used herein, the term "database" does not necessarily imply unity of structure. For example, two or more separate databases, when considered together, still constitute a "database"."DNA" - deoxyribonucleic acid - a polymer composed of two polynucleotide chains that coil around each other to form a double helix. The polymer carries genetic instructions for the development, functioning, growth and reproduction of all known organisms and many viruses. DNA and ribonucleic acid (RNA) are nucleic acids. Alongside proteins, lipids and complex carbohydrates (polysaccharides), nucleic acids are one of the four major types of macromolecules that are essential for all known forms of life (from Wikipedia)."dsDNA" - double-stranded DNA."dsRNA" - double-stranded RNA."LAMP" - loop-mediated isothermal amplification (LAMP) is a single-tube technique for the amplification of DNA for diagnostic purposes and a low-cost alternative to detect certain diseases. Isothermal amplification is carried out at a constant temperature, and does not require a thermal cycler as in a PCR system."Logic" - dedicated hardware, configured hardware, processors, processor systems, firmware, or software configured to execute logic functionality."Nucleotide" - an organic molecule composed of a nitrogenous base, a pentose sugar and a phosphate. It serves as a building block for deoxyribonucleic acid (DNA) and ribonucleic acid (RNA). Nucleotides are composed of three subunit molecules: a nucleobase, a five-carbon sugar (ribose or deoxyribose), and a phosphate group consisting of one to three phosphates. The four nucleobases in DNA are guanine (G), adenine (A), cytosine (C), and thymine (T); in RNA, uracil (U) is used in place of thymine."PAM" - a protospacer adjacent motif sequence is part of the pathogen target, which starts with TTTV, where, for example in CRISPR-Casl2a technology V is A, C, or G. This sequence is different for other CRISPR technologies such as Cas 9, 13 and 14."PCR" - polymerase chain reaction - a method used to rapidly make millions or billions of copies of a specific DNA sample. Where this document mentions PCR, an implementation may use any similar amplification technology, thermocycling or isothermal, including RPA, LAMP, bridge amplification, DNA nanoball generation, multiple displacement amplification (MDA), rolling circle amplification (RCA), nucleic acid sequence-based amplification (NASBA), Helicase-dependent amplification (HDA), transcription-mediated amplification (TMA), crosspriming amplification (CPA), and any other amplification techniques."PCR blueprint" - a nucleotide sequence including a forward primer, reverse primer, and a probe, for example as generated by software tools such as Primer3, or by similar logic. A concatenated blueprint includes intermediate base pairs that link the probe region to theflanking forward primer and reverse primer binding sites. A split blueprint doesn't include the base pairs in between the probe and primers, so that the probe region is more flexible."Polymerase" - an enzyme that copies or builds DNA or RIMA by linking smaller nucleotide units."Reporter" sequence or gene - a short synthetic single-stranded DNA sequence that encodes a protein whose expression can be easily measured and used to track the activity of a gene or promoter. A reporter gene may be detectable through fluorescence, luminescence, or enzymatic activity."RNA" - ribonucleic acid, a single stranded molecule made up of nucleotides which plays a central role in carrying out an instruction in genetic code."RNP" - ribonucleoprotein - a complex formed by the Casl2a protein and its guide RNA (crRNA)."RPA" - recombinase polymerase amplification, a single tube, isothermal alternative to PCR. By adding a reverse transcriptase enzyme to an RPA reaction, it can detect RNA as well as DNA, without the need for a separate step to produce complementary DNA. This combined method is called reverse transcript RPA (RT-RPA)."ssDNA" - single-stranded DNA."ssRNA" - single-stranded RNA.IMPLEMENTATIONSEnvironment
[0040] FIG. 1 illustrates an overview of an example system 100 in which a molecular assay development platform is used to identify at least one blueprint for a target pathogen. FIG. 1 includes a molecular assay development platform 110 and external molecular datasets 120. The external molecular datasets 120 can be stored in one or more databases. FIG. 1 also includes external tools, which may include primer / probe design software or logic 440, such as Primer3 or logic implemented by Primer3, CRISPR design software or logic 542, such as ChopChop or CRISPOR, and specificity check software or logic 545, such as BLAST or logic implemented by BLAST, to avoid detection of additional organisms that show high sequence similarity to the target pathogen. In an implementation, molecular assay development platform 110 can use the internally implemented logic for PCR to process internal molecular datasets (not shown in FIG. 1) to design probes and primers and to check specificity, thereby eliminating the need for primer / probe design software or logic 540 and / or specificity check software or logic 545.
[0041] In one implementation, the technology can be used to design blueprints for a CRISPR-Casl2a system, where the blueprint sequence may be defined in the guide RNA(crRNA - CRISPR single-stranded RNA), a genome sequence of 20-24 nucleotides preceded by the Casl2a enzyme, and a reporter (fluorescence). Such systems can be used for nucleic acid detection. Nucleic acid-based technologies require only knowledge of the pathogen genome sequence to enable accurate and early diagnosis. In such an implementation, the blueprint can also be referred to as a CRISPR blueprint. The CRISPR-Casl2a system enables a simplified architecture for molecular assay development platform 110. It eliminates the need for external design tools such as primer / probe design software or logic 540, CRISPR design software or logic 542, and possibly specificity check software or logic 545. Further implementations may use other detection systems, including Cas9, Casl3, Casl4, and others.
[0042] Molecular assay development platform 110, external molecular datasets 120, primer / probe design software or logic 540, CRISPR design software or logic 542, and specificity check software or logic 545 are in communication with each other via one or more network(s) 181. Further details of the architecture of molecular assay development platform 110 are presented with reference to FIGS. 5 and 6.Detecting pathogens with PCR
[0043] FIG. 2 illustrates an example of the pathogen detection process 200 using Polymerase Chain Reaction (PCR). PCR is a technology that can quickly make many copies of a specific DNA sequence. Its capability to detect the specific DNA sequence also makes it suitable for diagnostic applications.
[0044] The polymerase enzyme is used to amplify the target defined by primers 203A-B. To detect the target, forward primer 203A and reverse primer 203B are moved along DNA under test that may include the targeted pathogen genome 201. The target region 206 of a nucleotide sequence is called an amplicon. The amplicon may be detected by use of a forward primer 203A that detects the amplicon's start sequence and a reverse primer 203B that detects the amplicon's end sequence. When primers 203A-B find target region 206, the polymerase enzyme generates identical copies called amplified amplicons 211.
[0045] The forward primer 203A and reverse primer 203B may be included in a PCR blueprint 215 that further includes a nucleotide sequence called a probe 205, which releases a fluorescent molecule to confirm detection.
[0046] PCR applies heat to separate the original DNA strand and its copy, and lowers the temperature to anneal the primers and probes. In successive temperature cycles ("PCR cycles"), the DNA polymerase generates more copies from the original DNA strand as well as from the amplified amplicons 211, each time doubling the total number of copies.
[0047] Molecular assay development platform 110 can determine one or more PCR blueprints 215. The amplicon for a PCR blueprint 215 (the exact sequence of primers and a probe) may be required to meet several criteria for use in pathogen detection process. Forexample, its length may need to be in a range of 100 to 150 base pairs, so that replication is sufficiently reliable. It may need to be stable at 60 degrees so that the test can be done in laboratory conditions. The primers cannot repeat inside the amplicon. The first five and the last three base pairs in the primers are the most important, along with the probe. There are thousands of possible pathogens, and the amplicon must be unique and detect only the target, and none of the other pathogens.
[0048] Conventional techniques include designing amplicons by hand, i.e., by a human who understands the science of assay design and manually performs the process. For example, a well-known PCR test for COVID-19 is based on this approach. Designing them is time-consuming work and requires significant expertise, but there are publicly available tools that can save time and make it easier. Some examples of such tools are:
[0049] Primer3: given a longer sequence (e.g. 400 bp), Primer3 software can identify potential forward-primer / probe / reverse-primer candidates in the sequence and give quality estimates across several dimensions on how well they might work. Primer3 may also be accessed through its web interface Primer3Plus.
[0050] BLAST (Basic Local Alignment Search Tool): given an amplicon candidate, BLAST will search all organisms in its database to find matching sequences. It can be focused on potential pathogens to limit the search time.
[0051] MUMmer: given two genomes, MUMmer will find alignments between them.
[0052] The technology automates the amplicon design process. The technology can use existing tools such as those listed above or other tools that perform similar functions and / or computations. Additionally, the systems and methods disclosed herein can partially or completely implement the functionality provided by external tools, eliminating the need for human intervention. The systems and methods disclosed herein can be used to systematically and efficiently identify and evaluate good amplicons. The technology can also be applied for assay design in other areas such as oncology. Any pathogens can be detected in this manner. The genomes of these pathogens mutate frequently, and therefore the amplicon designs need to be refreshed regularly over time. The technology provides reliable, comprehensive, and efficient methods of designing the blueprints of amplicons. The technology therefore allows use of PCR tests in a reliable and a practical manner for pathogen detection.Detecting pathogens with CRISPR
[0053] FIG. 3 illustrates an example of the pathogen detection process 300 using CRISPR- Casl2a technology. Casl2a is an RNA-guided endonuclease protein, and its primary function is to recognize and cleave double-stranded DNA (dsDNA) at the intended target sites within a targeted pathogen genome 301 (from a patient's sample). Ribonucleoprotein RNP 308 is a complex formed by the Casl2a protein 312 and attached CRISPR RNA (crRNA 311) that guidesto the target. CRISPR blueprint 315 includes the Protospacer Adjacent Motif PAM 306 sequence of 4 base pairs and the protospacer 307 sequence of 20-24 base pairs, found in pathogen genome 301. In Casl2a, the PAM 306 sequence must start with the nucleotides TTTV, where V is A, C, or G. The PAM sequence is different for other CRISPR technologies, like the Cas9, 13 and 14 systems.
[0054] When the RNP 308 sequence has found PAM 306 and protospacer 307, RNP 308 initiates indiscriminate collateral cleavage. The nucleotides in crRNA 311 matches the nucleotides in protospacer 307, except that the crRNA 311 has a uracil (U) nucleobase in every location where protospacer 307 has a thymine (T) nucleobase.
[0055] Implementations manage the identification of high-performance protospacers 307 and design of matching crRNA 311 by algorithms that find stronger bindings with guide RNA. The stronger the binding, the stronger the fluorescence signal. The reporter 305 sequence is also critical. A better design will yield a stronger fluorescence reading. Reporter 305 is a synthetic single-stranded DNA (ssDNA) that is labeled with a fluorophore (a molecule that emits fluorescence) on one end and a quencher (which suppresses fluorescence when target is not detected) on the other end. Once RNP 308 finds the PAM 306 and protospacer 307 sequence, RNP 308 starts collateral cleavage activity, which breaks down all surrounding DNA and RNA including the reporter 305. Reporter 305 then starts emitting light, which can be read by a low-cost flatbed fluorescence reader. If crRNA 311 and a nucleotide sequence following a PAM 306 sequence do not match, no fluorescence is emitted.
[0056] Reporter 305 is a small ssDNA sequence with a length of 3 to 25 nucleotides. The combination of nucleobases A, C, G, and T can have a large impact on signal strength during collateral cleavage. Generally, researchers use / FAM / TTATT / FQ for Casl2a, wherein FAM indicates the presence of a fluorescent dye, 6-carboxyfluorescein, attached to the DNA sequence. The "TTATT", or any other structure such as " I I I I I " or "TTTCTTT", is the core DNA sequence, and FQ signifies that a quencher molecule is also attached, forming a fluorophore- quencher (FQ) pair. The fluorescence signal can also be captured on a lateral flow strip 309 just by changing the reporter's FQ to biotin, everything else being the same.
[0057] Some implementations use a reporter with 3 to 25 cytosine (C) nucleobases, for example "CCCCCCCC", which may provide a higher signal strength due to a higher cleaving activity with cytosine nucleobases, and which may result in a better signal-to-noise ratio. To further increase signal and signa-to-noise ratio some implementations create a hairpin loop structure (or stem loop structure) at the FQ end by adding 2-7 complimentary nucleotide bases, for example CT on one end and AG on the other end of the linear structure of the reporter. For example, "CCCCCCCC" becomes "CTCCCCCCCCAG" to create the stem loop structure.
[0058] FIG. 4 illustrates an example method of amplification of an intended target using RPA isothermal amplification at room temperature. Temperatures up to 42°C are useful, and the method eliminates the use of equipment. The method provides amplification before the detection step. Casl2a is large protein, therefore 200 base pairs 402 are selected from pathogen genome 301, including the intended target comprising PAM 306 and protospacer 307. The intended target (CRISPR blueprint 315) is located somewhere in the middle of the 200 base pairs. Some implementations use somewhat fewer or more base pairs than 200. The forward primer 316 are located at the beginning and reverse primer 317 located at the end of the 200 base pairs each include 35 base pairs. RPA amplifies all 200 base pairs, including the intended target. Within a few minutes, RPA with reagents can generate millions of replicates of the 200 base pairs to achieve a high sensitivity. In another implementation, only the blueprint along with 25-250 base pairs on both sides of the target (for a total of 78 to 528 base pairs) can be amplified.
[0059] Some implementations that use RPA add a probe, enabling visual or automatic validation of the amplification process. When amplification occurs, a fluorescence reader can measure its strength.Architecture
[0060] FIG. 5 presents an example architecture of the system 100 to systematically design assays for pathogen detection using PCR. FIG. 5 presents architectural components of the molecular assay development platform 110 from FIG. 1. The components, databases and applications that are implemented as part of the molecular assay development platform are presented inside a dashed box. Molecular assay development platform 110 can access external molecular datasets 120. The external molecular datasets 120 store molecular data that can be used for pathogen detection. Molecular assay development platform 110 can also include internal molecular datasets (not shown in FIG. 5) that may only be accessible within molecular assay development platform 110. The platform may include a pre-processing component 505 to curate molecular data obtained from external molecular datasets 120. The pre-processing operations can remove low-quality molecular data so that only high-quality molecular data is used for assay development. Pre-processing component 505 may also include logic to select molecules with more base pairs when multiple molecules that represent the same DNA or RNA sequence are available. Longer molecules with more base pairs provide better opportunities for detecting all possible information regarding the target pathogen. This can aid the development of high-quality assays for the target pathogen. The pre-processed molecules are stored in a DNA and RNA curated database 140.
[0061] Molecular assay development platform 110 can use a wide variety of search algorithms to determine one or more blueprints (each comprising a forward primer, a reverse primer and a probe) to detect a target pathogen. Implementations can systematically searchthe target pathogen when designing a blueprint for it. In such systematic searches, all base pairs of a target pathogen are reviewed using a sliding window without missing a base pair. The length of the sliding window can range from 10 to 1000 base pairs or beyond. Search component 101 may use Bayesian optimization techniques to account for mutations in the target pathogen. It can use wobble positioning algorithms and evolutionary artificial intelligence (Al) techniques to search for a blueprint for the target pathogen, and any other search techniques and algorithms. The search techniques and algorithms disclosed herein are presented for illustration of the technology, not for its limitation. The search component 101, implementing the intelligent search algorithms, can interface with systems such as primer / probe design software or logic 540 and specificity check software or logic 545 as shown in FIG. 5. Additionally, search component 101 can interface with other types of systems such as a CRISPR-associated system for generation of CRISPR blueprints. Molecular assay development platform 110 provides the flexibility to incorporate new types of pathogen detection techniques using the infrastructure it provides. The platform can implement such techniques internally as part of the platform or provide interfaces with external tools to design blueprints for pathogen detection.
[0062] The blueprints determined using the intelligent search algorithms implemented by search component 101 are stored in a molecular knowledge database 150. In one implementation, molecular knowledge database 150 can store information about molecules that have been processed by molecular assay development platform 110 using a variety of search algorithms along with their respective properties. The implementation can categorize properties in two or more categories. It can associate tags with respective molecules for ease of searching. One or more blueprints determined by the search techniques can be stored with respective molecules. Tags and / or properties can identify a DNA or RNA molecule and associate with one or more blueprints for a particular pathogen. Therefore, the implementation can easily search molecular knowledge database 150 for various DNA and RNA molecules, pathogens, blueprints, etc.
[0063] The implementation can send a blueprint from search component 101 to an assay design molecular recognition sequence multiplexing application (which may be simply referred to as an assay design multiplexing application 510). Assay design multiplexing application 510 can test the performance of multiple blueprints used in a chemical process for different target pathogens. For example, up to three different blueprints for respective target pathogens can be multiplexed in a first single chemical reaction. Other applications 530 can be used for testing in other multiplexing scenarios. For example, one other application can include up to ten blueprints in a second single chemical reaction. Individual blueprints that may work well when operating independently on a target pathogen may interfere with each other and not produce desired results when used together. Therefore, assay design multiplexing application 510 and other applications 530 detect such interactions that mayreduce the efficiency of the process to detect target pathogen. The blueprints are then sent to a laboratory test apparatus 520 that can perform wet-lab tests to determine the quality of the blueprints. A blueprint that performs well in the wet-lab test is selected for pathogen detection. The blueprints that do not perform well may be discarded. A feedback loop (or a feedback signal) may inform search component 101 about the results of the wet-lab test. Search component 101 may use the feedback signal iteratively to improve the design and development of the blueprints. The wet-lab testing process implemented by laboratory test apparatus 520 can determine whether the blueprint is of good quality based on a signal generated by laboratory test apparatus 520 during the testing process. A good signal indicates a good quality assay (or blueprint). The implementation stores blueprints that meet the desired quality level in an assays database 515 for use in pathogen detection.
[0064] FIG. 6 presents an example architecture of the system 100 to systematically design assays for pathogen detection using CRISPR. FIG. 6 presents architectural components of the molecular assay development platform 110 from FIG. 1. The components, databases and applications that are implemented as part of molecular assay development platform 110 are presented inside a dashed box. Molecular assay development platform 110 can access external molecular datasets 120. The external molecular datasets 120 store genomic data that can be used for pathogen detection. Molecular assay development platform 110 can also include internal molecular datasets (not shown in FIG. 6) that may only be accessible within molecular assay development platform 110. The platform may include a pre-processing component 505 to curate molecular data obtained from external molecular datasets 120. The pre-processing operations can remove low-quality molecular data so that only high-quality molecular data is used for assay development. Pre-processing component 305 may also include logic to select molecules with more base pairs when multiple molecules that represent the same DNA or RNA sequence are available. Longer molecules with more base pairs provide better opportunities for detecting all possible information regarding the target pathogen. This can aid the development of high-quality assays for the target pathogen. The pre-processed molecules are stored in a DNA and RNA curated database 140.
[0065] Molecular assay development platform 110 can use a wide variety of search algorithms to determine one or more blueprints (each comprising a PAM and a Protospacer sequence) to detect a target pathogen using CRISPR systems. Implementations can systematically search the target pathogen when designing a blueprint for it. In such systematic searches, all base pairs of a target pathogen are reviewed using a sliding window that quickly searches for occurrences of a PAM (e.g., TTTV for Casl2a as described earlier in this document). For part or all of the PAMs found, the implementation then reviews a range from 100 to 300 base pairs or beyond, starting at the PAM, for its suitability as a CRISPR blueprint 315. Search component 101 may use Bayesian optimization techniques to account for mutations in the target pathogen. It can use wobble positioning algorithms andevolutionary artificial intelligence (Al) techniques to search for a blueprint for the target pathogen, and any other search techniques and algorithms. The search techniques and algorithms disclosed herein are presented for illustration of the technology, not for its limitation. Blueprint and assay design software or logic 610 uses curated database 140 for the evaluation of CRISPR blueprints and generation of assay designs.
[0066] The blueprints determined using the intelligent search algorithms (blueprint and assay design software or logic 610) are stored in a molecular knowledge database 150. In one implementation, molecular knowledge database 150 can store information about molecules that have been processed by molecular assay development platform 110 using a variety of search algorithms along with their respective properties. The implementation can categorize properties in two or more categories. It can associate tags with respective molecules for ease of searching. One or more blueprints determined by the search techniques can be stored with respective molecules. Tags and / or properties can identify a DNA or RNA molecule and associate with one or more blueprints for a particular pathogen. Therefore, the implementation can easily search molecular knowledge database 150 for various DNA and RNA molecules, pathogens, blueprints, etc.
[0067] Blueprint and assay design software or logic 610 not only evaluates and ranks candidate blueprints, but also designs the associated components; crRNA sequence, reporter, and reagents, that make up an assay. This design application can find hundreds of high value targets from a single pathogen that is made up of, say, 5 million base pairs of DNA, where a pathogen DNA is stored in curated genomic database 140. The assay design can be sent for test in laboratory test apparatus 520 and the results are fed back to blueprint and assay design software or logic 610 and, if acceptable (a good signal indicates a good quality assay), the implementation stores it in assays database 515. However, if an assay does not sufficiently meet expectations, then an implementation may redesign or iterate the assay, and / or perform further wet-lab testing. This iterative verification process allows blueprint and assay design software or logic 610 to apply machine learning to improve a particular assay design, and to improve its overall assay design process. Other apps 550 may multiplex different blueprints from different pathogens in a single assay, allowing for the detection of multiple pathogens in a single test.
[0068] Additionally, blueprint and assay design software or logic 610 may also automatically generate or find data for multiplication, e.g., it can find forward and reverse primers for RPA amplification as part of the assay design.Process Steps
[0069] FIG. 7 presents an example method 700 to systematically find and evaluate blueprints for pathogen detection. The flowchart in FIG. 7 presents high-level process steps for designing determining blueprints and designing assays. As with all flowcharts herein,operations can be combined, performed in parallel or performed in a different sequence without affecting the functions achieved. In some cases, a re-arrangement of operations will achieve the same results only if certain other changes are made as well. In other cases, a rearrangement of operations will achieve the same results only if certain conditions are satisfied. Furthermore, the flow charts herein show only operations that are pertinent to an understanding of the technology, and numerous additional operations for accomplishing other functions can be performed before, after and between those shown.
[0070] Method 700 includes the following operations.
[0071] 710 - Start.
[0072] 720 - Identifying feasible subsequences or high-specificity areas in the genome of the target pathogen. Feasible subsequences and high-specificity areas may include one or more candidate blueprints. One implementation may use a window function to scan the whole target genome and find subsequences of a specific length to meet criteria for a blueprint. For example, a PCR blueprint of 400 base pairs requires a window of 400 base pairs length. The window, when sliding over the target pathogen or organism, yields successive blueprint candidates that can then be evaluated for various qualities, including probe strength, lack of repeats, chance of mutations, and uniqueness, to yield a feasible subsequence. Another implementation may use a sliding window to find occurrences of a PAM for a specific Cas enzyme (e.g., Casl2a) in the target genome. Once a PAM occurrence has been identified, the PAM and 20-24 base pairs following the PAM become a feasible subsequence.
[0073] In yet another implementation, the genome of the target pathogen is matched against genomes of each of the likely non-target pathogens. A subsequence that has sufficient nucleotide or base pair mismatches with each of the likely non-target pathogens is a high- specificity area. A subsequence that has no or very few mismatches with a non-target pathogen cannot be used to distinguish a target pathogen from a non-target pathogen. A high-specificity area must have sufficient mismatches and also be large enough to successfully detect.
[0074] 730 - Generating candidate blueprints. This operation uses genomes remaining from 720 and identifies a blueprint (including forward primer, probe, and reverse primer for PCR, or PAM and protospacer for a CRISPR system) for each pathogen genome. In addition, this operation can also include assigning quality and specificity (or uniqueness) scores to blueprints.
[0075] 740 - Evaluating candidate blueprints and assigning specificity scores. The evaluation process may include matching a blueprint candidate with relevant non-target pathogens and identifying the number of mismatches between the two sequences to determine the specificity score. The higher the number of mismatches, the higher the uniqueness (and therefore specificity) of the blueprint.
[0076] 750 - Determining if the specificity score of a candidate blueprint is higher than a threshold.
[0077] 760 - In response to determining that the specificity score of the candidate blueprint is lower than the threshold, discarding the candidate blueprint.
[0078] 770 - In response to determining that the specificity score of the candidate blueprint is higher than the threshold, designing a candidate assay related to the candidate blueprint, determine a quality score for the candidate blueprint and the candidate assay, and rank the candidate blueprint among other candidate blueprints based on the quality score. A candidate assay, designed for a candidate blueprint, may impact the quality score related to the assay effectiveness. For example, if the assay (or blueprint) delivers a strong fluorescence signal, the quality score may increase. If the assay results in false positives or negatives, the quality score may decrease. In one implementation, the technology can use a large number of searches to construct candidate assays that are future proof, i.e., such candidate assays are likely to stay effective even when the target pathogen mutates, resulting in new variations. The disclosed technology can use Bayesian optimization and / or evolutionary optimization to design candidate assays that can be used to detect target pathogens even when the target pathogen mutates.
[0079] 780 - Storing the candidate blueprints and candidate assays, along with the quality and specificity score for use in a pathogen detection.
[0080] 790 - End of the method.
[0081] Implementations may include the following elements. The "unit-step-size sliding window traversal logic" explores every base pair in a genome for purposes of blueprint generation. In case of a CRISPR system, an implementation only evaluates sequences that start with a PAM. A PCR system may identify "split search query blueprints" without taking into account intermediate base pairs that link probes to the flanking forward primers and the reverse primers. The PCR system may then apply "early-stopping logic" to terminate the genome matching process once a blueprint matches a non-target organism or mismatches a target organism. A system may further apply a mutation sampling process to generate mutation-robust blueprints, as discussed further below.Unit-Step-Size Sliding Window Traversal Logic
[0082] FIG. 8 illustrates an example process 800 based on method 700. Logic implementation 800 works on genome 810 of a target pathogen, that includes a sequence of base pairs with one or more subsequences that identify the target pathogen. Unit-step sliding window logic 820 looks at genome 810, one predetermined window at a time, to record successive subsequences of base pairs in blueprint generation logic 830 (e.g., Primer3). Thus, for each window, it records one subsequence of base pairs. Because successive windows coverall base pairs in the target genome, logic implementation 800 reviews all candidate blueprints, given a selected blueprint length, primer length, and probe length. This is in contrast with conventional approaches that generate blueprints for only certain regions of the target genome (for example, regions that are clinically identified).
[0083] In the case molecular assay development platform 110 develops blueprints and assays for PCR, the size of the sliding window may equal the intended size of the blueprint, i.e., the number of bytes from the start position of forward primer 203A until the end position of reverse primer 203B. This way, the system can evaluate all potential blueprints of the intended size and select the best candidate(s) from among them.
[0084] In the case molecular assay development platform 110 develops blueprints and assays for CRISPR, the size of the sliding window may equal the size of the PAM associated with the Cas enzyme. For all PAMs it finds, the system can evaluate the respective protospacers of the intended size, i.e., the intended number of protospacer bytes following the PAMs. Again, this way, the system can evaluate all potential blueprints of the intended size and select the best candidate(s) from among them. PAMs may be, for example, four base pairs, and protospacers may be 20-24 base pairs. The implementation may furthermore review the 80-90 base pairs on each the left and the right to include all roughly 200 base pairs required for a detection reaction (in case of a Casl2a protein). Within the 200 base pairs, the PAM or the blueprint should not repeat.
[0085] Blueprint-to-pathogen comparison logic 840 compares the sequences found with target organisms and non-target organisms in genomes database 850 to find matches and near-matches with target organisms and mismatches with non-target organisms. From these, it calculates specificity scores. Discarding blueprints with low specificity (for example, based on the specificity scores), it stores remaining blueprints in mapped blueprint database 860 and stores their quality scores and specificity scores in metadata in database 870. An implementation may work as follows: match on target organism >> increment specificity score mismatch on non-target organism >> increment specificity score mismatch on target organism >> decrement specificity score match on non-target organism >> decrement specificity score.
[0086] An implementation may further design an assay for each remaining blueprint and assign quality scores, based on properties and / or performance of the blueprint and the assay.
[0087] Quality sorting logic 880 sorts the blueprints, for example in order of decreasing quality scores, and stores the sorted blueprints in quality blueprints database 890.Split Search Query Blueprint
[0088] FIG. 9 shows an implementation 900 of a split search query blueprint generation process 910 that may be used for evaluating blueprints and assays for PCR (CRISPR doesn't require split searches). Concatenated search query blueprints include intermediate base pairs that link the probes to the flanking forward primer 203A and the reverse primer 203B. During search, these intermediate base pairs can produce false positive and false negative matches. Furthermore, these intermediate base pairs also confine the blueprint search query to a fixed probe location / distance relative to the flanking forward primers and the reverse primers.
[0089] A split search query blueprint 915 does not include the intermediate base pairs between the forward primer 203A and the reverse primer 203B. Therefore, the resulting blueprint search results have blueprints that have fewer false positives and false negatives. Furthermore, the resulting blueprint search results have blueprints with probes that are invariant to their relative distance to the flanking forward primers and the reverse primers.
[0090] The example process 800 may use split search query blueprint 915 instead of (concatenated) PCR blueprints 215. Its blueprint-to-pathogen comparison logic 840 may then stop comparison when a blueprint matches with a non-target organism, or when a blueprint does not match a target organism.
[0091] This approach delivers even higher quality blueprints than other approaches.Mutated Genomes
[0092] FIG. 10 shows an implementation 1000 for the generation of mutated genomes for a target organism genome 1010 that includes sequences identifying a target pathogen. The mutation probability assignment logic 1020 assigns mutation probabilities to different regions of the target genome, for example, based on Bayesian probability. The mutation sampling logic 1030 applies mutations on regions of the target genome based on the mutation probability and thereby generates the mutated genomes of the target organism. These mutated genomes are simulated, i.e., they may not yet have been found in nature or discovered clinically.Mutation-Proof Blueprints
[0093] FIG. 11 shows an implementation 1100 of the generation of blueprints that are robust to potential future genome mutations. The mutated genomes database 1155 stores simulated mutations in the target genomes and non-target genomes such as may be generated by implementation 1000. Blueprint-to-pathogen comparison logic 840 (or 1140) not only compares the blueprints generated in blueprint generation logic 830 with the target and non-target genomes in genomes database 850 (or 1150) but also with the simulated mutated genomes in mutated genomes database 1155.METHOD
[0094] The disclosed technology performs the development of high-quality blueprints for pathogen detection in three stages. These include: (Stage 1) identifying feasible segments in a entire pathogen genome, (Stage 2) generating promising blueprint candidates, and (Stage 3) evaluating and ranking these candidates to find the best ones. In the description below, these example steps are based on MUMmer, Primer3, and BLAST software, but they can be implemented in any software that performs the same tasks. The disclosed technology provides systems and methods to apply existing techniques and tools in a systematic manner to make the blueprint design process efficient and comprehensive. Design of PCR blueprints can be achieved with the example architecture of FIG. 5. Design of CRISPR blueprints can be achieved with the example architecture of FIG. 6. Details of the three stages are presented below:
[0095] Stage 1: Identifying high-specificity areas: Using software such as MUMmer, the genome of the target pathogen, such as E. coli, is paired up with genomes of each of the possible non-target genomes, such as Shigella. MUMmer aligns these genomes, i.e. finds subsequences where they match. The disclosed technology can aggregate the alignments, forming a heatmap that shows where the most promising areas for blueprints are. That is, for each possible high-specificity area (i.e. a subsequence of e.g. 200 to 400 base pairs), the heatmap shows how many mismatches there are. The high-specificity areas that have fewer than a minimum number of mismatches (e.g. 20) are discarded, i.e. not considered further. This process is repeated for each possible non-target genome.
[0096] Similarly, when high-specificity areas are being reviewed that should match multiple similar target genomes, such as mutated variants of a respiratory disease virus, an implementation may compare the genomes to identify possible high-specificity areas that have no or very few mismatches. The identified areas may then be reviewed against those of non-target genomes where the number of mismatches should be as large as possible, or at least more than the minimum number of mismatches, to qualify for further evaluation.
[0097] The majority of the genome can be discarded in this manner. Only the high- specificity areas that remain in the end are passed on to Stage 2, making the blueprint design process much more efficient. The minimum mismatch parameter can be adjusted, resulting in faster results or a more comprehensive search.
[0098] The method in Stage 1 can apply to both PCR and CRISPR.
[0099] FIG. 12 illustrates example pseudocode 1200 for Stage 1.
[0100] Stage 2: Generating amplicon candidates: For a PCR system, each of the candidates identified in Stage 1 is given in turn as input to software such as Primer3, which then identifies possible blueprints. A blueprint includes of a forward primer, a probe, and a reverse primer. In addition, Primer3 returns quality metrics (scores) for each of these choices,which is used to sort the results from most to least promising. Only blueprints that meet a minimum score are kept; others are discarded. This threshold can be adjusted to make the search faster or more comprehensive.
[0101] The score only measures how well the blueprint would work for the target organism; it does not consider possible non-target organisms. The list of these candidates is passed on to Stage 3 for false-positive evaluation.
[0102] In a CRISPR-Cas system, only unique blueprints are selected. They are verified using the internal databases (FIG. 6).
[0103] FIG. 13 illustrates example pseudocode 1300 for Stage 2 for PCR. The code for a CRISPR-Cas system is different, but achieves the same results.
[0104] Stage 3: Evaluating false positives: This logic is the same for PCR and CRISPR- Cas systems. Each candidate amplicon and / or blueprint is evaluated based on how well it matches with other pathogen organisms (or non-target pathogens). Such matches are potential false positives in that the implementation would replicate a non-target pathogen molecule (i.e., DNA or RNA) in addition to the target pathogen. Software such as BLAST can be used for this phase. It returns alignments as well as penalties that quantify how useful the mismatches may be. If penalties are higher than a threshold, the blueprint / amplicon should not be considered: the likelihood of false positives is too high. Otherwise, a combined quality and specificity score is formed (i.e. numerical combination of e.g. Primer3 score and BLAST penalties), and the amplicon is added in the list of solutions, sorted by the combined score.
[0105] FIG. 14 illustrates example pseudocode 1400 for Stage 3.
[0106] Optionally, the resulting solutions can be further improved through various local search methods, i.e. shifting or editing the blueprint sequences slightly, either manually through expert knowledge, or algorithmically. The algorithms may be continually improved using feedback from wet-lab testing (laboratory test apparatus 520 results) to assay design multiplexing application 510 or blueprint and assay design software or logic 610.Other Implementations
[0107] The disclosed technology can be implemented using other techniques such as Bayesian optimization or evolutionary optimization to develop high-quality blueprints for pathogen detection. The following two implementations can also scale up the process.
[0108] In one implementation, Bayesian optimization can be used to identify the most promising high-specificity areas. Initially all high-specificity area locations have equal probability. As blueprints are discovered and evaluated, these associated probabilities are updated through a Gaussian process. The resulting model can then be used to sample the space of high-specificity areas optimally, making it possible to discover high-quality blueprints while evaluating only a fraction of possible high-specificity areas.
[0109] In another implementation, evolutionary optimization is used to design the assay. In this manner, the disclosed technology can be scaled to an even larger context where Bayesian optimization may not be used to cover the entire genome. Each individual target in the population includes start and end locations of the target genome. Each high-specificity area can then be evaluated using various algorithms, resulting in fitness that drives the selection. Variation is created through mutation and crossover of the start and end locations.
[0110] The two approaches (or implementations presented above) can also be combined so that the fitness scores determine the partial Bayesian optimization, which is then used in turn to inform mutation and crossover. In this manner, evolution discovers the broader areas where Bayesian optimization can be focused.
[0111] The comprehensive search method as described above may be sufficient for finding blueprint candidates for a variety of pathogens such as E. coli and shigella, etc. The search can be run in a matter of a hours, which can represent a reasonable computational cost for developing a new assay.
[0112] In another implementation, the disclosed technology can apply artificial intelligence (Al) and / or machine learning techniques or models to scale up the disclosed method. For example, the disclosed technology can scale up the disclosed method in at least two ways: (1) using other genomic search applications that include, e.g., the entire human genome may help in achieving a scale-up of three orders of magnitude. Such applications may be developed based on characterizing genetic disorders. And (2) using applications where the search has to run orders of magnitude more times to get the result.
[0113] In one implementation, the disclosed technology can use a large number of searches to construct assays that are future proof, i.e., such assays are likely to stay effective even when the target pathogen mutates, resulting in new variations. These assays are more reliable and can have a longer shelf life, making it possible to manufacture and distribute them in larger quantities and to broader markets.
[0114] In such an implementation, the technology can comprise two components. The first component can comprise an Al or machine learning model that predicts where the mutations are most likely to occur. The output from such models can be used to build a mutation heatmap. There is considerable knowledge about mutation rates that is already available, and it can be used as a starting point. This model can be further augmented by training Al and / or machine learning models using historical datasets of past mutations.
[0115] The second component can comprise an intelligent search (e.g., performed by the intelligent search engine) that can be extended to take likely mutations into account. This extension may require running the search many times. If the effects of mutations were independent, it would be possible to use the mutation rate heatmap as a surrogate model during search, that is, write an objective function based on probabilities of base-pairs at eachlocation. However, there are interactions between locations that may not be covered by such an objective function. For instance, repetitive DNA sequences do not work as effectively for designing blueprints as nonrepetitive sequences; It is not enough to evaluate statistical properties of possible mutations; they need to be actually created and evaluated comprehensively.
[0116] Once the most likely high-specificity areas are identified using artificial intelligence (Al) or machine learning models, mutated versions of these high-specificity areas can be created by sampling the heatmap model. The possible blueprints for each sample can be found using the intelligent search as described before and Al and / or machine learning models may be used to focus the search so that it can be done in a reasonable amount of time. The necessary compute can also be adjusted by adjusting the amount of sampling and the degree of focus, in other words, by trading accuracy and coverage for compute.
[0117] In one implementation, when generating blueprints that match many mutations, it is possible to take advantage of technology that creates nucleotides at a given location in a given frequency. For instance, a nucleotide "A" can be placed in a location 70% of the time and a nucleotide "G" can be placed in that location 30% of the time. An assay based with such a mix of blueprints can then be effective against the most likely mutations on that location (albeit proportionately less strongly).
[0118] The distribution result of blueprints can then be used to select a set of blueprints that works best with the most likely mutations. The larger the set, the more likely it may cover a future mutation. The complexity of the assay is proportional to the coverage and therefore the robustness of the assay against future mutations.OUTPUT
[0119] The result (or output) from the above process can be a list of promising candidate blueprints for detecting given pathogens, ready for testing in the laboratory. The method is systematic and can be configured to evaluate all base pairs in a target pathogen, therefore performing a comprehensive search. It can also be configured to consider only the most promising alternatives, thus making the search more efficient. The disclosed technology therefore provides a comprehensive and practical approach to designing superior blueprints for pathogen detection.COMPUTER SYSTEM
[0120] FIG. 15 shows an example computer system 1500 that can be used to implement the disclosed technology. FIG. 15 presents a simplified block diagram of a computer 1510, or network node, which can be used to implement the functions of molecular assay development platform 110. Computer system 1500 includes a processor subsystem 1514 which communicates with a number of peripheral devices via bus subsystem 1512. These peripheraldevices may include a storage subsystem 1524, comprising a memory subsystem 1526 and a file storage subsystem 1528, user interface input devices 1522, user interface output devices 1520, and a communication module 1516. The input and output devices allow user interaction with computer 1510. Communication module 1516 provides physical and communication protocol support for interfaces to outside networks, including an interface to network(s) 181, and is coupled via network(s) 181 to corresponding communication modules in other computer systems. Network(s) 181 may comprise many interconnected computer systems and communication links. These communication links may be wireline links, optical links, wireless links, or any other mechanisms for communication of information, but typically it is an IP-based communication network, at least at its extremities. While in one implementation, network(s) 181 includes the Internet, in other implementations, network(s) 181 may be any suitable computer network.
[0121] The physical hardware component of network interfaces is sometimes referred to as network interface cards (NICs), although they need not be in the form of cards: for instance, they could be in the form of integrated circuits (ICs) and connectors fitted directly onto a motherboard, or in the form of macrocells fabricated on a single integrated circuit chip with other components of the computer system.
[0122] User interface input devices 1522 may include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated into the display, audio input devices such as voice recognition systems, microphones, and other types of tangible input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information into the computer system or onto computer network.
[0123] User interface output devices 1520 may include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide non-visual display such as via audio output devices. In general, use of the term "output device" is intended to include all possible types of devices and ways to output information from the computer system to the user or to another machine or computer system.
[0124] Storage subsystem 1524 stores the basic programming and data constructs that provide the functionality of certain implementations of the disclosed technology.
[0125] Storage subsystem 1524, when used for implementation of server nodes, comprises a product including a non-transitory computer readable medium storing a machine- readable data structure including a spatial event map which locates events in a workspace, wherein the spatial event map includes a log of events, entries in the log having a location ofa graphical target of the event in the workspace and a time. Also, storage subsystem 1524 comprises a product including executable instructions for performing the procedures described herein associated with the server node.
[0126] Storage subsystem 1524, when used for implementation of client-nodes, comprises a product including a non-transitory computer readable medium storing a machine readable data structure including a spatial event map in the form of a cached copy as explained below, which locates events in a workspace, wherein the spatial event map includes a log of events, entries in the log having a location of a graphical target of the event in the workspace and a time. Also, storage subsystem 1524 comprises a product including executable instructions for performing the procedures described herein associated with the client node.
[0127] For example, the various modules implementing the functionality of certain implementations of the disclosed technology may be stored in storage subsystem 1524. These software modules are generally executed by processor subsystem 1514.
[0128] Memory subsystem 1526 typically includes a number of memories including a main random-access memory (RAM 1530) for storage of instructions and data during program execution and a read only memory (ROM 1532) in which fixed instructions are stored. File storage subsystem 1528 provides persistent storage for program and data files, and may include a hard disk drive, a floppy disk drive along with associated removable media, a CD ROM drive, an optical drive, or removable media cartridges. The databases and modules implementing the functionality of certain implementations of the disclosed technology may have been provided on a computer readable medium such as one or more CD-ROMs and may be stored by file storage subsystem 1528. The memory subsystem 1526 includes, among other things, computer instructions which, when executed by the processor subsystem 1514, cause the computer system to operate or perform functions as described herein. As used herein, processes and software that are said to run in or on the "host" or the "computer," execute on processor subsystem 1514 in response to computer instructions and data in memory subsystem 1526 including any other local or remote storage for such instructions and data.
[0129] Bus subsystem 1512 provides a mechanism for letting the various components and subsystems of a computer system communicate with each other as intended. Although bus subsystem 1512 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0130] The computer 1510 itself can be of varying types including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, or any other data processing system or user device. In one implementation, a computer system includes several computer systems. Due to the ever-changing nature of computers and networks, the description of computer 1510 depicted in FIG. 15 is intended only as a specific example for purposes of illustrating the preferred implementations of the disclosed technology. Many other configurations of the computer system are possible having more or less components than the computer system depicted in FIG. 15. The same components and variations can also make up each of the other devices and / or engines and the databases in the environment of FIG. 1, as well as the devices and / or engines and the databases as shown in FIGS. 5-6.PARTICULAR IMPLEMENTATIONS
[0131] Described implementations of the subject matter can include one or more features, alone or in combination, as described in the following clauses.Clause 1. A method of systematically determining a blueprint comprising three sequences of nucleotides for a subsequence of nucleotide bases representing a genome of a target pathogen, including: identifying one or more feasible subsequences of nucleotide bases in the subsequence of nucleotide bases representing the genome of the target pathogen; identifying, for the one or more feasible subsequences, a plurality of amplicons wherein an amplicon is a portion of a feasible subsequence of the plurality of feasible subsequences; matching the respective amplicons with at least one non-target pathogen, and assigning a specificity score to the respective amplicons based on mismatches of nucleotides of the respective amplicons with nucleotides of the at least one non-target pathogen; selecting an amplicon with a respective specificity score above a minimum-specificity threshold; identifying the blueprint for the selected amplicon wherein the blueprint comprises three sequences of nucleotide bases that are used for generating the selected amplicon; and storing the blueprint for the selected amplicon for use in detecting the target pathogen in a biological sample.Clause 2. The method of clause 1, further including determining a respective quality score for the selected amplicon.Clause 3. The method of clause 1 or clause 2, wherein identifying the blueprint further includes determining, for the selected amplicon, at least three sequences of nucleotide bases that correspond to three associated sequences of nucleotide bases in the amplicon.Clause 4. The method of any of the clauses 1 to 3, wherein the at least three sequences of nucleotide bases include a forward primer, a reverse primer, and a probe.Clause 5. The method of any of the clauses 1 to 4, wherein the identifying the one or more feasible subsequences of nucleotide bases in a sequence of nucleotide bases further includes: repeatedly matching the subsequence of nucleotide bases representing the genome of the target pathogen with a second subsequence of nucleotide bases representing a genome of a non-target pathogen, wherein each matching is performed with a different alignment of the second subsequence with the subsequence of nucleotide bases representing the genome of the target pathogen; identifying, for each alignment, subsequences in the second sequence of nucleotide bases that match with subsequences in the sequence of nucleotide bases representing the genome of the target pathogen; aggregating the matched subsequences in the second subsequence of nucleotide bases; generating a mismatch heatmap identifying portions of the sequence of nucleotide bases representing the genome of the target pathogen that do not match with the subsequences in the second subsequence; and selecting subsequences in the sequence of nucleotide bases representing the genome of the target pathogen with respective mismatches greater than a minimum-mismatch threshold for use as amplicons.Clause 6. The method of any of the clauses 1 to 5, wherein the mismatch heatmap identifies a number of mismatches for a subsequence having a length between 80 base pairs to 1000 base pairs.Clause 7. The method of any of the clauses 1 to 6, further including: using Bayesian optimization to identify subsequences in the sequence of nucleotide bases; and using the identified subsequences for identification of amplicons.Clause 8. The method of any of the clauses 1 to 7, further including using evolutionary optimization to determine respective fitness scores for the one or more subsequences in the sequence of nucleotide bases representing the genome of the target pathogen and selecting subsequences in dependence on their respective fitness scores.Clause 9. The method of any of the clauses 1 to 8, further including:predicting, using a trained machine learning model, mutations in the plurality of amplicons, wherein the predicted mutations in the plurality of amplicons are represented as a mutation heatmap and wherein the trained machine learning model is trained using historical datasets of past mutations; generating a plurality of mutations to the subsequence of nucleotide bases representing the genome of the target pathogen; determining a plurality of amplicons each including at least a portion of the plurality of mutations; wherein amplicons in the plurality of amplicons can include proportions of possible alternative nucleotides at specific locations; and storing the plurality of amplicons as a future-proof assay for detection of variants of the target pathogen created by mutations in the target pathogen.Clause 10. A system including one or more processors coupled to memory, the memory loaded with computer instructions to systematically determine a blueprint comprising three sequences of nucleotides for a subsequence of nucleotide bases representing a genome of a target pathogen, the instructions, when executed on the processors, implementing actions comprising: identifying one or more feasible subsequences of nucleotide bases in the subsequence of nucleotide bases representing the genome of the target pathogen; identifying, for the one or more feasible subsequences, a plurality of amplicons wherein an amplicon is a portion of a feasible subsequence of the plurality of feasible subsequences; matching the respective amplicons with at least one non-target pathogen, and assigning a specificity score to the respective amplicons based on mismatches of nucleotides of the respective amplicons with nucleotides of the at least one non-target pathogen; selecting an amplicon with a respective specificity score above a minimum-specificity threshold; identifying the blueprint for the selected amplicon wherein the blueprint comprises three sequences of nucleotide bases that are used for generating the selected amplicon; and storing the blueprint for the selected amplicon for use in detecting the target pathogen in a biological sample.Clause 11. The system of clause 10, wherein identifying of the blueprint further implements actions comprising determining, for the selected amplicon, at least three sequences of nucleotide bases that correspond to three associated sequences of nucleotide bases.Clause 12. The system of clause 10 or clause 11, wherein the at least three sequences of nucleotide bases include a forward primer, a reverse primer, and a probe.Clause 13. The system of any of the clauses 10 to 12, wherein the identifying the one or more feasible subsequences of nucleotide bases in a sequence of nucleotide bases further implements actions comprising: repeatedly matching the subsequence of nucleotide bases representing the genome of the target pathogen with a second subsequence of nucleotide bases representing a genome of a non-target pathogen, wherein each matching is performed with a different alignment of the second subsequence with the subsequence of nucleotide bases representing the genome of the target pathogen; identifying, for each alignment, subsequences in the second sequence of nucleotide bases that match with subsequences in the sequence of nucleotide bases representing the genome of the target pathogen; aggregating the matched subsequences in the second subsequence of nucleotide bases; generating a mismatch heatmap identifying portions of the sequence of nucleotide bases representing the genome of the target pathogen that do not match with the subsequences in the second subsequence; and selecting subsequences in the sequence of nucleotide bases representing the genome of the target pathogen with respective mismatches greater than a minimum-mismatch threshold for use as amplicons.Clause 14. The system of any of the clauses 10 to 13, further implementing actions comprising: using Bayesian optimization to identify subsequences in the sequence of nucleotide bases; and using the identified subsequences for identification of amplicons.Clause 15. The system of any of the clauses 10 to 14, further implementing actions comprising using evolutionary optimization to determine respective fitness scores for the one or more subsequences in the sequence of nucleotide bases representing the genome of the target pathogen and selecting subsequences in dependence on their respective fitness scores.Clause 16. A non-transitory computer-readable storage medium storing computer program instructions to systematically determine a blueprint comprising three sequences of nucleotides for a subsequence of nucleotide bases representing a genome of a target pathogen, whereinthe computer program instructions, when executed on a processor, implement actions comprising: identifying one or more feasible subsequences of nucleotide bases in the subsequence of nucleotide bases representing the genome of the target pathogen; identifying, for the one or more feasible subsequences, a plurality of amplicons wherein an amplicon is a portion of a feasible subsequence of the plurality of feasible subsequences; matching the respective amplicons with at least one non-target pathogen, and assigning a specificity score to the respective amplicons based on mismatches of nucleotides of the respective amplicons with nucleotides of the at least one non-target pathogen; selecting an amplicon with a respective specificity score above a minimum-specificity threshold; identifying the blueprint for the selected amplicon wherein the blueprint comprises three sequences of nucleotide bases that are used for generating the selected amplicon; and storing the blueprint for the selected amplicon for use in detecting the target pathogen in a biological sample.Clause 17. The non-transitory computer-readable storage medium of clause 16, wherein identifying the blueprint further includes determining, for the selected amplicon, at least three sequences of nucleotide bases that correspond to three associated sequences of nucleotide bases in the amplicon.Clause 18. The non-transitory computer-readable storage medium of clause 16 or clause 17, wherein the at least three sequences of nucleotide bases include a forward primer, a reverse primer, and a probe.Clause 19. The non-transitory computer-readable storage medium of any of the clauses 16 to 18, wherein identifying the one or more feasible subsequences of nucleotide bases in a sequence of nucleotide bases further includes: repeatedly matching the subsequence of nucleotide bases representing the genome of the target pathogen with a second subsequence of nucleotide bases representing a genome of a non-target pathogen, wherein each matching is performed with a different alignment of the second subsequence with the subsequence of nucleotide bases representing the genome of the target pathogen; identifying, for each alignment, subsequences in the second sequence of nucleotide bases that match with subsequences in the sequence of nucleotide bases representing the genome of the target pathogen;aggregating the matched subsequences in the second subsequence of nucleotide bases; generating a mismatch heatmap identifying portions of the sequence of nucleotide bases representing the genome of the target pathogen that do not match with the subsequences in the second subsequence; and selecting subsequences in the sequence of nucleotide bases representing the genome of the target pathogen with respective mismatches greater than a minimum-mismatch threshold for use as amplicons.Clause 20. The non-transitory computer-readable storage medium of any of the clauses 16 to 19, further comprising: using Bayesian optimization to identify subsequences in the sequence of nucleotide bases; and using the identified subsequences for identification of amplicons.Clause 21. Logic configured to execute functionality comprising: comparing a target genome with one or more non-target genomes to generate a heatmap of mismatches for each of the one or more non-target genomes; on the target genome, identifying one or more high-specificity areas, wherein a high- specificity area is a sequence of nucleotides with a density of mismatches higher than a minimum mismatch density for all non-target genomes; using a sliding window to identify, in one or more of the one or more high-specificity areas, multiple candidate blueprints; for each of the multiple candidate blueprints, designing an associated candidate assay; determining a quality score for each candidate blueprint and each associated candidate assay; and ranking the multiple candidate blueprints based on their quality scores.Clause 22. The logic of clause 21, wherein using a sliding window to identify multiple candidate blueprints comprises: determining if a nucleotide sequence in the sliding window includes an occurrence of a protospacer adjacent motif (PAM) associated with a Cas enzyme; determining if a candidate protospacer in a nucleotide sequence following the PAM does not include another occurrence of the PAM; in response to determining that the candidate protospacer in the nucleotide sequence following the PAM does include another occurrence of the PAM, discarding the occurrence ofthe PAM and the candidate protospacer; and in response to determining that the candidate protospacer in the nucleotide sequence following the PAM does not include another occurrence of the PAM, adding the PAM and the candidate protospacer to a list of candidate blueprints.Clause 23. The logic of clause 21 or clause 22, wherein designing an associated candidate assay includes generating a ribonucleoprotein (RNP) including the Cas enzyme and guide RIMA based on the candidate protospacer.Clause 24. The logic of any of the clauses 21 to 23, wherein designing an associated candidate assay includes generating a reporter DNA sequence of a format / FAM / Cn / FQ, wherein FAM indicates a presence of a fluorescent dye, Cn is a sequence of n cytosine nucleobases with n in a range of 3 to 25, and FQ signifies a presence of a quencher molecule.Clause 25. The logic of any of the clauses 21 to 24, wherein the quencher molecule is biotin.Clause 26. The logic of any of the clauses 21 to 25, wherein using a sliding window to identify multiple candidate blueprints further comprises: determining if a nucleotide sequence of 25 to 250 base pairs preceding the PAM and a nucleotide sequence of 25 to 250 base pairs following the candidate protospacer include areas suitable as a forward primer and a reverse primer, respectively, for recombinase polymerase amplification (RPA).Clause 27. A system for systematic design of blueprints for pathogen detection, comprising: blueprint and assay design logic; logic to remove low-quality molecular data from a molecular dataset to obtain curated molecular data; a database to store the curated molecular data; a molecular knowledge database to store and retrieve information about molecules processed by the blueprint and assay design logic; and an assays database to store candidate assays; wherein the blueprint and assay design logic is configured to: compare a target genome with one or more non-target genomes to generate a heatmap of mismatches for each of the one or more non-target genomes;on the target genome, identify one or more high-specificity areas, wherein a high- specificity area is a sequence of nucleotides with a density of mismatches higher than a minimum mismatch density for all non-target genomes; use a sliding window to identify, in one or more of the one or more high-specificity areas, multiple candidate blueprints; for each of the multiple candidate blueprints, design an associated candidate assay; determine a quality score for each candidate blueprint and each associated candidate assay; and rank the multiple candidate blueprints based on their quality scores.Clause 28. The system of clause 27, wherein using a sliding window to identify multiple candidate blueprints comprises: determining if a nucleotide sequence in the sliding window includes an occurrence of a protospacer adjacent motif (PAM) associated with a Cas enzyme; determining if a candidate protospacer in a nucleotide sequence following the PAM does not include another occurrence of the PAM; in response to determining that the candidate protospacer in the nucleotide sequence following the PAM does include another occurrence of the PAM, discarding the occurrence of the PAM and the candidate protospacer; and in response to determining that the candidate protospacer in the nucleotide sequence following the PAM does not include another occurrence of the PAM, adding the PAM and the candidate protospacer to a list of candidate blueprints.Clause 29. The system of clause 27 or clause 28, wherein designing an associated candidate assay includes generating a ribonucleoprotein (RNP) including the Cas enzyme and guide RIMA based on the candidate protospacer.Clause 30. The system of any of the clauses 27 to 29, wherein designing an associated candidate assay includes generating a reporter DNA sequence of a format / FAM / Cn / FQ, wherein FAM indicates a presence of a fluorescent dye, Cn is a sequence of n cytosine nucleobases with n in a range of 3 to 25, and FQ signifies a presence of a quencher molecule.Clause 31. The system of any of the clauses 27 to 30, wherein the quencher molecule is biotin.Clause 32. The system of any of the clauses 27 to 31, wherein using a sliding window to identify multiple candidate blueprints further comprises: determining if a nucleotide sequence of 25 to 250 base pairs preceding the PAM and a nucleotide sequence of 25 to 250 base pairs following the candidate protospacer include areas suitable as a forward primer and a reverse primer, respectively, for recombinase polymerase amplification (RPA).Clause 33. A non-transitory computer-readable storage medium storing computer program instructions to design blueprints for pathogen detection, wherein the computer program instructions, when executed on a processor, implement actions comprising: comparing a target genome with one or more non-target genomes to generate a heatmap of mismatches for each of the one or more non-target genomes; on the target genome, identifying one or more high-specificity areas, wherein a high- specificity area is a sequence of nucleotides with a density of mismatches higher than a minimum mismatch density for all non-target genomes; using a sliding window to identify, in one or more of the one or more high-specificity areas, multiple candidate blueprints; for each of the multiple candidate blueprints, designing an associated candidate assay; determining a quality score for each candidate blueprint and each associated candidate assay; and ranking the multiple candidate blueprints based on their quality scores.Clause 34. The non-transitory computer-readable storage medium of clause 33, wherein using a sliding window to identify multiple candidate blueprints comprises: determining if a nucleotide sequence in the sliding window includes an occurrence of a protospacer adjacent motif (PAM) associated with a Cas enzyme; determining if a candidate protospacer in a nucleotide sequence following the PAM does not include another occurrence of the PAM; in response to determining that the candidate protospacer in the nucleotide sequence following the PAM does include another occurrence of the PAM, discarding the occurrence of the PAM and the candidate protospacer; and in response to determining that the candidate protospacer in the nucleotide sequence following the PAM does not include another occurrence of the PAM, adding the PAM and the candidate protospacer to a list of candidate blueprints.Clause 35. The non-transitory computer-readable storage medium of clause 33 or clause 34, wherein designing an associated candidate assay includes generating a ribonucleoprotein (RNP) including the Cas enzyme and guide RIMA based on the candidate protospacer.Clause 36. The non-transitory computer-readable storage medium of any of the clauses 33 to 35, wherein designing an associated candidate assay includes generating a reporter DNA sequence of a format / FAM / Cn / FQ, wherein FAM indicates a presence of a fluorescent dye, Cn is a sequence of n cytosine nucleobases with n in a range of 3 to 25, and FQ signifies a presence of a quencher molecule.Clause 37. The non-transitory computer-readable storage medium of any of the clauses 33 to 36, wherein the quencher molecule is biotin.Clause 38. The non-transitory computer-readable storage medium of any of the clauses 33 to 37, wherein using a sliding window to identify multiple candidate blueprints further comprises: determining if a nucleotide sequence of 25 to 250 base pairs preceding the PAM and a nucleotide sequence of 25 to 250 base pairs following the candidate protospacer include areas suitable as a forward primer and a reverse primer, respectively, for recombinase polymerase amplification (RPA).CONSIDERATIONS
[0132] The applicant hereby discloses in isolation each individual feature described herein and any combination of two or more such features, to the extent that such features or combinations are capable of being carried out based on the present specification as a whole in light of the common general knowledge of a person skilled in the art, irrespective of whether such features or combinations of features solve any problems disclosed herein, and without limitation to the scope of the claims. The applicant indicates that aspects of the present technology may include of any such feature or combination of features. In view of the foregoing description, it will be evident to a person skilled in the art that various modifications may be made within the scope of the technology.
[0133] The foregoing description of preferred implementations of the present technology has been provided for the purposes of illustration and description. It is not intended to be exhaustive or to limit the technology to the precise forms disclosed. Obviously, many modifications and variations will be apparent to practitioners skilled in this art. In particular, and without limitation, any and all variations described, suggested by the Background section of this patent application or by the material incorporated by reference are specifically incorporated by reference into the description herein of implementations of the technology.In addition, any and all variations described, suggested or incorporated by reference herein with respect to any one implementation are also to be considered taught with respect to all other implementations. The implementations described herein were chosen and described in order to best explain the principles of the technology and its practical application, thereby enabling others skilled in the art to understand the technology for various implementations and with various modifications as are suited to the particular use contemplated. It is intended that the scope of the technology be defined by the following claims and their equivalents.
[0134] All features disclosed in the specification, including the claims, abstract, and drawings, and all the steps in any method or process disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. Each feature disclosed in the specification, including the claims, abstract, and drawings, can be replaced by alternative features serving the same, equivalent, or similar purpose, unless expressly stated otherwise.
[0135] Although the description has been described with respect to specific implementations thereof, these specific implementations are merely illustrative, and not restrictive. For instance, many of the operations can be implemented on a printed circuit board (PCB) using off-the-shelf devices, in a System-on-Chip (SoC), application-specific integrated circuit (ASIC), programmable processor, a coarse-grained reconfigurable architecture (CGRA), or in a programmable logic device such as a field-programmable gate array (FPGA), obviating the need for at least part of any dedicated hardware. Implementations may be as a single chip, or as a multi-chip module (MCM) packaging multiple semiconductor dies in a single package. All such variations and modifications are to be considered within the ambit of the disclosed technology the nature of which is to be determined from the foregoing description.
[0136] Any suitable programming language can be used to implement the routines of specific implementations including C, C++, Java, JavaScript, Python, compiled languages, interpreted languages and scripts, assembly language, machine language, etc. Different programming techniques can be employed such as procedural or object oriented. Methods embodied in routines can execute on a single processor device or on a multiple processor system. Although the steps, operations, or computations may be presented in a specific order, this order may be changed in different specific implementations. In some specific implementations, multiple steps shown as sequential in this specification can be performed at the same time.
[0137] Specific implementations may be implemented in a tangible, non-transitory computer-readable storage medium for use by or in connection with the instruction execution system, apparatus, board, or device. Specific implementations can be implemented in the form of control logic in software or hardware or a combination of both. The control logic, when executed by one or more processors, may be operable to perform that which is described inspecific implementations. For example, a tangible non-transitory medium such as a hardware storage device can be used to store the control logic, which can include executable instructions.
[0138] One or more implementations of the technology or elements thereof can be implemented in the form of a computer product, including a non-transitory computer-readable storage medium with computer usable program code for performing any indicated method steps and / or any configuration file for one or more processors to execute a high-level program. Furthermore, one or more implementations of the technology or elements thereof can be implemented in the form of an apparatus including a memory and at least one processor that is coupled to the memory and operative to perform exemplary method steps, and / or a processor that is operative to execute a high-level program based on a configuration file. Yet further, in another aspect, one or more implementations of the technology or elements thereof can be implemented in the form of means for carrying out one or more of the method steps described herein and / or executing a high-level program described herein. Such means can include (i) hardware module(s); (ii) software module(s) executing on one or more hardware processors; (iii) bit files for configuration of a processor array; or (iv) a combination of aforementioned items.
[0139] It will also be appreciated that one or more of the elements depicted in the drawings / figures can also be implemented in a more separated or integrated manner, or even removed or rendered as inoperable in certain cases, as is useful in accordance with a particular application.
[0140] Thus, while specific implementations have been described herein, latitudes of modification, various changes, and substitutions are intended in the foregoing disclosures, and it will be appreciated that in some instances some features of specific implementations will be employed without a corresponding use of other features without departing from the scope and spirit as set forth. Therefore, many modifications may be made to adapt a particular situation or material to the essential scope and spirit.
Claims
CLAIMS1. A method of systematically determining a blueprint comprising three sequences of nucleotides for a subsequence of nucleotide bases representing a genome of a target pathogen, including: identifying one or more feasible subsequences of nucleotide bases in the subsequence of nucleotide bases representing the genome of the target pathogen; identifying, for the one or more feasible subsequences, a plurality of amplicons wherein an amplicon is a portion of a feasible subsequence of the plurality of feasible subsequences; matching the respective amplicons with at least one non-target pathogen, and assigning a specificity score to the respective amplicons based on mismatches of nucleotides of the respective amplicons with nucleotides of the at least one non-target pathogen; selecting an amplicon with a respective specificity score above a minimum-specificity threshold; identifying the blueprint for the selected amplicon wherein the blueprint comprises three sequences of nucleotide bases that are used for generating the selected amplicon; and storing the blueprint for the selected amplicon for use in detecting the target pathogen in a biological sample.
2. The method of claim 1, further including determining a respective quality score for the selected amplicon.
3. The method of claim 1, wherein identifying the blueprint further includes determining, for the selected amplicon, at least three sequences of nucleotide bases that correspond to three associated sequences of nucleotide bases in the amplicon.
4. The method of claim 3, wherein the at least three sequences of nucleotide bases include a forward primer, a reverse primer, and a probe.
5. The method of claim 1, wherein the identifying the one or more feasible subsequences of nucleotide bases in a sequence of nucleotide bases further includes: repeatedly matching the subsequence of nucleotide bases representing the genome of the target pathogen with a second subsequence of nucleotide bases representing a genome of a non-target pathogen, wherein each matching is performed with a different alignment of the second subsequence with the subsequence of nucleotide bases representing the genomeof the target pathogen; identifying, for each alignment, subsequences in the second sequence of nucleotide bases that match with subsequences in the sequence of nucleotide bases representing the genome of the target pathogen; aggregating the matched subsequences in the second subsequence of nucleotide bases; generating a mismatch heatmap identifying portions of the sequence of nucleotide bases representing the genome of the target pathogen that do not match with the subsequences in the second subsequence; and selecting subsequences in the sequence of nucleotide bases representing the genome of the target pathogen with respective mismatches greater than a minimum-mismatch threshold for use as amplicons.
6. The method of claim 5, wherein the mismatch heatmap identifies a number of mismatches for a subsequence having a length between 80 base pairs to 1000 base pairs.
7. The method of claim 1, further including: using Bayesian optimization to identify subsequences in the sequence of nucleotide bases; and using the identified subsequences for identification of amplicons.
8. The method of claim 1, further including using evolutionary optimization to determine respective fitness scores for the one or more subsequences in the sequence of nucleotide bases representing the genome of the target pathogen and selecting subsequences in dependence on their respective fitness scores.
9. The method of claim 1, further including: predicting, using a trained machine learning model, mutations in the plurality of amplicons, wherein the predicted mutations in the plurality of amplicons are represented as a mutation heatmap and wherein the trained machine learning model is trained using historical datasets of past mutations; generating a plurality of mutations to the subsequence of nucleotide bases representing the genome of the target pathogen; determining a plurality of amplicons each including at least a portion of the plurality of mutations; wherein amplicons in the plurality of amplicons can include proportions of possible alternative nucleotides at specific locations; andstoring the plurality of amplicons as a future-proof assay for detection of variants of the target pathogen created by mutations in the target pathogen.
10. A system including one or more processors coupled to memory, the memory loaded with computer instructions to systematically determine a blueprint comprising three sequences of nucleotides for a subsequence of nucleotide bases representing a genome of a target pathogen, the instructions, when executed on the processors, implementing actions comprising: identifying one or more feasible subsequences of nucleotide bases in the subsequence of nucleotide bases representing the genome of the target pathogen; identifying, for the one or more feasible subsequences, a plurality of amplicons wherein an amplicon is a portion of a feasible subsequence of the plurality of feasible subsequences; matching the respective amplicons with at least one non-target pathogen, and assigning a specificity score to the respective amplicons based on mismatches of nucleotides of the respective amplicons with nucleotides of the at least one non-target pathogen; selecting an amplicon with a respective specificity score above a minimum-specificity threshold; identifying the blueprint for the selected amplicon wherein the blueprint comprises three sequences of nucleotide bases that are used for generating the selected amplicon; and storing the blueprint for the selected amplicon for use in detecting the target pathogen in a biological sample.
11. The system of claim 10, wherein identifying of the blueprint further implements actions comprising determining, for the selected amplicon, at least three sequences of nucleotide bases that correspond to three associated sequences of nucleotide bases.
12. The system of claim 11, wherein the at least three sequences of nucleotide bases include a forward primer, a reverse primer, and a probe.
13. The system of claim 10, wherein the identifying the one or more feasible subsequences of nucleotide bases in a sequence of nucleotide bases further implements actions comprising: repeatedly matching the subsequence of nucleotide bases representing the genome of the target pathogen with a second subsequence of nucleotide bases representing a genome of a non-target pathogen, wherein each matching is performed with a different alignment of the second subsequence with the subsequence of nucleotide bases representing the genome of the target pathogen;identifying, for each alignment, subsequences in the second sequence of nucleotide bases that match with subsequences in the sequence of nucleotide bases representing the genome of the target pathogen; aggregating the matched subsequences in the second subsequence of nucleotide bases; generating a mismatch heatmap identifying portions of the sequence of nucleotide bases representing the genome of the target pathogen that do not match with the subsequences in the second subsequence; and selecting subsequences in the sequence of nucleotide bases representing the genome of the target pathogen with respective mismatches greater than a minimum-mismatch threshold for use as amplicons.
14. The system of claim 10, further implementing actions comprising: using Bayesian optimization to identify subsequences in the sequence of nucleotide bases; and using the identified subsequences for identification of amplicons.
15. The system of claim 10, further implementing actions comprising using evolutionary optimization to determine respective fitness scores for the one or more subsequences in the sequence of nucleotide bases representing the genome of the target pathogen and selecting subsequences in dependence on their respective fitness scores.
16. A non-transitory computer-readable storage medium storing computer program instructions to systematically determine a blueprint comprising three sequences of nucleotides for a subsequence of nucleotide bases representing a genome of a target pathogen, wherein the computer program instructions, when executed on a processor, implement actions comprising: identifying one or more feasible subsequences of nucleotide bases in the subsequence of nucleotide bases representing the genome of the target pathogen; identifying, for the one or more feasible subsequences, a plurality of amplicons wherein an amplicon is a portion of a feasible subsequence of the plurality of feasible subsequences; matching the respective amplicons with at least one non-target pathogen, and assigning a specificity score to the respective amplicons based on mismatches of nucleotides of the respective amplicons with nucleotides of the at least one non-target pathogen; selecting an amplicon with a respective specificity score above a minimum-specificity threshold; identifying the blueprint for the selected amplicon wherein the blueprint comprises threesequences of nucleotide bases that are used for generating the selected amplicon; and storing the blueprint for the selected amplicon for use in detecting the target pathogen in a biological sample.
17. The non-transitory computer-readable storage medium of claim 16, wherein identifying the blueprint further includes determining, for the selected amplicon, at least three sequences of nucleotide bases that correspond to three associated sequences of nucleotide bases in the amplicon.
18. The non-transitory computer-readable storage medium of claim 17, wherein the at least three sequences of nucleotide bases include a forward primer, a reverse primer, and a probe.
19. The non-transitory computer-readable storage medium of claim 16, wherein identifying the one or more feasible subsequences of nucleotide bases in a sequence of nucleotide bases further includes: repeatedly matching the subsequence of nucleotide bases representing the genome of the target pathogen with a second subsequence of nucleotide bases representing a genome of a non-target pathogen, wherein each matching is performed with a different alignment of the second subsequence with the subsequence of nucleotide bases representing the genome of the target pathogen; identifying, for each alignment, subsequences in the second sequence of nucleotide bases that match with subsequences in the sequence of nucleotide bases representing the genome of the target pathogen; aggregating the matched subsequences in the second subsequence of nucleotide bases; generating a mismatch heatmap identifying portions of the sequence of nucleotide bases representing the genome of the target pathogen that do not match with the subsequences in the second subsequence; and selecting subsequences in the sequence of nucleotide bases representing the genome of the target pathogen with respective mismatches greater than a minimum-mismatch threshold for use as amplicons.
20. The non-transitory computer-readable storage medium of claim 16, further comprising: using Bayesian optimization to identify subsequences in the sequence of nucleotide bases; andusing the identified subsequences for identification of amplicons.
21. Logic configured to execute functionality comprising: comparing a target genome with one or more non-target genomes to generate a heatmap of mismatches for each of the one or more non-target genomes; on the target genome, identifying one or more high-specificity areas, wherein a high- specificity area is a sequence of nucleotides with a density of mismatches higher than a minimum mismatch density for all non-target genomes; using a sliding window to identify, in one or more of the one or more high-specificity areas, multiple candidate blueprints; for each of the multiple candidate blueprints, designing an associated candidate assay; determining a quality score for each candidate blueprint and each associated candidate assay; and ranking the multiple candidate blueprints based on their quality scores.
22. The logic of claim 21, wherein using a sliding window to identify multiple candidate blueprints comprises: determining if a nucleotide sequence in the sliding window includes an occurrence of a protospacer adjacent motif (PAM) associated with a Cas enzyme; determining if a candidate protospacer in a nucleotide sequence following the PAM does not include another occurrence of the PAM; in response to determining that the candidate protospacer in the nucleotide sequence following the PAM does include another occurrence of the PAM, discarding the occurrence of the PAM and the candidate protospacer; and in response to determining that the candidate protospacer in the nucleotide sequence following the PAM does not include another occurrence of the PAM, adding the PAM and the candidate protospacer to a list of candidate blueprints.
23. The logic of claim 22, wherein designing an associated candidate assay includes generating a ribonucleoprotein (RNP) including the Cas enzyme and guide RIMA based on the candidate protospacer.
24. The logic of claim 22, wherein designing an associated candidate assay includes generating a reporter DNA sequence of a format / FAM / Cn / FQ, wherein FAM indicates a presence of a fluorescent dye, Cn is a sequence of n cytosine nucleobases with n in a range of 3 to 25, and FQ signifies a presence of a quencher molecule.
25. The logic of claim 24, wherein the quencher molecule is biotin.
26. The logic of claim 22, wherein using a sliding window to identify multiple candidate blueprints further comprises: determining if a nucleotide sequence of 25 to 250 base pairs preceding the PAM and a nucleotide sequence of 25 to 250 base pairs following the candidate protospacer include areas suitable as a forward primer and a reverse primer, respectively, for recombinase polymerase amplification (RPA).
27. A system for systematic design of blueprints for pathogen detection, comprising: blueprint and assay design logic; logic to remove low-quality molecular data from a molecular dataset to obtain curated molecular data; a database to store the curated molecular data; a molecular knowledge database to store and retrieve information about molecules processed by the blueprint and assay design logic; and an assays database to store candidate assays; wherein the blueprint and assay design logic is configured to: compare a target genome with one or more non-target genomes to generate a heatmap of mismatches for each of the one or more non-target genomes; on the target genome, identify one or more high-specificity areas, wherein a high- specificity area is a sequence of nucleotides with a density of mismatches higher than a minimum mismatch density for all non-target genomes; use a sliding window to identify, in one or more of the one or more high-specificity areas, multiple candidate blueprints; for each of the multiple candidate blueprints, design an associated candidate assay; determine a quality score for each candidate blueprint and each associated candidate assay; and rank the multiple candidate blueprints based on their quality scores.
28. The system of claim 27, wherein using a sliding window to identify multiple candidate blueprints comprises: determining if a nucleotide sequence in the sliding window includes an occurrence of aprotospacer adjacent motif (PAM) associated with a Cas enzyme; determining if a candidate protospacer in a nucleotide sequence following the PAM does not include another occurrence of the PAM; in response to determining that the candidate protospacer in the nucleotide sequence following the PAM does include another occurrence of the PAM, discarding the occurrence of the PAM and the candidate protospacer; and in response to determining that the candidate protospacer in the nucleotide sequence following the PAM does not include another occurrence of the PAM, adding the PAM and the candidate protospacer to a list of candidate blueprints.
29. The system of claim 28, wherein designing an associated candidate assay includes generating a ribonucleoprotein (RNP) including the Cas enzyme and guide RIMA based on the candidate protospacer.
30. The system of claim 28, wherein designing an associated candidate assay includes generating a reporter DNA sequence of a format / FAM / Cn / FQ, wherein FAM indicates a presence of a fluorescent dye, Cn is a sequence of n cytosine nucleobases with n in a range of 3 to 25, and FQ signifies a presence of a quencher molecule.
31. The system of claim 30, wherein the quencher molecule is biotin.
32. The system of claim 28, wherein using a sliding window to identify multiple candidate blueprints further comprises: determining if a nucleotide sequence of 25 to 250 base pairs preceding the PAM and a nucleotide sequence of 25 to 250 base pairs following the candidate protospacer include areas suitable as a forward primer and a reverse primer, respectively, for recombinase polymerase amplification (RPA).
33. A non-transitory computer-readable storage medium storing computer program instructions to design blueprints for pathogen detection, wherein the computer program instructions, when executed on a processor, implement actions comprising: comparing a target genome with one or more non-target genomes to generate a heatmap of mismatches for each of the one or more non-target genomes; on the target genome, identifying one or more high-specificity areas, wherein a high- specificity area is a sequence of nucleotides with a density of mismatches higher than a minimum mismatch density for all non-target genomes; using a sliding window to identify, in one or more of the one or more high-specificityareas, multiple candidate blueprints; for each of the multiple candidate blueprints, designing an associated candidate assay; determining a quality score for each candidate blueprint and each associated candidate assay; and ranking the multiple candidate blueprints based on their quality scores.
34. The non-transitory computer-readable storage medium of claim 33, wherein using a sliding window to identify multiple candidate blueprints comprises: determining if a nucleotide sequence in the sliding window includes an occurrence of a protospacer adjacent motif (PAM) associated with a Cas enzyme; determining if a candidate protospacer in a nucleotide sequence following the PAM does not include another occurrence of the PAM; in response to determining that the candidate protospacer in the nucleotide sequence following the PAM does include another occurrence of the PAM, discarding the occurrence of the PAM and the candidate protospacer; and in response to determining that the candidate protospacer in the nucleotide sequence following the PAM does not include another occurrence of the PAM, adding the PAM and the candidate protospacer to a list of candidate blueprints.
35. The non-transitory computer-readable storage medium of claim 34, wherein designing an associated candidate assay includes generating a ribonucleoprotein (RNP) including the Cas enzyme and guide RNA based on the candidate protospacer.
36. The non-transitory computer-readable storage medium of claim 34, wherein designing an associated candidate assay includes generating a reporter DNA sequence of a format / FAM / Cn / FQ, wherein FAM indicates a presence of a fluorescent dye, Cn is a sequence of n cytosine nucleobases with n in a range of 3 to 25, and FQ signifies a presence of a quencher molecule.
37. The non-transitory computer-readable storage medium of claim 36, wherein the quencher molecule is biotin.
38. The non-transitory computer-readable storage medium of claim 34, wherein using a sliding window to identify multiple candidate blueprints further comprises: determining if a nucleotide sequence of 25 to 250 base pairs preceding the PAM and a nucleotide sequence of 25 to 250 base pairs following the candidate protospacer include areassuitable as a forward primer and a reverse primer, respectively, for recombinase polymerase amplification (RPA).
Citation Information
Patent Citations
Method for designing primers for multiplex PCR
US20190221287A1
Methods and compositions for amplicon concatenation
US20210189384A1