Method for functionally screening biological sequence fragments
The method addresses the challenge of false positives in nucleic acid screening by fragmenting sequences and using a test database with cryptographic hash functions to ensure accurate detection and prevention of harmful sequences, enabling automated screening.
Patent Information
- Application Number
- JP2022544739
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-01-23
- Filing Date
- 2021-01-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-01-23
AI Technical Summary
Current screening methods for nucleic acid synthesis orders rely on similarity search algorithms that cannot effectively identify small pieces of nucleic acid capable of being assembled into harmful sequences, leading to false positives and requiring expert curation, making automated screening impossible.
A method involving preselecting a biomolecule, fragmenting its sequence into predetermined lengths, and detecting matches within a test sequence database to evaluate and prevent the synthesis of sequences with unrelated functions, using cryptographic hash functions to protect sequence identities and ensure accurate detection.
Minimizes false positives by reliably identifying functional sequences while preventing the synthesis of harmful sequences, enabling automated screening without human intervention.
Smart Images

Figure 0007704761000035 
Figure 0007704761000036 
Figure 0007704761000037
Abstract
Description
Related Applications
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 965,138, filed on January 23, 2020, the disclosure of which is hereby incorporated by reference in its entirety.
Technical Field
[0002] The present invention relates in part to a method for detecting a biological sequence corresponding to a particular biological function while minimizing the inaccurate detection of sequences having unrelated functions.
Background Art
[0003] The government strongly recommends that the industry screen nucleic acid synthesis orders to prevent the construction of biological weapons [Diggans, J. and E. Leproust. Frontiers in Bioengineering and Biotechnology 7 (April): 86 (2019)]. Many companies are voluntarily screening for pathogens identified by national and international groups [see International Gene Synthesis Consortium. (2017) “Harmonized Screening Protocol V2 / / genesynthesisconsortium.org / wp-content / uploads / IGSCHarmonizedProtocol11-21-17.pdf]. Current screening methods rely on similarity search algorithms [Altschul et al. Journal of Molecular Biology 215 (3): 403-10 (1990)] to identify sequences similar to those from well-known biological weapons. These algorithms cannot screen for small pieces of nucleic acid that can be assembled into larger pieces. Many harmless sequences are similar enough to be identified as harmful by similarity searches, resulting in false positives that require expert curation and making automated screening impossible. Automated screening methods that do not necessarily rely on experts to curate false positives and can be applied to benchtop nucleic acid synthesizers and assemblers are not available.
Summary of the Invention
[0004] According to an aspect of the present invention, there is provided a method for evaluating a biological sequence capable of performing a preselected function, the method comprising: (a) preselecting a biomolecule capable of performing the function of interest; (b) creating a test sequence database comprising a plurality of sequence fragments of the preselected biomolecule, wherein the preselected sequence fragments are of a predetermined length; (c) fragmenting the sequence of one or more test biomolecules to a length equivalent to the predetermined length of the sequence fragments of the preselected biomolecule in the test sequence database; (d) detecting the presence or absence of a sequence match between at least one fragment of the fragmented test biomolecule and at least one of the plurality of sequence fragments of the preselected biological sequence, and (e) taking an action in response to the detection in (d), wherein the detection in (d) provides an evaluation of the test biomolecule. In some embodiments, the step of taking an action in response to (d) comprises: preventing the synthesis of the test biomolecule, permitting the synthesis of the test biomolecule, sequencing one or more polynucleotide molecules, including one or more of DNA sequencing, DNA molecule design, polypeptide sequencing, and further sequence identification steps. In certain embodiments, the method further comprises identifying, in the test sequence database, one or more sequence fragments of the preselected biological sequence that respectively match one or more sequence fragments of a second biomolecule having a biological function unrelated to the biological function of interest of the preselected biomolecule, and removing the identified sequence fragment(s) from the test sequence database. In certain embodiments, when the presence of a sequence match is detected in (d), the action comprises preventing the synthesis of the test biomolecule.In some embodiments, the means for creating a test array database comprises: (a) screening a plurality of sequence fragments of a preselected biological sequence molecule against at least one control array database, the control array database comprising a plurality of control sequence fragments of at least one molecule capable of performing a function unrelated to the target function of the preselected biomolecule; (b) identifying the presence of a match between a sequence fragment among the plurality of sequence fragments of the preselected biomolecule and a sequence fragment in the control array database that is identified as a fragment of a biomolecule capable of performing a function unrelated to the target function of the preselected biomolecule; and (c) removing from the test array database sequence fragments of the preselected biological sequence that are identified as matching sequence fragments of a biological sequence identified as capable of performing a function unrelated to the target function of the molecule capable of performing the target function. In some embodiments, the preselected biomolecule is a polynucleotide. In certain embodiments, the sequence of the biomolecule is the full-length nucleic acid sequence of the polynucleotide or a portion of the full-length nucleic acid sequence of the polynucleotide. In certain embodiments, the full-length nucleic acid sequence encodes a protein. In some embodiments, the predetermined length of the sequence fragments of the preselected polynucleotide molecule is 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 or more nucleotides. In some embodiments, the preselected biomolecule comprises a polypeptide. In some embodiments, the amino acid sequence of the preselected biomolecule is the full-length amino acid sequence of the polypeptide or a portion of the full-length amino acid sequence of the polypeptide.In some embodiments, the predetermined length of a sequence fragment of a preselected polypeptide molecule is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95 or more amino acids. In certain embodiments, the plurality of sequence fragments of a preselected biomolecule: (1) includes all or an important portion or at least one essential fragment of a possible fragment of the biomolecule that can perform the function of interest, and (2) does not include sequences found in biomolecules that can perform a function unrelated to the function of the preselected biomolecule. In some embodiments, the control sequence database includes a plurality of control sequence fragments of at least one molecule that can perform a function unrelated to the function of interest of the preselected biomolecule. In some embodiments, the control sequence database includes a plurality of control sequence fragments of at least one molecule that cannot perform the function of interest of the preselected biomolecule. In certain embodiments, the predetermined length is the same for all fragments of the preselected biomolecule. In some embodiments, the predetermined length of the sequence fragments of the preselected biomolecule includes a plurality of lengths. In some embodiments, the test sequence database includes one or more sequence fragments randomly or pseudo-randomly selected from the sequences of molecules known to be able to perform a function different from the function of interest of the preselected molecule. In certain embodiments, the randomly or pseudo-randomly selected sequence fragments are biased towards sequence regions with greater homology to functionally or phylogenetically related sequences. In certain embodiments, the test sequence database further includes sequences that are functional equivalents of a plurality of sequence fragments of the preselected biomolecule. In some embodiments, the means for identifying functional equivalents includes computer-based means. In some embodiments, the means for identifying functional equivalents includes experimental means.In some embodiments, the computer means for selecting functionally equivalent elements included in the test array database includes the step of using a classifier based on experimental data for evaluating the accuracy of the computer means. In certain embodiments, the means for selecting functionally equivalent elements included in the test array database includes incorporating a minimum number of arrays calculated to achieve a predetermined certainty of successfully preventing the test array from slipping through detection. In some embodiments, the means for selecting functionally equivalent elements included in the test array database includes a random selection method or a pseudo-random selection method. In certain embodiments, the identity of all sequence fragments of one or both of the test array database and the test biomolecule is protected. In some embodiments, the means of protection includes the application of a cryptographic hash function, which deterministically maps array data to a fixed-size bit string using a one-way function. In some embodiments, the application of the cryptographic hash function cannot be reversed without a brute-force search of all possible array inputs to the test array database. In some embodiments, the application of the cryptographic hash function further includes the use of one or more information keys that must be accessed in order to attempt a brute-force search. In certain embodiments, the application of the cryptographic hash function requires keys from multiple independent sources that must cooperate to calculate the hash without any server accessing the array data. In certain embodiments, the independent sources include independent computer servers. In some embodiments, the method also includes the step of splitting the created test array database into two or more partial test array databases, and the created test array database used to detect the presence or absence of a sequence match is one, two or more partial test array databases. In some embodiments, when a sequence match is detected between a partial test array database and one or more fragments of the test biomolecule, the method further includes the step of detecting the presence or absence of the sequence using another two or more partial test array databases.In some embodiments, the test array database includes a portion of a larger database of array fragments such that the fragments included in the test array database can be rotated frequently or by the discovery of a match.
[0005] According to another aspect of the present invention, a method for identifying a biological sequence capable of performing a preselected function is provided. The method includes: (a) preselecting a biomolecule, wherein the preselected biomolecule is capable of performing the function of interest; (b) creating a test sequence database containing a plurality of sequence fragments of the preselected biomolecule, wherein the preselected sequence fragments are of a predetermined length; (c) fragmenting the sequence of one or more test biomolecules into lengths equivalent to the predetermined length of the sequence fragments of the preselected biomolecule in the test sequence database; and (d) detecting the presence or absence of a sequence match between at least one fragment of the fragmented test biomolecule and at least one of the plurality of sequence fragments of the preselected biological sequence. The means for creating the test sequence database includes: (i) screening a plurality of sequence fragments of a preselected biological sequence molecule against at least one control sequence database, wherein the control sequence database contains a plurality of control sequence fragments of at least one molecule capable of performing a function unrelated to the function of interest of the preselected biomolecule; (ii) identifying the presence of a match between a sequence fragment among the plurality of sequence fragments of the preselected biomolecule and a sequence fragment in the control sequence database that is identified as a fragment of a biomolecule capable of performing a function unrelated to the function of interest of the preselected biomolecule; and (iii) removing from the test sequence database the sequence fragments of the preselected biological sequence that are identified as matching the sequence fragments of the biological sequence identified as capable of performing a function unrelated to the function of interest of the molecule of interest. In some embodiments, the preselected biomolecule is a polynucleotide. In certain embodiments, the sequence of the preselected biomolecule is the full-length nucleic acid sequence of the polynucleotide or a part of the full-length nucleic acid sequence of the polynucleotide. In some embodiments, the full-length nucleic acid encodes a protein.In some embodiments, the predetermined length of a sequence fragment of a preselected polynucleotide molecule is 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 or more nucleotides. In certain embodiments, the preselected biomolecule comprises a polypeptide molecule. In some embodiments, the sequence of the preselected biomolecule is the full-length amino acid sequence of the polypeptide or a part of the full-length amino acid sequence of the polypeptide. In some embodiments, the predetermined length of a sequence fragment of a preselected polynucleotide molecule is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95 or more amino acids. In some embodiments, a plurality of sequence fragments of a preselected biomolecule: (1) include all or important parts of possible fragments of the biomolecule that can perform the desired function, and (2) do not include sequences found in biomolecules that can perform a function unrelated to the function of the preselected biomolecule. In certain embodiments, the control sequence database includes a plurality of control sequence fragments of at least one molecule that can perform a function unrelated to the desired function of the preselected biomolecule. In some embodiments, the control sequence database includes a plurality of control sequence fragments of at least one molecule that cannot perform the desired function of the preselected biomolecule. In some embodiments, a single length is the predetermined length of a sequence fragment of a preselected biomolecule. In certain embodiments, there are a plurality of predetermined lengths of sequence fragments of a preselected biomolecule. In some embodiments, the test sequence database includes one or more sequence fragments randomly selected from the sequences of molecules known to be able to perform a function different from the desired function of the preselected molecule.In some embodiments, the randomly selected sequence fragments are biased towards the stored regions. In certain embodiments, the test database further includes sequences that are functional equivalents of a plurality of sequence fragments of a preselected biomolecule. In some embodiments, the identities of all sequence fragments of one or both of the test sequence database and the test biomolecule are protected. In certain embodiments, the means of protection includes the application of a cryptographic hash function, which deterministically maps the sequence data to a fixed-size bit string using a one-way function. In some embodiments, the application of the cryptographic hash function cannot be reversed without a brute-force search of all possible sequence inputs to the test sequence database. In some embodiments, the application of the cryptographic hash function further includes the use of one or more information keys that must be accessed to attempt a brute-force search. In certain embodiments, the application of the cryptographic hash function requires keys from multiple independent sources that must cooperate to calculate the hash without any server accessing the sequence data. In certain embodiments, the independent sources include independent computer servers. In certain embodiments, the method also includes the step of splitting the created test sequence database into two or more partial test sequence databases, and the created test sequence database used to detect the presence or absence of a sequence match is one, two, or more partial test sequence databases. In some embodiments, when a sequence match is detected between a partial test sequence database and one or more fragments of the test biomolecule, the method further includes the step of detecting the presence or absence of the sequence using another two or more partial test sequence databases. In some embodiments, the test sequence database includes a portion of a larger database of sequence fragments such that the fragments included in the test sequence database can be rotated frequently or by the discovery of a match.In some embodiments, the method also includes taking an action upon detection, and the action-taking step includes steps of preventing the synthesis of the test biomolecule, permitting the synthesis of the test biomolecule, sequencing one or more polynucleotide molecules, DNA sequencing, DNA molecule design, determining the amino acid sequence of a polypeptide, and including one or more of further sequence identification steps. In certain embodiments, when the presence of a sequence match is detected, the action includes the step of preventing the synthesis of the test biomolecule. In some embodiments, the method includes identifying, in a created test sequence database, one or more sequence fragments of one or more preselected biological sequences that respectively match one or more sequence fragments of a second biomolecule having a biological function unrelated to the intended biological function of the preselected biomolecule, and excluding the identified sequence fragment(s) from the test sequence database.
[0006] In another aspect of the invention, there is provided a test sequence database created according to any embodiment of any of the foregoing methods. In another aspect of the invention, there is provided a method of evaluating a biological sequence using an embodiment of the foregoing test sequence database. In certain embodiments, the evaluating step includes determining whether to permit or prevent the synthesis of the evaluated biological sequence.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
[0008] Aspects of the present invention include, in part, methods and systems for reliably and efficiently detecting sequences corresponding to a preselected biological function, also referred to herein as "functional sequences", while minimizing the detection of functionally unrelated sequences, also referred to herein as "unrelated sequences". In some embodiments, the methods of the present invention include detecting nucleic acid functional sequences. In some embodiments, the methods of the present invention include detecting polypeptide functional sequences. As used herein, the term "polypeptide" is used interchangeably with the term "protein". Embodiments of the detection system of the present invention may include the assay sequence databases described herein.
[0009] The methods of the present invention can be used to detect nucleic acid or peptide sequences corresponding to particularly important biological functions while minimizing the chance of incorrectly identifying sequences that do not correspond to that function, such that the sequences encoding that function can be reliably identified.
[0010] Certain embodiments of the present invention are useful for preventing those interested in sequences considered undesirable for synthesis (also referred to herein as "adversaries") from avoiding detection. Randomly selecting fragments from functional sequences prevents adversaries from knowing which fragments are being screened and allows adversaries to introduce mutations throughout the test sequences in an attempt to escape detection. If adversaries do not include sufficient mutations in a particular fragment, their sequences may match one of the computationally determined functional variants included in the assay sequence databases of the present invention. The more fragments and the more computationally determined functional variants of these fragments that are included, the greater the likelihood of detection. If adversaries include too many mutations throughout their test sequences, they will no longer perform the desired function [Gray et al. Genetics 207(1):53-61(2017); Jackson et al. PloS One 12(4):e0164905(2017) and Pokusaeva et al. PLoS Genetics 15 (4):e1008079(2019)].
[0011] Inspection array database In some embodiments of the present invention, the detection method includes searching for and / or identifying an array that matches an array database. In some embodiments, the array database, also referred to herein as an "inspection array database," includes a plurality of sequence fragments of a preselected biomolecule. In some embodiments, at least a portion of the preselected biomolecules are selected because they can perform a desired function. The preselected biomolecule may be a polypeptide molecule or a polynucleotide molecule, and the preselected biomolecule can perform a desired function. Non-limiting examples of biomolecules that can perform a desired function include: viruses that can infect humans from humans, such as, but not limited to, sequences corresponding to the Ebola virus; and toxins that can kill mammalian cells at very low doses, such as, but not limited to, sequences encoding ricin. Additional biomolecules that can perform a desired function are well known in the art, and such sequences can be included in embodiments of the methods of the present invention.
[0012] The test array database may be created in such a manner as to include a plurality of array fragments of the sequences of preselected biomolecules, such fragments may also be referred to herein as "preselected array fragments". In some embodiments of the present invention, the preselected array fragments in the test array database are of a predetermined length. In embodiments where the preselected biomolecule is a polynucleotide, the predetermined length of the preselected array fragment includes all integers between 15 and 300, including 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300 or more nucleotides. In embodiments where the preselected biomolecule is a polypeptide, the predetermined length of the preselected array fragment includes all integers between 7 and 150, including 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 100, 110, 120, 130, 140, 150 or more amino acids. In certain embodiments, the test array database includes preselected array fragments of the same predetermined length. In some embodiments, the test array database includes preselected array fragments of different predetermined lengths.
[0013] In certain embodiments of the present invention, a plurality of sequence fragments of a preselected biomolecule include all or important portions of possible fragments of a biomolecule that can perform a desired function. In certain embodiments of the methods of the present invention, a plurality of sequence fragments of a preselected biomolecule do not include sequences found in biomolecules that can perform functions unrelated to the function of the preselected biomolecule. As used herein, the term "plurality" means more than one, for example it can mean at least 2, 3, 4, 5, 6, 7, 8, 9, 10 or more.
[0014] Figures 1 and 2 provide diagrams showing certain embodiments of the methods of the present invention. Figure 1 illustrates a method by which nucleic acid and peptide sequences can be broken down into pieces of a certain length to detect exact matches within a database of pieces unique to potential biological weapons. In some embodiments of the present invention, the database can include well-known sequence fragments from potential biological weapons and / or computationally generated functional equivalents, but does not include fragments that match sequences that are functionally unrelated from public databases. Figure 2 shows how, in certain embodiments of the present invention, the sequences being screened and the contents of the database can be hashed to enable screening while avoiding providing any information in clear text.
[0015] Certain means for creating a database of sequences for testing In some embodiments of the present invention, the means for creating a test sequence database includes screening a plurality of sequence fragments of a preselected biological sequence molecule against at least one control sequence database, where the control sequence database includes a plurality of control sequence fragments of at least one molecule that can perform a function unrelated to the target function of the preselected biomolecule. The means for creating a test sequence database may also include identifying the presence of a match between a sequence fragment among the plurality of sequence fragments of the preselected biomolecule and a sequence fragment in the control sequence database that is identified as a fragment of a biomolecule that can perform a function unrelated to the target function of the preselected biomolecule. As used herein, the term "control sequence database" means a database that includes one or more groups of sequences unrelated to what is being searched for by the detection system, and its acceptance into the test sequence database would result in false positive matches. Non-limiting examples include all plasmid sequences in GenBank, the European Nucleotide Archive, and the Addgene repository requested from at least 25 institutions.
[0016] Furthermore, the means for creating a test sequence database also includes excluding from the test sequence database one or more sequence fragments of the preselected biological sequence that are identified as matching sequence fragments of a biological sequence identified as being able to perform a function unrelated to the target function of the molecule that can perform the target function. As used herein, the term "biological sequence" refers to a molecule found in a biological system, and non-limiting examples of biological sequences are DNA sequences, RNA sequences, gene sequences, polynucleotide sequences, protein sequences, polypeptide sequences, amino acid sequences, and nucleic acid sequences.
[0017] In some embodiments of the method and system of the present invention, the test array database includes randomly selected fragments of functional arrays. In some cases, a ranked list of arrays predicted to be functionally equivalent to the arrays included in the test array database is computed on a computer using methods well known in the art (see, e.g., Bromberg, Y., & B. Rost Nucleic Acids Res. 35, 3823-3835 (2017); Miller et al. Sci. Rep. 7, 41329 (2017); Miller, M. et al. Nucleic Acids Res. 47, e142 (2019); Choi, Y. et al. PLoS One 7(10), e46688 (2012); Hopf, T. A. et al. Nat. Biotechnol. 35, 128-135 (2017); Gray, V. E. et al. Cell Syst. 6, 116-124.e3 (2018); and Riesselman, A. J. et al. Nat. Methods 15, 816-822 (2018), the contents of each of which are hereby incorporated by reference in their entirety). The minimum number of the list of functionally equivalent arrays, non-limiting examples of which are: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more, may be included in the test array database for detection purposes. For example, such functionally equivalent arrays may be included as control arrays.
[0018] In some embodiments of the method of the present invention, the equivalent functional arrays computed on one or more computers are selected randomly or in a biased random manner from a ranked list of arrays predicted to be functionally equivalent and included in the test array database of the present invention, using methods well known in the art. In some embodiments, the randomly selected fragments and the computationally equivalent functional fragments are pre-screened for matches to well-known unrelated arrays present in a database, non-limiting of which is GenBank, to ensure that fragments that are erroneously associated with unrelated arrays are not included in the test array database.
[0019] Some embodiments of the present invention include a pre-screening step in which sequences unrelated to the sequence of a pre-selected biomolecule are examined. In some embodiments of the present invention, prior to inputting into the test sequence database, the rate at which unrelated sequences from a set of well-known sequences included in the pre-screening step are incorrectly identified is 0%. The rate at which unrelated sequences that are not well-known or not included in the pre-screening step are accidentally and incorrectly identified varies depending on the length and number of fragments included in the database, and the rate of incorrect identification per fragment corresponds to 1 per total number of nucleic acid or peptide sequences of a defined length. The use of the devices, systems and methods of the present invention can reliably identify true functional sequences at rates of 90%, 95%, 99%, 99.9% or 100% including all percentages within the indicated range, where the exact rate depends on the number of randomly selected fragments and equivalent functional sequence fragments included in the test sequence database.
[0020] Figures 3A - B illustrate non-limiting examples of creating a database of the present invention. In this example, a database of sequences corresponding to potential biological weapons is created for an exact match to the nucleic acid sequences under consideration. Figure 3A provides a schematic diagram showing the use of nucleic acid or peptide fragments of a predetermined size for computer calculation of functionally equivalent fragments in a list ranked by probability of function. Figure 3B shows an item from such a ranked list, indicating that the item can be included up to a random adversarial threshold in the database in a deterministic, random, pseudo-random or biased random manner.
[0021] In another example of an embodiment of the present invention, FIGS. 4A - B show a graphical representation of how a random selection of fragments in a pattern biased towards optionally conserved regions from potential biological weapons can be used to reliably detect sequences corresponding to functionally similar biological weapons. As used herein, the term "biased towards conserved regions" refers to the selection of sequences that exhibit a higher level of homology to the sequences of the relevant genes and organisms. Such homology is often associated with greater functional importance. FIG. 4A shows five fragments and computationally calculated functionally equivalent variants along with a description of simple or elaborate "attacks" that attempt to evade detection by introducing mutations across the entire sequence of the biological weapon. FIG. 4B illustrates the failure of attempts to evade screening because the adversary does not know which of these fragments or how many functionally equivalent variants (also referred to herein as functional variants) are included in the database and attempts to avoid rendering the function non - functional by including too many mutations. As used herein with reference to a first nucleotide or polypeptide sub - sequence, the term "functionally equivalent" means a second nucleotide or polypeptide sub - sequence that can replace the first nucleotide or polypeptide sub - sequence without imposing a substantial cost on the function of the overall sequence, biomolecule, or molecule encoded by the sequence.
[0022] Databases created using embodiments of the method of the present invention are expected to enable the identification of potentially harmful nucleotide and / or polypeptide sequences and provide opportunities to prevent their production. When screening sequences using a database created using embodiments of the present invention, the likelihood of false - positive results is very low. For example, FIG. 5 shows the number of false positives predicted per year for the estimated level of DNA synthesis worldwide over time, assuming a database size of approximately one billion fragments, nucleic acid sequences of 57 base pairs, or peptide sequences of 19 amino acids.
[0023] Exclusion list The terms "excluded array" and "exclusion list" are used herein with reference to arrays that are specifically permitted for use by an individual and / or a research institute. Thus, when requesting the synthesis of an array, the individual or research institute requesting the array can provide the synthesis facility with a list of arrays that the individual and / or research institute owns and / or is permitted to use. By way of non-limiting example, a research institute may be permitted to handle work involving array "X" if it is necessary for the research institute to use it in developing a treatment or vaccine for an organism that includes array "X" but is considered a harmful array. Other individuals and / or research institutes are not permitted to synthesize or use array "X", but it will be considered an excluded array for the permitted research institute and will be listed in the research institute's exclusion list.
[0024] FIG. 6 shows an overview of an embodiment of the present invention and illustrates a flowchart of how the exclusion list functions. For example, typically a research institute needs to obtain approval from their institutional biosafety committee or other authorities to use a particular pathogen. Certain embodiments of a fully automated screening system created using the methods of the present invention will recognize that such a research institute is permitted to obtain DNA corresponding to genes and genomes that they are approved to use without any human intervention. In a non-limiting example, each gene and genome listed in the biosafety committee approval report has an associated GenBank® ID, as do all genes and genomes. The exclusion list contains all GenBank IDs of genes and genomes that the research institute is explicitly permitted to use. This information can be used to identify the genes and genomes permitted in the screening system of the present invention. In some embodiments of the system of the present invention, each research institute attempting to synthesize a sequence needs to send their exclusion list to the screening system of the present invention, which uses an oblivious multiparty server system to hash each GenBank ID once and then hash again using the institute's unique ID as a salt. This ensures that all research institutes have different, unique hashes, keeps the list confidential, and prevents an adversary from having a copy of the database and knowledge of which hashes correspond to which genes or genomes and thus which research institutes are working on which genes and genomes. Thereby, the adversary is prevented from using this information to determine what is present in the database. If a harmful gene or genome is entered into the database protected by the operator of the screening system of the present invention, the corresponding GenBank ID is hashed using an oblivious multiparty server system once and associated with each hashed sequence fragment of the gene or genome in the database.When the user issues an order for the system of the present invention to detect the presence of a DNA fragment in the hazard database, the database hashes the associated (once hashed) GenBank® ID using the customer's laboratory ID and then checks to see if the resulting hash matches the (similarly hashed) input on the customer's exclusion list. If so, the order passes through and the requested sequence is synthesized. Otherwise, the system of the present invention rejects the order and records the incident.
[0025] Conformance cost In some embodiments of the method of the present invention, it includes evaluating how a fragment of a gene or gene product sequence has different fitness costs when mutated. Figure 7 shows how some fragment windows across a gene can be changed to almost anything, other fragment windows allow a few substitutions, but otherwise substitutions disrupt the function of the gene or gene product, while still other fragment windows of the sequence show a tendency for most individual mutations to impose a small cost that increases as more mutations are added. The random adversarial threshold approach described herein with respect to the method of the present invention uses a classifier algorithm to predict the function of mutants for each fragment window and operates with a large number of wild-type sequences and predicted mutant sequences included in a database. To obtain a functional version of a harmful gene or genome protected by the system of the present invention, an adversary must choose to issue a synthetic order that includes either the wild-type sequence or a variant for each window of the gene or genome. For an order to pass through, none of these all speculations should exist in the database. Since the adversary does not know which windows are protected, they must speculate on variants for all windows and bear the potential fitness cost and risk of discovery at each window. The overall cost of the system must not exceed a minimum threshold for the resulting sequence to be functional. The greater the proportion of variants protected within a given window and the more windows that are protected, the lower the likelihood that an order will escape detection. Since the approach can be fully automated, it can also function for next-generation desktop DNA synthesizers and central suppliers.
[0026] Certain applications of the method Certain embodiments of the present invention are useful for biosecurity applications. For example, without intending to limit, the methods of the present invention can be used to detect functional sequences corresponding to biological weapons in DNA synthesis orders to prevent such synthesis and reject these orders. Another non-limiting implementation of embodiments of the methods of the present invention involves detecting functional sequences corresponding to biological weapons in DNA sequencing results. Another non-limiting implementation of embodiments of the present invention involves detecting functional sequences corresponding to biological weapons from a set of sequences input into a DNA design and analysis software program.
[0027] Certain aspects of the present invention enable highly efficient calculations of whether a sequence is functional. A time corresponding to O(log(N)) is considered the pinnacle of optimal algorithms, where in the context of the present invention N, N corresponds to the number of fragments in a database (Cormen, T.H., et al., 2009. Introduction to Algorithms. MIT Press.). Some data structures enable exact match searches in a time corresponding to O(1); since the present invention relies on exact match searches, certain embodiments of the present invention enable similar efficiency.
[0028] Certain embodiments of the present invention enable automated screening for functional sequences without human intervention. For example, by embodiments of the methods of the present invention, a nucleic acid synthesizer or peptide synthesizer can be programmed to automatically screen and reject sequence synthesis orders that contain functional sequences derived from a prohibited list of biological and toxin genes, such as those from the U.S. Federal Select Agent Program (FSAP) for unified export controls and the Australia Group treaty.
[0029] Molecules and Sequences for Screening In some embodiments of the present invention, the test sequence is screened against a test sequence database in the method or system of the present invention. The term "screened against" means "compared to". In a non-limiting example, the target sequence for synthesis is the test sequence, which is screened using the method and / or system of the present invention. Screening against the test sequence database of the present invention can provide information that helps determine actions taken with respect to the test sequence, such as, but not limited to: permitting or preventing the sequence from being synthesized. Other actions that can be informed by the results of applying embodiments of the method of the present invention to the test sequence include, but are not limited to: sequencing one or more polynucleotide molecules, DNA sequencing, DNA molecule design, and further sequence identification steps. Various means for evaluating DNA molecule design, DNA and / or protein sequence sequencing are well known in the art and can be applied as part of the actions taken based at least in part on the information obtained from the use of embodiments of the test database of the present invention.
[0030] In some embodiments of the present invention, the test biomolecule is fragmented into one or more of a plurality of partial and all possible overlapping pieces that are 1 base pair or 1 amino acid shifted and of a desired length in comparison to equal-sized pieces of the relevant sequence. The fragmented sequences of one or more test biomolecules are of a length equal to the predetermined length of the sequence fragments of a preselected biomolecule in the test sequence database. A test biomolecule is a molecule that is evaluated / tested using the test sequence database of the present invention. For example, but not intended to be limiting, a test biomolecule can be a polynucleotide synthesized by an individual or a research institute or produced by a service provider or a synthesizer.
[0031] Sequence identity protection In some methods and systems of the present invention, the identity of each sequence fragment of one or both of the test sequence database and the test biological molecule is protected. As used herein, the term "protected" means that the user of the method of the system of the present invention cannot identify the sequence of the fragment or the test biological molecule. For example, the sequence fragments to be screened can be "hashed" using methods well known to those skilled in the art to create a one-to-one information mapping that is not easily reversed. Including fragments that are equally hashed from related sequences in the test sequence database enables reliable database searches and detections without disclosing the identity of the sequences. Various means well known in the art for protecting the identity of sequences can be used. Non-limiting examples of means for protection include the application of cryptographic hash functions, where the cryptographic hash function deterministically maps the sequence data to a fixed-size bit string using a one-way function (see, for example, Cormen, T.H., et al., 2009. Introduction to Algorithms. MIT Press). Cryptographic hash functions are used in the art and methods well known in the art can be used including cryptographic hash functions in the methods and systems of the present invention. In some embodiments of the present invention, the cryptographic hash function is selected and applied such that it cannot be reversed or decoded without a brute-force search of all possible sequence inputs to the test sequence database. In some embodiments of the present invention, the cryptographic function applied also includes the use of one or more information keys that must be accessed to attempt a brute-force search. The inclusion of such one or more information keys restricts what a user can do to access the identity of each sequence fragment of one or both of the test sequence database and the test biological molecule. It is understood that additional means for protecting the identity of sequence fragments can also be used in conjunction with embodiments of the methods of the present invention.For example, see Yao, A.C. 27th Annual Symposium on Foundations of Computer Science (sfcs 1986), Toronto, ON, Canada, 1986, pp. 162 - 167 and I. Damgard, I.’89 Proceedings, Lecture Notes in Computer Science Vol. 435, G. Brassard, ed, Springer - Verlag, 1990, pp. 416 - 427, each of which is incorporated herein by reference in its entirety.
[0032] Examples
Example
[0033] Construct an inspection array database by selecting all possible fragments from a prohibited list of biological and toxin genes, such as those from the U.S. Federal Select Agent Program (FSAP) and the Australia Group treaty. Prescreen the database against a well - known database such as GenBank to remove all those that match sequences that are functionally unrelated.
[0034] The DNA synthesis provider fragments the sequences from the customer order into all possible overlapping pieces of equal size to those in the database. Translate the fragments in all possible reading frames to produce equivalent peptides. Compare the fragments from the customer order with those in the database in an automated fashion.
[0035] Results The synthesis provider can screen all orders for fragments that exactly match those from the prohibited list, with or without including a few false positives corresponding to unrelated sequences. The screening can be done in a fully automated fashion, avoiding the burden on experts.
Example
[0036] The inspection array database is constructed by randomly selecting fragments from the prohibited lists of biological and toxin genes, such as those from the U.S. Federal Select Agent Program (FSAP) and the Australia Group treaty, optionally biased towards highly conserved regions. Functional equivalents of these fragments are calculated computationally using a prediction program or algorithm, and random numbers are included in the database. The database is pre-screened against a well-known database, such as GenBank, to remove any fragments that match functionally unrelated sequences.
[0037] The DNA synthesis provider fragments the sequences from the customer order into all possible overlapping pieces of equal size to those in the database. The fragments are translated in all possible reading frames to produce equivalent peptides. The fragments from the customer order are compared in an automated fashion to those in the database to detect orders that produce functional equivalents of prohibited biological or toxin genes.
[0038] Results The synthesis provider can screen all orders for fragments that are functionally equivalent to those from the prohibited list, with or without including a few false positives corresponding to functionally unrelated sequences. The screening can be performed in a fully automated fashion.
Example
[0039] The inspection array database is constructed by randomly selecting fragments from the prohibited lists of biological and toxin genes, e.g., those from the U.S. Federal Select Agent Program (FSAP) and the Australia Group treaty, optionally biased towards highly conserved regions. Functional equivalents of these fragments are computed computationally using prediction programs or algorithms, and random numbers are included in the database. The database is pre-screened against a well-known database, e.g., GenBank, to remove all fragments that match functionally unrelated sequences.
[0040] The DNA synthesis provider assigns an information key to each customer interested in protecting orders from industrial espionage. Customer orders are fragmented into all possible overlapping pieces of equal size to those in the database, translated in all possible reading frames to produce equivalent peptides, and all results are hashed using the key. The provider similarly hashes all sequences in the database. Fragments from customer orders are compared in an automated fashion to those in the database to detect orders that produce functional equivalents of prohibited biological or toxin genes without sharing the customer order.
[0041] Results The synthesis provider can screen all orders for fragments that are functionally equivalent to those from the prohibited list, with or without including a few false positives corresponding to functionally unrelated sequences. The screening can be done in a fully automated fashion. The screening can be done in a fully automated fashion without requiring the customer to provide the synthesis provider with the order in clear text, protecting the customer from industrial espionage.
Example
[0042] The test array database is constructed by randomly selecting fragments from a list of prohibited biological and toxin genes, e.g., those from the U.S. Federal Select Agent Program (FSAP) and the Australia Group treaty, optionally biased towards highly conserved regions. Functional equivalents of these fragments are computed computationally using a prediction program or algorithm, and random numbers are included in the database. The database is pre-screened against a well-known database, e.g., GenBank, to remove any fragments that match non-functional sequences.
[0043] The DNA synthesis provider fragments the sequencing results from the customer sample into all possible overlapping pieces of equal size to those in the database. The fragments are translated in all possible reading frames to produce equivalent peptides. Fragments from the customer order are compared in an automated fashion to those in the database to detect customers who can produce materials that can generate functional equivalents of prohibited biological or toxin genes.
[0044] Results The sequencing provider can screen all sequencing results for fragments that are functionally equivalent to those from the prohibited list, with or without including a few false positives corresponding to non-functional sequences. The screening can be done in a fully automated fashion.
Example
[0045] The inspection array database is constructed by randomly selecting fragments from the prohibited lists of biological and toxin genes, e.g., those from the U.S. Federal Select Agent Program (FSAP) and the Australia Group treaty, optionally biased towards highly conserved regions. Functional equivalents of these fragments are computationally calculated using prediction programs or algorithms, and random numbers are included in the database. The database is pre-screened against a well-known database, e.g., GenBank, to remove any fragments that match non-functional sequences.
[0046] The DNA design software provider fragments the sequences entered by the customer into all possible overlapping pieces of equal size to those in the database. The fragments are translated in all possible reading frames to produce equivalent peptides. Fragments from the customer order are compared in an automated fashion to those in the database to detect customers who could inadvertently or deliberately design engineered constructs with functions equivalent to prohibited biological or toxin genes.
[0047] Results The design software provider can screen all designs for fragments that are functionally equivalent to those from the prohibited list, with or without including a few false positives corresponding to non-functional sequences. The screening can be done in a fully automated fashion.
Example
[0048] The study was conducted using an embodiment of the array screening method of the present invention. This experiment examined a random sample of 10,000 variants of the window PQSVECRPFVFGAGKPYEF (SEQ ID NO: 23) within the PIII gene of the M13 bacteriophage, which is an example of a virus that infects Escherichia coli and is harmless to humans. M13 was used in the study as a representative virus. In this case, a library of variant sequences was created, sequenced, and subjected to iterative rounds of infection to select for variants that retained function. After each round of selection, the survivors were sequenced to quantify the frequency of change for each variant. The classifier used FUNTRP and BLOSUM62 to obtain a fitness estimate in arbitrary units, while for the purpose of determining the ground truth, the phage was considered to be sufficiently fit to be "harmful" if its measured proportion in the large phage population exceeded a certain threshold before and after growth in the bacterial culture.
[0049]
Number
[0050] Experimental data on the effects of substitution variants of a 19 - amino - acid window within a protein and a 42 - base - pair sequence within a functional nucleic acid sequence were obtained by evaluating the genome of bacteriophage M13 that infects Escherichia coli using the prediction tools funtrp (for protein sequences) and nucleic acid conservation (for nucleic acid sequences in the viral origin of replication and packaging signals). Fifteen stretches of 42 - or 57 - base pairs were identified for experimental investigation in the packaging signal, positive and negative replication origins, and genes I, II, III, and IV. An oligonucleotide library of 220,000 sequences containing variants at positions predicted by funtrp or by structural analysis (for nucleic acids) was constructed to evaluate the accuracy of variant prediction. These libraries of variants were cloned into a phagemid, which is a plasmid containing a copy of the M13 replication origin and genes encoding related proteins from the M13 virus (see Figure 8). A helper plasmid containing an M13 phage with a replication origin and packaging signal disrupted by the insertion of the p15a plasmid origin and the kanamycin - resistance gene was constructed. Each helper plasmid had all protein - coding genes intact (for nucleic acid research) or was missing either gene I, gene II, gene III, or gene IV. For example, a helper plasmid missing gene I could be complemented by a phagemid library encoding gene I variants such that the phagemid library produced M13 particles encoding the phagemid library rather than the helper plasmid. Mixing these with recipient Escherichia coli and effectively selecting recipient cells successfully infected with phagemids carrying the phagemid selects for phagemids that can complement the missing gene. The degree of enrichment or non - enrichment compared to the wild - type sequence corresponds to the fitness of the variant with respect to virus production and infection.
[0051] A library of gene III variants for the amino acid sequence PQSVECRPFVFGAGKPYEF (SEQ ID NO: 23) was first cloned into DH5 alpha cells and sequenced by MiSeq (see Figure 9) to measure the initial library diversity (NGS point 1, initial library). Next, they were transformed into cells carrying a helper plasmid lacking gene III and sequenced again (NGS point 2, pre-selection). The resulting cells were grown, M13 particles were purified and sequenced (NGS point 3, phage release), and then mixed with recipient cells carrying different antimicrobial resistance markers. The resulting cells were grown, selected for both the phagemid and the recipient marker, and sequenced (NGS point 4, post 1-selection). The selection was repeated two more times to obtain additional enrichment data (NGS points 5 and 6, post 2-selection and post 3-selection). The library sequencing coverage was approximately 40x at 100% coverage for the two pre-selection samples.
[0052] The selection resulted in 4-digit enrichment / depletion in each direction (Figure 10), indicating that variants were actually selected by several criteria. Analysis of the results of all single-mutant variants of a particular amino acid residue revealed that alanine at position 13 tolerated any mutation, while proline at position 1 did not tolerate any (Figures 11 and 12).
[0053] A variant was considered to be fit if the ratio of its measured proportion in a larger population before and after selection (NGS point 1 compared to NGS point 4 or point 6) exceeded a certain threshold for various thresholds. These distributions of fit and unfit sequences from the library were used as an empirical dataset to evaluate the predictions.
[0054] To predict function, funtrp analysis (Figure 13) of whether each position is likely to accept substitutions at zero, moderate, or high fitness costs was combined with the BLOSUM62 matrix that defines common substitutions for amino acids at these positions to create predictions of functional variants for all variants in the library. These were compared against empirical results and various threshold boundaries to establish ROC curves that evaluate the true positive rate and false positive rate for the combined funtrp+BLOSUM62 classifier. Notably, the ROC curves for NGS points 1 through 4 and points 1 through 6 are nearly identical, suggesting that only one round of selection is needed.
[0055] Figure 14 shows a graph of ROC curves created from experimental data. The graph is a receiver operating characteristic (ROC) curve for weak classifiers based on the tools FUNTRP and BLOSUM62, evaluated against biologically ground truth data obtained from experimental studies. The ROC curve captures the trade-off between type I (false positive) and type II (false negative) errors for a yes / no classifier. The false positive rate (horizontal axis) is the proportion of variants not classified as fits, and the true positive rate (vertical axis) is the proportion of fit variants correctly classified. The curve is parameterized by the fitness threshold in arbitrary units, the classifier used to separate positive from negative based on an estimate of its noisy fitness level, s ∈ (-∞, ∞). Figure 15 is a graph showing another receiver operating characteristic (ROC) curve for the same classifier (as in Figure 14) including enrichment / depletion scores compared for NGS point 1 and NGS point 4. The ROC curve similarly demonstrates that iterative selection was not necessary.
[0056] Similar evaluations can be performed for other classifiers and for other variant libraries, as needed, to improve predictions. For nucleic acids, the prediction can combine conservation analysis of each position with structural analysis that calculates changes in the folding energy of the associated RNA secondary structure caused by the mutation.
[0057] Notably, the ROC curve for funtrp+BLOSUM62 was sufficient to predict 90% of the functional sequences from the library at the cost that half of the sequences were false positives. That is, considering 10,000 functional sequences, the ROC curve indicates that the classifier can predict 18,000 and cover 9,000 out of 10,000 well. If such sequences are included in the database at multiple positions across the hazard, the probability of detection is very high, and the probability for an adversary to obtain a functional sequence undetected is very low considering the cost of including variants sufficient to have a chance of escaping detection.
Example
[0058] A test array database is constructed by randomly selecting fragments from biological and toxin gene prohibited lists, e.g., from the U.S. Federal Select Agent Program (FSAP) and the Australia Group treaty, optionally biased towards highly conserved regions. The functional equivalents of these fragments are computationally calculated using a prediction program or algorithm, and random numbers are included in the database. The database is pre-screened against a well-known database, e.g., GenBank, to remove any fragments that match non-functional sequences.
[0059] An adversary attempts to synthesize a prohibited gene or a functional version of the genome. How can the probability of their success or failure be determined? The following describes the analysis of the screening method of the present invention implemented. The term "secureDNA system" refers to an embodiment of the screening method system of the present invention. The system is used to identify "harmful" sequences and / or sequences that are potential functional variants of harmful sequences. In the following description, the term "individual" means an organism, such as a virus, bacterium, or other organism. Using the following method, nucleotide sequences derived from an individual were evaluated to determine whether they are functional sequences, for example, if they are contained in an organism, whether the organism can survive and replicate. In this example, the term "adversary" means a person or group interested in or synthesizing a sequence regarded as a harmful polynucleotide sequence. In the following description, the term "defender" means an operator of the system of the present invention who attempts to prevent unauthorized persons and groups from synthesizing or otherwise accessing harmful polynucleotide sequences.
[0060] The SecureDNA system succeeds in DNA screening when it prevents all adversaries from assembling sequences encoding functional biohazards. The most dangerous types of hazards are self-replicating pathogens that can spread exponentially without human assistance. Functional sequences for such replicating pathogens are defined as DNA sequences that have sufficient fitness to survive and replicate in a shared environment and that become increasingly prevalent in the absence of human intervention, for example, novel pandemic viruses. Fitness is formulated in a number of ways: the probability that an object survives to reproduce, the expected number of offspring of the object, or any of these normalized for some relevant population. In any case, a probability-like real number in [0,1] is a sufficient representation of fitness, below which there is some minimum fitness f min for which it can be inferred that there exists. If the maximum fitness for all hazard variants that can be synthesized despite SecureDNA is less than f min then the SecureDNA system succeeds.
[0061] SecureDNA uses random adversarial threshold (RAT) screening to search for fragments of hazards and plausible functional variants. A variant is a DNA or amino acid sequence window that differs from the wild-type sequence (e.g., the sequence of the actual pathogen is found online) at one or more positions at the same locus. Each hazard consists of many loci where any variant is possible within any window within any locus. The conditional distribution F(v) is defined as the fitness or functionality of a given variant v of a hazard, where v is a triple (h, l, s v ), h: hazard identity or index; l: window index within the genome; s v : the exact variant sequence. By definition, the fitness of the wild type at any locus, F((h, l, s h:l )) is 1. Since the sequence variant s v is typically unique to both the hazard and the window, it can always be said that F((h, l, s v )) > ∈, including a slight abuse of notation, and F(s v ) = F(v). The total number of windows across the coding sequence of the hazard is denoted as N. Although complex interactions between variants were expected, it was estimated that the fitness of a hazard containing multiple variants, at least multiplicative compounding within a small fitness adjustment from the wild type, i.e., was at most the product of the fitnesses of the individual variants. Individual studies using the SecureDNA system carefully select windows to screen within each hazard and variant included using predictive software available in the art. Experiments were performed using funtrp [Miller, M., et al., Nucleic Acids Research, 2019, Vol. 47, No. 21, e142] in combination with the BLOSUM62 matrix [Eddy, S. Nat Biotechnol 22, 1035 - 1036 (2004)] of amino acid substitution probabilities.
[0062]
Number
[0063] In this example, it is conservatively assumed that the adversary has an oracle that can perfectly predict the fitness of any given variant, i.e., the adversary knows F(v). There are very few currently available methods for estimating F(v), and thus the information presented includes some interpretation of the impact of the significant inaccuracies in this estimate on the realistic state of the evaluation, which is the realistic state regarding the evaluation.
[0064] 2. Approximation of Destructive Changes The actual fitness distribution can only be measured empirically, and in part; a given experiment attempts to evaluate millions of variants in a single window, and thus only for biomolecules suitable for measurement. The study enabled a rough estimate of the fitness distribution for the most essential and evolutionarily conserved windows: for example, for a total of 737,280 variants whose sequences do not completely disrupt the function with F(v)>0.5 hazard in a specific window, substitutions with 7, 7, 5, 5, 4, 3, 3, 1, and 1 alternative residues are tolerated, allowing substitutions with a fitness cost below medium for 9 out of 19 amino acid residues. This situation was conservatively approximated by assigning a value of 1 to all of these and a value of 0 to all other variants.
[0065]
Number
[0066] For example, if one amino acid plays a decisive topological or affinity role in that protein, such that only a small set of replacements will result in functional proteins and hazards, this approximation was good. It was assumed that an individual's selection of decisive regions for screening could meet the conditions for this approximation.
[0067] The system involves selecting k different windows, i.e., k serves as an indicator of the w hazard genome. V w is made the set of all variants (h, w, s v ).
[0068]
Number
[0069] Assuming no preference among functional variants and assuming all have equal and effective fitness, the probability of the "evasion" phenomenon, denoted as "E", where an adversary randomly selects functional variants that do not exist at all k positions in the database is
[0070]
Number
[0071] (1) is as follows. The notable boundary in the probability of evasion is derived from the arithmetic - geometric mean inequality
[0072]
Number
[0073] (2) is as follows. More roughly, the underlying boundary can also be introduced in terms of the maximum scope of application α over all k windows max : P(E) ≤ 1 - α max (3) If the inventors can establish even one strong guarantee in the scope of application from 3, the inventors can rely on the maximum scope of application provided by any one window to the boundary P(E).
[0074]
Number
[0075] However, the defender does not have complete knowledge of F b (v).
[0076]
Number
[0077] If only a weak guarantee can be established in the average application range, it can be understood from (2) that the means for correction is to include more windows, i.e., increase k. A stronger bound that also exploits the fact that the positions of the k windows are unknown to the adversary is investigated in Section 4.1 below. First, the following considerations relate to the impact of the uncertainty in the defender's estimate of Fb(v) when the defender's choice of the k windows is considered known to the adversary.
[0078] 3. Error trade-off
[0079]
Number
[0080] The ROC curve accurately captures the trade-off between type I and type II errors. Selecting a point on the curve, called the operating point, based on a selection criterion constitutes a clear compromise that can be chosen in a principled way.
[0081] There are many ways to quantify and optimize across the ROC curve. A useful example is to define the costs C tp , C fp , C fn and C tn as the costs of true positive, false positive, false negative, and true negative test results, respectively, in a game-theoretic sense.
[0082]
Number
[0083] In fact, neither the adversary nor the defender will select variants outside a certain Hamming distance r before the variant becomes too different from the wild type for the variant to function. r is an empirical biological parameter. One possible expression for q is
[0084] [Number]
[0085] and can be.
[0086] [Number]
[0087] This Hamming ball capacity can be understood a priori as the size of a reasonable set of likely functional variants. Once s w , opt is selected, the application range α w is α w = tp w (s w , ort ) and is given by. This approach is attractive because it 1. Incorporates experimental data that coherently compares predicted fitness with actual fitness for some specific fitness estimation methods by constructing an empirical ROC curve based on, for example, next-generation sequencing (NGS) data from a population of harmless viruses in a control experiment, and 2. Coherently incorporates interpretable cost parameters that capture the resulting trade-off and offers. In this system, there is a trade-off between the probability of escaping screening by the size of the RAT database and the overall rate of accidentally classifying a random sequence as a hazard.
[0088]
Number
[0089] g should be as low as possible due to the accelerating increase in the total amount of DNA synthesized annually. Any inclusion in the database incurs the same cost C with respect to the overall false alarm rate of random misclassification tp = C fp := C p The cost of a true negative is zero (C tn = 0). The contact criterion from (4) is
[0090]
Number
[0091] will be. Incidentally, for the relationship with g, C p is inversely proportional to |S|, which is an exponential function of the window length. The window is as long as possible without enabling the easy assembly of longer DNA sequences from short sequences that cannot be screened due to being shorter than the window length, which is approximately 50 base pairs and is an inherent physical property of DNA. This constraint is the reason why C p cannot be arbitrarily low.
[0092] The cost of a false negative C fn is still under discussion. C fn is related to the expected exploitability of false negatives by an adversary to increase P(E), which can be the subject of detailed analysis. In particular, it depends on the scope of application of the database and its current size. For now, it is treated as an external parameter to confirm its effect.
[0093] To illustrate an example, simplify the assumption that all k windows have the same ROC, drop the subscript w, and from Equation 1 P(E)=(1-α) k This example shows how the quality of the classifier captured by the ROC curve affects the optimal selection of parameters, particularly k.
[0094]
Number
[0095] The maximum Hamming distance was determined to be 6 before additional changes would likely render the function inoperative. The capacity of the Hamming sphere of strings of length 19 from 20 alphabets with radius 6 is
[0096]
Number
[0097] is. q ≒ 4.8x10 -10 is the ratio (Equation 5). The value is C fn = 10 8 C p When set to, that is, ignoring the inclusion of functional variants in the database, it is 100 million times more costly than including additional items in the database (imaginable from the scale of the effect of successful hazard synthesis).
[0098]
Number
[0099] Currently, there is no closed - form or data for the classifier ROC curve, but an intuition can be obtained about the relationship between the "quality" of the classifier and the boundary that can be placed on P(E). Qualitatively, a "high - quality" classifier makes a clean separation between functional and non - functional variants. It has an ROC curve with a steep slope near fp = 0 and near fp = 1, reaching high towards the point (fp, tp)=(0, 1). The area under the ROC curve (AUC) of it is close to 1. It may have a slope in the range
[0100]
Number
[0101] and can have a slope in the range [2 / 3, 3 / 2]. In contrast, the ROC curve of a "low - quality" classifier is closer to the line tp = fp, and its AUC is close to.5, meaning that it does not exceed chance and may have a slope in the range [2 / 3, 3 / 2].
[0102] Assume that a high - quality classifier reaches a target slope of 21 at fp(s opt ) = 3x10 -6 ; tp(s opt ) =.95. The interpretation is that an application range of α = 95% is achieved by covering.0003% of the Hamming sphere of reasonable variants near the wild - type sequence corresponding to a database size of about one million, and it would be an ideal compromise that gives the specified balance of costs C p and C fn by definition.
[0103] The target boundary in terms of the probability of escape is set at P(E)=.001, and on average, the attacker creates a functional hazard only once in 1000 full orders. Using a high - quality classifier, the number k of windows to be covered is
[0104]
Number
[0105] It should be. Assume that a low-quality classifier can be used instead. This classifier
[0106]
Number
[0107] does not have s opt The interpretation is that the low-quality classifier cannot be used to achieve an optimal compromise between the given balance of costs. Instead, s opt is selected to result in a maximum practical database size of 10 million corresponding to fp(s opt ) = 3.2x10 -5 . Since the ROC curve is close to the line tp(s opt ) = fp(s opt ), the true positive rate is shown as 6x10 -4 and cannot be higher. The number k of windows to be covered to achieve the same boundary P(E) =.001 is
[0108]
Number
[0109] and would not be achieved except for the replicating pathogen with the largest genome. Even in that case, it cannot be achieved with this database size. What was important in the exercise described above was the following: · The ROC curve of the classifier for the fitness approximation of disruptive changes could be empirically measured, plotted, and analyzed for any dataset that compares the experimentally measured fitness with an easily obtainable computer-based tool for predicting protein or DNA functionality. · Explicit costs may be associated with including or not including variants in the database, and these can set the ideal operating point for each classifier. · The use of weak classifiers was possible, but more windows (larger k) were needed to compensate for their insufficient true positive rate. k was a function of quality, evaluated only by the ROC and cost settings. · There is a complex interaction between the cost setting and the optimal operating point. For example, k directly affects C p through the number of encryption operations, which is directly translated into DOPRF calls, and the expected exploitable vulnerability of a single false negative by an adversary to increase P(E) directly affects C fn This interaction between these parameters is expected to be convex and have a solution that is easy to handle using numerical methods. · There is a minimum average classifier strength required to accurately bound the probability of missing a target when screening by cost constraints.
[0110] 4. Randomly cover variants As the defender approaches complete knowledge of F(V), they can deterministically choose which windows to protect because they require at most arbitrarily close to zero, the fewest database inputs to show the bounds of P(E), and most functional variants can be recovered. Once these completely protectable windows are covered (only when determined by the windows selected by this criterion), it is understood that nothing is gained by including any other windows. An adversary with oracle knowledge of F(V) knows this and can focus their attention only on these areas, developing better fitness predictions to find counterintuitive functional variants that are less likely to be screened. Paradoxically, the simpler the hazard screening is due to the small number of functional variants, the easier it is for an adversary to develop a good estimate of fitness to evade screening, as long as the database construction only focuses on deterministically covering a particular window.
[0111] By randomly selecting windows non-deterministically, a randomization defender strategy can be used to increase the predicted amount of work that an adversary with oracle knowledge of F(V) should do to an infeasible level.
[0112] 4.1 Making the Adversary Modify Even More Windows This section describes how to make the boundaries from Section 2 (above in this specification) stronger.
[0113]
Number
[0114] In Section 2, an implicit assumption was made that the adversary actually knew which k windows to modify, but in reality these are not actually known to the adversary. One approach is to describe a variant collection that the adversary sends as
[0115]
Number
[0116] (assuming no window overlaps). The adversary has a modified l value compared to the threat it is intended to protect against
[0117]
Number
[0118] The first observation is that l ≥ k is required: if fewer than exactly k windows are modified by the adversary, for each such window the original array is in D and the adversary will always be caught.
[0119] Next, an adversary that modifies all N windows was considered. The chance of successfully passing the inspection using the "conforming" array is
[0120]
Number
[0121] and wherein
[0122]
Number
[0123] is the fitness of the actual array according to the aforementioned fitness function of the present inventors.
[0124]
Number
[0125] is
[0126]
Number
[0127] is the probability that cannot be captured by the RAT. Set l = N, and then
[0128]
Number
[0129] can be bounded in exactly the same way as in the case of Section 2. However, additionally, the success of an adversary without divine inspiration of fitness is affected by a fitness term that is likely to be 0 if all arrays have to be modified.
[0130] Towards establishing the boundary for k ≦ l < N, A l, l modifications are described by an event selected by an adversary such that all k windows protected by the adversary are contained. Further, let P(E) be the probability that none of the previous k windows have been detected, and let E be the related event. Since the wild - type sequence is surely in the database, and the adversary must pass at least all k tests and simultaneously correctly identify the correct k out of N windows using l modifications from the wild - type,
[0131]
Number
[0132] it should be. Thus
[0133]
Number
[0134] it is, and the probability of the adversary's success is
[0135]
Number
[0136] which becomes, where
[0137]
Number
[0138] is a sequence containing l modifications. The standard approach to the upper bound of the adversary's success probability is to find the maximum with respect to l and then appropriately select k, which requires that the aforementioned function be differentiable [in particular
[0139]
Number
[0140] Section 2 already gives a clear boundary to P(E), so it is necessary to analyze other terms. In summary, Example 7 provides a mathematical assessment of the extreme challenges even a well - positioned adversary faces when attempting to synthesize the sequences protected by the described system of the present invention. The provided assessment provides insights into the effectiveness of the screening method of the present invention. The greater the proportion of functional variants for a particular window existing in the database, the lower the probability of evading detection. The greater the number of windows to be protected, the lower the probability of evading detection. Assuming that an adversary can fully predict the function of the generated sequences (which is another, even more difficult problem compared to the unsolved problem of fully predicting the fitness of variants for a particular window), these results suggest that a real adversary who risks making the sequence non - functional has only a very small chance of success as long as a sufficient number of sequences are included in the database.
[0141] Equivalents Although some embodiments of the present invention are described and illustrated herein, those skilled in the art can readily envision various other means and / or structures for performing the functions and / or obtaining the results and / or one or more of the benefits described herein, and each such change and / or variation is to be regarded as within the scope of the present invention. More generally, those skilled in the art will understand that all parameters, sizes, materials, and arrangements described herein are meant to be exemplary, and that the actual parameters, sizes, materials, and / or arrangements depend on the specific one or more applications in which the teachings of the present invention are used. Those skilled in the art can recognize or confirm, using only routine experimentation, numerous equivalents to the specific embodiments of the invention described herein. Accordingly, the foregoing embodiments are presented by way of example only and are within the scope of the appended claims and their equivalents; it is understood that the invention can be practiced otherwise than as specifically described and claimed herein. The present invention is directed to the individual characteristics, systems, articles, materials, and / or methods described herein. Additionally, any combination of two or more of such characteristics, systems, articles, materials, and / or methods is included within the scope of the present invention, provided such characteristics, systems, articles, materials, and / or methods do not conflict with each other.
[0142] All definitions defined and used herein are to be understood as governing dictionary definitions, definitions in incorporated by reference documents, and / or ordinary meanings of defined terms. As used herein and in the claims, the indefinite articles "a" and "an" are to be understood to mean "at least one" unless clearly indicated otherwise.
[0143] As used herein and in the claims, the phrase "and / or" is to be understood as meaning "either or both" of the connected elements, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Other elements, whether or not specifically identified in relation to a specifically identified element, may optionally be present in addition to the elements specifically identified by the "and / or" clause, unless otherwise clearly indicated to the contrary.
[0144] All references, patents, and patent applications and publications cited or referenced in this application are hereby incorporated by reference in their entirety. Claims at the time of international filing [Claim 1] A method for evaluating a biological sequence capable of performing a preselected function, comprising: (a) preselecting a biomolecule capable of performing the function of interest; (b) creating a test sequence database containing a plurality of sequence fragments of the preselected biomolecule, wherein the preselected sequence fragments are of a predetermined length; (c) fragmenting the sequence of one or more test biomolecules to a length equivalent to the predetermined length of the sequence fragments of the preselected biomolecule in the test sequence database; (d) detecting the presence or absence of a sequence match between at least one fragment of the fragmented test biomolecule and at least one of the plurality of sequence fragments of the preselected biological sequence, and (e) taking an action in response to the detection in (d), wherein the detection in (d) provides an evaluation of the test biomolecule A method comprising. [Claim 2] The method according to claim 1, wherein taking an action in response to (d) includes one or more of the following: preventing the synthesis of the test biomolecule, permitting the synthesis of the test biomolecule, sequencing one or more polynucleotide molecules, DNA sequencing, DNA molecule design, polypeptide sequencing, and further sequence identification steps. [Claim 3] The method according to claim 1, further comprising identifying one or more sequence fragments of a preselected biological sequence that respectively match one or more sequence fragments of a second biomolecule having a biological function unrelated to the biological function of interest of the preselected biomolecule in the test sequence database, and removing the identified sequence fragment(s) from the test sequence database. [Claim 4] The method according to claim 1, wherein when the presence of a sequence match is detected in (d), the action includes preventing the synthesis of the test biomolecule. [Claim 5] The means for creating the test sequence database is: (a) Screening a plurality of sequence fragments of a preselected biological sequence molecule against at least one control sequence database, wherein the control sequence database includes a plurality of control sequence fragments of at least one molecule capable of performing a function unrelated to the target function of the preselected biological molecule; (b) Identifying the presence of a match between a sequence fragment among the plurality of sequence fragments of the preselected biological molecule and a sequence fragment in the control sequence database that is identified as a fragment of a biological molecule capable of performing a function unrelated to the target function of the preselected biological molecule; and (c) Removing from the test sequence database the sequence fragment of the preselected biological sequence that is identified as matching a sequence fragment of a biological sequence identified as capable of performing a function unrelated to the target function of the molecule capable of performing the target function. The method according to any one of claims 1, comprising the above steps. [Item 6] The method according to any one of claims 1, wherein the preselected biological molecule is a polynucleotide. [Item 7] The method according to claim 6, wherein the sequence of the biological molecule is the full-length nucleic acid sequence of the polynucleotide or a part of the full-length nucleic acid sequence of the polynucleotide. [Item 8] The method according to claim 7, wherein the full-length nucleic acid sequence encodes a protein. [Item 9] The predetermined length of the sequence fragment of the preselected polynucleotide molecule is 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 or more nucleotides. The method according to any one of claims 6-8. [Item 10] The method according to any one of claims 1-6, wherein the preselected biological molecule includes a polypeptide. [Item 11] The method according to claim 10, wherein the amino acid sequence of the preselected biological molecule is the full-length amino acid sequence of the polypeptide or a part of the full-length amino acid sequence of the polypeptide. [Item 12] The method according to claim 10, wherein the predetermined length of the sequence fragment of the preselected polypeptide molecule is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95 or more amino acids. [Item 13] A plurality of sequence fragments of a preselected biomolecule: (1) include all or an important part or at least one essential fragment of a possible fragment of a biomolecule capable of performing the intended function, and (2) do not include sequences found in biomolecules that perform functions unrelated to the function of the preselected biomolecule. The method according to claim 1. [Item 14] The method according to claim 2, wherein the control sequence database includes a plurality of control sequence fragments of at least one molecule that performs a purpose function unrelated to the purpose function of the preselected biomolecule. [Item 15] The method according to claim 2, wherein the control sequence database includes a plurality of control sequence fragments of at least one molecule that cannot perform the purpose function of the preselected biomolecule. [Item 16] The method according to claim 1, wherein the predetermined length is the same for all fragments of the preselected biomolecule. [Item 17] The method according to claim 1, wherein the predetermined length of the sequence fragment of the preselected biomolecule includes a plurality of lengths. [Item 18] The method according to claim 1, wherein the test sequence database includes one or more sequence fragments randomly or pseudo-randomly selected from the sequences of molecules known to be capable of performing functions different from the purpose function of the preselected molecule. [Item 19] The method according to claim 18, wherein the randomly or pseudo-randomly selected sequence fragments are biased towards sequence regions having greater homology to functionally or phylogenetically related sequences. [Item 20] The method according to claim 1, wherein the test sequence database further includes a sequence that is a functional equivalent of a plurality of sequence fragments of the preselected biomolecule. [Item 21] The method according to claim 20, wherein the means for identifying functional equivalents includes computer-based means. [Item 22] The method according to claim 20, wherein the means for identifying functional equivalents includes experimental means. [Item 23] The method according to any one of claims 20 to 22, wherein the computer means for selecting functionally equivalent elements included in the test array database includes the step of using a classifier based on experimental data for evaluating the accuracy of the computer means. [Item 24] The method according to claim 20, wherein the means for selecting functionally equivalent elements included in the test array database includes incorporating the minimum number of arrays calculated to achieve a predetermined accuracy that successfully prevents the test array from leaking from detection. [Item 25] The method according to any one of claims 20 to 22, wherein the means for selecting functionally equivalent elements included in the test array database includes a random selection method or a pseudo-random selection method. [Item 26] The method according to claim 1, wherein the identity of all array fragments of one or both of the test array database and the test biomolecule is protected. [Item 27] The method according to claim 26, wherein the means for protection includes the application of a cryptographic hash function, and the cryptographic hash function deterministically maps the array data to a fixed-size bit string using a one-way function. [Item 28] The method according to claim 27, wherein the application of the cryptographic hash function cannot be reversed without a brute-force search of all possible array inputs to the test array database. [Item 29] The method according to claim 28, wherein the application of the cryptographic hash function further includes the use of one or more information keys that must be accessed to attempt a brute-force search. [Item 30] The method according to claim 29, wherein the application of the cryptographic hash function requires keys from a plurality of independent sources that must cooperate to calculate the hash without any server accessing the array data. [Item 31] The method according to claim 30, wherein the independent sources include independent computer servers. [Item 32] The method according to claim 1, further including the step of dividing the created test array database into two or more sub-test array databases, and the created test array database used to detect the presence or absence of an array match is one, two, or more sub-test array databases. [Item 33] The method according to claim 32, further comprising the step of detecting the presence or absence of an array using two or more other partial test array databases when an array match is detected between a partial test array database and one or more fragments of a test biomolecule. [Item 34] The method according to claim 1, wherein the test array database includes a part of a larger database of array fragments such that the fragments included in the test array database can be rotated frequently or by finding a match. [Item 35] A method for identifying a biological sequence that can perform a preselected function, (a) a step of preselecting a biomolecule, the preselected biomolecule being capable of performing a target function; (b) a step of creating a test array database including a plurality of array fragments of a preselected biomolecule, the preselected array fragments having a predetermined length; (c) fragmenting the sequence of one or more test biomolecules to a length equivalent to the predetermined length of the array fragments of the preselected biomolecule in the test array database; and (d) detecting the presence or absence of an array match between at least one fragment of the fragmented test biomolecule and at least one of the plurality of array fragments of a preselected biological sequence comprising; The means for creating the test array database is: (i) a step of screening a plurality of array fragments of a preselected biological sequence molecule against at least one control array database, the control array database including a plurality of control array fragments of at least one molecule capable of performing a function unrelated to the target function of the preselected biomolecule; (ii) identifying the presence of a match between an array fragment among the plurality of array fragments of the preselected biomolecule and an array fragment in the control array database that is identified as a fragment of a biomolecule capable of performing a function unrelated to the target function of the preselected biomolecule; and (iii) removing from the test array database the array fragments of the preselected biological sequence that are identified as matching the array fragments of the biological sequence identified as capable of performing a function unrelated to the target function of the molecule capable of performing the target function comprising a method. [Item 36] The method according to claim 35, wherein the preselected biomolecule is a polynucleotide. [Item 37] The method according to claim 36, wherein the sequence of the preselected biomolecule is the full-length nucleic acid sequence of the polynucleotide or a part of the full-length nucleic acid sequence of the polynucleotide. [Item 38] The method according to claim 37, wherein the full-length nucleic acid encodes a protein. [Item 39] The method according to claim 36, wherein the predetermined length of the sequence fragment of the preselected polynucleotide molecule is 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 or more nucleotides. [Item 40] The method according to claim 35, wherein the preselected biomolecule includes a polypeptide molecule. [Item 41] The method according to claim 40, wherein the sequence of the preselected biomolecule is the full-length amino acid sequence of the polypeptide or a part of the full-length amino acid sequence of the polypeptide. [Item 42] The method according to claim 40 or 41, wherein the predetermined length of the sequence fragment of the preselected polynucleotide molecule is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 2, 93, 94, 95 or more amino acids. [Item 43] The method according to claim 35, wherein a plurality of sequence fragments of the preselected biomolecule: (1) include all or important parts of possible fragments of the biomolecule that can perform the desired function, and (2) do not include sequences found in biomolecules that can perform a function unrelated to the function of the preselected biomolecule. [Item 44] The method according to claim 35, wherein the control sequence database includes a plurality of control sequence fragments of at least one molecule that can perform a function unrelated to the desired function of the preselected biomolecule. [Item 45] The method according to claim 35, wherein the control sequence database includes a plurality of control sequence fragments of at least one molecule that cannot perform the desired function of the preselected biomolecule. [Item 46] The method according to claim 35, wherein the single length is a predetermined length of a sequence fragment of a biomolecule selected in advance. [Item 47] The method according to claim 35, wherein there are a plurality of predetermined lengths of sequence fragments of a biomolecule selected in advance. [Item 48] The method according to claim 35, wherein the test sequence database includes one or more sequence fragments randomly selected from the sequences of molecules that are known to be able to perform functions different from the target function of the preselected molecule. [Item 49] The method according to claim 48, wherein the randomly selected sequence fragments are biased towards conserved regions. [Item 50] The method according to claim 35, wherein the test database further includes sequences that are functional equivalents of a plurality of sequence fragments of a preselected biomolecule. [Item 51] The method according to claim 35, wherein the identity of all sequence fragments of one or both of the test sequence database and the test biomolecule is protected. [Item 52] [Item 53] The method according to claim 51, wherein the means of protection includes the application of a cryptographic hash function, and the cryptographic hash function deterministically maps the array data to a fixed-size bit string using a one-way function. [Item 54] The method according to claim 53, wherein the application of the cryptographic hash function cannot be reversed without a brute-force search of all possible sequence inputs to the test sequence database. [Item 55] The method according to claim 53, wherein the application of the cryptographic hash function further includes the use of one or more information keys that must be accessed in order to attempt a brute-force search. [Item 56] The method according to claim 51, wherein the application of the cryptographic hash function requires keys from a plurality of independent sources that must cooperate to calculate the hash without any server accessing the array data. [Item 57] The method according to claim 55, wherein the independent source includes an independent computer server. [Item 58] Further including the step of dividing the created test sequence database into two or more partial test sequence databases, and the creation used to detect the presence or absence of an array match [Item 59] The method according to claim 35, wherein the created test sequence database is one, two or more partial test sequence databases. [Item 58] The method according to claim 57, further comprising the step of detecting the presence or absence of an array using two or more other partial test array databases when an array match is detected between a partial test array database and one or more fragments of a test biomolecule. [Item 59] The method according to claim 35, wherein the test array database includes a part of a larger database of array fragments such that the fragments included in the test array database can be rotated frequently or by the discovery of a match. [Item 60] The method according to claim 35, comprising the step of taking an action upon receiving the detection in claim 35(d), and the step of taking an action further includes one or more of the following: the step of preventing the synthesis of a test biomolecule, the step of permitting the synthesis of a test biomolecule, the step of sequencing one or more polynucleotide molecules, DNA sequencing, DNA molecule design, the step of determining the amino acid sequence of a polypeptide, and a further sequence identification step. [Item 61] The method according to claim 35, wherein when the presence of an array match is detected in claim 35(d), the action includes the step of preventing the synthesis of a test biomolecule. [Item 62] The method according to claim 35, further comprising the step of identifying one or more sequence fragments of a preselected biological sequence that respectively match one or more sequence fragments of a second biomolecule having a biological function unrelated to the target biological function of a preselected biomolecule in the created test array database, and the step of removing the identified sequence fragment(s) from the test array database. [Item 63] A test array database created by the method according to any one of claims 1 to 34. [Item 64] A test array database created by the method according to claim 1. [Item 65] A method for evaluating a biological sequence using the test array database according to claim 63. [Item 66] The method according to claim 65, wherein the step of evaluating includes the step of determining whether to permit or prevent the synthesis of the evaluated biological sequence. [Item 67] A test array database created using the method according to any one of claims 35 to 62. [Item 68] A test array database created using the method according to claim 35. [Item 69] A method for evaluating a biological sequence using the inspection array database according to claim 67. [Item 70] The method according to claim 69, wherein the evaluating step includes determining whether to permit or prevent the synthesis of the evaluated biological sequence.
Claims
1. A method for evaluating a biological sequence capable of performing a preselected function, comprising: (a) preselecting a biomolecule, wherein the preselected biomolecule is capable of performing the function of interest; (b) creating a test sequence database comprising a plurality of sequence fragments of the preselected biomolecule, wherein the plurality of sequence fragments are of a predetermined length; (c) fragmenting the sequence of the test biomolecule into lengths equivalent to the predetermined lengths of the plurality of sequence fragments of the preselected biomolecule in the test sequence database; (d) detecting, in an automated manner, the presence or absence of a sequence match between at least one fragment of the fragmented test biomolecule and at least one of the plurality of sequence fragments of the preselected biomolecule in the test sequence database, and (e) taking an action in response to the detection in (d), wherein taking the action comprises one or more of the following steps: preventing the synthesis of the test biomolecule, permitting the synthesis of the test biomolecule, sequencing one or more polynucleotide molecules, DNA sequencing, DNA molecule design, polypeptide sequencing, and further sequence identification steps A method comprising the above steps.
2. The method of claim 1, further comprising identifying one or more of the plurality of sequence fragments of the preselected biomolecule that respectively match one or more sequence fragments of a second biomolecule having a biological function unrelated to the function of interest of the preselected biomolecule in the test sequence database, and removing the one or more identified sequence fragments from the test sequence database.
3. The method of claim 1, wherein the step of taking the action in (e) comprises preventing the synthesis of the test biomolecule.
4. The step of creating the test sequence database in (b) comprises: (i) screening a plurality of sequence fragments of the preselected biomolecule against at least one control sequence database, wherein the at least one control sequence database comprises a plurality of control sequence fragments of at least one molecule capable of performing a function unrelated to the function of interest of the preselected biomolecule; (ii) identifying a match between a sequence fragment among the plurality of sequence fragments of the preselected biomolecule and a control sequence fragment in the at least one control sequence database; and (iii) removing from the test sequence database the sequence fragments of the preselected biomolecules identified as matching the control sequence fragments The method according to any one of claims 1 to 3, comprising:
5. The method according to claim 1, wherein the test sequence database comprises one or more sequence fragments randomly or pseudo-randomly selected from sequences of molecules known to be able to perform a function different from the target function of the preselected biomolecule.
6. The method according to claim 5, wherein the one or more sequence fragments randomly or pseudo-randomly selected are biased towards sequence regions having greater homology to sequences that are functionally or phylogenetically related.
7. The method according to claim 1, wherein the test sequence database further comprises sequences that are functional equivalents of multiple sequence fragments of the preselected biomolecule.
8. The method according to claim 7, wherein identifying the functional equivalents comprises computer-implemented means.
9. The method according to claim 7, wherein identifying the functional equivalents comprises experimental means.
10. The method according to claim 8, wherein the computer-implemented means for identifying the functional equivalents comprises using a classifier based on experimental data for evaluating the accuracy of the computer-implemented means.
11. The method according to claim 7, wherein the functional equivalents included in the test sequence database comprise the minimum number of sequences calculated to achieve a predetermined accuracy that successfully prevents the test sequences from leaking from detection.
12. The method according to any one of claims 7 to 9, wherein selecting the functional equivalents included in the test sequence database comprises a random selection method or a pseudo-random selection method.
13. The method according to claim 1, further comprising dividing the test sequence database into two or more partial test sequence databases, and detecting the presence or absence of a sequence match in (d) using one of the two or more partial test sequence databases.
14. The method according to claim 13, further comprising detecting the presence or absence of one or more fragments of the test biomolecule in another one of the two or more partial test sequence databases when a sequence match is detected between one of the two or more partial test sequence databases and one or more fragments of the test biomolecule.
15. The method according to claim 1, wherein the test array database includes a portion of a larger database of array fragments such that the fragments contained in the test array database can be rotated frequently or by the discovery of a match.
16. A method for identifying a biological sequence that can perform a preselected function, comprising: (a) preselecting a biomolecule, wherein the preselected biomolecule can perform the target function; (b) creating a test array database containing a plurality of array fragments of the preselected biomolecule, wherein the plurality of array fragments have a predetermined length, and the step comprises: (i) screening a plurality of array fragments of the preselected biomolecule against at least one control array database, wherein the at least one control array database contains a plurality of control array fragments of at least one molecule that can perform a function unrelated to the target function of the preselected biomolecule; (ii) identifying a match between an array fragment among the plurality of array fragments of the preselected biomolecule and a control array fragment in the at least one control array database; and (iii) removing from the test array database the array fragments of the preselected biomolecule identified as matching the control array fragments to create a test array database; (c) fragmenting the sequence of the test biomolecule into lengths equivalent to the predetermined lengths of the plurality of array fragments of the preselected biomolecule in the test array database; and (d) detecting, in an automated manner, the presence or absence of a sequence match between at least one fragment of the fragmented test biomolecule and at least one of the plurality of array fragments of the preselected biomolecule in the test array database A method.
17. The method according to claim 1 or 16, wherein the preselected biomolecule is a polynucleotide.
18. The method according to claim 17, wherein the sequence of the preselected biomolecule is the full-length nucleic acid sequence of a polynucleotide or a part of the full-length nucleic acid sequence of a polynucleotide.
19. The method according to claim 18, wherein the full-length nucleic acid encodes a protein. **Claim 20**: The method according to claim 17, wherein the predetermined lengths of the plurality of sequence fragments of the polynucleotide are 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 or more nucleotides. **Claim 21** The method according to claim 1, 16, or 17, wherein the preselected biomolecule comprises a polypeptide molecule. **Claim 22** The method according to claim 21, wherein the sequence of the preselected biomolecule is the full-length amino acid sequence of a polypeptide or a part of the full-length amino acid sequence of a polypeptide molecule. **Claim 23**: The method according to claim 21 or 22, wherein the predetermined lengths of the plurality of sequence fragments of the polypeptide molecule are 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95 or more amino acids. **Claim 24** The method according to claim 1 or 16, wherein the plurality of sequence fragments of the preselected biomolecule: (1) includes all or important parts of possible fragments of the preselected biomolecule that can perform the desired function, and (2) does not include sequences found in biomolecules that can perform a function unrelated to the function of the preselected biomolecule. **Claim 25**: The method according to claim 4 or 16, wherein at least one control sequence database includes a plurality of control sequence fragments of at least one molecule that cannot perform the desired function of the preselected biomolecule. **Claim 26** The method according to claim 1 or 16, wherein the predetermined length of the plurality of sequence fragments of the biomolecule with a single preselected length. **Claim 27** The method according to claim 1 or 16, wherein there are a plurality of predetermined lengths of the plurality of sequence fragments of the preselected biomolecule. **Claim 28** The method according to claim 16, wherein the test array database includes one or more array fragments randomly selected from the arrays of molecules known to be capable of performing functions different from the intended function of a preselected biomolecule.
29. The method according to claim 28, wherein one or more randomly selected array fragments are biased towards conserved regions.
30. The method according to claim 16, wherein the test database further includes sequences that are functional equivalents of a plurality of array fragments of a preselected biomolecule.
31. The method according to claim 1 or 16, wherein the identity of all array fragments of one or both of the test array database and the test biomolecule is protected.
32. The method according to claim 31, further comprising applying a cryptographic hash function that deterministically maps array data to a fixed-size bit string using a one-way function.
33. The method according to claim 32, wherein the application of the cryptographic hash function cannot be reversed without a brute-force search of all possible array inputs to the test array database.
34. The method according to claim 33, wherein the application of the cryptographic hash function further includes the use of one or more information keys that must be accessed to attempt a brute-force search.
35. The method according to claim 32, wherein the application of the cryptographic hash function requires keys from multiple independent sources that must cooperate to calculate the hash without any server accessing the array data.
36. The method according to claim 35, wherein the multiple independent sources include independent computer servers.
37. The method according to claim 16, further comprising splitting the test array database into two or more partial test array databases, and detecting the presence or absence of an array match in (d) using one of the two or more partial test array databases.
38. The method according to claim 37, further comprising detecting the presence or absence of one or more fragments of the test biomolecule in another one of the two or more partial test array databases when an array match is detected between one of the two or more partial test array databases and one or more fragments of the test biomolecule.
39. The method of claim 16, wherein the test array database comprises a portion of a larger database of array fragments such that fragments contained in the test array database can be rotated frequently or by the discovery of matches. **Claim 40**: The method of claim 16, comprising the step of taking an action in response to the detection in (d), the step of taking an action further comprising one or more of the following: the step of preventing the synthesis of the test biomolecule, the step of permitting the synthesis of the test biomolecule, the step of sequencing one or more polynucleotide molecules, DNA sequencing, DNA molecule design, the step of determining the amino acid sequence of a polypeptide, and a further sequence identification step. **Claim 41** The method of claim 40, wherein the step of taking an action comprises the step of preventing the synthesis of the test biomolecule when an array match is detected in (d). **Claim 42**: The method of claim 16, further comprising the step of identifying one or more of a plurality of sequence fragments of a preselected biomolecule that each match one or more sequence fragments of a second biomolecule having a biological function unrelated to the intended function of the preselected biomolecule in the test array database, and the step of removing the one or more sequence fragments identified from the test array database.