System and method for increasing synthesized protein stability
A 3DCNN predicts destabilizing residues in proteins, addressing the inefficiencies of traditional methods by enhancing protein stability and enabling effective high-throughput engineering for industrial use.
Patent Information
- Application Number
- JP2025071233
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-05-02
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-13
AI Technical Summary
Existing protein engineering methods struggle to efficiently identify and stabilize destabilizing residues in proteins, leading to incomplete understanding of protein sequence/structure/function relationships, especially for properties like stability and folding, which hinders high-throughput protein variant screening and industrial application adaptation.
A three-dimensional convolutional neural network (3DCNN) is trained on amino acid sequences and crystal structures to predict destabilizing residues, allowing for targeted mutagenesis and improved protein stability through machine learning, without requiring prior knowledge of specific protein features.
The 3DCNN accurately identifies residues for mutagenesis, enhancing protein stability and enabling more efficient protein engineering, improving protein properties such as folding and stability, thus facilitating industrial applications.
Smart Images

Figure 2025118694000031 
Figure 2025118694000032 
Figure 2025118694000033
Abstract
Description
[Technical Field]
[0001] Related Applications This application is a division of "System and Method" filed on May 2, 2019. for Increasing Synthesized Protein Stab This application claims the benefit of and priority to U.S. Provisional Patent Application No. 62 / 841,906, entitled "Energy Efficiency." No. 6,299,499, filed on Oct. 1, 2003, which is incorporated herein by reference in its entirety.
[0002] STATEMENT OF FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT This invention was made possible by a grant from the National Institutes of Health. Grant number R43 NS105463 awarded, and Air Force Office Grant number F awarded by the ce of Scientific Research The Government has granted a patent application to the present invention under Patent Application No. A9550-14-1-0089. have certain rights. [Background technology]
[0003] Protein engineering is a transformative approach in biotechnology and biomedicine. These include conferring novel functionality to existing proteins or synthesizing proteins in non-native environments. The goal is either to make proteins more sustainable or to improve the quality of life of the organism. The design consideration that gives rise to this is the overall stability of the protein. Gain-of-function mutations that expand the role of proteins through design or directed evolution are often Introduced at a thermodynamic cost. Most natural proteins are only marginally stable. Therefore, it has been shown that increasing stability before selection promotes protein evolvability. Therefore, functional mutations that destabilize proteins to the point of unfolding may be overlooked. There is a possibility.
[0004] A significant barrier to the conversion of useful naturally occurring biocatalysts to industrial applications is fundamentally Adaptation of proteins to different environmental conditions, temperatures, and solvents. Increasing productivity can alleviate many of these pressures and allow for higher yields and lower costs. Therefore, stabilization is the key to many protein engineering efforts. is essential to the success of
[0005] There are many ways to engineer proteins, all of which generally involve the creation of protein variants. How quickly and accurately can it be measured and how efficiently the protein variant status is determined? Mutagenesis polymerase chain reaction (PCR) represents a compromise between the amount of DNA that can be sampled and the amount of DNA that can be sampled. Techniques such as RI require minimal knowledge of the relationship between sequence and function, but Nevertheless, high-throughput methods are still required to isolate large libraries of protein variants. Depends on the screen or selection. Uses structural data and computational approaches. These can be used to narrow the search space and simultaneously reduce the amount of downstream characterization. Tools are increasingly needed for proteins where desired properties are difficult to measure, especially on a large scale. However, due to an incomplete understanding of protein sequence / structure / function relationships, Different computational tools for protein engineering are often quite different or This is especially true for properties such as stability and folding. However, these are often many small fragments distributed throughout the entire protein sequence. This is the result of a complex interaction.
[0006] Typically, the computational method involves a computationally intensive folding scheme. By performing simulations, we identify residues that destabilize proteins. The level of detail involved in these simulations varies, drawing on quantum mechanics (MOE). Some go so far as to describe molecular interactions by comparing the two, while others use more coarse-grained methods (R To a first approximation, coarse-grained approaches Find gaps in the structure (RosettaVIP) or perform fast local free energy calculations (f oldX) or find residues that are evolutionary outliers (PROSS). Problematic residues are then identified by either hydrophobic packing or evolutionary constraints. Returning to the census suggests better matching residues. The effects of these substitutions on the quality of the mutants were estimated through energy simulations. Overall, this process (residue identification, substitution proposal, refolding, and free cleavage) Energy calculations) can take from a few hours to a few days.
[0007] Machine learning does not require prior knowledge of specific protein features or time-consuming manual work. This is an attractive alternative as it does not require inspection and assignment of individual structural features. Recently, Torng and Altman (Torng, incorporated herein by reference) ng et al.,“3D deep convolutional neural networks for amino acid environment simi laarity analysis,”BMC Bioinformatics,18:3 02, 2017) is a method for identifying amino acids given information about the surrounding protein microenvironment. By predicting the uniformity, we propose a three-dimensional convolutional neural network (3DCNN). This paper describes a general framework for applying this neural network to protein structure analysis. The network achieved a 42% prediction accuracy in assigning amino acids to wild-type sequences. and other computational methods that rely on identifying pre-assigned structure-based features. Furthermore, given the structural data of the model protein T4 lysozyme, As such, 3D CNNs typically use wild-type sequences where mutations are known to be destabilizing. Given the structures of these known destabilizing mutants, we predict the destabilizing residues relative to the wild-type residues. showed a strong preference for Summary of the Invention
[0008] The proteome is a collection of proteins with a number of properties, including folding, stability, catalytic properties, and binding specificity. Considering that several unrelated or even opposing phenotypes must be simultaneously exhibited, This suggests that amino acids that are structural outliers at positions away from the active site may contribute to folding and It is reasonable that this may affect the quality and stability of the product, but not its functionality. It uses artificial intelligence to learn the consensus microenvironments of different amino acids and scan the entire structure. Improved protein engineering techniques to scan and identify residues that deviate from the structural consensus There is a need in the art for: those residues that are considered to have low wild-type probability , is believed to be a locus of instability and has therefore been subjected to mutagenesis and stability engineering. Implementations of the systems and methods discussed herein are good candidates for such improvements. To provide improved protein engineering techniques.
[0009] In one embodiment, a neural network is trained to improve protein properties. The computer-implemented method includes collecting a set of amino acid sequences from a database; and compiling a set of three-dimensional crystal structures with chemical environments for a set of acids. , translating the chemical environment into a voxelized matrix and sub- The neural network is trained on the set and the target task is Identifying candidate residues to be mutated in proteins and using a neural network to identify candidate residues and identifying predicted amino acid residues to replace the mutant protein, The protein exhibits improved properties over the target protein. In one embodiment, the method comprises: Position, partial charge, beta factor, secondary structure, aromaticity, electron density, polarity, and their combinations The spatial arrangement of features selected from the group consisting of a combination of at least one of the three-dimensional crystal structures The method further includes adding the two together.
[0010] In one embodiment, the method comprises adjusting the set of amino acid sequences to reflect their inherent frequencies. In one embodiment, the method further comprises selecting an amino acid sequence from a random position in the sequence. The method further comprises sampling at least 50% of the amino acids in the set of amino acid sequences. In one embodiment, the method further comprises: training a second, independent neural network on the same set and and identifying candidate and predicted residues based on the results of the network. In one embodiment, the property is stability, maturation, folding, or a combination thereof. That is it.
[0011] In another aspect, a system for improving a protein property includes a processor and instructions and a non-transitory computer-readable medium having stored thereon instructions executed by a processor. providing a target protein comprising a sequence of residues when each three-dimensional model is generated; provides a set of three-dimensional models surrounding amino acids and a set of protein property values for each and estimating a set of parameters at various points of each three-dimensional model. and the neural network with the three-dimensional model, parameters, and protein property values. and training the neural network to identify candidate residues to be mutated in the target protein. and using a neural network to identify predicted amino acid residues to replace the candidate residues. and performing steps including the steps of identifying a group and producing a mutant protein. The protein exhibits improved properties over the target protein.
[0012] In one embodiment, the protein property is stability. Recompile at least one amino acid sequence of the folded amino acid sequence In one embodiment, the steps include reconstructing the three-dimensional model and generating an updated three-dimensional model. Before folding, at least one amino acid sequence of the folded amino acid sequence has This includes adding the spatial arrangement of features.
[0013] In another aspect, the present invention provides a method for the preparation of secBFP2 fragments containing T18, S28, S32, S42, S52, S62, S72, S82, S92, S102, S112, S124, S132, S142, S152, S162, S172, S182, S192, S194, S196, S198, S199 Choose from Y96, S114, V124, T127, D151, N173, and R198 and secBFP2 variants having one or more mutations at another residue that is In one embodiment, the protein is a protein selected from the group consisting of SEQ ID NOs: 2 to 28. In one embodiment, the secBFP2 variant comprises one of the amino acid sequences The cBFP2 variant is a variant of one of the amino acid sequences of SEQ ID NO: 2 to SEQ ID NO: 28. In one embodiment, the secBFP2 variant is selected from the group consisting of SEQ ID NO: 2 to SEQ ID NO: 2 In one embodiment, the BFP comprises a fusion protein comprising one of the amino acid sequences of , and a fragment of one of the amino acid sequences of SEQ ID NO: 2 to SEQ ID NO: 28.
[0014] In another aspect, the present invention provides a nucleic acid encoding a protein comprising a secBFP2 variant. In one embodiment, the nucleotide sequence is a nucleic acid molecule comprising SEQ ID NO: Amino acid sequences set forth in SEQ ID NOs: 2 to 28, variants thereof, and fusion proteins thereof; or a fragment thereof. In one embodiment, the molecule is a plasmid. In one embodiment, the molecule is an expression vector. In one embodiment, the nucleic acid molecule is a heterologous protein coding In another aspect, the present invention provides a method for the preparation of a recombinant vector comprising: A composition comprising the above protein, a composition comprising the above nucleic acid molecule, a composition comprising the above protein A kit or a nucleic acid molecule as described above. [Brief explanation of the drawings]
[0015] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with the drawing(s) will be provided upon request and payment of the necessary fee. Provided by the office after payment.
[0016] The above and other objects and features are intended to be understood as being within the scope of the present description and understanding of the present invention. The following description of the preferred embodiments is provided to illustrate, by way of example only, the principles of the present invention, and is not to be construed as limiting the present invention. and in the drawings, like numbers represent like elements. [Figure 1A] FIG. 1 is a diagram of a computer-implemented neural network implementation for increasing synthetic protein properties. [Figure 1B] 1 is a flowchart of an implementation of a method for determining amino acid residues at the center of a microenvironment. [Figure 1C] 1 is a flowchart of an implementation of a method for increasing synthetic protein properties during testing. [Figure 1D] FIG. 1 is a block diagram of an implementation of a neural network for increasing synthetic protein properties during training. [Figure 1E] FIG. 1 is a block diagram of an implementation of a convolutional neural network for increasing synthetic protein properties. [Figure 2A] 1 is a graph of experimental results of an implementation of a method and system for increasing synthetic protein properties. [Figure 2B] 10 is another graph of experimental results of an implementation of a method and system for increasing synthetic protein properties. [Figure 3A] 10 is another graph of experimental results of an implementation of a method and system for increasing synthetic protein properties. [Figure 3B] Photographs of proteins synthesized with modifications suggested by implementation of the system to increase synthetic protein properties. [Figure 4A] 10 is another graph of experimental results of an implementation of a method and system for increasing synthetic protein properties. [Figure 4B]FIG. 1 is a diagram of suggested protein modifications suggested by implementation of the system for increasing synthetic protein properties. [Figure 5] 1 is a set of photographs of experimental results of the implementation of a system for increasing synthetic protein properties. [Figure 6] 10 is a graph of experimental results of implementation of a system for increasing synthetic protein properties. [Figure 7] 10 is a graph of experimental results of implementation of a system for increasing synthetic protein properties. [Figure 8] 1 is a graph showing the fold change in fluorescence of 17 blue fluorescent protein variants relative to the wild-type protein. [Figure 9] 1 is a graph showing the fold change in fluorescence of blue fluorescent protein variants relative to the wild-type protein. [Figure 10] 1 provides exemplary images of the fluorescence of the blue fluorescent protein variant "Bluebonnet," which contains the S28A, S114T, N173H, and T127L mutations, compared to the parent protein and other blue fluorescent proteins. [Figure 11A] FIG. 1 is a block diagram illustrating an implementation of a system for increasing synthetic protein properties. [Figure 11B] FIG. 1 is a block diagram illustrating an implementation of a system for increasing synthetic protein properties. DETAILED DESCRIPTION OF THE INVENTION
[0017] The drawings and description of the invention are for illustrative purposes only and are not to be construed as limiting the scope of the invention. While the present invention has been simplified, for clarity, many of the features found in related systems and methods have been omitted. It should be understood that the present invention excludes other elements. Recognize that other elements and / or steps may be desirable and / or necessary to However, such elements and steps are well known in the art. and because they do not facilitate a better understanding of the invention, such elements and steps No discussion of such techniques is provided herein. The disclosure herein is intended to provide a basis for such techniques known to those skilled in the art. All such changes and modifications to elements and methods are covered.
[0018] Unless otherwise defined, all technical and scientific terms used herein are defined by the It has the same meaning as commonly understood by a person skilled in the art to which the invention belongs. Any methods and materials similar or equivalent to those described herein may be used in the practice or testing of the present invention. Exemplary methods and materials that can be used in the study are described.
[0019] As used herein, each of the following terms has the meaning associated with it in this section: It has.
[0020] The articles "a" and "an" are used herein to refer to one of the grammatical objects of the article, or It is used to refer to one or more (i.e., at least one). "An element" means one element or more than one element.
[0021] As used herein when referring to a measurable value, such as an amount, a time duration, etc., "about" means "approximately" or "approximately" " indicates ±20%, ±10%, ±5%, ±1%, and ±0.1% variation from the specified value. The term "variation" is intended to encompass movement and therefore variation is appropriate.
[0022] The term "nucleic acid molecule" or "polynucleotide" refers to a molecule in either single-stranded or double-stranded form. refers to a deoxyribonucleotide or ribonucleotide polymer in either Unless otherwise indicated, it is understood that nucleotides function in a manner similar to naturally occurring nucleotides. and polynucleotides containing known analogues of naturally occurring nucleotides, which can be used to When a nucleic acid molecule is represented by a DNA sequence, this means that "U" (uridine) is replaced by "T" It is understood that the present invention also includes RNA molecules having corresponding RNA sequences that replace "(thymidine)" ( It will be done.
[0023] The term "recombinant nucleic acid molecule" refers to a molecule containing two or more linked polynucleotide sequences. A recombinant nucleic acid molecule is a nucleic acid molecule that is not naturally occurring. It may be produced by recombinant nucleic acid or by chemical synthesis methods. The molecule may be a fusion protein, e.g., a polypeptide of interest, as discussed herein, linked to the polypeptide. Encoding fluorescent protein variants suggested by the systems and methods described herein The term "recombinant host cell" refers to a cell that contains a recombinant nucleic acid molecule. Thus, recombinant host cells contain "genes" that are not found within the native (non-recombinant) form of the cell. " can be used to express a polypeptide from the
[0024] Reference to a polynucleotide that "encodes" a polypeptide refers to transcription of the polynucleotide. and the resulting mRNA, upon translation, produces a polypeptide. A coding polynucleotide is a coding sequence that is identical to the mRNA. The encoding polynucleotide is considered to include both the original strand and its complementary strand. It is recognized that a sequence is considered to include degenerate nucleotide sequences that encode the same amino acid residues. The nucleotide sequence encoding the polypeptide may contain introns. It may include polynucleotides as well as coding exons.
[0025] The term "expression control sequence" refers to a sequence that controls the transcription or translation of a polynucleotide or It refers to a nucleotide sequence that controls the localization of an operably linked polypeptide. The expression control sequences regulate the transcription, and optionally translation (i.e., i.e., transcriptional or translational regulatory elements, respectively), or cellular specification of the encoded polypeptide A molecule is "operably linked" when it controls or regulates the localization of a molecule to a compartment. Expression control sequences include promoters, enhancers, transcription terminators, and start codons ( ATG), intron excision and splicing to maintain the correct reading frame signal, stop codon, ribosome binding site, or targeting a polypeptide to a specific location a targeting sequence, such as a cellular compartmentalization signal (which directs the polypeptide to the cytosol, nucleus, , plasma membrane, endoplasmic reticulum, mitochondrial membrane or matrix, chloroplast membrane or chloroplast It can be targeted to the lumen, the intermediate trans-Golgi sacculus, lysosomes, or endosomes. The cell compartmentalization domain may be, for example, a human type II membrane anchor protein. Amino acid residues 1-81 of protein galactosyltransferase, or cytochrome c oxygen The peptides containing amino acid residues 1 to 12 of the presequence of subunit IV of cis-aminobutyric acid enzyme were (Also, Hancock et al., EMBO J. 10:4033-40 39,1991, Buss et al.,Mol.Cell.Biol.8:3960 See also U.S. Pat. Nos. 5,776,689, 1988, 5,776,689, each of which is incorporated herein by reference. , which is incorporated herein by reference).
[0026] When used to describe a chimeric protein, "operably linked" or The terms "operably linked" or "operably associated" or similar terms are used interchangeably. A polypeptide refers to polypeptide sequences placed in a physical and functional relationship to each other. In some embodiments, the function of the polypeptide components of the chimeric molecule is enhanced relative to their functional activity alone. For example, the system and method discussed herein may be implemented in a similar manner to the system and method described herein. The fluorescent protein to be detected can be fused to a polypeptide of interest. In this case, the fusion molecule The target polypeptide retains its original biological activity. In some embodiments of the systems and methods discussed herein, fluorescent tags are used. The activity of either the protein or the protein of interest is compared to their activity alone. Such fusion can also be reduced by the systems and methods discussed herein. It can be used together with
[0027] The term "label" refers to any method of detecting a substance by, for example, visual inspection, spectroscopy, or photochemical reaction, biochemical reaction, A composition that can be detected by immunochemical or chemical reaction, with or without an instrument. Useful labels include, for example, phosphorus-32, fluorescent dyes, fluorescent proteins, and electron-dense reagents, enzymes (such as those commonly used in ELISA), small molecules, e.g., biotin , digoxigenin, or antisera or antibodies are available which may be monoclonal antibodies. Other haptens or peptides may be included. The fluorescent protein variants suggested by the implementation of quality, but nevertheless by means other than its own fluorescence, e.g., radioactivity A nuclide label or peptide tag can be incorporated into a protein to enhance its expression, e.g. Detectable by facilitating the identification and isolation of expressed proteins, respectively It will be appreciated that the systems and methods discussed herein may be labeled to Labels useful for purposes of implementing the methods generally include radioactive signals, fluorescent light, enzymatic activity, and the like. any of which may be, for example, a fluorescent signal in the sample. It can be used to quantify the amount of a protein variant.
[0028] The term "polypeptide" or "protein" refers to a polymer of two or more amino acid residues. These terms refer to a mer in which one or more amino acid residues are substituted with the corresponding naturally occurring amino acid. Amino acid polymers that are artificial chemical analogues of amino acids, as well as naturally occurring amino acid polymers The term "recombinant protein" applies to proteins derived from recombinant DNA molecules. A protein produced by expression of a nucleotide sequence encoding the amino acid sequence of a protein. Refers to...
[0029] The terms "isolated" and "purified" refer to substances in their natural state in nature. Refers to a substance that is substantially or essentially free from components that normally accompany it. The properties are generally determined by analytical methods such as polyacrylamide gel electrophoresis and high performance liquid chromatography. The polynucleotide or polypeptide is determined using analytical techniques. A substance is considered isolated if it is the predominant species present in the substance. Protein or nucleic acid molecules represent over 80% of the macromolecular species present in a preparation and often Represents more than 90% of all macromolecular species present, typically more than 95% of the macromolecular species, and especially When tested using conventional methods for determining the purity of such molecules, A polypeptide or polynucleotide purified to essential homogeneity such that it is the only species present. It's a punchline.
[0030] The term "naturally occurring" refers to a protein, nucleic acid molecule, cell, or molecule that does not occur in nature. It is used to refer to other substances that are present in living organisms, such as polypeptides present in living organisms, including viruses. A naturally occurring substance is a sequence of a nucleic acid or polynucleotide sequence as it occurs in nature. It may be in a form, for example, modified by the hand of man to be in isolated form.
[0031] The term "antibody" refers to a gene encoding an immunoglobulin gene(s), or an antigen-binding fragment thereof. refers to polypeptides substantially encoded by the The immunoglobulin genes recognized are kappa, lambda, Alpha, gamma, delta, epsilon, and mu constant region genes, as well as numerous Immunoglobulin variable region genes are included. Antibodies exist as complete immunoglobulins. However, antigen-binding fragments of antibodies are similarly characterized, which can be isolated by digestion with peptidases. Such antibodies can be produced by recombinant DNA technology. Antigen-binding fragments include, for example, Fv, Fab', and F(ab)'2 fragments. As used herein, the term "antibody" refers to an antibody produced by modification of a whole antibody. antibody fragments synthesized either by recombinant DNA technology or de novo using recombinant DNA technology. The term "immunoassay" refers to an assay that utilizes an antibody to specifically bind to an analyte. Immunoassays use the specific binding properties of specific antibodies to isolate and target the analyte. The method is characterized by the following:
[0032] Used in conjunction with two or more polynucleotide sequences or two or more polypeptide sequences When used, the term "identical" means that the sequences are the same when aligned for maximum correspondence. When percentage sequence identity is used in reference to polypeptides, it refers to the residues in the sequence that One or more otherwise non-identical residue positions may differ by conservative amino acid substitutions. , the first amino acid residues have similar chemical properties, such as similar charge or hydrophobic or hydrophilic properties. The amino acid residue is substituted for another amino acid residue that has the desired functional properties and therefore improves the function of the polypeptide. It is recognized that conservative substitutions do not alter the functional properties of polypeptides. The percent sequence identity may be adjusted upwards to correct for the conservative nature of the substitution. Such adjustments can be made, for example, to convert conservative substitutions into partial mismatches rather than full mismatches. This is done by scoring as a match, thereby increasing the percentage of sequence identity. Thus, for example, identical amino acids may be given a score of 1, and non-conservative substitutions may be given a score of 2. is given a score of zero, conservative substitutions are given a score between zero and one. Scoring of conservative substitutions is done, for example, by Meyers and Miller, mp.Appl.Biol.Sci.4:11-17,1988, Smith and Waterman, Adv. Appl. Math. 2:482, 1981, Needle man and Wunsch, J. Mol. Biol. 48:443, 1970, Pe. arson and Lipman,Proc.Natl.Acad.Sci.,USA 85:2444 (1988), Higgins and Sharp, Gene 73 :237-244, 1988, Higgins and Sharp, CABIOS 5 :151-153;1989, Corpet et al., Nucl. Acids R es.16:10881-10890,1988, Huang,et al.,Comp .Appl.Biol.Sci.8:155-165,1992, Pearson et. al., Meth. Mol. Biol., 24:307-331, 1994 (these Each is calculated using the algorithms discussed in Alignment can also be performed by simple visual inspection and manual alignment of sequences. It can be implemented.
[0033] When used in reference to a particular polynucleotide sequence, "conservatively modified variations" The term refers to different polynucleotides that encode the same or essentially the same amino acid sequence. A polynucleotide refers to an amino acid sequence that is identical to an essentially identical sequence. The degeneracy of the genetic code allows for a large number of functionally identical polynucleotides will encode any given polypeptide. For example, the codons CGU, CGC, CGA, CGG, AGA, and AGG all code for the amino acid arginine. At every position where an arginine is specified by a codon, the codon It can be changed to any of the corresponding codons listed without altering the peptide. Nucleotide sequence variations such as Therefore, the fluorescent protein variants can be considered as a kind of "mutation" Each polynucleotide sequence disclosed herein as a whole contains all possible sequences. It will be appreciated that this also accounts for the lent mutation. Also, the only codon for methionine is usually Polynucleotides, except for AUG, which is the only codon for tryptophan, and UUG, which is usually the only codon for tryptophan. Each codon in the peptide is modified by standard techniques to yield a functionally identical molecule. It will also be recognized that the sequence of the encoded polypeptide may be altered. Each silent variation of a polynucleotide that does not disrupt the transcription is implicitly described herein. In particular, a single amino acid or a small percentage of amino acids (typically less than 5%) in the encoded sequence Individual substitutions, deletions, or additions that alter, add, or delete less than 1% of the total DNA fragments (less than 1%, generally less than 1%) Conservatively modified variations may be considered, provided that the alterations are with chemically similar amino acids. It will be appreciated that the term "functionally similar" is used in conjunction with "functionally similar" amino acid substitutions. Conservative amino acid substitutions that provide the following amino acids may include the following six groups: Each contains amino acids that are considered conservative substitutions for one another: 1) Alanine (Ala, A), serine (Ser, S), threonine (Thr, T), 2) Aspartic acid (Asp, D), glutamic acid (Glu, E), 3) Asparagine (Asn, N), Glutamine (Gln, Q), 4) Arginine (Arg, R), Lysine (Lys, K), 5) Isoleucine (Ile, I), Leucine (Leu, L), Methionine (Met, M) ), valine (Val, V), and 6) Phenylalanine (Phe, F), Tyrosine (Tyr, Y), Tryptophan (T rp, W).
[0034] Amino acid or nucleotide sequences are compared with each other or over a given comparison window. Two or more amino acid sequences are considered to be identical if they share at least 80% sequence identity with a reference sequence. or two or more nucleotide sequences are "substantially identical" or "substantially similar" Thus, a substantially similar sequence is, for example, at least 85% of the sequence identity, at least 90% sequence identity, at least 95% sequence identity, or at least Both contain sequences with 99% sequence identity.
[0035] If the complement of a subject nucleotide sequence is substantially identical to a reference nucleotide sequence, then the complement is A given nucleotide sequence is considered to be "substantially complementary" to a reference nucleotide sequence. It will be done.
[0036] Fluorescent molecules are coupled via fluorescence resonance energy transfer (FR) with donor and acceptor molecules. It is useful in FRET. To optimize detectability, several factors must be balanced: The emission spectrum is maximally aligned with the excitation spectrum of the acceptor to maximize the overlap integral. The quantum yield of the donor moiety and the absorbance coefficient of the acceptor should overlap as much as possible. The number is as close as possible to maximize RO, which represents the distance at which the energy transfer efficiency is 50%. However, the fluorescence resulting from direct excitation of the acceptor is higher than that of the FR The donor and acceptor fluorescence may be difficult to distinguish from the fluorescence resulting from ET. The excitation spectrum of the donor is such that the donor can be efficiently excited without directly exciting the acceptor. There should be as little overlap as possible so that a wavelength region that can be used can be found. Similarly, the emission spectra of the donor and acceptor are such that the two emissions are clearly distinguishable. There should be as little overlap as possible so that the emission from the acceptor is the only readout. If it is to be measured either as a fraction or as part of the emission ratio A high fluorescence quantum yield of the acceptor moiety is desirable. One factor that should be considered when selecting a fluorescent protein is the efficiency of fluorescence resonance energy transfer (FER) between the fluorescent protein and the fluorescent protein. Preferably, the efficiency of FRET between the donor and acceptor is at least 10%, more preferably More preferably, it is at least 50%, and even more preferably, it is at least 80%.
[0037] The term "fluorescence properties" refers to the molar extinction coefficient at the appropriate excitation wavelength, the fluorescence quantum efficiency, and the excitation wavelength. The shape of the excitation or emission spectrum, the maximum excitation wavelength and the maximum emission wavelength, 2 The ratio of excitation amplitudes at two different wavelengths, the ratio of emission amplitudes at two different wavelengths, The spectral variation of the fluorescent protein is compared with that of the wild-type or parent fluorescent protein. A measurable difference in any one of these characteristics between an ant or its variants is useful. A measurable difference is the amount of any quantitative fluorescent property, e.g., the amount of fluorescence at a particular wavelength. The amount of fluorescein can be determined by determining the integral of the fluorescence across the emission spectrum. Determining the ratio of excitation or emission amplitudes at two different wavelengths (respectively, "excitation" and "emission"). The "emission amplitude ratio calculation" and "emission amplitude ratio calculation") are particularly advantageous because they are ratio calculation processes. The sensor provides an internal reference and determines the absolute brightness of the excitation source, the sensitivity of the detector, and the light scattering by the sample. As used herein, the term "interaction" refers to a process for adjusting a concentration of a compound to compensate for variations in the concentration of a compound. The term "fluorescent protein" refers to a chemically tagged protein whose fluorescence is due to the chemical tag. proteins, and emission peaks in the ultraviolet wavelengths (i.e., below about 400 nm) are considered fluorescent proteins for purposes of implementing the systems and methods discussed herein. Fluorescence occurs only in the presence of certain amino acids, such as tryptophan or tyrosine. Any polypeptide capable of emitting fluorescence when excited with appropriate electromagnetic radiation, except for polypeptides Generally, the term "protein" refers to a protein of interest. Fluorescent proteins useful for detecting or for use in implementing the methods discussed herein Proteins are proteins that derive their fluorescence from the autocatalytic formation of a chromophore. Proteins may be naturally occurring or engineered (i.e., variants or When used in reference to fluorescent proteins, "mutant" may contain an amino acid sequence. The term "protein" or "variant" refers to a protein that differs from a reference protein.
[0038] The term "blue fluorescent protein" is used herein to refer to a protein that emits blue fluorescence. The term "blue fluorescent protein" or "BFP" is used extensively throughout this specification. Used in the broadest sense, specifically mTagBFP, secBFP2, and from any species blue fluorescent proteins, as well as their variants (which have the ability to emit blue fluorescence) (for as long as you retain it).
[0039] The terms "mutant" or "variant" refer to a mutant of the corresponding wild-type or parent fluorescent protein. As used herein in reference to fluorescent proteins containing mutations to the We have demonstrated mutant fluorescent proteins with different fluorescent properties compared to the corresponding wild-type fluorescent proteins. To clarify this, we will discuss "spectral variants" or "spectral mutants" of fluorescent proteins. and are referred to herein.
[0040] Throughout this disclosure, various aspects of the implementation of the systems and methods discussed herein are , may be presented in range format. The range format is merely for convenience and brevity. and should not be construed as an inflexible limitation on the scope of the invention. Thus, the description of a range should be interpreted as including all possible subranges, if any. and each individual value within that range. For example, 1 to A range such as 6 is a subrange such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, etc. Ranges, as well as individual numbers within those ranges, e.g., 1, 2, 2.7, 3, 4, 5, 5.3, 6 and any whole or partial increments therebetween. This applies regardless of the extent of the scope.
[0041] In some aspects of the systems and methods discussed herein, the The software for executing the instructions may be stored on a non-transitory computer-readable medium. The software, when executed on a processor, implements the methods discussed herein. Perform some or all of the steps in
[0042] Aspects of the systems and methods discussed herein may be implemented in computer software. Particular embodiments relate to algorithms written in particular programming languages. or specific operating system or computing platform Although the systems and methods discussed herein may be described as being implemented on a The implementation of the calls and methods may not be implemented in any particular computing language, platform, or other medium. It is understood that the present invention is not limited to these combinations. The software that runs it is C, C++, C#, Objective-C, Java, J JavaScript, Python, PHP, Perl, Ruby, or Visual Basic written and compiled in any programming language, including but not limited to, The elements of the systems and methods discussed herein may be implemented as servers, , cloud instances, workstations, thin clients, mobile devices, Embedded microcontroller, television, or any other suitable computing Any acceptable computing platform, including but not limited to a It is further understood that this may be performed on a form.
[0043] Some implementations of the systems discussed herein may run on a computing device. The software described herein is described as software that is on a specific computing device (e.g., a dedicated server or workstation) Although it may be disclosed as working, the software may be portable in nature. In addition, software running on dedicated servers can also be used on desktop or mobile devices. Chairs, laptops, tablets, smartphones, watches, and wearable electronic devices or other wireless digital / cellular phones, televisions, cloud instances, embedded microcontroller, thin client device, or any other suitable computer The systems and methods discussed herein may be used in any of a wide variety of devices, including computer devices. and may be performed for the purpose of implementing the method.
[0044] Similarly, some of the implementations of the systems discussed herein may involve various wireless or wired computers. The systems discussed herein are described as communicating over a computer network. For purposes of implementing the system and method, the terms "network," "networking," and "networking" are used interchangeably. The term "working" refers to wired Ethernet, fiber optic connections, and various 802.11 standards. Wireless connectivity, including any of the following: 3G, 4G / LTE, or 5G networks Ruler WAN Infrastructure, Bluetooth®, Bluetooth Bluetooth Low Energy (BLE), or Zigbee (registered trademark) ) a communications link, or any means by which one electronic device can communicate with another electronic device It is understood that the present invention encompasses any other method of the present invention. The networking part of the system implementation is based on a virtual private network (VPN). ) may be implemented on
[0045] Aspects of the implementation of the systems and methods discussed herein include machine learning algorithms, It relates to machine learning engines, or neural networks. Neural networks are Based on various attributes of proteins, e.g., the atomic environment of amino acids within known proteins It may be trained to provide suggested sequences for one or more amino acids in a protein based on their attributes. In some embodiments, the attributes may include atom type, electrostatic, base The resulting amino acid sequence may include: The amino acids may be judged according to one or more quality indicators, and the weights of the attributes are determined by the quality indicators. In this way, the neural network can be optimized to maximize its performance. It can be trained to predict and optimize any quality metric that can be measured objectively. Examples of quality metrics on which neural networks can be trained include accuracy of wild-type amino acids, known stable amino acids, and The accuracy of the stabilized / destabilized positions, amino acid groups, and any other suitable type of In some embodiments, the neural network may include a multi-task neural network. It may have risk capabilities and allow for simultaneous prediction and optimization of multiple quality indicators.
[0046] In embodiments that implement such neural networks, queries can be generated in a variety of ways: The query can be performed thermally or via a desired parameter, e.g., a melting curve. Enhanced protein stability can be chemically embodied using guanidine or urea denaturation To do this, we ask the neural network to identify the amino acids in a given protein. The neural network of the implementation of the systems and methods discussed herein may be Its predicted identity (as assessed by the neural network) is comparable to its natural identity. identifies one or more amino acid residues in different proteins, thereby providing improved protein Proteins can be generated by mutating natural amino acid residues to predicted amino acid residues. As contemplated herein, predicted amino acid residues may be any naturally occurring or It may also be a non-natural (eg, artificial or synthetic) amino acid.
[0047] In some embodiments, the neural network is Training the neural network using the desired parameter values associated with the residues. By updating the neural network in this way, the optimal The ability of the neural network to suggest amino acid residues can be improved. In this study, training a neural network involves generating proteins that are mutated at predicted amino acid residues. This may involve using values of desired parameters associated with the protein. For example, In some embodiments, training the neural network comprises: The desired parameter values are predicted and the predicted values are compared with the parameters associated with known amino acids. Comparing it to the corresponding values and training a neural network based on the results of the comparison If the predicted value is the same as or substantially similar to the known value, the new value is The local network may be updated minimally or not at all. If the predicted value differs from the known amino acid value, the neural network tries to resolve this discrepancy. The neural network can be effectively updated to better compensate. Whether it's a retrained neural network or a new one, it will generate additional amino acids. Acid may be suggested.
[0048] The technology of the present application relates to increasing protein stability, which is different from other types of Protein parameters or attributes of the protein, e.g., half-life, activity, resistance to degradation, solubility, thermal stability Qualitative analysis, post-translational modification, enhanced pH tolerance, shortened maturation time, nucleic acid binding, protein-protein interaction Non-limiting applications of these techniques are possible, as they may be applied to crystalline, hydrophobic, or combinations thereof. It should be understood that the purpose of this is to train neural networks. Depending on the type of data received, the neural network can identify different types of proteins, It can be optimized for protein-protein interactions and / or protein attributes. In this way, neural networks can be trained to identify peptides, also called proteins. This allows for improved identification of possible amino acid sequences. Performing the neural network analysis may include inputting an initial amino acid sequence for the protein. The network may have been previously trained using different amino acid sequences. The query to the network is performed to find proteins with higher stability than the initial amino acid sequence. It may be related to the proposed amino acid sequence. Each residue of the proposed amino acid sequence The proposed amino acid sequence showing specific amino acids for may be received.
[0049] Inputting a sequence with a discrete representation, neural network with a continuous representation and sequentially input it as input to the neural network. By discretizing the output before serving it to the neural network, The techniques described herein related to performing machine learning may be applied to other machine learning applications. Such techniques can be particularly useful in applications where a final output having a discrete representation is desired. Such techniques use data that relate discrete attributes to the characteristics of a set of discrete attributes. By applying the model generated by the trained neural network, It can be generalized to specify a set of discrete attributes. In this case, the discrete attributes may include different amino acids.
[0050] In some embodiments, the model includes data resulting from molecular simulations. An initial sequence having, but not limited to, discrete attributes located at each position in the sequence can be used as input. Each of the discrete attributes in the initial series may be one of a plurality of discrete attributes. Querying a neural network involves inputting an initial set of discrete attributes. and an output series of discrete features having different levels of the feature from the initial series. and generating attributes in response to querying the neural network. In response, the output sequence and the different discrete attributes for each position in the output sequence are The associated values may be received from a neural network. The value of a discrete attribute is a new value for the level of the characteristic if the discrete attribute is selected for the position. The data may correspond to predictions of the neural network and form a continuous data set of values. The output sequence may be spread over discrete attributes of In some embodiments, the output sequence of discrete versions Identifying the value of the discrete attribute for each position in the series involves choosing, for each position in the series, the value of the discrete attribute for that position. The proposed set of discrete attributes may include selecting the discrete attribute having the highest value from the set of discrete attributes. The property may be received as an output specifying the discrete version.
[0051] In some embodiments, the iterative process performs a neural network analysis on the output sequence. querying the network, receiving the output sequence, and Additional iterations of the iterative process are formed by identifying the The process may include inputting a series of discrete versions of the output from the iterations. A sequence is generated when the current output sequence matches the previous output sequence from the previous iteration. may stop at
[0052] In some embodiments, a neutron detector is used to identify amino acid sequences with multiple quality indicators. than the desired value for a single quality metric, including for training neural networks. Rather, it is desirable to have a desired value for multiple quality metrics (e.g., a value higher than the values of another array). Such techniques allow the identification of proteins with different properties. This can be particularly useful in applications where identification of the proposed amino acid sequence for a protein is desirable. In implementations of such techniques, the training data is used to train the neural network. The data may include data relating to different properties for each of the amino acid sequences obtained. The models generated by training neural networks are composed of different combinations of properties. In some embodiments, the parameter The data may represent the weight between the first feature and the second feature, which is used to determine the proposed amino acid sequence. used to balance the likelihood that a sequence will have the first characteristic compared to the second characteristic. In some embodiments, training the neural network may include and assigning scores for the properties of the proposed amino acid sequence. It can be used to estimate values for the parameters of the model used to make the predictions. In some such embodiments, the training data comprises atomic sequences associated with amino acid sequences. The neural network may include a microenvironment, which is used to train the neural network. If so, generate a model that will be used to predict the proposed amino acid sequence. Training a neural network can involve assigning scores to parameters. All values can be estimated using the scores.
[0053] Biological applications of convolutional neural networks are relatively rare. Proteins are not analyzed as amino acid sequences, but rather as a set of proteins that are analyzed to determine their three-dimensional structure. One embodiment of the implementation of the methods discussed herein is A three-dimensional convolutional neural network characterizing the unique chemical environment of each amino acid in The same neural network is then trained to find the best fit for a given environment. The neural network described herein can predict the amino acid that , 1.6 million amino acid environments across 19,000 phylogenetically distant protein structures After training, the network achieved an in-sample accuracy of 80.0%. The out-of-sample accuracy was 72.5%, which is an improvement of about 20-30% (about 40%) over the current state of the art. out-of-sample precision).
[0054] Sites with large discrepancies between predicted and observed amino acids are considered to be stable. and targets for manipulating protein characteristics such as folding maturation. The systems and methods described herein are directed to three biological instances: beta-lactamases; enzyme antibiotic marker, blue fluorescent protein from coral, and yeast Candida We experimentally characterized the phosphomannose isomerase from C. albicans and identified its neural Predictions from the network demonstrate improved protein function and stability in vivo. These results predict new biological tools at the intersection of AI and molecular biology .
[0055] In one embodiment, the implementation of the methods discussed herein is performed using a neural network, e.g. For example, the neural network model published by Torng and Altman, referenced above, The implementation of the systems and methods discussed herein utilizes the following: As the experimental results discussed below show, published neural network designs are substantially The original Torng and Altman set was 32,760 training 3696 training protein families and 1601 test structures and 194 test protein families.
[0056] Implementations of the systems and methods discussed herein address the problem of protein stabilization. To achieve this, we build on the Torng and Altman framework. In the example, the protein crystal structure is treated like a three-dimensional image. There are many observations about individual amino acids and their atomic environments. This allows for a consistent frame of reference to be centered on one of these amino acids. , oxygen, nitrogen, sulfur, and carbon atoms in a 20 x 20 x 20 Å box is separated, removing all atoms associated with the central amino acid. This set of compatible amino acids is used for a three-dimensional convolutional neural network. This trained neural network , experimentally introduced destabilizing mutations can be detected.
[0057] Implementation of the systems and methods discussed herein may be used to identify novel stabilizing mutations. The improvements described herein improve the quality of predictions without known destabilizing factors. Not only can it justify mutations, but it can also identify unknown destabilizing residues and suggest stabilizing mutations. Make it sufficient.
[0058] In some implementations, the systems and methods discussed herein may include detecting input tampering. Such an implementation allows for the identification of wild-type amino acids located in favorable environments on proteins. This may narrow the sequence space to residues with very low wild-type probability. The improvements provided by implementations of the systems and methods discussed herein may be combined When combined, significant advances have been made in identifying candidate protein residues for improving overall utility. This can be described as several individual improvements that result in a significantly improved model.
[0059] Figure 1A shows a computer-implemented neural network for increasing synthetic protein properties. 1 is a diagram of an implementation of the work. Some properties of proteins that an engineer may wish to modify include: Maturation kinetics, thermal stability, K m、 K cat cation for proper folding or In 101, the protein is For each residue in the protein, the microenvironment may be translated, and a three-dimensional model of the protein and its Other methods for generating 3D models include: There is also a method where the unknown protein model is taken from a known protein structure. A fragment set constructed from a pool of known protein segments with amino acid sequences If a match is found, a segment match or a known protein model is selected ("Template" The residues of the amino acid sequence are mapped to residues in the template sequence (alignment) Constraints on various distances, angles, and dihedral angles in the sequence determine the alignment with the template structure. Comparison based on the achievement of spatial constraints when constraint violations are minimized. Protein modeling is one example. When a three-dimensional model of a protein crystal structure is generated, , a corresponding microenvironment associated with the structure is generated.
[0060] In some embodiments, the three-dimensional model merely illustrates or represents a protein without a microenvironment. The three-dimensional model may be displayed in some implementations as a three-dimensional array. In one example, the coordinates of the three-dimensional model are stored in a three-dimensional array. In an embodiment, the three-dimensional image may be generated from a three-dimensional model, The image data in the array may be mapped to the original array. A pixel can represent an addressable element of an image in two-dimensional space. Thus, a voxel represents an addressable element in three-dimensional space.
[0061] In some implementations, image features are combined through 3D convolutional and max-pooling layers. The three-dimensional filter in the three-dimensional convolution layer is a 20 amino acid microarray. We search for recurrent spatial patterns that best capture local biochemical features to separate microenvironments. The max pooling layer performs downsampling on the input and parallelizes the network. We will further discuss convolutional neural network architectures below. This will be further considered.
[0062] The first convolutional layer 121 detects low-level features through filters. Neural networks use convolutions to highlight features in a dataset. In the convolutional layer of a neural network, a filter is applied to a three-dimensional array. , to generate a feature map. In a convolutional layer, a filter is a function of the input and the filter elements. The input is stored as a feature map. In an embodiment, a 3x3x3 filter may be applied to the three-dimensional image.
[0063] The convolution filters and feature maps from the images are denoted by 102. In an embodiment of the present invention, a reference frame may be created around a central amino acid in the image, Features may be extracted around that central amino acid. Convolution of Images and Filters The feature map created from the filter summarizes the presence of filter-specific features in the image. Increasing the number of filters used increases the number of features that can be tracked. 100 filters were applied to create an 18x18x18 feature map. In the implementation, other numbers of filters may be used. The resulting feature map is then filtered to obtain the feature To account for non-linear patterns, an activation function may be passed through.
[0064] In some implementations, a normalized linear function with the formula f(x)=max(0,x) is used to calculate the activation A normalized linear activation function may be applied to the feature maps as a normalization function. The linear behavior makes this function easy to optimize, and then the neural network It allows us to achieve high prediction accuracy. Also, the normalized linear activation function can handle any negative It outputs zero for any given input, meaning it is not a truly linear function. The output of the convolutional layer in a convolutional neural network is a feature map. The values in the group may be passed through a normalized linear activation function.
[0065] The second convolutional layer is illustrated at 122. Increasing the number of convolutional layers increases the number of possible tracks. The complexity of the features can increase. The convolutional layers in 122 use different In some embodiments, the filters incorporate 100 filters. In order to ensure the accuracy of In some embodiments, a different filter may be incorporated into the second convolutional layer. In some embodiments, atoms associated with the central amino acid may be filtered out.
[0066] In some implementations, a smaller dataset of dimensions 16x16x16 is (In other implementations, other dimensions may be utilized, or more or fewer (A number of filters are applied.) The dot product of the convolutions in the second convolutional layer is The dataset 103 is a set of duplicates from the original protein images 101. It includes a feature map that tracks coarse features.
[0067] In some implementations, the first pooling layer of dimensions 2x2x2 is implemented as 123. A pooling layer may be implemented to downsample the data. A ring window may be applied to the feature map. The filtering layer outputs the maximum value of the data within the window and downsamples the data within the window. Max pooling emphasizes the most prominent features in the pooling window. In another embodiment, the pooling layer outputs the average value of the data within the window.
[0068] The downsampled data at 104 is split into 200 independent 8x8x8 arrays. By downsampling the data, the neural network can Having large amounts of data is discussed further below. It is advantageous to allow the network to fine-tune the accuracy of its weights so that It is possible, but with large amounts of data, neural networks will take a significant amount of processing time. Downsampling the data can reduce the number of computers required in the network. This can be important in neural networks to reduce computer computation. 2×2 pooling layer123, and downsampled data of dimensions 8×8×8. Although shown with the data, other implementations may use other sizes of pooling windows and downlinks. Sampled data may be used.
[0069] In some implementations, the subsequent convolutional layer 124 consists of 200 independent 2x2x2 filters. We use a filter to reprocess the downsampled data and extract features in the new feature map. In contrast to the 3x3x3, the smaller filter 2x2x2 To account for the downsampled data, a convolutional layer is implemented in 124. The depth of the convolution filter is determined by the depth of the data to facilitate the dot product matrix multiplication. In other implementations, other sizes or dimensions may be used, as discussed above. A modal filter may be used.
[0070] The convolutional layer 124 and the feature maps from the image are shown at 105. The feature map created from the convolution of the filtered data and the filter is The implementation shown in Figure 105 summarizes the presence of unique features of the filter. We have a 7×7 array. The dot product from the convolution further reduces the size of the data.
[0071] The convolutional layer 125 extracts 400 pixels from the low-resolution dataset 105 as shown. Using additional filters, such as by using an independent 2x2x2 filter, More complex features can be extracted. Increasing the number of filters applied to the image increases the number of features that are tracked. This data is downsampled from the pooling layer 123, Due to the substantial size reduction, more filters are applied in this convolutional layer. extract and analyze protein 101 images without the need for extensive processing or memory requirements. It can be called and emphasized.
[0072] The feature maps from the convolutional layer 125 are shown at 106. The feature map created from the convolution of the data and the filter is the filter-specific feature map in the image. The implementation shown in Figure 106 has 400 independent 6x6x6 arrays. Although there are ray, in various implementations other numbers or sizes of arrays may be used. The dot product from the convolution further reduces the size of the data.
[0073] In some implementations, it has dimensions 2x2x2 (or any other suitable dimension size). A second pooling layer is implemented at 126 to further downsample the data. In some embodiments, the same type of pooling layer is implemented in the first pooling layer. It may be implemented with a second pooling layer as shown in the figure. Depending on the type of pooling layer, , the pooling window used to downsample the data is determined. For example, max pooling layers may be implemented at 123 and 126. In some embodiments, different pooling layers may be implemented in a convolutional neural network. For example, a max pooling layer may be implemented at 123 and an average pooling layer at 126. The max pooling layer may be implemented in the most prominent part of the pooling window. While emphasizing prominent features, the average pooling layer outputs the average value of the data in the window. do.
[0074] In the illustrated implementation, the downsampled data at 107 is Although a separate 3x3x3 array is shown, other numbers or dimensions of arrays may be utilized. Having a large amount of data allows the network to determine its weights, as discussed further below. This can be advantageous to allow for fine-tuning of accuracy, but the large amount of data can make it difficult to The neural network can consume significant processing time. Ringing is a technique used to reduce the computational complexity required in the network. This can be useful in a mobile network.
[0075] Reducing the size of the data may result in further flattening of the data in some implementations. This means that the data can be arranged in a one-dimensional vector, rather than being fully connected. fully connected layers are flattened for the purposes of matrix multiplication that occurs in the connected layers. The layer 127 may receive a flattened one-dimensional vector of length 10800 (e.g. , from the 400x3x3x3 array in step 107, but the vector is (These may have different lengths.) In this case, each number in a one-dimensional vector is applied to a neuron. The neuron sums the inputs and , applying an activation function. In some embodiments, the activation function is a rectified linear function. In alternative embodiments, the activation function may be a hyperbolic tangent or a sigmoid function. stomach.
[0076] In the illustrated implementation, the first fully connected layer 127 is 108 kb long and has a length of 10800 kb. outputs a one-dimensional vector in (although other lengths may be used, as discussed above). The vector output by a fully connected layer represents a vector of real numbers. In some embodiments, the real number may be output and classified. Real numbers are used in subsequent fully connected neural networks to improve the accuracy of convolutional neural networks. The input may be further input to the selected layer.
[0077] In this embodiment, the output of the first fully connected layer 108 is shown at 128 The output of the first fully connected layer 108 is Since it is already a one-dimensional vector, it is flattened before being input to the subsequent fully connected layer. In some embodiments, to improve the accuracy of the neural network, The number of additional fully connected layers is determined by the number of new may be limited by the processing power of the computer running the network; or Adding a fully connected layer requires additional computational resources to process the additional fully connected layer. It may be limited by a small increase in accuracy compared to the increase in computation time.
[0078] In the illustrated implementation, the second fully connected layer 128 has a length of 100 at 109. It outputs a one-dimensional vector of zeros (although other lengths may be used). The vector output by the layer represents a vector of real numbers. , the real numbers may be output and classified. In other embodiments, the real numbers may be To improve the accuracy of the neural network, the inputs are further fed into the subsequent fully connected layers. Good too.
[0079] In some implementations, the output of the fully connected layer 109 is The softmax classifier uses the softmax function or normalization The exponential function is used to convert real inputs into normalized In an alternative embodiment, a sigmoid function is used to convert the The output of the neural network may be classified. The sigmoid function is The softmax function is a multi-class sigmoid function.
[0080] At 110, the output of the softmax layer is a set of 20 identified amino acids. The probability of improving the properties of a target protein (but with more or fewer amino acids) (This output is used to train additional convolutional neural networks.) Additional queries are added to allow the query to perform different queries given a predicted amino acid sequence. The output 110 may be a target tangent. They may also be used directly as predicted amino acids that improve protein properties.
[0081] Figure 1B shows the flow of the implementation of the method for determining the amino acid residue at the center of the microenvironment. This chart shows how a neural network classifies an output given a particular input. To learn the algorithm, the neural network is trained on known input / output pairs. Once the neural network has learned how to classify known input / output pairs, Once the algorithm is learned, the neural network predicts what the classified output should be. In this embodiment, the neural network During testing, neural networks are trained to predict amino acids at the center of a microenvironment. The network is provided with an amino acid sequence, analyzes the microenvironment surrounding the amino acid, and identifies naturally occurring amino acids. Neural networks can predict amino acid residues different from the amino acid residues. The acid is to mutate natural amino acid residues to predicted amino acid residues to produce improved proteins. It shows that it can be generated by
[0082] In step 130, in some implementations, a neural network is trained. A diverse protein sample set may be compiled or constructed for use in the analysis. The more diverse the sample set, the better the neural network will be at classifying it. For example, neural networks can be more robust than input / output networks during the first iteration of training. During the next training iteration, the input / output pairs are compared to the learning If the input / output pairs are similar to the ones the neural network has learned, It should work simply because the data is similar, not because the network is robust. A variety of input / output pairs can then be used for the third iteration. When inputs to the network are The similarity of the first two input / output pairs is likely to be much larger than it would have been if The property allows the neural network to distinguish similar input / output pairs in the first two iterations. It may fine-tune itself to learn, which is a way of "overtraining" the network. This can be called "doing something."
[0083] Alternatively, the second iteration of training may have distinct input / output pairs compared to the input / output pairs of the first iteration. When using output pairs, the neural network can classify a wider range of input / output pairs. During testing, the output is not known, so the network Ideally, the network should be able to classify a wide range of input / output pairs.
[0084] Therefore, in some implementations of step 130, The training dataset consists of proteins that are all phylogenetically divergent across a certain threshold. In various embodiments, the data set is constructed from at least 20%, 30%, 40%, %, or 50% phylogenetically divergent proteins. Filtering removes highly similar / duplicate proteins that may occur multiple times in the training set. Such improvements can be achieved by over-sampled proteins. This can reduce the bias that exists in the current state of the art.
[0085] In some embodiments, individual proteins in the training dataset are identified as proteins lacking annotation. These protein database (PDB) structures were modified by adding hydrogen atoms. In one embodiment, the addition of hydrogen atoms is performed using a software converter, e.g., pdb2pqr In another embodiment, the atoms are selected based on their binding capacity and the DNA backbone. Further separation is provided by the inclusion of other atoms, such as phosphorus, in the lattice.
[0086] In some embodiments, individual proteins in the training set are analyzed by partial charges, beta factors, Additional characteristics of proteins, including but not limited to, secondary structure, aromaticity, and polarity The protein model is modified by adding biophysical channels to take into account the Decorated.
[0087] In some embodiments, the high resolution model and the low resolution model of the same protein are If it can coexist in the protein database, the training data may be removed. According to some implementations of the methods discussed in the literature, the relevant structures have a resolution below a threshold. All genes are grouped together into groups with sequence similarity above a certain percentage threshold. As used herein, "resolution" typically refers to a resolution in angstroms. This refers to the resolution of the electron density map of a molecule, measured in Å. It is possible to resolve lower distances between points, meaning that more features of the molecular structure are visible. Therefore, molecular models with "lower" resolution are more sensitive to molecular models with "higher" resolution. In one example, the related structures, as well as the resolution below 2.5 Å and All genes with at least 50% sequence similarity were grouped together and The available structures with low resolution are selected for use in the training model, and the higher Lower resolution (lower quality) molecular models are removed.
[0088] In some embodiments, amino acid sampling is performed using equal numbers of all 20 amino acids. Expression was normalized to its abundance in the PDB relative to cysteine. In some embodiments, the amino acid sampling may be normalized to natural occurrence. In this case, amino acid sampling may be normalized to natural occurrence within a given species. Cysteine can be artificially assigned a high probability at any given position. Amino acids were modified in the data sample. Cysteine is the rarest amino acid observed in the PDB. amino acids, and therefore more abundant amino acids are undersampled and likely to be occupied. It is possible that the diversity of potential protein microenvironments was incompletely represented. Modifying the cysteine amino acids in the sample resulted in a significant increase in accuracy over the wild type. For each amino acid, the accuracy ranged from 96.7% to 32.8% (see Figure 2A). stomach).
[0089] In step 131, the amino acids in the protein are randomly selected from the amino acid sequence. In one embodiment, up to 50% of the amino acids in a protein may be sampled. Proteins were sampled unless they were large, in which case 100% of each protein was sampled. In another embodiment, the upper limit is 1000 amino acids per individual protein. The disclosed sampling method was to extract the outermost amino acid sequence of the protein. Remove bias in the dataset for residues.
[0090] In step 132, the three-dimensional model of the protein crystal structure is constructed by dividing each amino acid sequence comprising the structure. For example, microenvironments can be created with acid-associated microenvironments to generate three-dimensional models. Some methods, including others, involve the matching of an unknown protein model to a known protein. A fragment set constructed from a pool of candidate fragments taken from the structure of a known protein Segment matches where the segment matches an amino acid sequence, or a known protein model A template is selected (a "template"), and the residues in the amino acid sequence are matched to the residues in the template sequence. (alignment) and constraints on various distances, angles, and dihedrals in the sequence. is derived from an alignment with a template structure, minimizing violation of constraints. Comparative protein modeling based on the achievement of intermolecular constraints is one example. Once the three-dimensional model is generated, the microenvironments associated with each amino acid comprising the structure are also generated. One drawback of existing protein structure databases is that they require a large number of copies of each protein as new proteins are added. To create a 3D structure, different methods are used. Different methods may introduce different biases or artifacts that can affect the accuracy of the model. You can add new entities to the structure using the latest and same version of the same method. This ensures that the training structure is not an artifact or error present in older versions. This will ensure that the chemical composition will vary.
[0091] In step 133, the three-dimensional model generated from step 132 is In one example, the coordinates of the three-dimensional model may be stored in a three-dimensional array. In some embodiments, the three-dimensional image may be generated from a three-dimensional model. The dimensional image may be mapped into a three-dimensional array. The image data in the array may be represented as a voxel. The pixels are the addressable regions of the image in two-dimensional space. As such, voxels represent addressable elements in three-dimensional space.
[0092] In step 134, the image is fed to a convolutional layer in a convolutional neural network. The convolutional layer detects image features through filters. In a simplified example, a high-pass filter is used to detect the presence of a particular feature in a signal. The high pass filter detects the presence of high frequency signals. The output of the high pass filter is the signal with high frequencies. Similarly, image filters are designed to track specific features in an image. The more filters that are applied to an image, the more features that can be tracked.
[0093] In step 135, the image is convolved with a filter in a convolution layer to generate a In a convolutional layer, the filter extracts the filter-specific features of the input and the filter. The input is stored as a feature map by sliding over the element-wise dot product of the data.
[0094] The decision at 136 depends on whether there are more filters. As mentioned above, the more filters implemented, the more features that can be tracked in the image. Each filter is independently convolved with the image to create an independent feature map. If more filters are to be convolved with the image, steps 134 and 135 are repeated. When all of the filters have been convolved with the image, the process returns to step 13. Proceed to 7. In some embodiments, the feature maps are concatenated together and applied to the image. In other embodiments, the feature map may be as deep as the number of filters. may be processed one at a time.
[0095] In step 137, the activation function is The activation function is applied to the feature maps of the layer. It allows for the detection of nonlinear patterns in maps. The formula f(x)=max(0,x) A regularized linear function having the following formula may be applied to the feature maps: It behaves linearly for positive values, making this function easy to optimize, and then neural It allows the network to achieve higher accuracy. It also uses a regularized linear activation function. outputs zero for any negative input, meaning that it is not a truly linear function. Therefore, the output of a convolutional layer in a convolutional neural network is a feature map. The values in the feature maps are passed through a normalized linear activation function.
[0096] The decision at 138 depends on whether there are further convolutional layers. Increasing the number can increase the complexity of features that can be tracked. If so, a new filter may be applied to the image and the process repeats steps 134-138. In some embodiments, the filter may be used to ensure accuracy of the tracked features. In an alternative embodiment, different The filter may be incorporated into the second convolutional layer. If there are no further convolutional layers, The process proceeds to step 139.
[0097] In step 139, a pooling layer downsamples the data. A filtering window may be applied to the feature maps. The layer outputs the maximum value of the data in the window and downsamples the data in the window. Max pooling emphasizes the most prominent features in the pooling window. In this embodiment, the pooling layer outputs the average value of the data within the window.
[0098] The decision at 140 depends on whether there are further convolutional layers. Increasing the number can increase the complexity of features that can be tracked. If so, a new filter may be applied to the image and the process repeats steps 134-140. In some embodiments, the filter may be used to ensure accuracy of the tracked features. In an alternative embodiment, a different filter may be incorporated into the second convolutional layer. The repeated iterations from 34 to 138 and 134 to 140 demonstrate the flexibility and and provides increased complexity. Without further convolutional layers, the process proceeds as follows: Go to 141.
[0099] In step 141, in some implementations, the downsampled data is averaged. This means that the data is arranged in a one-dimensional vector. , are flattened for the purposes of matrix multiplication that occurs in fully connected layers.
[0100] In step 142, in some implementations, the flattened one-dimensional vector is The inputs are fed into the fully connected layers of a convolutional neural network. In a fully connected layer of a network, each number in a one-dimensional vector is applied to a neuron as an input. Neurons sum the inputs and apply an activation function. In an alternative embodiment, the activation function is a hyperbolic function. It may be a tangent or sigmoid function.
[0101] In some embodiments, the output of the first set of neurons in the fully connected layer is , may be input to another set of neurons via weights. The number of hidden layers in a fully connected network can be calculated using the In other words, the number of hidden layers in a neural network is determined by the number of neurons. It can change adaptively as the network learns how to classify the outputs.
[0102] In step 143, in some implementations, the network may include a fully connected network. Neurons are connected to other neurons by weights, which are the weights of some neurons. The strength of each neuron is adjusted to strengthen the effect of some neurons and weaken the effect of others. The training allows the neural network to better classify the output. The weights connecting the neurons determine how the neural network classifies the input or "training" In some embodiments, the neural network number of neurons may be removed. In other words, in the neural network The number of neurons that are active determines how quickly the neural network learns to classify the output. It adapts and changes as it learns.
[0103] The decision at 144 depends on whether there are additional fully connected layers. In some embodiments, the output of one fully connected layer is fed to a second fully connected layer. In some embodiments, to improve the accuracy of the neural network, To achieve this, additional fully connected layers are implemented. The number of additional fully connected layers is determined by the number of may be limited by the processing power of the computer running the network; or ,Adding a fully connected layer requires more computation time to process the additional fully connected layer. This may be limited by a small increase in accuracy compared to the increase in computation time. In this case, the output of one fully connected layer may be sufficient to classify an image. If there is a fully connected layer of This is repeated so that the neurons connected to each other are fed through the same network. If there are no connected layers, the process proceeds to step 145 .
[0104] In step 145, in some implementations, the fully connected layer is a vector of real numbers. In some embodiments, real numbers may be output and classified. In an embodiment, the output of the fully connected layer is input to a softmax classifier. A softmax classifier uses a softmax function or a normalized exponential function to compute real-valued Transforms the inputs of the eigenvalues into normalized probability distributions for the predicted output classes. In this state, a sigmoid function is used to classify the output of a convolutional neural network. The sigmoid function can be used when there is one class. The number is a multi-class sigmoid function. In some embodiments, the neural network The output of the work represents predicted amino acid residues at the center of chemical microenvironments.
[0105] For example, a neural network outputs a vector of length 20 containing 20 real numbers. The vector may have a length of 20, which means that there are 20 possible amino acids. The value in the vector is at the center of the microenvironment. The real numbers in the vector are passed through a softmax classifier to represent the likelihood of an amino acid. .
[0106] In step 146, in some implementations, the predicted amino acid residue is For example, the true amino acid vector is a vector of length 20. A single "1" indicates a natural amino acid in the center of a chemical environment, and a vector Other values in the vector hold "0".
[0107] In neural networks, learning involves comparing known input / output pairs during training. This type of learning is called supervised learning. The difference between predicted and known values is determined. The information is then backpropagated through the neural network. The weights are then used to calculate the error signal. This method of training a neural network is called backpropagation. It is called.
[0108] In step 147, in some implementations, the weights are updated via steepest descent. Equation 1 below shows how the weights are adjusted at each iteration n.
number
[0109] Gradient descent is an optimization technique that minimizes an objective function. In other words, gradient descent is , it is possible to adjust the unknown parameters in the direction of steepest descent. The weight values that optimize the classification accuracy of the network are unknown. It is an unknown parameter that is adjusted in the direction of the steep decline.
[0110] In some embodiments, the objective function may be a cross-entropy error function. Minimizing the difference entropy error function is the optimal solution for the probability distribution of predicted amino acid vectors and natural amino acid sequences. In some embodiments, the probability distribution of the amino acid vectors is minimized. , the objective function may be a squared error function. Minimizing the squared error objective function is It represents minimizing the instantaneous error of a neuron.
[0111] During each training iteration, the weights are adjusted to approach their optimal values. Depending on the neuron's position, different formulas are used to determine how the weights adjust to the objective function. Equation 2 below determines whether the weight between neuron i and neuron j is We show how it is adjusted for the cross-entropy error function.
number
[0112] In some embodiments, the weights may be adjusted whenever a modification to the weights is determined. This type of training may be called online or incremental training. One advantage of parallel training is that it allows neural networks to track small changes in the input. In some embodiments, the weights are used to determine the input / output characteristics of the neural network. It may be modified after receiving a batch of output pairs. This type of training is called batch training. One advantage of batch training is that it trains the neural network to optimized weight values. In this embodiment, the neural network is 1.6 million amino acid and microenvironment pairs were trained. In step 148, a counter is incremented. The neural network completes one round of batch training when the counter reaches 20. In other words, the neural network generates its output based on 20 input / output pairs. Evaluating itself completes one round of training.
[0113] The decision at 149 depends on whether the current batch of training samples is complete. If the number of training samples required to fill one batch is achieved, the network , proceed to step 150. As discussed above, one batch of training requires 20 inputs / outputs. A pair of forces is required. The number of samples required to fill one batch is not achieved. If not, the neural network repeats steps 134 to 149.
[0114] In step 150, the weight modifications temporarily stored in step 147 are performed. The weight values are calculated by adding a new batch of 20 input / output pairs to the newly modified weights The value is evaluated using the sum of the corrections.
[0115] The decision at 151 depends on whether the maximum number of training iterations has been reached. Completing one round of completes one training iteration. In some situations, the weights The weights never reach their optimum value because they keep fluctuating around it. Therefore, in some embodiments, the neural network A maximum number of iterations may be set to prevent the network from training indefinitely.
[0116] If the maximum number of iterations has not been reached, the neural network created in step 130 Retrain the network using different input / output pairs from the retrieved data samples. The iteration counter is used to determine when the neural network has completed one batch of training. After this, it is incremented in step 153.
[0117] When the maximum number of iterations is reached, the neural network may store the values of the weights. Step 152 indicates storing the weight values. These weights are used for training by the network. These are the weights that are trained and then used when testing the neural network. Therefore, it is stored in memory.
[0118] If the number of repeats is not reached, the error between the predicted amino acid residue and the known natural amino acid residue may be evaluated. This evaluation is performed in step 154. In some situations In this case, the error between the predicted value and the known natural value is so small that the error is acceptable. In these situations, the neural network does not need to continue training. The weight values that yield such small error rates are stored and used in subsequent tests. In some embodiments, the neural network may be The network either predicts one output very well or predicts one output very well. A small error over a few iterations to ensure that the model has not learned how to predict The network must be trained to maintain a small error over several iterations. By requesting a network, the network is more likely to classify a diverse range of inputs correctly. If the error between the predicted and known values is still too large, the neural network It can continue to train itself and repeat steps 131 to 154. In many implementations, During the iterations of steps 131-154, the neural network The neural network is trained using the
[0119] Figure 1C is a flowchart of the implementation of the method for enhancing synthetic protein properties during testing. In step 160, the weights stored from the training scenario are used in step 172. These weights are set as the weights of the fully connected layers in the input The weights are broad and It has been trained on a diverse set of inputs, so unknown inputs need to be classified. Used for.
[0120] In step 161, in some implementations, unknown proteins are randomly sampled. In one embodiment, up to 50% of the amino acids in the protein are are sampled unless they are large, in which case no more than 100 are sampled from each protein. In another embodiment, the upper limit is 2 amino acids per individual protein. 00 amino acids. The disclosed sampling method is for residues on the exterior of the protein. Remove bias from the dataset.
[0121] In step 162, the three-dimensional model of the protein crystal structure is constructed by dividing each amino acid sequence comprising the structure. Several methods for generating three-dimensional models can be used to create acid-associated microenvironments. There are other methods, but the unknown protein model is derived from a known protein structure. The fragment set is constructed from a pool of candidate fragments taken from known protein segments. Segment matches where the segment matches the amino acid sequence, or known protein models are selected. A template is selected and the residues in the amino acid sequence are mapped to residues in the template sequence. The alignment is then performed and constraints on various distances, angles, and dihedrals in the sequence are applied to the template. spatial constraints derived from alignment with rate structures, where violation of the constraints is minimized Comparative protein modeling based on the achievement of the three-dimensional model of protein crystal structure is one example. When a model is generated, the microenvironment associated with each amino acid, including its structure, is also generated. One obstacle to protein structure databases is the need to reconstruct crystal structures as new proteins are added. The difference is that different methods are used to create three-dimensional structures. The methods may introduce different biases or artifacts that may affect the accuracy of the model. By rebuilding the structure using the latest and same version of the same method, ,The training structure is chemically correct, rather than artifacts or errors present in older versions. It is certain that the composition will vary.
[0122] In step 163, the three-dimensional model generated from step 162 is converted into a three-dimensional array. In one example, the coordinates of the three-dimensional model may be stored in a three-dimensional array. In some embodiments, the three-dimensional image may be generated from a three-dimensional model. The dimensional image may be mapped into a three-dimensional array. The image data in the array may be represented as a voxel. The pixels are the addressable regions of the image in two-dimensional space. As such, voxels represent addressable elements in three-dimensional space.
[0123] In step 164, the image is fed to a convolutional layer in a convolutional neural network. The convolutional layer detects image features through filters. The filters are , which are designed to detect the presence of specific features in an image. The high-pass filter detects the presence of high-frequency signals. The output of the high-pass filter Similarly, image filters are filters that are designed to track specific features in an image. The more filters that are applied to an image, the more features that can be tracked. .
[0124] In step 165, the image is convolved with a filter in a convolution layer to generate a In a convolutional layer, the filter extracts the filter-specific features of the input and the filter. The input is stored as a feature map by sliding over the element-wise dot product of the data.
[0125] The decision at 166 depends on whether there are more filters. As mentioned above, the more filters implemented, the more features that can be tracked in the image. Each filter is independently convolved with the image to create an independent feature map. If more filters are to be convolved with the image, steps 164 and 165 are repeated. If all of the filters have been convolved with the image, the process continues at step 167. In some embodiments, the feature maps are concatenated together and applied to the image. A feature map may be created that is as deep as the number of filters. , may be processed one at a time.
[0126] In step 167, in some implementations, the activation function is The activation function is applied to the feature maps of the convolutional layers of the neural network. The function f(x) allows the detection of nonlinear patterns in the extracted feature maps. A rectified linear function with σ = max(0,x) is applied to the feature maps as the activation function. The normalized linear activation function behaves linearly for positive values, and optimizing this function This makes it easier for neural networks to achieve high prediction accuracy. Also, the normalized linear activation function outputs zero for any negative input, This means that is not a truly linear function. Therefore, convolutional neural networks The output of the convolutional layer in is a feature map, and the values in the feature map are calculated by the normalized linear activation function Numbers can be passed.
[0127] The decision at 168 depends on whether there are further convolutional layers. Increasing the number can increase the complexity of features that can be tracked. If so, a new filter may be applied to the image and steps 164-168 may be repeated. In some embodiments, the filter is a first convolution to ensure accuracy of the tracked features. In an alternative embodiment, a different filter is used in the second convolution. If there are no further convolutional layers, the process continues with step Go to 169.
[0128] In step 169, a pooling layer downsamples the data. A filtering window may be applied to the feature maps. The layer outputs the maximum value of the data in the window and downsamples the data in the window. Max pooling emphasizes the most prominent features in the pooling window. In this embodiment, the pooling layer outputs the average value of the data within the window.
[0129] The decision at 170 depends on whether there are further convolutional layers. Increasing the number can increase the complexity of features that can be tracked. If so, a new filter may be applied to the image and steps 164-170 may be repeated. In some embodiments, the filter is a first convolution to ensure accuracy of the tracked features. In an alternative embodiment, a different filter is used in the second convolution. If there are no further convolutional layers, the process continues with step Go to 171.
[0130] In step 171, in some implementations, the downsampled data is averaged. This means that the data is arranged in a one-dimensional vector. , are flattened for the purposes of matrix multiplication that occurs in fully connected layers.
[0131] In step 172, in some implementations, the flattened one-dimensional vector is The inputs are fed into the fully connected layers of a convolutional neural network. In a fully connected layer of a network, each number in a one-dimensional vector is applied to a neuron. A neuron sums the inputs and applies an activation function. In some embodiments, the activation function In an alternative embodiment, the activation function is a hyperbolic tangent or a sinusoidal function. It may also be a gmoid function.
[0132] In step 173, in some implementations, the network may include a fully connected network. Neurons are multiplied by weights. Weights in a fully connected network are multiplied by a step These weights are the weights initialized in 160. These weights are used to accurately distinguish the unknown inputs. The weights should be large and varied so that it is likely possible to classify It is used when unknown inputs are evaluated because it has been trained over a set.
[0133] The decision at 174 depends on whether there are additional fully connected layers. In some embodiments, the output of one fully connected layer is fed to a second fully connected layer. In some embodiments, to improve the accuracy of the neural network, To achieve this, additional fully connected layers are implemented. The number of additional fully connected layers is determined by the number of may be limited by the processing power of the computer running the network; or ,Adding a fully connected layer requires more computation time to process the additional fully connected layer. This may be limited by a small increase in accuracy compared to the increase in computation time. In this case, the output of one fully connected layer may be sufficient to classify an image. If there is a fully connected layer of This is repeated so that the neurons connected to each other are fed through the same network. If there are no connected layers, the process proceeds to step 175 .
[0134] In step 175, the fully connected layer outputs a vector of real numbers. In some embodiments, the real numbers may be output and sorted. The output of the connected layers is input to a softmax classifier, which Using a softmax function or a normalized exponential function, we convert real inputs into predicted In another embodiment, the sigmoid function is used to convert the The sigmoid function may be used to classify the output of a convolutional neural network. The softmax function can be used when there is one class. In some embodiments, the output of the neural network is a tangent function. The table shows candidate residues and amino acid residues predicted to improve protein quality indicators.
[0135] In step 176, the synthetic protein is generated according to the output of the neural network. Synthetic proteins can be generated as computational tools that run neural networks. The neural network is then run on a computing device. by another computing device with which you are communicating, Candidate amino acid residues identified by a genetic algorithm or by a neural network, and It can be generated by another entity making substitutions according to predicted amino acid residues. For example, some In an embodiment, the synthetic protein is generated by a neural network and / or Neural networks or computing that runs neural networks One or more substitutions are made according to the predicted amino acid residues and candidate residues identified by the device orientation. In some embodiments, the neural network may predict amino acid residues that are the same as natural amino acid residues. Neural networks can predict amino acid residues that are different from the natural amino acid residues. The predicted amino acids of the neural network are improved to predict natural amino acid residues. This indicates that the protein can be produced by mutating the amino acid residues. Proteins can be generated according to the output of a neural network.
[0136] Figure 1D shows a block diagram of a neural network during training, according to some implementations. The input is fed to the neural network at 180. As discussed above, In other words, neural networks can accept a variety of inputs. In embodiments, the neural network accepts amino acid sequences or residues. In an embodiment, the neural network is a set of discrete attributes located at each position in the set. A sequence of amino acids may be received.
[0137] In the block diagram, 181 represents a neural network that changes over time. As mentioned above, during training, the neural network adaptively adapts each iteration of new inputs / outputs. The weights are updated according to an error signal calculated by the difference between the predicted output and the known output. The neural network is adaptively updated because it is updated as the learning progresses.
[0138] In the block diagram, 182 indicates whether the output predicted by the neural network satisfies the query. For example, a neural network can identify specific amino acid residues that can be modified. In these situations, neural networks may be queried and trained to determine The output of the work may be amino acid residues, and the amino acid residues may be selected to have improved properties. In another embodiment, the ATPase inhibitors may be used to synthesize new proteins. The output of the network may be amino acid residues that can be used as substitutions, where the substitutions are: It may also be used to synthesize new proteins with improved properties. In this form, the neural network generates proteins with parameters different from the initial amino acid sequence. A query may be made against the proposed amino acid sequence for the protein. In this case, the output of the neural network is a specific amino acid sequence for each residue in the amino acid sequence. It may also be an amino acid sequence representing an acid.
[0139] In the block diagram, 186 represents the desired value. This type of training is In order to train a network, the inputs that correspond to the outputs must be known, which is called supervised training. During training, the neural network tries to produce results as close as possible to the desired values. You are asked to exert effort.
[0140] The desired value 186 and the output value from the neural network 182 are calculated at 185 The difference between the output value and the desired value is determined and passed through the neural network. This error signal 183 is propagated again, and the neural network learns from this error. As shown in Equations 1 and 2 above, the weights are based on the error signal. It is updated accordingly.
[0141] Figure 1E is a block diagram of a convolutional neural network, with some implementations. In the block diagram, 190 represents a convolution layer. The convolution layer converts an image into a Filters are designed to detect the presence of specific features in an image. In a simplified example, a high-pass filter detects the presence of high-frequency signals. The output of a frequency filter is the part of the signal that has high frequencies. Similarly, an image filter The more filters that are applied to an image, the more efficient the image becomes. The more features that can be tracked, the more likely it is that a person is using a mobile device.
[0142] In some implementations, the image is convolved with a filter in a convolutional layer to In the convolutional layer, the filter extracts the unique features of the input and the filter. The input is stored as a feature map by sliding over the element-wise dot product. The activation function is , which is applied to the feature maps of the convolutional layers of a convolutional neural network. The neural network detects nonlinear patterns in the extracted feature maps. A normalized linear function with the formula f(x)=max(0,x) is applied to the feature maps. A normalized linear activation function behaves linearly for positive values, and this function can be expressed as It makes optimization easy, and then the neural network achieves high prediction accuracy. Also, the normalized linear activation function outputs zero for any negative input. , which means it is not a truly linear function. Therefore, convolutional neural networks The output of the convolutional layer in the network is a feature map, and the values in the feature map are normalized linearly active. In other embodiments, a sigmoid function or a hyperbolic tangent function is used. It can be applied as an activation function.
[0143] The extracted feature map, which is acted upon by the activation function, is then calculated by 191 As shown, the data may be fed into a pooling layer, which downscales the data. A pooling window may be applied to the feature maps. In an embodiment, the pooling layer outputs the maximum value of the data in the window, Max pooling is most noticeable in the pooling window. Highlight prominent features.
[0144] The downsampled pooled data is then convolved in some implementations. may be flattened before being input to the fully connected layer 192 of the neural network. .
[0145] In some embodiments, a fully connected layer has only one set of neurons. In an alternative embodiment, the fully connected layer may be composed of neurons in the first layer 193. A set of neurons in the hidden layer 194 may be included, and a set of neurons in the subsequent hidden layer 194 may be included. The number of hidden layers in the connected neural network can be reduced. The number of hidden layers in the network increases as the neuron network learns how to classify the output. It can change adaptively.
[0146] In a fully connected layer, the neurons in each of layers 193 and 194 are connected to each other. Neurons are connected by weights. During training, weights are assigned to some neurons. The effect of each neuron is adjusted to strengthen the effect of some neurons and weaken the effect of others. Adjusting the strength allows the neural network to better classify the output. In some embodiments, the number of neurons in the neural network is reduced. In other words, the number of neurons that are active in the neural network The number of iterations changes adaptively as the neural network learns how to classify the output. do.
[0147] After training, the error between the predicted values and the known values can be very small, so the error is It can be deemed acceptable and the neural network does not need to continue training. In these situations, the weight values that resulted in such small error rates are stored and used for subsequent trials. In some embodiments, the neural network may be A network can predict one output very well, or predict one output very well. for a few iterations to ensure that the model has not learned how to predict incorrectly. A small error rate must be met. If we ask the network to do this, the network may be able to classify the various input ranges appropriately. Sexuality becomes higher.
[0148] In the block diagram, 195 represents the output of the neural network. The output of is a vector of real numbers. In some embodiments, the real numbers are output and sorted. In an alternative embodiment, the output of the fully connected layer may be used as a softmax classifier. is entered into
[0149] In the block diagram, 196 represents a softmax classifier layer. Using a softmax function or a normalized exponential function, we convert real inputs into predicted In another embodiment, the sigmoid function is used to convert the The sigmoid function may be used to classify the output of a convolutional neural network. The softmax function can be used when there is one class. In some embodiments, the output of the neural network is a tangent function. The table shows candidate residues and amino acid residues predicted to improve protein quality indicators. In embodiments, the output of the neural network is a specific sequence for each residue in the amino acid sequence. It may also be an amino acid sequence showing the amino acids
[0150] In some embodiments, problematic residues are identified and analyzed using multiple independently trained neural networks. By combining predictions from neural networks, new residues are proposed. By identifying residues based on an independently trained neural network, , neural networks emerge during training, and for any individual neural network Bias due to inherent idiosyncrasies can be eliminated. Many independent neural networks Averaging the work eliminates the quirks associated with any individual neural network.
[0151] Various improvements to the existing algorithm have cumulatively improved accuracy. As shown, the various improvements, taken together, in one embodiment, improve the model of wild-type amino acid prediction. We increased the accuracy of our algorithm from approximately 40% to over 70% across all amino acids.
[0152] Engineered proteins Implementation of the systems and methods discussed herein may involve the use of native or parent proteins. one or more that modify a desired trait or characteristic of a protein compared to the desired trait or characteristic The present invention further provides or identifies compositions comprising engineered proteins comprising mutations of In some embodiments, the information generated or identified by the implementation of the systems and methods discussed herein may be used to The modified proteins are suitable for three-dimensional folding in the implementation of the systems and methods discussed herein. One or more objects predicted by a neural network (3DCNN) prediction pipeline containing one or more mutations in amino acid residues to confer a desired trait or characteristic to the protein. The 3DCNN prediction pipeline is used to analyze the mutations in the predicted residues. The present invention relates to a method for generating a digital image, a digital video signal, or a digital audio signal, which is generated by implementing the systems and methods discussed herein. The engineered proteins identified herein are referred to as 3DCNN engineered proteins. It is called crystalline.
[0153] 3DCs generated or identified by implementation of the systems and methods discussed herein Examples of traits or properties that can be modified in NN-engineered proteins include stability, parental These include compatibility, activity, half-life, fluorescence properties, and susceptibility to photobleaching. Not limited.
[0154] 3DCs generated or identified by implementation of the systems and methods discussed herein NN engineered proteins can be made using chemical methods, e.g., 3DCN. Proteins engineered with N have been synthesized using solid-phase techniques (Roberge JY et al. (1999 5) Science 269:202-204) and cleaved from the resin. Purification can be achieved by preparative high performance liquid chromatography. Automated synthesis can be performed, for example, in a manufacturing facility. Peptide synthesis was performed on an ABI 431 A peptide synthesizer (PerkinElmer) according to the instructions provided by the manufacturer. This can be achieved using a fluoroscopy (e.g., a fluoroscopy kit) or a fluoroscopy kit (
[0155] Alternatively, proteins engineered in 3DCNNs can be synthesized by translation of coding nucleic acid sequences. may be produced by recombinant means or by truncation from a longer protein sequence. The composition of proteins engineered in 3DCNNs can be confirmed by amino acid analysis or sequencing. It may be recognized.
[0156] 3DCs generated or identified by implementation of the systems and methods discussed herein NN-engineered protein variants are those in which (i) one or more of the amino acid residues It may be substituted with a conserved or non-conserved amino acid residue (preferably a conserved amino acid residue), and such such substituted amino acid residues may or may not be encoded by the genetic code, (ii) one or more modified amino acid residues, e.g., residues modified by the attachment of a substituent group; (iii) protein fragments engineered with 3DCNN, and / or (iv ) The proteins engineered in 3DCNN are fused with another protein or polypeptide. The fragments may be proteins from the original 3DCNN engineered protein sequence. This includes polypeptides produced via proteolytic cleavage (including multi-site proteolysis). Variants may be post-translationally or chemically modified. Such variants are within the scope of the present invention. It is deemed to be within the scope of those skilled in the art from the teachings herein.
[0157] As is known in the art, "similarity" between two polypeptides is measured by the similarity of the two polypeptides. The amino acid sequence of the polypeptide and its conserved amino acid substitutions are compared with the sequence of a second polypeptide. A variant is a sequence that differs from the original sequence and is determined by comparing the sequence of interest. Residues per segment of interest that differ from the original sequence by less than 40% of the residues per segment differ from the original sequence by less than 25% of residues and by less than 10% of residues per segment of interest. or differ from the original protein sequence by only a few residues per segment of interest. and at the same time, the functionality of the original sequence and / or ubiquitin or ubiquitinated proteins. Polypeptides that are sufficiently homologous to the original sequence to retain the ability to bind to proteins The term "antisense oligonucleotide" is defined to include a sequence of amino acids. At least 60%, 65%, 70%, 72%, and 74% of the original amino acid sequence ,76%,78%,80%,90%,91%,92%,93%,94%,95%,96% , 97%, 98%, or 99% similar or identical amino acid sequences. The identity between two amino acid sequences can be determined preferably using the BLASTP algorithm. [BLAST Manual, Altschul, S., et al., NCBI NLM NIH Bethesda,Md.20894,Altschul,S.,et al., J. Mol. Biol. 215:403-410 (1990)]. and is determined by.
[0158] 3DCs generated or identified by implementation of the systems and methods discussed herein NN engineered proteins can be post-translationally modified. For example, the sequences discussed herein can be modified by the Post-translational modifications included within the scope of implementation of the system and method include signal peptide cleavage. , glycosylation, acetylation, isoprenylation, proteolysis, myristoylation, Some modifications include protein folding, and proteolytic processing. Alternatively, the processing event requires the introduction of additional biological machinery, e.g., a signal peptide. Processing events such as cleavage and core glycosylation were observed in canine microsomal membranes or Xenop by adding US egg extract (U.S. Patent No. 6,103,489) to the standard translation reaction. and inspected.
[0159] 3DCs generated or identified by implementation of the systems and methods discussed herein NN-engineered proteins can incorporate unnatural amino acids either post-translationally or during translation. The present invention may include unnatural amino acids formed by introducing unnatural amino acids during protein translation. Various approaches to introduce amino acids are available. For example, suppressor Special tRNAs, such as suppressor tRNAs, which are tRNAs with specific properties, It is used in the process of selective non-natural amino acid substitution (SNAAR). mRNA and RNA sequences that act to target unnatural amino acids to specific sites during protein synthesis. and a unique codon on the suppressor tRNA (described in WO90 / 05785). However, suppressor tRNAs are present in the protein translation system. It must not be recognizable by aminoacyl-tRNA synthase. Specifically modifying amino acids without significantly altering the functional activity of aminoacylated tRNAs Using an enzyme reaction, the tRNA molecule is aminoacylated to form the unnatural amino acid. These reactions are called post-aminoacylation modifications. For example, the cognate tRNA ( tRNA LYSThe epsilon-amino group of lysine linked to ) provides amine-specific photoaffinity It may be modified with a label.
[0160] 3DCs generated or identified by implementation of the systems and methods discussed herein NN engineered proteins can be synthesized by combining them with other proteins, such as ribosomal proteins, to prepare fusion proteins. This may be, for example, an N- or C-terminal fusion protein. This may be achieved by synthesis of a fusion protein, provided that the resulting fusion protein is a 3DCNN. provided that the functionality of the engineered protein is retained.
[0161] 3DCNN-engineered protein mimics In some embodiments, the subject compositions comprise peptides of 3DCNN-engineered proteins. Peptidomimetics are compounds based on or containing peptides and proteins. Compounds derived from the above. The peptidomimetics that are constructed or identified typically incorporate unnatural amino acids, conformational constraints, etc. Structural modification of protein sequences manipulated by known 3DCNNs using electrostatic substitution, etc. The subject peptidomimetics can be obtained by combining structures between peptides and non-peptide synthetic structures. constitute a spatial continuum, and thus peptidomimetics delineate the pharmacophore and Translating peptides into non-peptide compounds with the activity of 3DCNN-engineered proteins This can be useful to help
[0162] Furthermore, as will be apparent from the present disclosure, the proteins engineered in the subject 3DCNN Such peptidomimetics may be non-hydrolyzable (e.g., For example, stability to proteases or other physiological conditions that degrade the corresponding peptide. improved specificity and / or potency, and improved intracellular localization of peptidomimetics For illustrative purposes, the systems discussed herein may have attributes such as increased cell permeability for Peptide analogs generated or identified by implementation of the systems and methods include, for example, benzo[a]pyridin-3-one. Diazepines (e.g., Freidinger et al. in Peptides:C hemistry and biology, GR Marshall ed., ES COM Publisher: Leiden, Netherlands, 1988 substituted gama lactam ring ring)(Garvey et al.in Peptides:Chemistry and Biology,GRMarshall ed.,ESCOM Publ isher:Leiden,Netherlands,1988,p123), C-7 model Mimics (Huffman et al. in Peptides: Chemistry a nd Biology,GRMarshall ed.,ESCOM Publicis her: Leiden, Netherlands, 1988, p.105), keto-methionine Ren pseudopeptides (Ewenson et al. (1986) J Med Chem 2 9:295, and Ewenson et al. in Peptides:Struc ture and Function(Proceedings of the 9th American Peptide Symposium)Pierce Chemi cal Co. Rockland, Ill., 1985), β-turn dipeptide core ( Nagai et al. (1985) Tetrahedron Lett 26:64 7, and Sato et al. (1986) J Chem Soc Perkin Trans 1:1231), β-amino alcohols (Gordon et al. (1 985)Biochem Biophys Res Commun 126:419, and Dann et al. (1986) Biochem Biophys Res C ommun 134:71), diaminoketones (Natarajan et al. (1 984)Biochem Biophys Res Commun 124:141), and methyleneamino modification (Roark et al. in Peptides: Ch emistry and Biology, GR Marshall ed., ESC OM Publisher:Leiden,Netherlands,1988,p13 4) can be generated using Session III: Analytical ic and synthetic methods,in Peptides: Chemistry and Biology,GR Marshall ed.,E SCOM Publisher: Leiden, Netherlands, 1988) Please refer to.
[0163] Various approaches can be implemented to generate engineered protein peptidomimetics in 3DCNNs. In addition to suitable side chain substitutions, implementations of the systems and methods discussed herein may also be used to modify peptide secondary The use of conformationally constrained mimetics of the structure for the amide bond of a peptide is contemplated. A number of surrogates have been developed. Frequently used surrogates for the amide bond include: The following groups: (i) trans-olefins, (ii) fluoroalkenes, (iii) methylene alkenes (iv) phosphonamides, and (v) sulfonamides.
[0164] nucleic acid In one embodiment, implementations of the systems and methods discussed herein are used to generating or isolating a nucleic acid comprising a nucleotide sequence encoding a protein engineered with NN; It can be identified.
[0165] Alternatively, the nucleotide sequence encoding the protein engineered in the 3DCNN can be obtained The polynucleotides to be analyzed may be polynucleotides according to implementation of the systems and methods discussed herein. Sequence variations relative to the original nucleotide sequence, provided that they encode a peptide, e.g. It may contain one or more nucleotide substitutions, insertions, and / or deletions. Therefore, the implementation of the systems and methods discussed herein can be used to The nucleotide sequence is substantially identical to that of the protein engineered by the 3DCNN. The nucleotide sequence corresponding to the target gene can be generated or identified.
[0166] As used herein, a nucleotide sequence means a sequence of a nucleotide that is At least 60%, at least 70%, at least 85%, at least 95%, at least At least 96%, at least 97%, at least 98%, or at least 99% of the nucleotides Any of the nucleotide sequences described herein may be used if they have a degree of identity to the nucleotide sequence. The nucleotides encoding the proteins engineered in 3DCNN are "substantially identical" to the Nucleotide sequences that are substantially homologous to the nucleotide sequence are typically, for example, conservative substitutions. or by introducing non-conservative substitutions, based on the information contained in the nucleotide sequence. Based on the above, generated or determined by the implementation of the systems and methods discussed herein The polypeptide may be isolated from the organism that produces it. Other examples of possible modifications include insertion of one or more nucleotides at either end of the sequence, Addition of a nucleotide or deletion of one or more nucleotides at any end or within the sequence The identity between two nucleotide sequences is preferably determined using the BLASTN algorithm. Rhythm [BLAST Manual, Altschul, S., et al., NCBI NLM NIH Bethesda,Md.20894,Altschul,S.,e t al., J. Mol. Biol. 215:403-410(1990)]. It is determined by:
[0167] In another aspect, implementations of the systems and methods discussed herein may be used to Constructs containing nucleotide sequences encoding N-engineered proteins or derivatives thereof In certain embodiments, the construct may be transcribed, and optionally translated, to generate or identify a target gene. The construct is operably linked to the systems and regulatory elements discussed herein. operably linked to the expression of a nucleotide sequence produced or identified by implementing the method. Regulatory sequences can be incorporated, thus forming an expression cassette.
[0168] Proteins engineered with 3DCNN or chimeric 3DCNN , may be prepared using recombinant DNA methods. The nucleic acid molecules encoding proteins or proteins engineered in chimeric 3DCNNs are then synthesized in 3D Good results for proteins engineered with CNN or chimeric 3DCNN It may be incorporated into an appropriate expression vector to ensure expression.
[0169] Thus, in another aspect, implementations of the systems and methods discussed herein may be used. and the nuclei generated or identified by implementation of the systems and methods discussed herein. A vector containing the nucleic acid sequence or construct may be generated or identified. , depending on the host cell into which it is subsequently introduced. In certain embodiments, The vectors generated or identified by implementation of the systems and methods are expression vectors. Suitable host cells include a wide variety of prokaryotic and eukaryotic host cells. In certain embodiments, Expression vectors include viral vectors, bacterial vectors, and mammalian cell vectors. Prokaryotic and / or eukaryotic vector-based systems are selected from the group consisting of: and the like, for use with implementations of the systems and methods discussed in Many such systems can be used commercially to produce cognate polypeptides. is readily and widely available.
[0170] Additionally, the expression vector may be provided to the cell in the form of a viral vector. Viruses that are useful as vectors include retroviruses, adenoviruses, and adeno-associated viruses. viruses, herpes viruses, and lentiviruses. Generally, a suitable vector will contain an origin of replication, a promoter, or a promoter sequence that is functional in at least one organism. A motor sequence, convenient restriction endonuclease sites, and one or more selectable markers (See, for example, WO 01 / 96584, WO 01 / 29058, and U.S. (See US Patent No. 6,326,193).
[0171] Suitable vectors for inserting the polynucleotide include expression vectors in prokaryotes, e.g., p UC18, pUC19, Bluescript and its derivatives, mp18, mp19, pBR322, pMB9, ColE1, pCR1, RP4, phage, and "shuttle" vectors, such as pSA3 and pAT28; expression vectors in yeast, such as 2min Clonal plasmid type vectors, integrating vectors, YEP vectors, centromeric vectors plasmids; expression vectors in insect cells, such as the pAC series and pVL Vectors; expression vectors in plants, e.g., pIBI, pEarleyGate, pAV A, pCAMBIA, pGSA, pGWB, pMDC, pMY, pORE series, etc.; and viral vector-based expression vectors in eukaryotic cells (adenovirus, adenovirus, viruses related to viruses, such as retroviruses and, in particular, lentiviruses); and non-viral vectors, such as pSilencer 4.1-CMV (Ambion ), pcDNA3, pcDNA3.1 / hyg, pHMCV / Zeo, pCR3.1, p EFI / His, pIND / GS, pRc / HCMV2, pSV40 / Zeo2, pTR ACER-HCMV, pUB6 / V5-His, pVAX1, pZeoSV2, pCI, pSVL and PKSV-10, pBPV-1, pML2d, and pTDT1 derived It is a vector.
[0172] Illustratively, a vector into which a nucleic acid sequence is introduced may be introduced into a host cell to induce transcription of the host cell. The plasmid may or may not be integrated into the genome of the host. Nucleotide sequences generated or identified by implementation of the systems and methods discussed in Illustrative, non-limiting examples of gene constructs into which the gene construct may be inserted include those for expression in eukaryotic cells. Examples include tet-on inducible vectors for expression.
[0173] In certain embodiments, the vector is a vector useful for transforming animal cells. .
[0174] The recombinant expression vectors can also be used to express engineered proteins or chimeric 3DCNNs. Increased expression of proteins engineered with NN, proteins engineered with 3DCNN, or Encode the portion of the protein engineered in the chimeric 3DCNN that improves solubility, or engineered in 3DCNNs by acting as ligands for affinity purification and / or Nucleic acid molecules that aid in the purification of proteins or proteins engineered in chimeric 3DCNNs For example, proteolytic cleavage sites may be included in the 3DCNN engineered proteins. After the fusion protein is purified, the fusion protein is inserted into the 3DCNN. Allows for the separation of engineered proteins or chimeric proteins engineered in 3DCNNs Examples of fusion expression vectors include glutathione S-transferase (GST) , maltose E binding protein, or protein A fused to the recombinant protein, respectively. pGEX (Amrad Corp., Melbourne, Australia) , pMAL (New England Biolabs, Beverly, Mass.) and pRIT5 (Pharmacia, Piscataway, NJ). can be.
[0175] Additional promoter elements, i.e., enhancers, regulate the frequency of transcription initiation. Typically, these are located in the region 30–110 bp upstream of the initiation site, but many promoters Promoters have recently been shown to contain functional elements downstream of the initiation site. The spacing between motor elements is often flexible, resulting in promoter function varying depending on the element They are retained when they invert or move relative to each other. In the mutant, the spacing between promoter elements increases to 50 bp apart before activity begins to decline. Depending on the promoter, the individual elements can act either cooperatively or independently. It appears to be able to function to activate transcription.
[0176] The promoter is a sequence of 5 non-coding sequences located upstream of the coding segment and / or exons. to a gene or polynucleotide sequence, such as may be obtained by isolating the gene sequence. It may also be a promoter with which it is naturally associated. Such a promoter is referred to as "endogenous." Similarly, an enhancer may be located either downstream or upstream of the sequence. It may be an enhancer naturally associated with the polynucleotide sequence. The coding polynucleotide segment is coupled to a recombinant or heterologous promoter. (This refers to a promoter that is not normally associated with a polynucleotide sequence in its natural environment. ) under the control of a recombinant or heterologous enhancer. Also included are enhancers that are not normally associated with a polynucleotide sequence in its natural environment. Such promoters or enhancers may be promoters or enhancers of other genes. enhancers, as well as any other prokaryotic, viral, or eukaryotic cell. promoters or enhancers that are not "naturally occurring," i.e., promoters containing different elements of different transcriptional regulatory regions and / or mutations that alter expression The promoter and enhancer nucleic acid sequences may also include promoters or enhancers. In addition to synthetically producing the sequences, the sequences may be synthesized in the context of the compositions disclosed herein. Produced using nucleic acid amplification techniques, including recombinant cloning and / or PCR™ (U.S. Patent No. 4,683,202, U.S. Patent No. 5,928,906). to regulate transcription and / or expression of sequences within non-nuclear organelles such as mitochondria and chloroplasts. It is contemplated that directing control sequences may also be used.
[0177] Naturally, the cell type, organelle, and organism chosen for expression will determine A promoter and / or enhancer that effectively directs expression of the DNA segment It is important to use a promoter that can be used to express, for example, recombinant proteins and and / or the introduced DNA segment is advantageous for large-scale production of the peptide. Constitutive, tissue-specific, inducible, and cytogenetic expression of β-glucan is possible under conditions appropriate to direct high levels of β-glucan. and / or may be useful. The promoter may be heterologous or endogenous.
[0178] The promoter sequences exemplified in the experimental examples presented herein are those encoding immediate early cytomegaloviruses. The promoter sequence is a CMV promoter sequence. A potent polypeptide capable of promoting high-level expression of any polynucleotide sequence linked to it. However, the first Simian virus 40 (SV40) promoter sequence phase promoter, mouse mammary tumor virus (MMTV), human immunodeficiency virus (HIV) ) long terminal repeat (LTR) promoter, Moloney virus promoter, avian leukosis Viral promoter, Epstein-Barr virus immediate early promoter, Rous meat tumor virus promoters, as well as human gene promoters, including but not limited to: However, the actin promoter, myosin promoter, hemoglobin promoter, and and other constitutive promoters, including but not limited to the muscle creatine promoter. Additionally, implementations of the systems and methods discussed herein may also be used. The present invention is not limited to the use of constitutive promoters; inducible promoters are also contemplated herein. Such systems and methods may be generated or identified through implementation of the systems and methods described herein. The use of inducible promoters generated or identified through the system or method provides for the expression of such promoters. and turning on expression of the operably linked polynucleotide sequence when expression is desired. or to provide a molecular switch capable of turning off expression when expression is not desired. Examples of inducible promoters include the metallothionine promoter, the glucocorticoid promoter, and the The promoters of the tetracycline, progesterone, and Additionally, the systems and methods discussed herein include, but are not limited to: Implementation of the method may allow for the use of tissue-specific promoters, which promoters may be: It is active only in the desired tissue. Examples of tissue-specific promoters include the HER-2 promoter. These include, but are not limited to, motor and PSA-related promoter sequences.
[0179] In one embodiment, expression of the nucleic acid is externally controlled. For example, in one embodiment, expression using the doxycycline Tet-On system or other inducible or repressible expression systems It is controlled externally.
[0180] Recombinant expression vectors also allow for the selection of transformed or transfected host cells. The gene may contain a selectable marker gene that facilitates the selection of the desired gene. It confers resistance to certain drugs, such as G418 and hygromycin, and β-galactosidases. acetyltransferase, chloramphenicol acetyltransferase, firefly luciferase, or an immunoglobulin or a portion thereof, e.g., an immunoglobulin, preferably an IgG A selectable marker is a gene that encodes a protein such as the Fc portion of a target gene. The nucleic acid may be introduced on a vector separate from the nucleic acid to be introduced.
[0181] The reporter gene identifies potentially transfected cells and identifies the regulatory sequences. Reporter genes are generally used to assess functionality in recipient organisms or or tissues, and the expression is not present in or expressed by them, e.g. encodes a protein that is exhibited by some easily detectable property, such as enzymatic activity The reporter gene expression is determined by the amount of DNA transferred into the recipient cells. Assays are performed at suitable time points.
[0182] Exemplary reporter genes include luciferase, beta-galactosidase, chloramphenicol, and guanidine. Phenicol acetyltransferase, secreted alkaline phosphatase, or green Genes encoding fluorescent proteins, including but not limited to fluorescent protein genes (e.g., Ui-Tei et al., 2000 FEBS Lett. 4 79:79-82).
[0183] In one embodiment, the present invention provides a method for detecting a cellular signal generated or generated by an implementation of the systems and methods discussed herein. The proteins engineered in 3DCNN that are identified are reporter genes and are suitable for expression. For example, in one embodiment, a compound produced or included in such a system or method may be The proteins engineered in the 3DCNN that are identified or identified are blue fluorescent proteins with increased fluorescence activity. In such embodiments, the implementation of the systems and methods discussed herein Nucleotides encoding proteins engineered by 3DCNNs generated or identified by the device A nucleotide sequence may be incorporated into the expression system to allow detection of the heterologous protein sequence. stomach.
[0184] The recombinant expression vector may be introduced into a host cell to produce a recombinant cell. The cells may be prokaryotic, eukaryotic, or prokaryotic. Vectors produced or identified by the method can be used to transform, for example, eukaryotic cells, such as yeast cells. cells, Saccharomyces cerevisiae, or mammalian cells, e.g. , epithelial kidney 293 cells or U2OS cells, or prokaryotic cells, such as bacteria, Esch Transformation of Escherichia coli or Bacillus subtilis Nucleic acids can be purified by co-precipitation with calcium phosphate or calcium chloride, DEAE-de transfection, lipofectin, electroporation, or microinjection The vector can be introduced into a cell using conventional techniques such as injection. Suitable methods for transfecting and transfecting the endonuclease-dependent nucleotide sequences are described by Sambrook et al. (Mo Lecular Cloning: A Laboratory Manual, 2nd Edition,Cold Spring Harbor Laboratory pr ess (1989)), and other experimental texts.
[0185] For example, the information generated or identified by the implementation of the systems and methods discussed herein Proteins engineered with 3DCNN or chimeric 3DCNN are , bacterial cells, e.g., E. coli, insect cells (using baculovirus), yeast cells The expression can be in endothelial cells, endothelial cells, or mammalian cells. Other suitable host cells are described in Goeddel, Gen. e Expression Technology:Methods in Enzym ology 185, Academic Press, San Diego, Calif. .(1991).
[0186] Modified blue fluorescent protein In one embodiment, implementations of the systems and methods discussed herein may be used to: In certain embodiments, the composition may comprise a stable BFP2 variant protein. The present invention relates to a secBFP2 variant protein comprising one or more mutations that enhance the identification of secBFP2. In certain embodiments, the secBFP2 variant protein has a specificity similar to that of wild-type secBFP2. This results in enhanced stability, enhanced fluorescence, enhanced half-life, and slower photobleaching. Indicates one or more of the following.
[0187] In some embodiments, the secBFP2 variant protein comprises one or more mutations. For example, in some embodiments, the secBFP2 variant The protein is expressed at T18, S28, Y96, and S in relation to the full-length wild-type secBFP2. One selected from 114, V124, T127, D151, N173, and R198 In one embodiment, the secBFP2 comprises one or more mutations in the above residues. Long wild-type secBFP2 Contains the amino acid sequence of JPEG2025118694000003.jpg40170.
[0188] In certain embodiments, within the secBFP2 variant proteins described herein The designations of the mutations are relative to SEQ ID NO: 1. For example, secBFP containing a mutation in T18 The term "variant protein" refers to secBFP2, but not to the full-length wild-type secBFP2 (sequence number 2). It has a mutation at the residue that correlates with the threonine at position 18 in column number 1).
[0189] In some embodiments, the secBFP2 variant protein is a full-length wild-type secBFP2 variant. In relation to cBFP2, T18X, S28X, Y96X, S114X, V124X, T1 27X, D151X, N173X, and R198X (where X is any amino acid) Some embodiments include secBFP2 having one or more mutations selected from the group consisting of: In this study, secBFP2 variant proteins were compared with full-length wild-type secBFP2. , T18W, T18V, T18E, S28A, Y96F, S114V, S114T, V1 24T, V124Y, V124W, T127P, T127L, T127R, T127D, D151G, N173T, N173H, N173R, N173S, R198V, and R 198L.
[0190] In one embodiment, the secBFP2 variant protein comprises a T18X mutation (wherein X In one embodiment, the secBFP2 comprises a secBFP2 containing any amino acid. The P2 variant protein contains the T18W mutation, relative to the full-length wild-type secBFP2. Contains secBFP2 containing the T18V mutation or the T18E mutation.
[0191] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000004.jpg40170, or any variant or fragment thereof.
[0192] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000005.jpg43170, or any variant or fragment thereof.
[0193] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000006.jpg41170, or any variant or fragment thereof.
[0194] In one embodiment, the secBFP2 variant protein has a S28X mutation (wherein X In one embodiment, the secBFP2 comprises a secBFP2 containing any amino acid. The P2 variant protein contains the S28A mutation relative to full-length wild-type secBFP2. Contains secBFP2.
[0195] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000007.jpg40170, or any variant or fragment thereof.
[0196] In one embodiment, the secBFP2 variant protein has a T96X mutation (wherein X In one embodiment, the secBFP2 comprises a secBFP2 containing any amino acid. The P2 variant protein contains the Y96F mutation relative to full-length wild-type secBFP2. Contains secBFP2.
[0197] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000008.jpg41170, or any variant or fragment thereof.
[0198] In one embodiment, the secBFP2 variant protein comprises a S114X mutation, X is any amino acid. The FP2 variant protein contains the S114V mutation relative to the full-length wild-type secBFP2. Contains secBFP2 with a mutation or S114T mutation.
[0199] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000009.jpg41170, or any variant or fragment thereof.
[0200] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000010.jpg41170, or any variant or fragment thereof.
[0201] In one embodiment, the secBFP2 variant protein comprises a V124X mutation, X is any amino acid. The FP2 variant protein contains the V124T mutation relative to the full-length wild-type secBFP2. secBFP2 containing a mutation, V124Y mutation, or V124W mutation.
[0202] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000011.jpg40170, or any variant or fragment thereof.
[0203] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000012.jpg42170, or any variant or fragment thereof.
[0204] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000013.jpg41170, or any variant or fragment thereof.
[0205] In one embodiment, the secBFP2 variant protein comprises a T127X mutation, X is any amino acid. The FP2 variant protein contains the T127P mutation relative to the full-length wild-type secBFP2. secBFP2 containing the mutation T127L, T127R, or T127D nothing.
[0206] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000014.jpg41170, or any variant or fragment thereof.
[0207] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000015.jpg41170, or any variant or fragment thereof.
[0208] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000016.jpg42170, or any variant or fragment thereof.
[0209] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000017.jpg40170, or any variant or fragment thereof.
[0210] In one embodiment, the secBFP2 variant protein comprises a D151X mutation, X is any amino acid. The FP2 variant protein contains the D151G mutation relative to the full-length wild-type secBFP2. Contains secBFP2, which contains the mutation.
[0211] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000018.jpg41170, or any variant or fragment thereof.
[0212] In one embodiment, the secBFP2 variant protein comprises a N173X mutation, X is any amino acid. The FP2 variant protein contains the N173T mutation relative to the full-length wild-type secBFP2. secBFP2 containing the mutation N173H, N173R, or N173S nothing.
[0213] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000019.jpg41170, or any variant or fragment thereof.
[0214] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000020.jpg41170, or any variant or fragment thereof.
[0215] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000021.jpg41170, or any variant or fragment thereof.
[0216] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000022.jpg38170, or any variant or fragment thereof.
[0217] In one embodiment, the secBFP2 variant protein comprises a R198X mutation, X is any amino acid. The FP2 variant protein contains the R198V mutation relative to the full-length wild-type secBFP2. Contains secBFP2 containing a mutation or the R198L mutation.
[0218] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000023.jpg41170, or any variant or fragment thereof.
[0219] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000024.jpg40170, or any variant or fragment thereof.
[0220] In one embodiment, the secBFP2 variant protein is a full-length wild-type secBFP. In relation to 2, T18X, S28X, Y96X, S114X, V124X, T127X, D151X, N173X, and R198X mutations (where X is any amino acid) 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more In one embodiment, the secBFP2 comprises 8 or more, or all 9 of the secB The FP2 variant proteins are T18W, T18W, and T18W relative to the full-length wild-type secBFP2. 18V, T18E, S28A, Y96F, S114V, S114T, V124T, V12 4Y, V124W, T127P, T127L, T127R, T127D, D151G, N Of 173T, N173H, N173R, N173S, R198V, and R198L 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more , or secBFP2 containing nine or more.
[0221] In one embodiment, the secBFP2 variant protein is a full-length wild-type secBFP. In relation to 2, T18X, S28X, S114X, V124X, T127X, D151X , N173X, and R198X (where X is any amino acid) mutations. In one embodiment, the secBFP2 variant protein comprises a complete Long: T18W, S28A, S114V, V124T, T containing secBFP2 with the mutations 127P, D151G, N173T, and R198L. nothing.
[0222] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000025.jpg41170, or any variant or fragment thereof.
[0223] In one embodiment, the secBFP2 variant protein is a full-length wild-type secBFP. 2, S28X, S114X, T127X, and N173X (where X is In one embodiment, the secBFP2 includes a mutation in secBFP2 (which may be any amino acid). The FP2 variant protein is S28A, S28B, S28C, S28D, S28E, S28F ... and secBFP2 containing the 114T, T127L, and N173H mutations.
[0224] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000026.jpg40170, or any variant or fragment thereof.
[0225] In one embodiment, the secBFP2 variant protein is a full-length wild-type secBFP. 2, the mutations S28X and S114X (where X is any amino acid) In one embodiment, the secBFP2 variant protein comprises a secBFP2 containing a variant. secBFP2, containing the S28A and S114T mutations relative to the full-length wild-type secBFP2. Contains ecBFP2.
[0226] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000027.jpg41170, or any variant or fragment thereof.
[0227] In one embodiment, the secBFP2 variant protein is a full-length wild-type secBFP2. In the context of S28X, S114X, and N173X (where X is any amino acid) In one embodiment, the secBFP2 variant comprises a secBFP2 comprising a mutation of The secBFP2 protein contains S28A, S114T, and and secBFP2 containing the N173H mutation.
[0228] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000028.jpg41170, or any variant or fragment thereof.
[0229] In one embodiment, the secBFP2 variant protein is a full-length wild-type secBFP. 2, S28X, Y96X, S114X, and N173X (where X is any In one embodiment, the secBFP2 includes a mutation at any amino acid. The P2 variant protein contains S28A, Y9 relative to the full-length wild-type secBFP2. and secBFP2 containing the 6F, S114T, and N173H mutations.
[0230] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000029.jpg41170, or any variant or fragment thereof.
[0231] In one embodiment, the secBFP2 variant protein is a full-length wild-type secBFP. In relation to 2, S28X, Y96X, S114X, T127X, and N173X (here and X is any amino acid. , secBFP2 variant proteins, relative to full-length wild-type secBFP2, S secBF containing the mutations 28A, Y96F, S114T, T127L, and N173H Includes P2.
[0232] For example, in one embodiment, the secBFP2 variant protein is JPEG2025118694000030.jpg40170, or any variant or fragment thereof.
[0233] In one embodiment, the present invention provides a method for detecting a cellular signal generated or generated by an implementation of the systems and methods discussed herein. The identified composition comprises a nucleotide sequence encoding a secBFP2 variant protein. In various embodiments, the nucleic acid molecule comprises an isolated nucleic acid molecule comprising a sequence selected from the group consisting of SEQ ID NO:2 to SEQ ID NO:3. 28, or a variant thereof. Contains fragments or fragments of
[0234] Fluorescent protein variants operably linked to one or more polypeptides of interest Also provided is a fusion protein comprising the polypeptide of the fusion protein, wherein the polypeptide is separated by a peptide bond. Alternatively, the fluorescent protein variants may be linked via a linker molecule. In one embodiment, the fusion protein may be linked to one or more polypeptides that are operably linked to one or more polynucleotides encoding the target polypeptide A recombinant nucleic acid molecule containing a polynucleotide encoding a fluorescent protein variant. It is expressed from
[0235] The polypeptide of interest may be tagged with a peptide tag, such as a polyhistidine peptide, or or cellular polypeptides such as enzymes, G proteins, growth factor receptors, or transcription factors. and can associate to form a complex. In one embodiment, the fusion protein is , a tandem fluorescent protein variant construct, which contains a donor fluorescent protein variant Ant, acceptor fluorescent protein variants, and the donor and the acceptor The cyclized amino acid of the donor is a peptide linker portion that connects the donor to the cyclized amino acid. The donor and acceptor emit light at a fluorescence resonance energy when the donor is excited. The linker moiety exhibits photon transfer and does not substantially emit light to excite the donor. Thus, the information generated or identified by the implementation of the systems and methods discussed herein A fusion protein is a protein that contains two or more operably linked proteins, which may be directly or indirectly linked. and may further comprise one or more polypeptides of interest. obtain.
[0236] kit In some implementations, the kit includes a kit for implementing the systems and methods discussed herein. To facilitate and / or standardize the use of compositions provided or identified therein, and These various methods may be provided to facilitate the methods discussed herein. The implementing materials and reagents may be provided in a kit to facilitate carrying out the method. As used herein, the term "kit" refers to a process, assay, analysis, or is used in reference to a combination of items that facilitates operation.
[0237] The kits may contain chemical reagents (e.g., polypeptides or polynucleotides), as well as other Additionally, the kits discussed herein may also include components, e.g., for sample collection. Equipment and reagents for product collection and / or purification, Apparatus and reagents for bacterial cell transformation and eukaryotic cell transfection reagents, already transformed or transfected host cells, sample tubes, holders, Trays, racks, dishes, plates, kit user instructions, solutions, buffer solutions or other chemicals Includes suitable samples to be used for reagents, standardization, normalization, and / or control samples The kit can also be used for, but not limited to, convenient storage and safety. For convenient shipping, it can be packaged in a box with a lid.
[0238] In some embodiments, for example, the kits discussed herein may comprise any of the components discussed herein. Fluorescent proteins produced or identified by implementation of the systems and methods described herein. Fluorescent proteins produced or identified by implementing the systems and methods discussed in a polynucleotide vector (e.g., a plasmid) encoding the vector, suitable for propagation of the vector Bacterial cell lines, as well as reagents for the purification of the expressed fusion proteins, can be provided. In some embodiments, the kits discussed herein can be used to prepare a polymerase chain reaction (PCR) that is prone to oligomerization. Anthozoan Fluorescent Proteins to Generate Reduced Fluorescence Protein Variants The present invention can provide the reagents necessary to perform the mutagenesis of the present invention.
[0239] The kits may be generated or specified by implementation of the systems and methods discussed herein. One or more compositions to be used, e.g., one or more fluorescent proteins that may be part of a fusion protein. One or more polynucleotides encoding a photoprotein variant or polypeptide. The fluorescent protein variant may comprise an oligomer, such as a non-oligomerizing monomer. The fluorescent protein may be a mutant fluorescent protein with reduced propensity for co-merization, or a tandem dimeric fluorescent protein. The kit may be a photoprotein, and the kit may include a plurality of fluorescent protein variants, the plurality of which may be , multiple mutant fluorescent protein variants, or multiple tandem dimeric fluorescent proteins , or a combination thereof.
[0240] The kits discussed herein may also contain one or more recombinant nucleic acid molecules, It partially encodes fluorescent protein variants that may be the same or different. , for example, a restriction endonuclease recognition site or a recombinase recognition site, or any an operably linked second polypeptide containing or encoding the polypeptide of interest; The kit may further comprise nucleotides. Generated or identified by implementations of the systems and methods discussed herein, including Instructions for using the composition can be included.
[0241] One of skill in the art can conveniently select one or more proteins with the desired fluorescent properties for a particular application. Such kits allow for the identification of multiple different fluorescent protein variants. Similarly, it may be particularly useful to provide a fluorescent protein encoding a different fluorescent protein variant. Kits containing multiple polynucleotides offer numerous advantages. The nucleotide contains a convenient restriction endonuclease or recombinase recognition site and thus, the regulatory element or the polypeptide of interest is encoded. operative linkage of the polynucleotide to the polynucleotide, or, if desired, a fluorescent tag. Two or more polynucleotides encoding protein variants are operably linked to each other. This can promote the following.
[0242] Use of fluorescent protein variants Fluorescent tags generated or identified by implementation of the systems and methods discussed herein Protein variants are useful in any method that uses a fluorescent protein. Fluorescent protein variants, including monomeric, dimeric, and tandem dimeric fluorescent proteins is detected in a detection assay, such as an immunoassay or hybridization assay. fluorescent protein bands for use in fluorescent proteins or to track protein movement in cells fluorescent matrices, including binding the reagent to an antibody, polynucleotide, or other receptor; They are useful as fluorescent markers in many methods where markers are already used. For trace studies, the first (or other) polynucleotide encoding the fluorescent protein variant The nucleotide is linked to a second (or other) polynucleotide that encodes a protein of interest. If desired, the construct may be inserted into an expression vector. The protein of interest is a protein whose localization is determined by the fluorescent protein component of the fusion protein. Fluorescence without the concern that the artifacts are caused by oligomerization In one embodiment of this method, two proteins of interest can be localized based on Independently, they are fused to two fluorescent protein variants with different fluorescent properties.
[0243] Fluorescent tags generated or identified by implementation of the systems and methods discussed herein Protein variants are useful in systems for detecting induction of transcription. Nucleotides encoding merized monomeric, dimeric, or tandem dimeric fluorescent proteins The sequence may be linked to a promoter or other expression control sequence of interest that may be contained in the present vector. The construct can be transfected into cells and fused to a promoter (or other regulatory The induction of the nodal element is measured by detecting the presence or amount of fluorescence, thereby The method allows for the observation of the responsiveness of signal transduction pathways from receptors to promoters. It is possible.
[0244] Fluorescent tags generated or identified by implementation of the systems and methods discussed herein Protein variants are also useful in applications involving FRET, where a fluorescent donor and detecting events as a function of the movement of the acceptors toward or away from each other. One or both of the donor / acceptor pairs may be fluorescent protein variants. Such donor / acceptor pairs can be substituted by the donor excitation peak and the emission peak. Provides wide separation between the peaks of the donor emission spectrum and the acceptor excitation spectrum provides good overlap between the
[0245] Using FRET, a donor and acceptor were attached to the substrate on either side of the cleavage site. Cleavage of the substrate can be detected. Upon cleavage of the substrate, the donor / acceptor pair Such an assay can be performed, for example, by physically separating the substrate from the sample. and determining a qualitative or quantitative change in FRET. (See, e.g., U.S. Pat. No. 5,741,657, incorporated herein by reference. Fluorescent protein variant donor / acceptor pairs are used to identify protein molecules. It may be part of a fusion protein bound by a peptide with a cleavage site (e.g. See, for example, U.S. Pat. No. 5,981,200, which is incorporated herein by reference. FRET can also be used to detect changes in electrical potential across a membrane. The acceptor and acceptor are arranged on either side of the membrane so that they move across the membrane in response to a voltage change. and thereby produce measurable FRET (see, e.g., See U.S. Patent No. 5,661,035, incorporated herein by reference.
[0246] In other embodiments, the present invention provides a method for detecting a virus generated or generated by an implementation of the systems and methods discussed herein. The fluorescent proteins identified are fluorescent receptors for protein kinase and phosphatase activity. Cansa, or Ca 2+ , Zn 2+ , cyclic 3′,5′-adenosine monophosphate, and cyclic 3 to create indicators of small ions and molecules such as ′,5′-guanosine monophosphate It is useful for
[0247] Fluorescence in a sample is typically measured using a fluorometer, which uses an excitation light having a first wavelength. The excitation radiation from the source passes through the excitation optics, which causes the excitation radiation to excite the sample. In response, the fluorescent protein variants in the sample emit light at a wavelength different from the excitation wavelength. The collection optics then collects the emitted light from the sample. The chair includes a temperature controller to maintain the sample at a specific temperature while it is being scanned. and holding multiple samples to position different wells to be exposed. The microtiter plate may have a multi-axis translation stage. The multi-axis translation stage, temperature controller, autofocus function, and power supply associated with the data collection The slave device may be managed by a suitably programmed digital computer, The computer also organizes the data collected during the assay into separate columns for presentation. This process can be miniaturized and automated in a high-throughput format. This may allow for screening of many thousands of compounds in a single system. Some ways to implement the "Say" are given in Lakowicz, "Principles of Fluorescence Spectroscopy”(Plenum Pre ss 1983), Herman, “Resonance energy transf. er microscopy”In“Fluorescence Microscopy of Living Cells in Culture”Part B,Meth. Cell Biol.30:219-243(ed.Taylor and Wang; Academic Press 1989), Turro, “Modern Molec ular Photochemistry”(Benjamin / Cummings P ubl. Co., Inc. 1978), pp. 296-361, and each of these is incorporated herein by reference.
[0248] Thus, the present disclosure also provides implementations of methods for identifying the presence of molecules in a sample. Such methods may be used, for example, to implement the systems and methods discussed herein. linking fluorescent protein variants generated or identified by the method to molecules and Detecting fluorescence from fluorescent protein variants in samples suspected of containing the protein The molecule to be detected may be a polypeptide, a polynucleotide, or can be, for example, an antibody, an enzyme, or any other molecule including a receptor, a fluorescent protein The variant may be a tandem dimeric fluorescent protein.
[0249] The sample being tested can be a biological sample, an environmental sample, or a sample containing a specific molecule. Any sample, including any other sample whose presence or absence is desired to be determined. Preferably, the sample comprises a cell or an extract thereof. The cell may be a human cell or an extract thereof. can be obtained from vertebrates, including mammals, or from invertebrates, and from plants or animals The cells may be obtained from a culture of such cells, e.g., a cell line, or or isolated from an organism. Thus, the cells may be contained in a tissue sample, which may , by any means commonly used to obtain tissue samples, e.g., from a living human The method may also be used to obtain intact living cells or freshly isolated tissue. When performed using live or organ samples, the presence of molecules of interest in living cells is identified. and thus can provide a means for determining, for example, the intracellular compartmentalization of molecules. To this end, implementation of the systems and methods discussed herein can The use of fluorescent protein variants generated or identified by the method described above is not limited to the use of oligomeric fluorescent proteins. This provides a substantial advantage in that the likelihood of aberrant identification or localization due to chromatin depletion is greatly minimized. Provides benefits.
[0250] The fluorescent protein variants are stable under the conditions to which the protein-molecule complexes are exposed. The fluorescent molecule may be linked to the fluorescent molecule directly or indirectly using any suitable linkage. Proteins and molecules are synthesized through chemical reactions between reactive groups present on the proteins and molecules. The linkage can be achieved by using a fluorescent protein and a molecule containing specific reactive groups. The fluorescent protein variant and the molecule can be linked by a linker moiety. Suitable conditions for the synthesis of hydroxybenzoates are selected depending on, for example, the chemical nature of the molecules and the type of linkage desired. It will be appreciated that when the molecule of interest is a polypeptide, fluorescent proteins may be used. A convenient means for linking protein variants and molecules is to link them together, e.g., as polypeptides. a tandem dimer fluorescent tag operably linked to a polynucleotide encoding a peptide molecule; and a fusion protein from a recombinant nucleic acid molecule comprising a polynucleotide encoding the protein. This is because it is expressed as a
[0251] Methods for identifying agents or conditions that modulate the activity of expression control sequences are also provided. Suitable methods include, for example, incorporating a fluorescent protein variant operably linked to an expression control sequence. A recombinant nucleic acid molecule comprising a polynucleotide encoding the polypeptide is prepared by cleaving the polynucleotide from an expression control sequence. exposure to agents or conditions suspected to be capable of modulating the expression of and detecting the fluorescence of the fluorescent protein variant upon such exposure. Such methods include, for example, the isolation of cellular factors involved in tissue-specific expression from regulatory elements. Chemical or other proteins, including cellular proteins, whose expression can be regulated from expression control sequences, are useful for identifying biological agents. Thus, expression control sequences include promoters, , enhancers, silencers, intron splicing recognition sites, polyadenylation sites The regulatory element may be a transcriptional regulatory element, such as a ribosome binding site, or a translational regulatory element, such as a ribosome binding site.
[0252] Fluorescent tags generated or identified by implementation of the systems and methods discussed herein Protein variants also provide a way to identify specific interactions between first and second molecules. Such methods are also useful in detecting specific phases of a first molecule and a second molecule. a donor first fluorescent protein variant linked to the first fluorescent protein variant under conditions that allow interaction; The first molecule is coupled to a second molecule linked to a second fluorescent protein variant of the acceptor. contacting, exciting the donor, and causing fluorescence or emission from the donor to the acceptor; Resonance energy transfer is detected, thereby determining the specific interaction between the first molecule and the second molecule. The conditions for such an interaction can be determined by determining whether the molecule is specifically It can be any condition that is predicted or suspected to be capable of interacting effectively with other substances. In particular, when the molecules being tested are cellular molecules, the conditions are generally physiological conditions. Therefore, the method uses conditions such as buffers, pH, and ionic strength that mimic physiological conditions. The method can be performed in vitro using a method for detecting leukemia virus infection, or the method can be performed in a cell or using a cell extract. It can be implemented using
[0253] Luminescence resonance energy transfer can be achieved by chemiluminescence, bioluminescence, lanthanide, or transition metal donors. The longer the red fluorescent protein, the more energy is transferred from the red fluorescent protein. Excitation wavelengths are available from a wider variety of donors than is possible with green fluorescent protein variants, and This allows for energy transfer over larger distances. Also, the longer emission wavelengths allow for Red light is detected more efficiently by solid-state photodetectors, much better than shorter wavelengths This is particularly useful for in vivo applications where chemiluminescent donors can penetrate tissues easily. Bioluminescence donors include phenol derivatives and peroxyoxalates. Quorin, Obelin, Firefly luciferase, Renilla luciferase, Bacterial luciferase Lanthanides include, but are not limited to, lanthanides ... As a donor, metal ions are linked to multiple ligand groups to protect them from the solvent water. Examples of suitable terbium chelates include, but are not limited to, terbium chelates containing UV-absorbing sensitizer chromophores. As transition metal donors, ruthenium and osmium in oligopyridine ligands are Chemiluminescent and bioluminescent donors include, but are not limited to, Metal-based systems do not require excitation light but are excited by the addition of a substrate, while metal-based systems require excitation light. Although it requires photoactivation, it offers a longer excited state lifetime and reduces unwanted background fluorescence. Facilitates time-gated detection to distinguish between light and scatter.
[0254] The first and second molecules are used to determine whether the proteins specifically interact. or cellular proteins being investigated to confirm such interactions. Such first and second cellular proteins may, for example, be characterized by their ability to oligomerize. These may be the same as when the protein is being tested in a It may be different from the specific binding partners that are being tested for in intracellular pathways. The first and second molecules can also be polynucleotides and polypeptides, e.g., known or or transcriptional regulatory element activity, and known or The first molecule may be a polypeptide being tested for transcription factor activity. For example, the first molecule may be a Transcriptional regulatory elements are tested for activity, which can be random or of known sequence. the second molecule may be a transcription factor; Such methods are useful for identifying novel transcriptional regulatory elements with desired activities. .
[0255] The present disclosure also provides implementations of methods for determining whether a sample contains an enzyme. Such methods include, for example, subjecting a sample to the systems and methods discussed herein. contacting the tandem fluorescent protein variants generated or identified by implementing the method; and exciting the donor and determining the fluorescence characteristics in the sample. To determine whether the presence of enzyme in the pull results in a change in the degree of Förster resonance energy transfer Similarly, the present disclosure provides a method for determining the activity of an enzyme in a cell. Such methods are useful for, for example, constructing tandem fluorescent protein variants. providing a cell expressing the product, wherein the peptide linker moiety is a donor and an acceptor; a donor comprising a cleavage recognition amino acid sequence specific to an enzyme that binds the donor; and determining the degree of fluorescence resonance energy transfer in the cell. The presence of enzyme activity results in a change in the degree of fluorescence resonance energy transfer. It can be implemented by: [Example]
[0256] Implementation of the systems and methods discussed herein may be further illustrated by reference to the following examples. These examples are provided for illustrative purposes only and are not intended to be limiting unless otherwise specified. Unless otherwise specified, no limitation is intended. The systems and methods should in no way be construed as being limited to the following examples. rather, any and all variations that become apparent as a result of the teachings provided herein. It should be interpreted as including form.
[0257] Without further explanation, one skilled in the art can, using the foregoing description and the following illustrative examples, It is believed that implementations of the systems and methods discussed herein can be made and used. The following examples are therefore illustrative of the systems and methods discussed herein. These and other aspects of the present disclosure are intended to specifically point out exemplary embodiments, and are not intended to limit in any way the remainder of the disclosure. should not be interpreted as
[0258] Example 1: Protein engineering using neural networks To empirically validate the neural network, we used three different model proteins. A first set of validated model proteins are selected, each representing a distinct protein engineering challenge. tem-1 beta-lactamase, mainly because: 1) its sensitivity to antibiotics; , as it is directly related to the overall stability of that protein, and 2) that protein However, it has been well characterized to reveal both stabilizing and destabilizing mutations. Next, we used a metalloproteinase to incorporate the non-standard amino acid L-DOPA. Protein phosphomannose isomerase was reused as a reporter to improve stability. While the enzyme's insufficient stability precludes its use to act as a reporter, The final protein example is the improvement of the blue fluorescent protein variant, secBFP2. Blue fluorescent proteins are well characterized, but they exhibit rapid photobleaching and slow photobleaching. Difficult maturation and folding, as well as relatively dim fluorescence, hinder more widespread use .
[0259] First, the wild-type amino acid is the residue that has been experimentally verified as the best residue at that position. We evaluated the true negative rate of the neural network by dividing the analysis into The effect of each individual amino acid change was quantified on organism fitness, tem-1 β-lactam The 263 mutations tested in tem-1 were tested using a previously published mutation scan of Of the positions, 136 sites had relative fitness values less than zero (i.e., the organism's fitness). sites that could not tolerate mutations away from the wild-type residue without loss of response This collection of 136 sites represents a complete collection of true negatives for tem-1 beta-lactamase. and for each individual change made to the neural network, the true negative sensitivity The final version accurately identified 92.6% of the 136 true negatives. This represents an increase of nearly 30% compared to the initial model. has an increasing ability to identify sites within proteins that are not amenable to mutation.
[0260] The results of the experiment are shown in Figures 3A and 3B. Figure 3A shows that improving BFP fluorescence The figure shows bar graphs of the regions predicted by the neural network and their extent. The rightmost bar 301 represents the wild-type transcription factor, each individually suggested by the neural network. The fluorescence observed by making specific combinations of amino acid substitutions into the protein A visual representation of the improvement is shown in Figure 3B. The modified blue fluorescent protein 302 , which shines much brighter than wild-type blue fluorescent protein 303.
[0261] Additional results are shown in Figures 4A and 4B. The bar graph in Figure 4B shows the We present an improvement of the neural network proposal for PMI. Each of the stabilizing mutations results in a 15% to 50% increase compared to the wild type, but when used in combination When the α-amino acid is added (bar 401), the improvement is additive, resulting in a significant improvement in stability of nearly 600%. Glass.
[0262] Venn diagram of Figure 4B shows 411 (blue fluorescent protein, pdb:3m24) and 412 (phospho Homannose isomerase (pdb:1pmi) is a protein that neural networks can use to identify other components. Computational protein stabilization techniques (Foldx PositionScan and Predict unique residue candidates not identified by Rosetta pmut scan This indicates that...
[0263] Figure 5 shows the TEM-1 β-lactamase inhibitors identified by the neural network. The mutant protein enables E. coli to grow at higher ampicillin concentrations than the ancestral protein. The results show that the single mutagenized β-lactamase mutants N52K, F60Y, 125 μg of E. coli expressing M182T, E197D, or A249V each WT cells can grow at concentrations of ampicillin higher than 1 / mL. E. coli expressing the tagged ancestral enzyme were unable to grow. E. coli (N52K, F) expressing a single enzyme variant containing all five of the mutations 60Y, M182T, E197D, and A249V, labeled "All" was able to grow at an ampicillin concentration of 3000 μg / mL. The neural network is used to generate catalytically relevant phenotypes, in this embodiment, the ability of E. coli to produce antibiotics. improved phenotype that allows it to exhibit higher resistance to the substance ampicillin .
[0264] Figure 6 shows that the neural network improved the thermal stability of blue fluorescent protein. In one example, after 10 minutes of heat challenge, the residual fluorescence was measured using the inducible protein Bluebo The amount of SecBFP2.1, the ancestor of nnet, was less than that of nnet. The blue fluorescent protein was diluted to 0.01 mg / mL with PBS pH 7.4 and 100 μL of Aliquot L of PCR strips in a thermal gradient using a thermal cycler. Fluorescence of thermally loaded variants and incubated at room temperature for 10 min. The control was measured using excitation and emission wavelengths of 402 nm and 457 nm, respectively. Fluorescence readings were normalized to the mean of the solutions incubated at room temperature. (e.g., a reading of 0.8 indicates that the heat-treated protein retains 80% of its native fluorescence.) As shown in Figure 6, Bluebonnet was approximately 84 Higher thermal stability compared to SecBFP2.1 across the entire temperature range from ℃ to approximately 100℃ For example, after a 10-minute heat shock at 100°C, the fluorescence was preserved by the ancestral protein. When not treated, it retained more than 20% of its original fluorescence.
[0265] Figure 7 shows that the neural network improved the chemical stability of blue fluorescent protein. In another example, the fluorescence half-life in guanidine melt is There is less information about the ancestral protein, SecBFP2.1, than there is about Bluebonnet. The purified blue fluorescent protein was diluted to 0.01 mg / ml in 6 M guanidine hydrochloride. Triplicate 100uL aliquots were placed in wells of a 96-well clear bottom black walled plate. These purified fluorescent proteins were added to the PBS solution and incubated at 25°C for 23 hours. for 30 min using excitation and emission wavelengths of 402 nm and 457 nm, respectively. The plate was agitated before each measurement. The fluorescence value measured at time zero was Normalize fluorescence throughout the remainder of the assay using (e.g., a reading of 0.8 is , indicating that the protein retained 80% of its initial fluorescence). As shown in Figure 7, Bluebonnet is for all times greater than time=0 up to time=approximately 24 hours. Over the entire time period, it showed higher chemical stability than SecBFP2.1.
[0266] Example 2: Bluebonnet, a brighter blue fluorescent protein When scientists look at how and where proteins move throughout the cell, requires specialized genetic tools. One of these tools is ultraviolet light, That is, fluorescent proteins are a family of proteins that fluoresce under blue fluorescent light. BFP (pdb:3m24) is by far the most commonly used red fluorescent protein. Although it is a high-quality derivative, it suffers from poor in vivo activity. Using a network pipeline, we calculated the fluorescence intensity when expressed in E. coli cells. We predicted the BFP variants that lead to an increase in the number of BFPs. We provide data showing that the chromosome predictions were tested for their ability to increase fluorescence (wild type Figure 9 shows that when beneficial mutations are combined, they produce 8 times more nucleotides than the wild type. Figure 10 provides data showing that a more than 2-fold increase in fluorescence was observed. Bluebonnet blue, containing a combination of 14T, T127L, and N173H mutations The increased fluorescence of the blue fluorescent protein was visible compared to the parent strain and other blue fluorescent proteins. Show that you can see.
[0267] Computer system diagram 11A and 11B are related to the implementation of the systems and methods discussed herein. 11A and 11B are block diagrams illustrating exemplary computer embodiments useful in 11A and 11B show a block diagram of an exemplary computer 1100. In particular, the computer 1100 includes a central processing unit 1102 and a main memory device 1104 . The computer 1100 may also include any other components, such as one or more input / output devices. 130a to 130n (generally referred to using the reference numeral 1130), coprocessor 1 106, and a cache in communication with the central processing unit 1102 and the coprocessor 1106. The system may include a flash memory 1140.
[0268] The central processing unit 1102 responds to and fetches data from the main memory 1104. In many embodiments, the central processing unit is an Int el Corporation(Mountain View,California) Manufactured by Motorola Corporation (Schaumb Manufactured by International B Business Machines (White Plains, New York) or Advanced Micro Devices (Sun Microprocessors such as those manufactured by Samsung Electronics Co., Ltd. (NYVALE, California) Provided by the device.
[0269] Similarly, the coprocessor 1106 responds to and fetches from the main memory 1104. In some embodiments, a coprocessor is any logic circuit that processes instructions received from a coprocessor. Sa1106 is Google (Mountain View, California) Tensor processing, artificial intelligence application-specific integrated circuits, such as those manufactured by It may also include a thermal processing unit (TPU).
[0270] The main memory 1104 stores data, and any memory location can be accessed by the main processor 1102. or can be accessed directly by the microprocessor of the coprocessor 1106. One or more memory chips, e.g., static random access memory, SRAM, Burst SRAM or SynchBurst SRAM (BSRAM), Dynamic Random Access Memory (DRAM), Fast Page Mode DRAM (FPM DRAM), Enhanced DRAM (EDRAM), Extended Data Output RAM ( EDO RAM), Extended Data Output DRAM (EDO DRAM), Burst Extended Data Power DRAM (BEDO DRAM), Enhanced DRAM (EDRAM), Synchronous DRAM (S DRAM), JEDEC SRAM, PC100 SDRAM, Double Data Rate SD RAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), SyncLink DRAM (SLDRAM), Direct Rambus DRAM (DRDRAM), or Strong Induction It can be a ferroelectric RAM (FRAM).
[0271] In the embodiment shown in FIG. 11A, the processor 1102 is connected to a system bus 1120 (hereinafter referred to as a The coprocessor 1104 communicates with the main memory 1104 via a processor 1106 (described in more detail below). 1106 communicates with the main memory 1104 via a system bus 1120. , the processor 1102 communicates directly with the main memory 1104 via a memory port, 11B illustrates an embodiment of a computer system 1100. For example, in FIG. 11B, main memory 110 4 may be DRAM. In some embodiments, the neural network The engine may involve a main memory to store the values of the trained weights. It may reside in a storage device.
[0272] 11A and 11B show that the main processor 1102 sometimes uses a "backside" bus 11 illustrates an embodiment in which the cache memory 1140 communicates directly with the cache memory 1140 via a secondary bus referred to as In some embodiments, the coprocessor 1106 communicates with the cache memory via the secondary bus. In other embodiments, main processor 1102 may communicate directly with system 1140. The system bus 1120 is used to communicate with the cache memory 1140. The coprocessor 1106 uses the system bus 1120 to access the cache memory 1140 Cache memory 1140 may typically communicate with the main memory 1104. It has the fastest response time and is typically provided by SRAM, BSRAM, or EDRAM. In some embodiments, the coprocessor is associated with a neural network. Tensor Processing Unit (TPU) or other co-processor to perform the computations The primary processor 11 may include a processor, for example, an application specific integrated circuit (ASIC). (This may be faster or more efficient than performing such calculations on a .O2.)
[0273] In the embodiment shown in FIG. 11A, the processor 1102 and the coprocessor 1106 , communicates with various I / O devices 1130 via a local system bus 1120. ESA VL bus, ISA bus, EISA bus, MicroChannel architecture (MC A) Bus, PCI bus, PCI-X bus, PCI-Express bus, or NuBu The central processing unit 1102 and the coprocessor 1106 are connected to the I / O bus using various buses, including the In an embodiment where the I / O device is a video display, In this state, the processor 1102 and / or coprocessor 1106 may communicate with the display using the Advanced Graphics Port (AGP). FIG. 11B shows the main processor 1102 using HyperTransport, Rap id I / O, or communicates directly with I / O device 1130b over InfiniBand FIG. 11B also illustrates an embodiment of a computer system 1100 for receiving local In a mixed bus and direct communication embodiment, the processor 1102 communicates with the I / O devices. I / O devices 1130b communicate directly with the local interconnect bus. Communicate with 30a.
[0274] A wide variety of I / O devices 1130 may be present in computer system 1100 . Input devices include keyboard, mouse, trackpad, trackball, and microphone. Output devices include video displays, Rays, speakers, inkjet printers, laser printers, and dye-sublimation printers The I / O devices may also include mass storage devices for the computer system 1100. , for example, hard disk drives, 3.5-inch, 5.25-inch disks or ZI Floppy disk drive for receiving floppy disks such as P disks, CD -ROM drive, CD-R / RW drive, DVD-ROM drive, various formats of Loop Drive, and Twintech Industry, Inc. (Los Altos USB Flash Drive manufactured by Amitos, California ve's line of devices, and Apple Computer, Inc. (Cupert iPod Shuffle devices manufactured by Iono, California A USB storage device such as a USB line may be provided.
[0275] In a further embodiment, the I / O device 1130 is connected to the system bus 1120 for external communication. Signal buses, such as USB bus, Apple Desktop bus, RS-232 serial Connections, SCSI bus, FireWire bus, FireWire 800 bus, Ether rnet bus, AppleTalk bus, Gigabit Ethernet bus, non-synchronized High-speed transfer mode bus, HIPPI bus, Super HIPPI bus, SerialPlu s-bus, SCI / LAMP bus, FibreChannel bus, or Serial Attached: A bridge between the small computer system interface bus It is also possible.
[0276] General-purpose desktop computers of the type shown in FIGS. 11A and 11B typically , an operating system that controls task scheduling and access to system resources. It operates under the control of a common operating system. Specifically, Microsoft Corp. (Redmond, Washington) MICROSOFT WINDOWS, Apple Computer ( MacOS, Intel manufactured by rnational Business Machines(Armonk,New Y ork), and Caldera Corp. (Salt A freely available operating system distributed by the University of Lake City, Utah. An example is Linux, an operating system.
[0277] The disclosures of any and all patents, patent applications, and publications cited herein are hereby incorporated by reference. The present invention is disclosed with reference to specific embodiments. However, other embodiments and modifications of the present invention may depart from the true spirit and scope of the present invention. It is clear that the invention can be devised by a person skilled in the art without departing from the scope of the appended claims. The scope of the present invention is intended to be construed to include all such embodiments and equivalent variations. do.
Claims
1. A protein comprising a secBFP2 variant having one or more mutations at one or more residues selected from T18, S28, Y96, V124, T127, D151, N173, and R198 relative to full-length wild-type secBFP2.
2. 2. The protein of claim 1, wherein the one or more mutations include two mutations at residues selected from T18, S28, Y96, V124, T127, D151, N173, and R198.
3. The protein of claim 2, wherein the two mutations are S28A and N173H.
4. 2. The protein of claim 1, wherein the one or more mutations include three mutations at residues selected from T18, S28, Y96, V124, T127, D151, N173, and R198.
5. The protein of claim 4, wherein the three mutations are S28A, N173H, and Y96F.
6. The protein of claim 4, wherein the three mutations are S28A, N173H, and T127L.
7. 2. The protein of claim 1, wherein the one or more mutations include four mutations at residues selected from T18, S28, Y96, V124, T127, D151, N173, and R198.
8. The protein of claim 7, wherein the four mutations are S28A, N173H, T127L, and Y96F.
9. A protein comprising an amino acid sequence selected from the group consisting of SEQ ID NO:2-SEQ ID NO:28; A variant of a protein comprising an amino acid sequence selected from the group consisting of SEQ ID NO: 2-SEQ ID NO: 28; A fusion protein comprising an amino acid sequence selected from the group consisting of SEQ ID NO:2-SEQ ID NO:28, and A fragment of a protein comprising an amino acid sequence selected from the group consisting of SEQ ID NO: 2 to SEQ ID NO: 28; The protein of claim 1, selected from the group consisting of:
10. A nucleic acid molecule comprising a nucleotide sequence encoding the protein of claim 1.
11. The nucleic acid molecule of claim 10, which is a plasmid.
12. The nucleic acid molecule of claim 10, which is an expression vector.
13. 13. The nucleic acid molecule of any one of claims 10 to 12, further comprising a multiple cloning site for insertion of a heterologous protein coding sequence.
14. A composition comprising the protein of claim 1.
15. A composition comprising a nucleic acid molecule according to any one of claims 10 to 13.
16. A kit comprising a nucleic acid molecule according to any one of claims 10 to 13.
Citation Information
Patent Citations
Methods and materials for giving teeth a white appearance
JP2013536245A
System and method for increasing synthesized protein stability
JP2024016257A
MONOMERIC VARIANTS OF THE TETRAMERIC eqFP611
US20120107244A1
Novel far red fluorescent protein
US20150225467A1