Method, device and storage medium for predicting enzyme

By acquiring the molecular information of the target product, constructing the enzyme's performance descriptor using mathematical fitting and machine learning, and determining the amino acid sequence of the target enzyme, the problem of time-consuming and laborious selection of catalytic enzymes in existing technologies is solved, achieving efficient and reliable enzyme catalytic reactions and reducing the generation of byproducts.

CN115188433BActive Publication Date: 2026-04-28SHANGHAI SYNTHEALL PHARM CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI SYNTHEALL PHARM CO LTD
Filing Date
2022-05-26
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies are time-consuming and laborious in selecting catalytic enzymes, relying on personal knowledge and experience, making it difficult to find the optimal catalytic enzyme and reaction conditions, resulting in low production efficiency of the target product and high output of by-reactants.

Method used

By acquiring the molecular information of the target product, constructing the enzyme's performance descriptor using mathematical fitting and machine learning, and determining the amino acid sequence of the target enzyme, efficient selection of catalytic enzymes can be achieved.

Benefits of technology

It improves the production efficiency of the target product, reduces the output and emission of by-reactants, and promotes the green development of enzyme catalysis technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115188433B_ABST
    Figure CN115188433B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of enzyme prediction method and device, and a kind of computer readable storage medium.The prediction method includes the following steps: obtaining the molecular information of target product;According to the molecular information, and the performance expectation of enzyme catalytic reaction to generate the target product, determine the performance descriptor of target enzyme for generating the target product;According to the performance descriptor, mathematical fitting is carried out to determine the matrix descriptor of the target enzyme;According to the matrix descriptor, determine the multiple amino acid residue descriptor involved in the target enzyme;And according to the multiple amino acid residue descriptor, determine the amino acid sequence of the target enzyme.Through executing these steps, the prediction method can efficiently and reliably select catalytic enzyme, to facilitate the efficient production of target product, and reduce the output and discharge of side reaction, to promote the further development of enzyme catalysis technology as a kind of green chemical technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to enzyme catalysis technology, and more particularly to an enzyme prediction method, an enzyme prediction device, and a computer-readable storage medium. Background Technology

[0002] Enzymes are large molecules in living organisms that perform catalytic functions. Their chemical composition is typically proteins, ribonucleic acid (RNA), or complexes of these with small organic molecules or metal ions. Enzyme-catalyzed reactions often reduce the number of steps compared to pure organic chemical synthesis, achieving higher atom economy and yield. Furthermore, enzymes themselves are degradable and obtainable from the biological world, making them renewable resources. Therefore, enzyme catalysis technology is widely regarded as a green chemistry technology in the industry.

[0003] However, enzyme-catalyzed reactions often involve both target reaction products and by-products, and enzymes with different sequences and structures exhibit vastly different catalytic effects on the target reaction and various by-products. Therefore, before selecting a catalytic enzyme based on the target product, technicians typically need to conduct extensive experiments to summarize the catalytic effects of different enzymes on the target reaction and various by-products. This process is not only time-consuming and labor-intensive, but also suffers from drawbacks such as heavy reliance on personal knowledge and experience, difficulty in exhaustively listing catalytic enzymes and reaction conditions, and difficulty in finding the optimal catalytic enzyme and optimal reaction conditions.

[0004] In order to overcome the above-mentioned defects in the existing technology, there is an urgent need in the field for an enzyme prediction technology to efficiently and reliably select catalytic enzymes, so as to facilitate the efficient production of target products and reduce the production and emission of by-reactants, thereby promoting the further development of enzyme catalysis technology as a green chemistry technology. Summary of the Invention

[0005] The following provides a brief overview of one or more aspects to offer a basic understanding of them. This overview is not an exhaustive summary of all conceived aspects, nor is it intended to identify key or decisive elements of all aspects, nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed descriptions that follow.

[0006] To overcome the aforementioned deficiencies in the existing technology, the present invention provides an enzyme prediction method, an enzyme prediction device, and a computer-readable storage medium, which can efficiently and reliably select catalytic enzymes to facilitate the efficient production of target products and reduce the production and emission of by-reactants, thereby promoting the further development of enzyme catalysis technology as a green chemistry technology.

[0007] Specifically, the enzyme prediction method according to the first aspect of the present invention includes the following steps: obtaining molecular information of a target product P; and determining a performance descriptor for a target enzyme that generates the target product P based on the molecular information and the performance expectation of the enzyme-catalyzed reaction to generate the target product P. According to the performance descriptor Perform mathematical fitting to determine the matrix descriptor of the target enzyme. According to the matrix descriptor Determine the multiple amino acid residue descriptors involved in the target enzyme. and based on the plurality of amino acid residue descriptors The amino acid sequence of the target enzyme was determined.

[0008] Furthermore, in some embodiments of the present invention, the performance expectation includes the expectation of the intrinsic value of the target enzyme and / or the expectation of the reaction conditions for the enzyme-catalyzed reaction.

[0009] Furthermore, in some embodiments of the present invention, the desired intrinsic values ​​of the target enzyme include the desired relative catalytic activity, inactivation rate constant, and specificity constant of the target enzyme, and the desired reaction conditions for the enzyme-catalyzed reaction include the desired pH value, initial substrate concentration, thermodynamic temperature, and the reaction conditions of the target product and the reference ligand molecule. The expected structural similarity. Based on the molecular information and the expected performance of the enzyme in the catalytic reaction to produce the target product P, a performance descriptor for the target enzyme is determined to generate the target product P. The steps include: determining the performance descriptor of the target enzyme based on the molecular information and the performance expectation.

[0010]

[0011] Where 'a' represents the relative catalytic activity, in units of s. -1 d is the deactivation rate constant, in seconds. -1 E is the specificity constant, dimensionless; pH is the negative logarithm of the hydrogen ion concentration in the reaction system, dimensionless; [S]0 is the initial substrate concentration, in g·L. -1 T represents the thermodynamic temperature in Kelvin; θ represents the temperature of the target product P relative to the reference ligand molecule. The structural similarity score.

[0012] Furthermore, in some embodiments of the present invention, the step of basing the data on the performance descriptor... Perform mathematical fitting to determine the matrix descriptor of the target enzyme. The steps include: based on the performance descriptor And the numerical approximate solution F of the predetermined quantitative structure-activity relationship function F. # Determine the matrix descriptor of the target enzyme. Wherein, the numerical approximate solution F # This is achieved by generating the reference ligand molecule. The sequence description tensor of at least one enzyme and performance description tensor The quantitative structure-activity relationship function between them was determined by mathematical fitting.

[0013] Furthermore, in some embodiments of the present invention, the numerical approximate solution F is determined. # The steps include: obtaining the reference ligand molecule. At least one similar ligand molecule P j The performance data, and the generation of each of the aforementioned similar ligand molecules P j The amino acid sequence of the sample enzyme $s i According to the aforementioned similar ligand molecules P j Based on the performance data, determine the corresponding performance description tensor. Based on the amino acid sequences of the enzymes in each sample $s i Determine the corresponding sequence description tensor And construct the sequence description tensor Regarding the performance description tensor Quantitative structure-activity relationship function The numerical approximate solution FQ of the quantitative structure-activity relationship function FQ is determined through machine learning training. # .

[0014] Furthermore, in some embodiments of the present invention, the step of obtaining the reference ligand molecule... At least one similar ligand molecule P j The performance data, and the generation of each of the aforementioned similar ligand molecules P j The amino acid sequence of the sample enzyme $s i The steps include: obtaining the reference ligand molecule. The molecular information is obtained; based on the molecular information, a corresponding ligand candidate set is determined, wherein the ligand candidate set includes at least one ligand node structure, and the ligand node structure includes at least a directed evolution data field and a comparison ligand field; the similarity score between the molecular information and the comparison ligand field of each ligand node structure in the ligand candidate set is determined to screen the reference ligand molecule. At least one similar ligand molecule P j The ligand node structure; and based on the directed evolution data fields of each of the ligand node structures, determine the generation of each of the similar ligand molecules P. jThe amino acid sequence of the sample enzyme $s i and each of the aforementioned similar ligand molecules P j Performance data.

[0015] Furthermore, in some embodiments of the present invention, the similar ligand molecule P j Performance data includes the generation of the similar ligand molecule P j The intrinsic value of the sample enzyme, and the generation of the similar ligand molecule P. j The reaction condition data for enzyme-catalyzed reactions. The data is based on the similar ligand molecules P... j Based on the performance data, determine the corresponding performance description tensor. The steps include: constructing the similar ligand molecule P based on the intrinsic value of the sample enzyme and the reaction condition data of the enzyme-catalyzed reaction. j performance descriptor According to the aforementioned similar ligand molecules P j performance descriptor Construct a performance description matrix for each of the sample enzymes. And a performance description matrix based on all sample enzymes. Construct the performance description tensor

[0016] Furthermore, in some embodiments of the present invention, the similar ligand molecule P is constructed based on the intrinsic value of the sample enzyme and the reaction condition data of the enzyme-catalyzed reaction. j performance descriptor The steps include: based on the similar ligand molecule P j Data on the change of substrate concentration over time in enzyme-catalyzed reactions were used to determine the corresponding sample enzyme's response to the similar ligand molecule P. j The relative catalytic activity *a* is determined; the inactivation rate constant *d* of the sample enzyme is determined based on the inactivation data of the sample enzyme in the enzyme-catalyzed reaction; the specificity constant *E* of the sample enzyme is determined based on the reaction rates of the target reaction and at least one side reaction in the enzyme-catalyzed reaction; and the pH value, initial substrate concentration, thermodynamic temperature, and similar ligand molecule *P* of the enzyme-catalyzed reaction are determined based on the relative catalytic activity *a*, the inactivation rate constant *d*, the specificity constant *E*, and the pH value, initial substrate concentration, thermodynamic temperature, and similar ligand molecule *P* of the enzyme-catalyzed reaction. j With the reference ligand molecule The structural similarity score is used to construct the similar ligand molecule P. j performance descriptor

[0017] Furthermore, in some embodiments of the present invention, the step of determining the amino acid sequence of each of the said sample enzymes is based on $s i Determine the corresponding sequence description tensor The steps include: calculating and generating the reference ligand molecule respectively. The amino acid sequence of the enzyme $s0 and the amino acid sequence of each of the sample enzymes $s i The similarity between them; based on a preset first similarity threshold, the reference ligand molecules are selected and generated. The enzyme sequence $s0 is similar to that of multiple sample enzymes; and the amino acid sequence $s0 of the multiple similar sample enzymes is based on the enzyme sequence $s0. i Determine the sequence description tensor

[0018] Furthermore, in some embodiments of the present invention, the step of determining the amino acid sequence of each of the said sample enzymes is based on $s i Determine the corresponding sequence description tensor The steps include: obtaining the amino acid sequence of the sample enzyme. i The side chain length l, side chain width w, and side chain acidity pK of each amino acid residue and placeholder are as follows: a The side chain rigidity r and side chain polarity μ are used to construct a description vector for the amino acid residues and the placeholders. Based on the amino acid sequence of the sample enzyme $s i Description vectors of each amino acid residue and each placeholder. Construct the amino acid sequence of the sample enzyme $s i Sequence description matrix And based on the amino acid sequences of the enzymes in each of the samples $s i Sequence description matrix Construct the sequence description tensor

[0019] Furthermore, in some embodiments of the present invention, a performance descriptor for the target enzyme that generates the target product P is determined based on the molecular information and the performance expectation for the enzyme-catalyzed reaction to generate the target product P. Previously, the prediction method further included the following steps: first molecular information of the target product P, and reference ligand molecules corresponding to the quantitative structure-activity relationship function F. The second molecular information is used to perform molecular similarity calculations to determine the similarity score between the first molecular information and the second molecular information; and in response to a judgment result that the similarity score is greater than or equal to a preset second similarity threshold, a performance descriptor for the target enzyme that generates the target product P is determined based on the molecular information and the performance expectation for the enzyme-catalyzed reaction to generate the target product P.

[0020] Furthermore, in some embodiments of the present invention, the prediction method further includes the following step: in response to the judgment result that the similarity score is less than the second similarity threshold, outputting a prompt that supplements the quantitative structure-activity relationship function F' corresponding to the target product P.

[0021] Furthermore, in some embodiments of the present invention, after determining the amino acid sequence of the target enzyme, the prediction method further includes the following steps: determining a corresponding sequence candidate set based on the amino acid sequence of the target enzyme, wherein the sequence candidate set includes at least one sequence node structure, and the sequence node structure includes at least a directed evolution data field and an alignment sequence field; determining the similarity score between the amino acid sequence of the target enzyme and the alignment sequence field of each sequence node structure in the sequence candidate set; selecting at least one sequence node structure similar to the target enzyme from the sequence candidate set according to a preset first similarity threshold; and outputting the directed evolution data field and / or alignment sequence field of each sequence node structure similar to the target enzyme.

[0022] Furthermore, the enzyme prediction apparatus provided according to a second aspect of the present invention includes a memory and a processor. The processor is connected to the memory and configured to implement the enzyme prediction method provided in the first aspect of the present invention.

[0023] Furthermore, the computer-readable storage medium provided according to the third aspect of the present invention stores computer instructions thereon. When the computer instructions are executed by a processor, the enzyme prediction method provided in the first aspect of the present invention is implemented. Attached Figure Description

[0024] The above-described features and advantages of the present invention will be better understood after reading the following detailed description of embodiments of the present disclosure in conjunction with the accompanying drawings. In the drawings, components are not necessarily drawn to scale, and components having similar related properties or features may have the same or similar reference numerals.

[0025] Figure 1 A flowchart illustrating an enzyme prediction method provided according to some embodiments of the present invention is shown.

[0026] Figure 2 A schematic diagram of the ECD structure of directed evolution data of an enzyme provided according to some embodiments of the present invention is shown.

[0027] Figure 3 A schematic diagram of a non-relational database of directed evolution data of enzymes provided according to some embodiments of the present invention is shown.

[0028] Figure 4A schematic flowchart of a molecular retrieval method for directed evolutionary data of enzymes provided according to some embodiments of the present invention is shown.

[0029] Figure 5 A schematic flowchart of a sequence retrieval method for directed evolution data of enzymes according to some embodiments of the present invention is shown. Detailed Implementation

[0030] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Although the description of the present invention is presented in conjunction with preferred embodiments, this does not mean that the features of the invention are limited to these embodiments. On the contrary, the purpose of describing the invention in conjunction with embodiments is to cover other options or modifications that may be derived based on the claims of the present invention. To provide a thorough understanding of the invention, many specific details will be included in the following description. The invention may also be implemented without using these details. Furthermore, to avoid confusion or obscuring the focus of the invention, some specific details will be omitted in the description.

[0031] It is understood that although terms such as "first," "second," and "third" may be used herein to describe various components, regions, layers, and / or parts, these components, regions, layers, and / or parts should not be limited by these terms, and these terms are only used to distinguish different components, regions, layers, and / or parts. Therefore, the first components, regions, layers, and / or parts discussed below may be referred to as second components, regions, layers, and / or parts without departing from some embodiments of the present invention.

[0032] As mentioned above, enzyme-catalyzed reactions often involve both target reaction products and by-products, and enzymes with different sequences and structures exhibit vastly different catalytic effects on the target reaction and various by-products. Therefore, before selecting a catalytic enzyme based on the target product, technicians typically need to conduct numerous experiments to summarize the catalytic effects of different enzymes on the target reaction and various by-products. This process is not only time-consuming and labor-intensive, but also suffers from drawbacks such as heavy reliance on personal knowledge and experience, difficulty in exhaustively listing catalytic enzymes and reaction conditions, and difficulty in finding the optimal catalytic enzyme and optimal reaction conditions.

[0033] To overcome the aforementioned deficiencies in the existing technology, the present invention provides an enzyme prediction method, an enzyme prediction device, and a computer-readable storage medium, which can efficiently and reliably select catalytic enzymes to facilitate the efficient production of target products and reduce the production and emission of by-reactants, thereby promoting the further development of enzyme catalysis technology as a green chemistry technology.

[0034] In some non-limiting embodiments, the prediction method provided in the first aspect of the present invention can be implemented by the prediction apparatus provided in the second aspect of the present invention. Specifically, the prediction apparatus is configured with a memory and a processor. The memory includes, but is not limited to, the computer-readable storage medium provided in the third aspect of the present invention, on which computer instructions are stored. The processor is connected to the memory and configured to execute the computer instructions stored in the memory to implement the enzyme prediction method provided in the first aspect of the present invention.

[0035] The working principle of the above-described prediction device and storage medium will be described below with reference to some embodiments of prediction methods. Those skilled in the art will understand that these prediction methods are merely non-limiting embodiments provided by this invention, intended to clearly demonstrate the main concepts of the invention and provide specific solutions convenient for public implementation, rather than limiting all functions or operating methods of the above-described prediction device and storage medium. Similarly, the prediction device and storage medium are also merely a non-limiting embodiment provided by this invention and do not constitute a limitation on the entities performing the steps in these prediction methods.

[0036] Please refer to Figure 1 , Figure 1 A flowchart illustrating an enzyme prediction method provided according to some embodiments of the present invention is shown.

[0037] like Figure 1 As shown, in some embodiments, the prediction method provided by the first aspect of the present invention can be implemented in two stages: offline training (left dashed line portion) and online prediction (right solid line portion). Correspondingly, the prediction device provided by the second aspect of the present invention can also be composed of an offline training device and an online prediction device. In some embodiments, the offline training device and the online prediction device can be integrated into the same hardware device. Optionally, in other embodiments, the offline training device and the online prediction device can be distributed across multiple different hardware devices, and these multiple different hardware devices cooperate with each other to implement the prediction method provided by the first aspect of the present invention.

[0038] Specifically, during the offline training phase, the training device can first construct a database of directed evolution data of enzymes, and then, based on reference ligand molecules... Small molecule retrieval is performed on the molecular information to determine the reference ligand molecule. At least one similar ligand molecule P j Then, the training device can be based on at least one similar ligand molecule P. jThe training set for the structure-activity relationship model F is constructed using the directed evolution data of the corresponding enzymes. Based on this training set, artificial intelligence (AI) training is performed to determine the numerical approximate solution Fi of the structure-activity relationship model Fi. # .

[0039] In some embodiments, the database for the directed evolution data of the above-mentioned enzymes can be the EvoCloud database, which is built based on the ECD (EvoCloudDroplet) structure. Please refer to... Figure 2 , Figure 2 A schematic diagram of the ECD structure of directed evolution data of an enzyme provided according to some embodiments of the present invention is shown.

[0040] like Figure 2 As shown, in the process of building the EvoCloud database, the builder can first collect multidimensional data on the directed evolution process of various enzymes from one or more channels, and then perform structured processing based on a pre-built data structure. The directed evolution data of the enzymes is then filled into the [A1~A12] fields according to the specified data types, thereby constructing an ECD structure for the directed evolution data of each enzyme. This integrates the multidimensional data involved in the directed evolution process of the enzyme, characterizes the catalytic performance of the enzyme, and achieves efficient storage and retrieval of this multidimensional data. It is understood that this description of the builder is only a non-limiting approach, and includes, but is not limited to, technicians performing the corresponding steps, and execution devices such as processors, controllers, and / or robotic arms executing the corresponding computer instructions.

[0041] Here, the data types and data structures involved in this ECD structure are described in the [A0] field below.

[0042] [A0-1] String: A general concept in electronic computer systems.

[0043] [A0-2] Integer: A general concept in electronic computer systems.

[0044] [A0-3] Floating-point number: A general concept in electronic computer systems.

[0045] [A0-4] Date and time: A general concept in electronic computer systems.

[0046] [A0-5]Charging Structure: Describes the method of adding a certain material in the reaction that measures the catalytic performance of an enzyme, and includes the following member fields:

[0047] [A0-5-1] Method field: Describes the feeding method. Its data type is string, and the allowed value is one of the string enumeration values ​​["continuous feeding", "one time charging", "other", "portionwise charging"], which respectively represent continuous feeding, one-time feeding, other, and batch feeding.

[0048] [A0-5-2] Speed ​​field: Describes the feeding speed; its data type is a physical quantity structure. If the method field is "continuous feeding", the unit field can be one of the string enumeration values ​​["L / h", "mL / h", "mL / min", "mL / s", "VVH", "VVM", "VVS"]. If the method field is "portionwise charging", the unit field can be one of the string enumeration values ​​["L / time", "mL / time", "V / time"].

[0049] [A0-5-3] The `timePoints` field describes the time points when feeding is performed. Its data type is an array of several physical quantity structures. If the `method` field is set to "continuous feeding" or "one time charging", the `timePoints` field contains only one member: the feeding start time. If the `method` field is set to "portionwise charging", the `timePoints` field can contain multiple members, each representing the time point of each feeding operation. The unit field for each member in the `timePoints` field can have a value from the string enumeration ["day", "h", "min", "s"].

[0050] [A0-6] Dilution structure: Describes the solution composition of a material used in a reaction that measures the catalytic performance of an enzyme, when the material is used in solution form. It includes the following member fields:

[0051] [A0-6-1] Solvent field: Describes the solvent or diluent of the solution. Its data type is ligand structure.

[0052] [A0-6-2] Loading field: Describes the amount of solvent or diluent used. Its data type is a physical quantity structure, and its unit field allows values ​​to be one of the string enumeration values ​​["L", "mL", "V", "μL"].

[0053] [A0-7] Ligand Structure: Describes information about a small molecule, such as its structure, composition, and index conforming to general rules. It contains the following member fields:

[0054] [A0-7-1] CAS field: Describes the CAS number of a substance. This number is assigned to each registered substance in the Chemical Abstracts Service (CASS), a chemical registry maintained by the American Chemical Society. This number is widely used to describe chemicals, especially those intended for sale and production. The data type is string.

[0055] [A0-7-2] inChI field: Describes a substance's InChI (International Chemical Identifier), a string jointly developed by the International Union of Pure and Applied Chemistry (IUPAC) and the National Institute of Standards and Technology (NIST) to uniquely identify a compound's IUPAC name. The data type is string.

[0056] [A0-7-3]iupacName field: Describes the scientific name of a substance according to the systematic nomenclature rules. This nomenclature is defined by the International Union of Pure and Applied Chemistry (IUPAC) and is a method for systematically naming chemical substances, specifying chemical terms from organic to inorganic, from molecules to polymers, and in various other aspects. The data type is string.

[0057] [A0-7-4]smiles field: A SMILES (Simplified Molecular InputLine Entry System) string describing a substance, a specification for explicitly describing molecular structure using ASCII strings. SMILES strings can be imported by most molecular editing software and converted into two-dimensional graphics or three-dimensional molecular models, and are one of the common solutions for processing the chemical composition of small molecules in current computer systems. The data type is string.

[0058] [A0-7-5] Structure field: A mol2 format text describing the chemical structure of a substance, including the elemental composition, interconnections, valence states, charges, and three-dimensional coordinates of the atoms contained in its molecule. Biochemical calculation software such as SYBYL and Discovery Studio typically use the mol2 format to store the chemical information of small molecules. The data type is string.

[0059] [A0-8] Mutation structure: Describes a difference in an amino acid or base sequence relative to its natural ancestor at some point, i.e., a mutation, and contains the following member fields:

[0060] [A0-8-1] MutationMotif field: Describes a fragment of the sequence at the location of the mutation. Its data type is string. If the type field is "nucleotide", the mutationMotif field allows one or more string enumeration values ​​["A", "C", "G", "T"] in one or three combinations, representing adenine deoxyribonucleotide, cytosine deoxyribonucleotide, guanine deoxyribonucleotide, and thymine deoxyribonucleotide, respectively. If the type field is "peptide", the mutationMotif field can be a string enumeration value ["A", "C", "D", "E", "F", "G", "H", "I", "K", "L", "M", "N", "P", "Q", "R", "S", "T", "V", "W", "Y"], representing alanine residue, cysteine ​​residue, aspartic acid residue, glutamic acid residue, phenylalanine residue, glycine residue, histidine residue, isoleucine residue, lysine residue, leucine residue, methionine residue, asparagine residue, proline residue, glutamine residue, arginine residue, serine residue, threonine residue, valine residue, tryptophan residue, and tyrosine residue.

[0061] [A0-8-2] Position field: Describes the position of the mutation location relative to the entire sequence. Its data type is integer, and positive integers are allowed.

[0062] [A0-8-3] TemplateMotif field: Describes a fragment of the sequence corresponding to the mutation site in the ancestor of nature. Its data type is string. If the type field is "nucleotide", the templateMotif field allows one or more string enumeration values ​​["A", "C", "G", "T"] in one or three combinations, representing adenine deoxyribonucleotide, cytosine deoxyribonucleotide, guanine deoxyribonucleotide, and thymine deoxyribonucleotide, respectively. If the type field is "peptide", the templateMotif field can have one of the string enumeration values ​​["A", "C", "D", "E", "F", "G", "H", "I", "K", "L", "M", "N", "P", "Q", "R", "S", "T", "V", "W", "Y"], representing alanine residue, cysteine ​​residue, aspartic acid residue, glutamic acid residue, phenylalanine residue, glycine residue, histidine residue, isoleucine residue, lysine residue, leucine residue, methionine residue, asparagine residue, proline residue, glutamine residue, arginine residue, serine residue, threonine residue, valine residue, tryptophan residue, and tyrosine residue.

[0063] [A0-8-4] Type field: Describes the type of mutation. Its data type is string, and the allowed value is one of the string enumeration values ​​["nucleotide", "peptide"], which represent mutations in the base sequence and mutations in the amino acid sequence, respectively.

[0064] [A0-9] Physical Quantity Structure: Describes a physical quantity. Scientifically, a physical quantity consists of a numerical value and a unit. Furthermore, in engineering, since absolute precision is impossible in any measurement or instrument setting, it is necessary to specify the minimum and maximum permissible deviations of the physical quantity from the target value as a process parameter.

[0065] [A0-9-1] Lower Limit (loweLimit) field: The lower limit of allowed values, and its data type is floating point.

[0066] [A0-9-2] Target Value field: The set target value, whose data type is floating point.

[0067] [A0-9-3] Unit field: The unit of the physical quantity, which is generally the basic unit specified by IUPAP or the product of their finite powers. Its data type is string.

[0068] [A0-9-4] UpperLimit field: The upper limit of allowed values, and its data type is floating point.

[0069] Furthermore, based on the definition of the [A0] field above, the [A1~A12] fields of the ECD structure can be defined as follows.

[0070] [A1] Alias ​​field: Describes an alternative name for the enzyme. Its data type is an array of strings. Enzymes can usually have multiple aliases, used for abbreviation in literature or as trade names, etc.

[0071] [A2] Wild Ancestor (acestorId) field: A unique identifier describing the wild ancestor of the enzyme in the storage system. Its data type is string.

[0072] [A3] Expression System field: Describes the enzyme's expression system. Its data type is an expression system structure, containing the following member fields:

[0073] [A3-1] Host field: Describes the expression host of the enzyme, and its data type is string. The source of the enzyme can be its natural source organism, i.e., the natural host, or an engineered host such as recombinant cells used for molecular cloning or overexpression.

[0074] [A3-2] Note field: Note information, its data type is string.

[0075] [A3-3] Vector field: Describes the DNA used to carry the enzyme gene; its data type is string. In different hosts, the enzyme gene can be integrated into the cell's chromosome or chromatin DNA, i.e., the genome, or it can be embedded in small DNA independent of the genome, such as plasmids or organelle DNA.

[0076] [A4]id field: Describes the enzyme's unique identifier in the storage system; its data type is string.

[0077] [A5] Name field: Describes the scientific name of the enzyme; its data type is string.

[0078] [A6] Note field: Note information, its data type is string.

[0079] [A7] Organism field: Describes the biological origin of the enzyme. Its data type is string. It is usually the organism from which the enzyme was first discovered or isolated, and follows the scientific name of the binomial nomenclature commonly used in modern biological taxonomy.

[0080] [A8] ParentId field: A unique identifier describing the evolutionary parent of the enzyme in the storage system. Its data type is string.

[0081] [A9] Catalytic Performance (performances) field: Describes the enzyme's catalytic performance data. Its data type is an array of several catalytic performance (performance) structures. Each catalytic performance (performance) structure represents one piece of enzyme reaction performance data and contains the following member fields:

[0082] [A9-1] The `conditions` field describes the conditions and parameters of the reaction that measure the catalytic performance of an enzyme. Its data type is a `conditions` structure, containing the following member fields:

[0083] [A9-1-1] Humidity field: Relative humidity, its data type is floating point, and the allowed value is 0 to 100%.

[0084] [A9-1-2]pH field: pH value, its data type is floating point, and the allowed value is 0 to 14.

[0085] [A9-1-3] Reaction Time field: Describes the reaction time that measures the catalytic performance of an enzyme. Its data type is a physical quantity structure, and its unit field allows one of the string enumeration values ​​["ms", "s", "min", "h", "day"].

[0086] [A9-1-4] Reactor field: Describes the container in which the reaction takes place to measure the catalytic performance of an enzyme. Its data type is a reactor structure, which contains the following member fields:

[0087] [A9-1-4-1] Agitation field: Describes the agitation method and speed. Its data type is an agitation structure, which contains the following member fields:

[0088] [A9-1-4-1-1]Magnitude field: Describes the magnitude of stirring, such as the size of the stirring magnet, the diameter of the mechanical stirring paddle, the flow rate of the air rise or the velocity of the crossflow, etc. Its data type is a physical quantity structure.

[0089] [A9-1-4-1-2] Method field: Describes the mixing method. Its data type is string, and the allowed value is one of the string enumeration values ​​["Air lift", "Cross current", "Linear shaking", "Magneticagitation", "Mechanic agitation", "Orbit shaking", "Other", "Vertex mixing"], which respectively represent air lift mixing, cross-flow mixing, reciprocating oscillation, magnetic stirring, mechanical stirring, circular oscillation, other, and vortex mixing.

[0090] [A9-1-4-1-3] Speed ​​field: describes the stirring speed. Its data type is a physical quantity structure, and its unit field allows values ​​to be one of the string enumeration values ​​["Hz", "m / s", "rad / s", "rpm"].

[0091] [A9-1-4-2] Diameter field: Describes the diameter of the reactor. Its data type is a physical quantity structure, and its unit field allows one of the string enumeration values ​​["cm", "dm", "m", "mm"].

[0092] [A9-1-4-3] Height field: Describes the height of the reactor. Its data type is a physical quantity structure, and its unit field allows values ​​to be one of the string enumeration values ​​["cm", "dm", "m", "mm"].

[0093] [A9-1-4-4] Shape field: Describes the shape of the reactor. Its data type is string, and the allowed value is one of the string enumeration values ​​["Eppendorf tube", "Glass vial", "Hydrogenation Reactor", "Jacket", "Microplate vial", "Round bottom flask", "T-flask", "Test tube", "Other"], representing centrifuge tube, glass bottle, hydrogenation reaction flask, jacketed reaction flask, microplate well, round bottom flask, conical flask, test tube, and others, respectively.

[0094] [A9-1-5] Temperature field: Describes the temperature at which the reaction takes place to measure the catalytic performance of the enzyme. Its data type is a physical quantity structure, and its unit field allows values ​​to be one of the string enumeration values ​​["℃", "K"].

[0095] [A9-2]id field: describes the experiment number, and its data type is string.

[0096] [A9-3] Product field: Describes the products and results of the reaction used to measure the catalytic performance of an enzyme. Its data type is a product structure, which contains the following member fields:

[0097] [A9-3-1] Conversion Ratio field: Describes the conversion rate of the reaction that measures the catalytic performance of the enzyme. Its data type is a floating point number, and the allowed value is a number from 0 to 1.

[0098] [A9-3-2]de field: Describes the diastereomer excess value of the reaction product used to measure the catalytic performance of an enzyme. Its data type is floating point, and it allows values ​​from 0 to 1.

[0099] [A9-3-3]dr field: describes the ratio of diastereomers of the products of the reaction that measures the catalytic performance of an enzyme. Its data type is floating point and allows values ​​greater than 0.

[0100] [A9-3-4]ee field: describes the enantiomer excess value of the reaction product used to measure the catalytic performance of an enzyme. Its data type is floating point, and the allowed value is a number from 0 to 1.

[0101] [A9-3-5] EnantioselectivityRatio field: Describes the enantioselectivity of the reaction for which the enzyme performs a catalytic performance. Its data type is a floating-point number, and values ​​greater than zero are allowed.

[0102] [A9-3-6]er field: Describes the enantiomeric ratio of the products of the reaction that measures the catalytic performance of the enzyme. Its data type is floating point and allows values ​​greater than 0.

[0103] [A9-3-7] The molecule field describes the product molecule of the reaction that measures the catalytic performance of an enzyme. Its data type is ligand structure.

[0104] [A9-3-8] Purity field: Describes the purity of the product of the reaction that measures the catalytic performance of the enzyme. Its data type is floating point, and it allows values ​​from 0 to 1.

[0105] [A9-3-9]re field: Describes the regioisomeric excess value of the product positional isomers of the reaction that measures the catalytic performance of an enzyme. Its data type is floating point, and it allows values ​​from 0 to 1.

[0106] [A9-3-10]rr field: describes the ratio of product positional isomers in a reaction that measures the catalytic performance of an enzyme. Its data type is floating point, and values ​​greater than 0 are allowed.

[0107] [A9-3-11] Isolated Yield field: Describes the isolated yield of the reaction used to measure the catalytic performance of the enzyme. Its data type is a floating-point number, and the allowed value is a number from 0 to 1.

[0108] [A9-3-12] Solution Yield field: Describes the in-situ yield of the reaction used to measure the catalytic performance of the enzyme. Its data type is a floating-point number, and the allowed value is a number from 0 to 1.

[0109] [A9-4] The `reagents` field describes the reagents involved in the reaction that measures the catalytic performance of the enzyme. Its data type is an array of `reagents` structures. Each `product` structure represents a reagent and contains the following member fields:

[0110] [A9-4-1] Charging field: describes the way the reagent is added to the reaction, and its data type is the charging structure.

[0111] [A9-4-2] Dilution field: Describes how a reagent is diluted when added in solution form. Its data type is a dilution structure.

[0112] [A9-4-3] Loading field: Describes the amount of reagent added. Its data type is physical quantity. Its unit field allows one of the string enumeration values ​​["eq.", "g", "L", "kg", "mg", "mL", "mmol", "mol", "V", "X", "μL"].

[0113] [A9-4-4] Molecule field: Describes the reagent molecule, and its data type is ligand structure.

[0114] [A9-5] The `substrates` field describes the substrate or main reactant used to measure the catalytic performance of an enzyme. Its data type is an array of `substrate` structures. Each `substrate` structure represents a substrate or main reactant and contains the following member fields:

[0115] [A9-5-1]Charging Field: Describes the way the substrate is added to the reaction in which the enzyme’s catalytic performance is measured. Its data type is the Charging structure.

[0116] [A9-5-2] Dilution field: Describes the dilution method when the substrate is added in solution during the reaction that measures the catalytic performance of an enzyme. Its data type is the dilution structure.

[0117] [A9-5-3] Loading field: Describes the amount of substrate added to the reaction that measures the catalytic performance of the enzyme. Its data type is a physical quantity structure, and its unit field allows one of the following string enumeration values: ["eq.", "g", "L", "kg", "mg", "mL", "mmol", "mol", "V", "X", "μL"].

[0118] [A9-5-4] The molecule field describes the substrate molecule of the reaction that measures the catalytic performance of an enzyme. Its data type is ligand structure.

[0119] [A10] References field: Describes information about references or cited materials. Its data type is an array of reference structures. Each reference structure represents one reference or cited material and contains the following member fields:

[0120] [A10-1] Citation field: Describes the source of the citation, such as journals, books, dissertations, and their volume, issue, and page numbers. Its data type is string.

[0121] [A10-2] Date field: Describes the earliest date and time of publication of the citation, and its data type is date and time.

[0122] [A10-3] Note field: Describes the notes in the citation, and its data type is string.

[0123] [A10-4] Reference URI field: A Uniform Resource Identifier that describes the access to the citation from the Internet or a storage system. Its data type is string.

[0124] [A10-5] Title field: Describes the title of the citation; its data type is string.

[0125] [A11] Sequences field: Describes the enzyme sequence. Its data type is an array of sequence structures. Each sequence structure represents one enzyme sequence and contains the following member fields:

[0126] [A11-1] GenBank Accession Field: Describes the GenBank Accession of the enzyme sequence, i.e., the accession number in the GenBank database. Its data type is string. GenBank is a DNA sequence database established by the National Center for Biotechnology Information (NCBI) in the United States. It obtains sequence data from public resources, primarily provided directly by researchers or from large-scale genome sequencing projects. To ensure the data is as complete as possible, GenBank has established cooperative relationships with EMBL (European EMBL-DNA Database) and DDBJ (Dual Degrees Data Bank of Japan) for data exchange.

[0127] [A11-2]gi field: describes the GI number of the enzyme sequence, i.e., the GenInfoIdentifier, and its data type is string.

[0128] [A11-3] Mutations field: Describes all the differences in the enzyme sequence relative to its natural ancestor, i.e., mutations. Its data type is an array of mutation structures. Each mutation structure represents a mutation.

[0129] [A11-4] Note field: Describes notes about the enzyme sequence. Its data type is string.

[0130] [A11-5] ReferenceURI field: Describes the source from which this enzyme sequence was first reported, such as a reference or the Uniform Resource Identifier (URI) of the corresponding entry in a public database. Its data type is string.

[0131] [A11-6] Sequence field: Describes the specific content of the enzyme sequence. Its data type is string, and the allowed values ​​must conform to the FASTA format.

[0132] [A11-7] SequenceURI field: Describes a Uniform Resource Identifier (URI) for accessing this enzyme sequence from the Internet or a storage system. Its data type is string.

[0133] [A11-8] Type field: Describes the type of the enzyme sequence. Its data type is string, and the allowed value is one of the string enumeration values ​​["Nucleotide", "Peptide"], which represent the base sequence and amino acid sequence, respectively.

[0134] [A11-9] The uniProt field: Describes the UniProt ID of the enzyme sequence; its data type is string. UniProt is an abbreviation for Universal Protein, the most comprehensive and resource-rich protein database. It integrates data from the Swiss-Prot, TrEMBL, and PIR-PSD databases. Its data mainly comes from protein sequences obtained after the completion of genome sequencing projects. It contains a wealth of information on the biological functions of proteins from the literature.

[0135] [A12] Structures field: Describes the enzyme structure. Its data type is an array of several structure objects. Each structure object represents a three-dimensional structure of an enzyme and contains the following member fields:

[0136] [A12-1] Ligands field: Describes the ligands in the enzyme structure, that is, the small molecular components of the enzyme structure other than the main chain that makes up the protein. These can typically be water molecules, water-soluble ions, water-soluble small organic solutes, or small organic molecules bound to the protein surface or interior, such as substrates, products, inhibitors, etc. Its data type is an array of several ligand structures. Each ligand structure represents one ligand in the enzyme structure.

[0137] [A12-2] Mutations field: Describes all the differences between the sequence corresponding to the enzyme structure and its natural ancestor, i.e., mutations. Its data type is an array of mutation structures. Each mutation structure represents a mutation.

[0138] [A12-3] Note field: Notes describing the enzyme structure; its data type is string.

[0139] [A12-4] Reference URI field: Describes the source from which this enzyme structure was first reported, such as a reference or the Uniform Resource Identifier of the corresponding entry in a public database. Its data type is string.

[0140] [A12-5] SequenceURI field: Describes the source of the first report of the enzyme sequence corresponding to this enzyme structure, such as the Uniform Resource Identifier (URI) of the corresponding entry in a reference or public database. Its data type is string.

[0141] [A12-6] StructureURI field: A Uniform Resource Identifier (URI) describing access to this enzyme structure from the Internet or storage system. Its data type is string.

[0142] [A12-7] Structure field: Describes the specific content of the enzyme sequence. Its data type is string, and the allowed values ​​must conform to the PDB format.

[0143] [A12-8] Category (type) field: Describes the category of enzyme structure. Its data type is string, and the allowed value is one of the string enumeration values ​​["CryoSEM", "NMR", "Other", "Predicted Model", "XRD"], which respectively represent cryo-electron microscopy, nuclear magnetic resonance, other, predicted structure model, and X-ray crystal diffraction structure.

[0144] In response to the completion of structured processing of directed evolution data of enzymes and the acquisition of enzyme ECD structures, the builder can store the directed evolution data of various enzymes into computer-readable storage media according to the architecture of the ECD structures to construct the EvoCloud database of directed evolution data of enzymes. Thus, this invention can integrate multi-dimensional data involved in the directed evolution process of enzymes, characterize the catalytic performance of enzymes, and achieve efficient storage of this multi-dimensional data.

[0145] Furthermore, the aforementioned ECD structure can use JSON (JavaScript Object Notation) as its container format, which is the format defined by the ECMA-404 standard (European Computer Manufacturers Association Standard 404). JSON is a widely used computer program data exchange format. Its serialization, deserialization, node insertion, node deletion, and node editing are directly supported by the latest mainstream high-level computer programming languages ​​such as ECMAScript (ECMA-262), C (ISO / IEC 9899:2011), C++ (ISO / IEC 14882), Java (ISO / IEC TR 13066), and C# (ECMA-334). It can be directly understood and processed by computer programs without additional data processing.

[0146] Furthermore, the ECD structure provided by this invention can be implemented using a dictionary-based non-relational database organization structure. Please refer to... Figure 3 , Figure 3 A schematic diagram of a non-relational database of directed evolution data of enzymes provided according to some embodiments of the present invention is shown.

[0147] like Figure 3As shown, in different computer programming languages ​​or database systems, a dictionary structure, also known as an associative array or map, is an abstract data structure containing multiple ordered pairs similar to (key, value). This data structure supports various common operations such as pair retrieval, adding pairs, deleting pairs, and modifying pairs. For example, in the pair retrieval operation, the operation parameter is the key to be searched, and the return value is the corresponding value. If the corresponding key-value pair does not exist, some implementations will raise an exception, while others will create and add a new key-value pair using the given key, where the "value" is the default value of its type (zero, empty container, etc.). As another example, in the add pair operation, the builder can add a new key-value pair and establish a mapping from the new key to the new value; the operation parameters are the key and value to be added. As yet another example, in the delete pair operation, the builder can remove a key-value pair and cancel the mapping from that key to that value; the operation parameter is the key to be deleted. For example, in the operation of modifying a pair, the builder can change the value of an existing key-value pair and map the original key to the new value, with the operation parameters being the key and the value.

[0148] In some embodiments, the builder may designate one or more of the [A1] to [A12] fields of the ECD structure as directed evolution data fields that record directed evolution data of the enzyme. Here, each field consists of an interrelated field key and field value, wherein the field key indicates the type of the field, and the field value indicates the specific content corresponding to the field type.

[0149] Because computer programming languages ​​such as JavaScript (i.e., ECMAScript, a computer programming language defined by the international standard ECMA-262) have built-in basic data types that provide support for dictionary structures, modern NoSQL (No-only Structural Query Language) database systems such as IndexedDB, Redis, and MongoDB directly support dictionary structures as a way to store data. Furthermore, CAM (Content-Addressable Memory) also implements hardware-level support for dictionary structures. Therefore, ECD structures stored in a dictionary format can be directly understood and processed by computer programs without additional data processing. In addition, because dictionary structures store data in a one-to-one key-value relationship, their retrieval efficiency is far higher than other relational storage methods, making them more suitable for the large-scale storage, indexing, and computational optimization design of enzyme-directed evolution data.

[0150] Furthermore, after completing the structured processing of the directed evolution data of enzymes and determining the ECD structures of various enzymes, this invention can also use these ECD structures to construct a retrieval candidate set of directed evolution data of enzymes, so that the training device can efficiently retrieve multi-dimensional data involved in the directed evolution process of sample enzymes based on the input retrieval information. Please refer to... Figure 4 , Figure 4 A schematic flowchart of a molecular retrieval method for directed evolutionary data of enzymes provided according to some embodiments of the present invention is shown.

[0151] like Figure 4 As shown, in constructing the search candidate set, the builder can first define at least one data structure based on the supported search information types. For example, for a scheme supporting ligand-based small molecule retrieval, the defined data structure correspondingly includes a ligand node structure. This ligand node structure includes at least a directed evolution data field and an alignment information field. Here, the data type of the directed evolution data field can be the aforementioned ECD structure, forming an ecd field to comprehensively integrate the multi-dimensional data involved in the directed evolution process of the enzyme, characterize the enzyme's catalytic performance, and facilitate efficient retrieval of this multi-dimensional data. The data type of the alignment information field is adapted to the type of search information and uses the SMILES format, forming an alignment ligand (smiles) field, thus explicitly describing the molecular structure of the ligand using ASCII strings.

[0152] The builder can then create a variable S in computer memory and define its data type as an array of at least one ligand node structure to construct the corresponding ligand candidate set to accommodate the ligand node structures of various candidate enzymes in the EvoCloud database.

[0153] Next, the builder can traverse the EvoCloud database based on the [A9-3] product field and the [A9-5] substrate field, and construct a ligand node structure for each member e[i].product.smiles and e[i].substrates[j].smiles that contains the [A9-3] field or the [A9-5] field, initialize its ecd field ecd = e[i].ecd, initialize its alignment ligand (smiles) field smiles = e[i].substrates[j].smiles or e[i].product.smiles, and add the ligand node structure to the end of array S and push it onto the stack to construct the ligand candidate set corresponding to the ligand retrieval function.

[0154] Furthermore, in some embodiments, the ligand node structure may preferably include a unique ID field. During the construction of the ligand candidate set, the builder can traverse the EvoCloud database based on the [A4]id field, creating a ligand node structure for each record, initializing its ecd field ecd = e[i].ecd, initializing its unique ID field id = e[i].id, initializing its alignment ligand (smiles) field smiles = e[i].substrates[j].smiles or e[i].product.smiles, and adding this ligand node structure to the end of array S and pushing it onto the stack to initially construct the ligand candidate set corresponding to the ligand retrieval function. Subsequently, for ligand node structures where the alignment ligand (smiles) field is empty (i.e. enzymes that do not contain the [A9-3] and [A9-5] fields in the EvoCloud database), the builder can preferably output a prompt message "supplement the content of the alignment ligand (smiles) field" to prompt technicians to synthesize, express, and experimentally verify the missing ligand information in the EvoCloud database.

[0155] Please continue to refer to this. Figure 4 After constructing the candidate ligand set, the training device can obtain the reference ligand molecule via a human-computer interaction interface or a communication interface. The molecular information 's' is used to determine the similarity score between the molecular information 's' and the alignment information fields of each data structure in the ligand candidate set. Based on a preset similarity threshold, the reference ligand molecule is selected from the candidate set to produce the reference ligand molecule. The system generates a data structure for at least one associated enzyme of the sample enzyme and outputs directed evolution data fields and / or alignment information fields for each data structure similar to the sample enzyme, thereby enabling efficient retrieval of directed evolution multidimensional data of the sample enzyme.

[0156] Specifically, when determining the similarity score between the molecular information s and the alignment information fields of each data structure in the ligand candidate set, the training device can first perform molecular similarity calculations between the molecular information s and the alignment ligand fields S[k].smiles of all members S[k] (i.e., the aforementioned ligand node structure ligandNode) in the ligand candidate set S, to determine the similarity score between the ligand molecular information s of the sample enzyme and each alignment ligand field S[k].smiles. Here, this molecular similarity calculation can be implemented based on various existing algorithms and programs such as Tanimoto, Topological Fingerprint, and Morgan Fingerprint, which will not be elaborated upon here.

[0157] If the similarity score is greater than or equal to a preset similarity threshold (e.g., 0.5), the training device can assign the similarity score to the similarity score field S[k].similarityScore of the corresponding ligand node structure S[k]. Conversely, if the similarity score is less than the preset similarity threshold (e.g., 0.5), the training device can delete this member S[k] from the ligand candidate set S. Thus, in response to completing the molecular similarity calculation for all members S[k] in the entire ligand candidate set S, the results of batch retrieval based on the input molecule information s can be generated in the final ligand candidate set S'. Afterwards, the training device can determine the reference ligand molecule to be produced based on each ligand node structure S[k] in the final ligand candidate set S'. The sample enzyme is compared with various similar enzymes, and its ligandNode structure is output with the ecd field, id field, similarity score field and / or alignment ligand (smiles) field to determine the P of each similar ligand molecule. j Performance data, and the generation of various similar ligand molecules P j The amino acid sequence of the sample enzyme $s i .

[0158] Based on the above description, by constructing a multi-dimensional data involved in the directed evolution process of enzymes, a structure representing the catalytic performance of enzymes (ECD), and a ligand node structure containing the ECD directed evolution data fields, this invention can integrate the multi-dimensional data involved in the directed evolution process of enzymes, characterize the catalytic performance of enzymes, and achieve efficient storage and retrieval of this multi-dimensional data, so as to enable training devices to efficiently and reliably expand the training samples in the training set.

[0159] Please continue to refer to this. Figure 1 In the search for and acquisition of various similar ligand molecules P j Performance data, and the generation of various similar ligand molecules P jThe amino acid sequence of the sample enzyme $s i Subsequently, the training device can train on various similar ligand molecules P j The performance data is digitized to determine the corresponding performance description tensor. And the amino acid sequences of the enzymes in each sample were analyzed. i Perform digitization to determine the corresponding sequence description tensor.

[0160] In the comparison of similar ligand molecules P j During the digitization of performance data, the training device can first determine the generation of similar ligand molecules P. j The intrinsic values ​​of the sample enzyme, and the P molecules generated from each similar ligand. j The reaction condition data for the enzyme-catalyzed reaction were used to construct similar ligand molecules P. j performance descriptor

[0161] Specifically, consider the following enzyme-catalyzed reaction:

[0162]

[0163] In the formula: S represents the substrate, and P represents the product. Thus, the change in substrate concentration over time can be expressed as the following differential equation:

[0164] d[S] / dt=-k·[S] [1-1-1]

[0165] In the formula: [S] represents the concentration, k is the reaction rate constant, and the negative sign indicates that the substrate decreases continuously during the reaction, which is directly proportional to the concentration of the active enzyme and the catalytic activity of the enzyme unit. Thus, we can obtain:

[0166] k = a·[E] / [S]0 [1-1-2]

[0167] In the formula: [E] is the concentration of active enzyme, [S0] is the initial substrate concentration, and a is the relative catalytic activity, i.e., the intrinsic value of the enzyme's catalytic activity.

[0168] Furthermore, consider the scenario where the enzyme is continuously deactivated during the reaction:

[0169]

[0170] In the formula: d is the enzyme inactivation rate constant, i.e., the intrinsic value of the enzyme's tolerance. Thus, the change in the concentration of the remaining active enzyme over time can be written as the following differential equation:

[0171] d[E] / dt=-d·[E]

[0172] In the formula: the negative sign indicates that the active enzyme decreases continuously during the reaction. Integrating the above formula by a definite integral yields:

[0173] [E] t / [E]0=e -dt [1-1-3]

[0174] Then, substituting equations [1-1-2] and [1-1-3] into equation [1-1-1] and performing a definite integral, we obtain:

[0175] c = [S] t / [S]0=(a / d)·([E]0 / [S0)·(e -dt -1) [1-1]

[0176] In the formula: c is the conversion rate of the reaction (dimensionless), and [E]0 / [S0] is the ratio of enzyme concentration to substrate concentration at the initial moment of the reaction (i.e., enzyme loading, dimensionless).

[0177] Furthermore, for any enzyme reaction, two different time points t1 and t2, or two different enzyme loads ([E]), can be considered. 0,1 / [S] 0,1 ),([E] 0,2 / [S] 0,2 Substituting the values ​​of α and β, and their corresponding conversion rates c1 and c2 into [1-1] yields two equations. Solving these two equations simultaneously allows us to determine the intrinsic catalytic activity value a and the intrinsic tolerance value d of the enzyme.

[0178] Furthermore, considering the intrinsic value E of enzyme catalytic specificity, the following enzyme-catalyzed reaction can be considered:

[0179]

[0180] In the formula: S and P represent the substrate and product of the target reaction, respectively. i and P i Let represent the substrate and product of the i-th side reaction, respectively. Thus, we have:

[0181]

[0182] In the formula: E is the enzyme specificity constant (dimensionless).

[0183] In particular, for stereoselective reactions (or chiral reactions), equation [1-2-1] can be simplified to:

[0184]

[0185] We can obtain:

[0186] E=ln{1-c(1+ee)} / ln{1-c(1-ee)} [1-2-4]

[0187] In the formula: E is the enantioselectivity (dimensionless), c is the conversion rate of the reaction (dimensionless), and ee is the enantioselectivity excess of the product (dimensionless).

[0188] Subsequently, for the reference ligand molecule P ligand molecules of various similar types j The training device can read the similarity score field S[k].similarityScore of the corresponding ligand node structure to determine the similarity of ligand molecules P. j Relative to the reference ligand molecule The structural similarity score θ is used to construct the following vector as the performance descriptor of the corresponding sample enzyme:

[0189]

[0190] In the formula: a represents the relative catalytic activity (unit: s). -1 ), where d is the enzyme inactivation rate constant (unit: s). -1 E is the enzyme specificity constant (dimensionless), pH is the negative logarithm of the hydrogen ion concentration in the reaction system (dimensionless), and [S]0 is the initial substrate concentration of the reaction (in g·L). -1 T is the thermodynamic temperature of the reaction (in K), and θ is the similarity of ligand molecules P. j With reference ligand molecule The structural similarity score is given. Here, the relative catalytic activity a, inactivation rate constant d, and specificity constant E are intrinsic values ​​of the enzyme, while pH, initial substrate concentration [S]0, and structural similarity score θ describe the reaction conditions of the enzyme-catalyzed reaction.

[0191] In constructing and generating various similar ligand molecules P j Sample enzyme performance descriptor Subsequently, the training device can be used to train P based on various similar ligand molecules. j performance descriptor Construct a performance description matrix for each sample enzyme i, i.e.:

[0192]

[0193] Then, based on the performance description matrix of enzyme i for all samples... Construct a higher-order tensor of order n×u×7, namely the performance description tensor of all sample enzyme i.

[0194]

[0195] In addition, regarding the amino acid sequences of the enzymes in each sample $s i During the digitization process, the amino acid sequence data of each sample enzyme i retrieved from the EvoCloud database using the above retrieval method can be recorded as follows:

[0196] $s1,$s2,…,$s i ,…,$s n

[0197] In the formula: Where, m i For sequence $s i The length of $a$ varies depending on the sequence. i,j ∈{“A”,”C”,”D”,”E”,”F”,”G”,”H”,”I”,”K”,”L”,”M”,”N”,”P”,”Q”,”R”,”S”,”T”,”V”,”W”,”Y”}, representing the sequence $s$ i The j-th amino acid residue is one of the following: alanine residue, cysteine ​​residue, aspartic acid residue, glutamic acid residue, phenylalanine residue, glycine residue, histidine residue, isoleucine residue, lysine residue, leucine residue, methionine residue, asparagine residue, proline residue, glutamine residue, arginine residue, serine residue, threonine residue, valine residue, tryptophan residue, or tyrosine residue.

[0198] In some embodiments, the training device may preferably be used to generate a reference ligand molecule. The amino acid sequence of the enzyme $s0 and the amino acid sequences of the enzymes in each sample $s i Multiple sequence alignment (MSA) calculations are performed between each amino acid sequence $s0 to compare it with the amino acid sequence $s of each enzyme i sample. i The similarity between them. Here, the MSA calculation can be implemented based on various existing algorithms and programs such as basic pairwise comparison BLAST, ClustalW, T-Coffee, Muscle, MAFFT, etc., which will not be elaborated here.

[0199] Subsequently, if the amino acid sequence $s0 matches the amino acid sequence $s of sample enzyme i... i If the similarity between samples is greater than or equal to a preset first similarity threshold (e.g., 0.3), the training device can retain the amino acid sequence $s of enzyme i in that sample. i Conversely, if the amino acid sequence $s0 matches the amino acid sequence $s of sample enzyme i... iIf the similarity between samples is less than a preset first similarity threshold (e.g., 0.3), the training device can filter out the amino acid sequence $s of enzyme i in that sample. i The process continues until a set of aligned sequences composed of multiple sample enzymes i whose amino acid sequences are all similar to the amino acid sequence $s0 is identified.

[0200] $s'1,$s'2,…,$s' i ,…,$s' n

[0201] In the formula: for any index number i, there exists a sequence $s' i With sequence $s i The only counterpart, $s' i =[$a' i,1 ,$a' i,2 ,…,$a' i,j ,…,$a' i,m ], all sequences $s' i All have the same sequence length m, and any $a' i,j ∈{““,”A”,”C”,”D”,”E”,”F”,”G”,”H”,”I”,”K”,”L”,”M”,”N”,”P”,”Q”,”R”,”S”,”T”,”V”,”W”,”Y”}, representing the sequence $s' i The j-th amino acid residue is a placeholder inserted to align homologous sequence fragments, or one of the following: alanine residue, cysteine ​​residue, aspartic acid residue, glutamic acid residue, phenylalanine residue, glycine residue, histidine residue, isoleucine residue, lysine residue, leucine residue, methionine residue, asparagine residue, proline residue, glutamine residue, arginine residue, serine residue, threonine residue, valine residue, tryptophan residue, or tyrosine residue.

[0202] Furthermore, any amino acid residue or placeholder $a can be represented by a vector. To describe this, that is, $a ∈ {"", "A", "C", "D", "E", "F", "G", "H", "I", "K", "L", "M", "N", "P", "Q", "R", "S", "T", "V", "W", "Y"} has a unique vector. As its descriptor:

[0203]

[0204] In the formula: l describes the side chain length, which is the number of atoms (excluding H) from the α-C of the amino acid residue to the furthest point in the side chain; w describes the side chain width, which is the number of branches in the side chain of the amino acid residue (here, the number of branches in the ring is defined as 2); pK aThe acidity of the side chain is described by the negative logarithm of the dissociation constant of the side chain according to the acidic ionization; r describes the rigidity of the side chain, taking 1 if the side chain can rotate freely (such as proline residues), and taking 0 if the side chain cannot rotate freely; μ describes the polarity of the side chain, taking 1 if it is polar or hydrophilic, and taking 0 if it is nonpolar or hydrophobic.

[0205] Table 2-2 below provides the l, w, pK values ​​for all placeholders and native amino acid residues. a The values ​​of r and μ are given.

[0206] Table 2-2 Descriptor vector values ​​for natural amino acid residues

[0207]

[0208]

[0209] By analyzing the amino acid sequences of enzymes in each sample $s i By performing MSA calculations on each pair of enzymes and filtering out amino acid sequences with low sequence similarity, this invention can further improve the similarity between enzyme samples from the perspective of enzyme sequence similarity, thereby improving the prediction accuracy of the trained structure-activity relationship function F.

[0210] Then, for any sequence $s' i Each of them can be constructed according to equation [2-1] to uniquely correspond to a matrix. As a sequence descriptor, that is:

[0211]

[0212] All sequence description matrices This can further construct a higher-order tensor of order n×m×5, which is the sequence description tensor of all sample enzyme i.

[0213]

[0214] Those skilled in the art will understand that the above is based on the aligned sequence $s' filtered by MSA calculation. i To construct sequence description tensors The proposed solution is merely a non-limiting implementation method provided by the present invention, intended to clearly demonstrate the main concept of the invention and provide a preferred solution that can improve prediction accuracy, rather than being used to limit the scope of protection of the present invention.

[0215] Alternatively, in other embodiments, given that the above-described ligand-based retrieval method has already eliminated reference ligand molecules... For dissimilar molecules with a molecular structure similarity lower than a preset similarity threshold (e.g., 0.5), the training device can skip the MSA calculation and screening steps mentioned above and directly use the amino acid sequence of each sample enzyme. i To construct sequence description tensors This also enables the construction of a structure-activity relationship function F to predict the basic function of catalytic enzymes.

[0216] Please continue to refer to this. Figure 1 After determining the performance description tensor of each sample enzyme, and sequence description tensor Then, the training device can construct the sequence description tensor. Regarding the performance description tensor Quantitative structure-activity relationship function The numerical approximate solution FQ of the quantitative structure-activity relationship function FQ is determined through machine learning training. # .

[0217] Specifically, in training numerical approximation solutions F # During the process, the training device can first assume that there is a mathematical relation F that theoretically satisfies the following equation:

[0218]

[0219] However, given the complexity of practical problems in enzyme-catalyzed reactions, it is often difficult to exhaustively provide a complete and precise analytical expression for the structure-activity relationship function F. Therefore, a training device can construct a neural network model based on the sequence description tensor of each enzyme sample in the training set. and performance description tensor Machine learning training is performed on the constructed neural network model to obtain an approximate numerical solution F of the structure-activity relationship function F. # ,Right now

[0220]

[0221] Thus, the prediction device provided in the second aspect of the present invention can be based on the approximate numerical solution F. # Mathematical fitting is performed to predict the amino acid sequence of the target enzyme suitable for producing the target product P.

[0222] Please continue to refer to this. Figure 1 The solid line on the right represents the online prediction stage of the target enzyme. Users can first provide molecular information about the target product P via the prediction device's human-machine interface or communication interface. In response to obtaining this molecular information, the prediction device can determine the performance descriptor of the target enzyme based on this information and the expected performance of the enzyme in catalyzing the reaction to produce the target product P. And based on this performance descriptor Perform mathematical fitting to determine the matrix descriptor of the target enzyme.

[0223] Specifically, the performance expectations for the enzyme-catalyzed reaction to produce the target product P include, but are not limited to, expectations of the intrinsic values ​​of the target enzyme, and / or expectations of the reaction conditions for the enzyme-catalyzed reaction. Further, the expectations of the intrinsic values ​​of the target enzyme include, but are not limited to, expectations of the relative specific catalytic activity a, expectations of the inactivation rate constant d, and expectations of the specificity constant E. The expectations of the reaction conditions for the enzyme-catalyzed reaction include, but are not limited to, expectations of the pH of the reaction system, expectations of the initial substrate concentration [S]0, expectations of the thermodynamic temperature T, and expectations of the ratio of the target product P to the reference ligand molecule. The expected structural similarity θ.

[0224] In some preferred embodiments, the performance descriptor of the target enzyme is determined based on molecular information and performance expectations. Previously, the prediction device could prioritize the first molecule information of the target product P and the reference ligand molecule corresponding to the quantitative structure-activity relationship function F. The second molecular information is used to perform molecular similarity calculations to determine the similarity score θ between the first and second molecular information. This molecular similarity calculation can be implemented using various existing algorithms and programs such as Tanimoto, Topological Fingerprint, and Morgan Fingerprint, which will not be elaborated upon here.

[0225] If the similarity score θ is less than a preset second similarity threshold (e.g., 0.5), the prediction device can determine that the target product P is similar to the reference ligand molecule. The molecular structures have low similarity, and the pre-fitted structure-activity relationship function F # If the prediction of the catalytic enzyme for the target product P is not suitable, the prediction will be terminated and a prompt will be output stating "Supplement the quantitative structure-activity relationship function F' for the corresponding target product P". Here, the second similarity threshold can be determined based on the expectation of structural similarity θ.

[0226] Conversely, if the similarity score θ is greater than or equal to a preset second similarity threshold (e.g., 0.5), the prediction device can determine that the target product P is similar to the reference ligand molecule. The molecular structures are highly similar, and the pre-fitted structure-activity relationship function F # The prediction of a suitable catalytic enzyme for the target product P is used to determine the performance descriptor of the target enzyme that generates the target product P based on the first molecule information and the aforementioned performance expectations for the enzyme-catalyzed reaction. Right now:

[0227]

[0228] Where 'a' represents the relative catalytic activity, in units of s. -1 d is the deactivation rate constant, in seconds. -1 E is the specificity constant, dimensionless; pH is the negative logarithm of the hydrogen ion concentration in the reaction system, dimensionless; [S]0 is the initial substrate concentration, in g·L. -1 T is the thermodynamic temperature, in K; θ is the ratio of the target product P to the reference ligand molecule. The structural similarity score.

[0229] Then, the prediction device can generate a performance descriptor for the target enzyme. Substituting into the above formula [3], that is

[0230]

[0231] To determine the corresponding matrix descriptor Then the obtained matrix descriptor Substituting into the above equations [2-3], we can determine the descriptors of the multiple amino acid residues $a involved in the target enzyme.

[0232] Then, the prediction device can use these multiple amino acid residue descriptors Substituting into the above equation [2-1], we can determine the side chain length l, side chain width w, and side chain acidity pK of the target enzyme. a And the rigidity of the side chain r, and thereby determine the amino acid sequence $s of the target enzyme.

[0233] Thus, the present invention can efficiently and reliably select a suitable catalytic enzyme for generating the target product P, thereby facilitating the efficient production of the target product P and reducing the production and emission of byproducts, thereby promoting the further development of enzyme catalysis technology as a green chemical technology.

[0234] Furthermore, in some embodiments of the present invention, in response to predicting and determining the amino acid sequence $s of the target enzyme, the prediction device may also preferably perform a sequence search based on the amino acid sequence $s to determine a variety of associated enzymes with similar sequences for the user to select.

[0235] Please refer to Figure 5 , Figure 5 A schematic flowchart of a sequence retrieval method for directed evolution data of enzymes according to some embodiments of the present invention is shown.

[0236] like Figure 5As shown, sequence retrieval of enzyme directed evolution data can be achieved based on a pre-constructed sequence candidate set. Specifically, during the construction of the sequence candidate set, the builder can first define at least one data structure according to the supported retrieval information type. For example, for a scheme supporting sequence-based retrieval, the defined data structure correspondingly includes a sequence node structure. This sequence node structure includes at least a directed evolution data field and an alignment information field. Here, the data type of the directed evolution data field can be selected from the aforementioned ECD structure, i.e., forming an ecd field, to comprehensively integrate the multi-dimensional data involved in the directed evolution process of the enzyme, characterize the catalytic performance of the enzyme, and facilitate efficient retrieval of this multi-dimensional data. The data type of the alignment information field is adapted to the type of retrieval information and uses the FASTA format, i.e., forming an alignment sequence field to represent the nucleic acid sequence or peptide sequence of the corresponding candidate enzyme.

[0237] Afterwards, the builder can create a variable S in computer memory and define its data type as an array consisting of at least one sequence node structure to construct the corresponding sequence candidate set to accommodate the sequence node structures of various candidate enzymes in the EvoCloud database.

[0238] Next, the builder can traverse the EvoCloud database based on the [A11] sequence field, construct a sequence node structure for each member e[i].sequences[j] that contains the [A11] field, initialize its ecd field ecd = e[i].ecd, initialize its alignment sequence field sequence = e[i].sequences[j], and add the sequence node structure to the end of array S and push it onto the stack to construct the sequence candidate set corresponding to the sequence retrieval function.

[0239] Furthermore, in some embodiments, the sequence node structure may preferably include a unique ID field. During the construction of the sequence candidate set, the builder can traverse the EvoCloud database based on the [A4]id field, creating a sequence node structure for each piece of data, initializing its ecd field ecd = e[i].ecd, initializing its unique ID field id = e[i].id, initializing its alignment sequence field sequence = e[i].sequences[j], and adding this sequence node structure to the end of array S and pushing it onto the stack to initially construct the sequence candidate set corresponding to the sequence retrieval function. Afterwards, for sequence node structures with an empty alignment sequence field (i.e., enzymes in the EvoCloud database that do not contain the [A11] field), the builder can preferably output a prompt message "Supplement the content of the alignment sequence field" to prompt technicians to synthesize, express, and experimentally verify the missing sequence information in the EvoCloud database.

[0240] Please continue to refer to this. Figure 5 After constructing the sequence candidate set, the prediction device can obtain the amino acid sequence $s of the target enzyme, determine the similarity score between the amino acid sequence $s and the alignment sequence field of each sequence node structure in the sequence candidate set, select at least one similar associated enzyme sequence node structure from the sequence candidate set according to the preset first similarity threshold, and output the directed evolution data field and / or alignment sequence field of each associated enzyme sequence node structure to achieve efficient retrieval of directed evolution multidimensional data of each associated enzyme.

[0241] Specifically, when determining the similarity score between the amino acid sequence $s and the alignment information fields of each data structure in the sequence candidate set, the prediction device can first perform multiple sequence alignment (MSA) calculations on the amino acid sequence $s and the alignment sequence field S[k].sequence of all members S[k] (i.e., the sequence node structure mentioned above) in the sequence candidate set S, to determine the similarity score between the sequence content s of the target enzyme and each alignment sequence field S[k].sequence. Here, the MSA calculation can be implemented based on various existing algorithms and programs such as basic pairwise comparison BLAST, ClustalW, T-Coffee, Muscle, and MAFFT, which will not be elaborated here.

[0242] If the similarity score is greater than or equal to a pre-set first similarity threshold (e.g., 0.3), the prediction device can assign the similarity score to the similarity score field S[k].identityScore of the corresponding sequence node structure S[k]. Conversely, if the similarity score is less than the pre-set first similarity threshold (e.g., 0.3), the prediction device can delete this member S[k] from the sequence candidate set S. Thus, in response to completing the MSA calculation of all members S[k] in the entire sequence candidate set S, the results of batch retrieval based on the input sequence s can be formed in the final sequence candidate set S'. Afterwards, the prediction device can determine each associated enzyme similar to the target enzyme based on each sequence node structure S[k] in the final sequence candidate set S', and output the ecd field, id field, similarity score field and / or alignment sequence field in its sequenceNode structure to achieve efficient retrieval of directed evolutionary multidimensional data of the target enzyme.

[0243] Based on the above description, by constructing a multi-dimensional data involved in the directed evolution process of enzymes, a structure representing the catalytic performance of enzymes (ECD), and a sequence node structure containing the ECD directed evolution data fields, this invention can integrate the multi-dimensional data involved in the directed evolution process of enzymes, characterize the catalytic performance of enzymes, and achieve efficient storage and retrieval of this multi-dimensional data. This allows for the efficient and reliable provision of associated enzymes with similar sequences to the target enzyme, thereby expanding the user's selection range.

[0244] Although the methods described above are illustrated and depicted as a series of actions for the sake of simplicity, it should be understood and appreciated that these methods are not limited by the order of the actions, as some actions may occur in a different order and / or concurrently with other actions from the illustrations and descriptions herein or not illustrated and described herein but which may be understood by those skilled in the art, according to one or more embodiments.

[0245] Those skilled in the art will understand that information, signals, and data can be represented using any of a variety of different techniques and arts. For example, the data, instructions, commands, information, signals, bits, symbols, and chips described throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or optical particles, or any combination thereof.

[0246] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in a generalized manner in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the invention.

[0247] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor such that the processor can read and write information to / from the storage medium. In an alternative, the storage medium may be integrated into the processor. The processor and storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and storage medium may reside as discrete components in the user terminal.

[0248] In one or more exemplary embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functionality may be stored or transmitted as one or more instructions or code on or through a computer-readable medium. A computer-readable medium includes both computer storage media and communication media, encompassing any medium that facilitates the transfer of a computer program from one location to another. A storage medium may be any available medium accessible to a computer. By way of example and not limitation, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and is accessible to a computer. Any connection is also legitimately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of a medium. As used in this article, disk and disc include compact discs (CDs), laser discs, optical discs, digital multi-purpose discs (DVDs), floppy disks, and Blu-ray discs. Disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of these should also be included within the scope of computer-readable media.

[0249] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for predicting enzymes, characterized in that, Includes the following steps: Obtain molecular information of the target product P; Based on the molecular information and the expected performance of the enzyme in catalyzing the reaction to produce the target product P, determine the performance descriptor of the target enzyme that produces the target product P. : in, The relative catalytic activity of the target enzyme is expressed in seconds (s). - ¹; The inactivation rate constant of the target enzyme is given in seconds. - ¹; is the specificity constant of the target enzyme, dimensionless; is the negative logarithm of the hydrogen ion concentration in the reaction system, dimensionless; Initial substrate concentration, in g·L - ¹; Temperature is thermodynamic temperature, in Kelvin (K). The target product P and the reference ligand molecule Structural similarity score; According to the performance descriptor Perform mathematical fitting to determine the matrix descriptor of the target enzyme. ; According to the matrix descriptor Determine multiple amino acid residue descriptors involved in the target enzyme. : in, The side chain length is described by the number of atoms (excluding H) from the α-C of an amino acid residue to the furthest point in the side chain. The width of the side chain is described by the number of branches in the side chain of the amino acid residues; The acidity of the side chain is described by the negative logarithm of the dissociation constant of the side chain when it is acidically ionized. Describe the stiffness of the side chain; take 1 if the side chain can rotate freely, and take 0 if the side chain cannot rotate freely. Describe the polarity of the side chain; assign 1 if it is polar or hydrophilic, and 0 if it is nonpolar or hydrophobic; and Based on the multiple amino acid residue descriptors The amino acid sequence of the target enzyme was determined.

2. The prediction method as described in claim 1, characterized in that, The performance expectations include expectations for the intrinsic values ​​of the target enzyme and / or expectations for the reaction conditions of the enzyme-catalyzed reaction.

3. The prediction method as described in claim 2, characterized in that, The desired intrinsic values ​​of the target enzyme include expectations for the relative catalytic activity, the inactivation rate constant, and the specificity constant. The desired reaction conditions for the enzyme-catalyzed reaction include expectations for the pH of the reaction system, the initial substrate concentration, the thermodynamic temperature, and the reaction between the target product and the reference ligand molecule. The expected structural similarity.

4. The prediction method as described in claim 3, characterized in that, According to the performance descriptor Perform mathematical fitting to determine the matrix descriptor of the target enzyme. The steps include: According to the performance descriptor and a pre-determined quantitative structure-activity relationship function F Numerical approximation solution F # Determine the matrix descriptor of the target enzyme. The numerical approximation solution F # This is achieved by generating the reference ligand molecule. The sequence description tensor of at least one enzyme and performance description tensor The quantitative structure-activity relationship function between them was determined by mathematical fitting.

5. The prediction method as described in claim 4, characterized in that, Determine the numerical approximate solution F # The steps include: Obtain the reference ligand molecule At least one similar ligand molecule The performance data, and the generation of each of the aforementioned similar ligand molecules. The amino acid sequence of the sample enzyme ; Based on the aforementioned similar ligand molecules Based on the performance data, determine the corresponding performance description tensor. ; Based on the amino acid sequences of each sample enzyme. Determine the corresponding sequence description tensor ;as well as Construct the sequence description tensor Regarding the performance description tensor Quantitative structure-activity relationship function The quantitative structure-activity relationship function is determined through machine learning training. F Numerical approximation solution F # .

6. The prediction method as described in claim 5, characterized in that, The acquisition of the reference ligand molecule At least one similar ligand molecule The performance data, and the generation of each of the aforementioned similar ligand molecules. The amino acid sequence of the sample enzyme The steps include: Obtain the reference ligand molecule Molecular information; Based on the molecular information, a corresponding ligand candidate set is determined, wherein the ligand candidate set includes at least one ligand node structure, and the ligand node structure includes at least a directed evolution data field and a comparison ligand field; The similarity score of the ligand field between the molecular information and each ligand node structure in the ligand candidate set is determined to screen the reference ligand molecule. At least one similar ligand molecule The ligand node structure; and Based on the directed evolution data fields of each ligand node structure, the generation of each of the aforementioned similar ligand molecules is determined. The amino acid sequence of the sample enzyme and each of the aforementioned similar ligand molecules Performance data.

7. The prediction method as described in claim 5, characterized in that, The similar ligand molecules Performance data includes the generation of the similar ligand molecules. The intrinsic value of the sample enzyme, and the generation of the similar ligand molecules. The reaction condition data for enzyme-catalyzed reactions, based on the aforementioned similar ligand molecules. Based on the performance data, determine the corresponding performance description tensor. The steps include: Based on the intrinsic values ​​of the sample enzyme and the reaction condition data of the enzyme-catalyzed reaction, the similar ligand molecule was constructed. performance descriptor ; Based on the aforementioned similar ligand molecules performance descriptor Construct a performance description matrix for each of the sample enzymes. ;as well as Based on the performance description matrix of all sample enzymes Construct the performance description tensor .

8. The prediction method as described in claim 7, characterized in that, The similar ligand molecule is constructed based on the intrinsic value of the sample enzyme and the reaction condition data of the enzyme-catalyzed reaction. performance descriptor The steps include: Based on the aforementioned similar ligand molecules Data on the change of substrate concentration over time in enzyme-catalyzed reactions were used to determine the corresponding sample enzyme's response to the similar ligand molecules. Relative catalytic activity ; Based on the inactivation data of the sample enzyme in the enzyme-catalyzed reaction, determine the inactivation rate constant of the sample enzyme. ; The specificity constant of the sample enzyme is determined based on the reaction rates of the target reaction and at least one side reaction in the enzyme-catalyzed reaction. ; Based on the relative catalytic activity The deactivation rate constant and the specificity constant The pH value, initial substrate concentration, thermodynamic temperature, and similar ligand molecules of the enzyme-catalyzed reaction. With the reference ligand molecule Structural similarity scores are used to construct the similar ligand molecules. performance descriptor .

9. The prediction method as described in claim 5, characterized in that, The amino acid sequence of each of the sample enzymes is used. Determine the corresponding sequence description tensor The steps include: The reference ligand molecule was calculated and generated separately. The amino acid sequence of the enzyme The amino acid sequences of the enzymes in each of the samples The similarity between them; The reference ligand molecule is selected and generated based on a preset first similarity threshold. The enzyme sequence Similar enzymes in multiple samples; and Based on the amino acid sequences of the similar multiple sample enzymes Determine the sequence description tensor .

10. The prediction method as described in claim 5 or 9, characterized in that, The amino acid sequence of each of the sample enzymes is used. Determine the corresponding sequence description tensor The steps include: Obtain the amino acid sequence of the sample enzyme. Side chain lengths of each amino acid residue and placeholder. Side chain width Side chain acidity Side chain rigidity and side chain polarity To construct a description vector for the amino acid residues and the placeholders. ; Based on the amino acid sequence of the sample enzyme Description vectors of each amino acid residue and each placeholder. Construct the amino acid sequence of the sample enzyme. Sequence description matrix ;as well as Based on the amino acid sequences of each sample enzyme. Sequence description matrix Construct the sequence description tensor .

11. The prediction method as described in claim 4, characterized in that, Based on the molecular information and the expected performance of the enzyme in catalyzing the reaction to produce the target product P, a performance descriptor for the target enzyme that produces the target product P is determined. Previously, the prediction method also included the following steps: The first molecule information of the target product P, and the corresponding quantitative structure-activity relationship function. F Reference ligand molecule The second molecular information is used to perform molecular similarity calculations to determine the similarity score between the first molecular information and the second molecular information; and In response to the judgment result that the similarity score is greater than or equal to a preset second similarity threshold, based on the molecular information and the performance expectation for the enzyme-catalyzed reaction to generate the target product P, a performance descriptor for the target enzyme generating the target product P is determined. .

12. The prediction method as described in claim 11, characterized in that, It also includes the following steps: In response to the judgment result that the similarity score is less than the second similarity threshold, a supplementary quantitative structure-activity relationship function corresponding to the target product P is output. F’ The prompt.

13. The prediction method as described in claim 1, characterized in that, After determining the amino acid sequence of the target enzyme, the prediction method further includes the following steps: Based on the amino acid sequence of the target enzyme, a corresponding sequence candidate set is determined, wherein the sequence candidate set includes at least one sequence node structure, and the sequence node structure includes at least a directed evolution data field and an alignment sequence field; Determine the similarity score between the amino acid sequence of the target enzyme and the alignment sequence field of each sequence node structure in the sequence candidate set; Based on a preset first similarity threshold, at least one sequence node structure similar to the target enzyme is screened from the sequence candidate set; and Output directed evolution data fields and / or alignment sequence fields for each sequence node structure similar to the target enzyme.

14. An enzyme prediction device, characterized in that, include: Memory, and A processor, connected to the memory, and configured to implement the enzyme prediction method as described in any one of claims 1 to 13.

15. A computer-readable storage medium storing computer instructions thereon, characterized in that, When the computer instructions are executed by the processor, the enzyme prediction method as described in any one of claims 1 to 13 is implemented.

Citation Information

Patent Citations

  • Directed evolution method and device of enzyme

    CN115171783A