Neural network-based enzyme modification method

Through the neural network-based enzyme transformation method, the protein activity prediction model is used to transform the active region of the enzyme, which solves the shortcomings in stability, tolerance and selectivity of the existing enzyme transformation methods, and achieves more efficient enzyme transformation screening.

CN119964632APending Publication Date: 2025-05-09ZHEJIANG UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510101129.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing enzyme modification methods have shortcomings in terms of stability, tolerance and selectivity, and are difficult to meet practical application needs.

Method used

The enzyme transformation method based on neural network is used to determine the final enzyme transformation scheme by obtaining the protein sequence database, determining the active region and target loop structure of the enzyme to be modified, performing multiple replacements of amino acid sequences and optimizing the activity prediction model.

Benefits of technology

It improves the screening efficiency of enzyme transformation, can more effectively improve the stability, tolerance and selectivity of enzymes, and meet practical application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964632A_ABST
    Figure CN119964632A_ABST
Patent Text Reader

Abstract

The invention relates to the field of enzyme engineering, and discloses an enzyme modification method based on a neural network. According to the method, the protein structure of the to-be-modified enzyme is reconstructed by using the pre-trained protein activity prediction model, the potential region with relatively high enzyme functional activity is determined, and the ring structure of the potential region with relatively high enzyme activity is rationally designed and modified by using the protein activity prediction model, so that the efficiency of screening the high-yield enzyme is improved; the activity multi-round optimization of the enzyme to be modified is carried out based on the amino acid sequence of the protein ring structure, and the screening efficiency of enzyme modification can be more effectively improved. In conclusion, a potential region with relatively high enzyme functional activity is determined through the provided deep learning strategy, and the enzyme is rationally designed and modified by using artificial intelligence to obtain an enzyme modification scheme, so that the efficiency of screening the effective high-yield enzyme is improved, and inconvenience caused by manual screening is avoided and reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of enzyme engineering, and in particular to an enzyme modification method based on neural network. Background Art

[0002] Enzymatic reactions have the advantages of high efficiency, specificity, and environmental protection. Therefore, enzymes are not only widely used in traditional fields such as chemical industry, food, and environment, but also play an irreplaceable role in emerging technologies and products such as gene editing, stem cell technology, and targeted drugs. A large number of studies have found that natural protein enzymes often cannot meet the needs of practical applications in terms of stability, tolerance, and selectivity. Therefore, optimizing and modifying enzyme molecules is not only the focus of protein science research, but also an urgent need for industrial production.

[0003] In the field of protein engineering, commonly used protein modification methods include directed evolution, semi-rational design and rational design. Directed evolution guides proteins to accumulate beneficial mutations through multiple rounds of repeated mutation, expression and screening. However, directed evolution introduces mutations in a random manner, resulting in a large number of mutants, which is not conducive to artificial screening. Semi-rational design selects several sites as modification targets based on prior knowledge such as crystal structure and catalytic mechanism, thereby improving the efficiency of modification. However, the success of semi-rational design is closely related to the richness of prior knowledge, which leads to considerable limitations in its application. Rational design attempts to obtain enzymes with desired properties by precisely regulating the structural space of proteins, but it is currently still limited by the high-precision acquisition of the spatial structure of enzyme molecules and the rational understanding of structure-function relationships and catalytic mechanisms.

[0004] The protein ring structure, often referred to as the protein loop structure, is an important component of the protein molecule. Summary of the invention

[0005] In order to solve the technical problem of enzyme modification, the present invention provides an enzyme modification method based on neural network.

[0006] The specific technical scheme of the present invention is: In a first aspect, the present invention provides an enzyme modification method based on a neural network, comprising the following steps: Acquire a protein sequence database, wherein the protein sequence database includes amino acid sequence information corresponding to each of a plurality of basic protein loop structures; Based on the protein activity prediction model and the original amino acid sequence information of the enzyme to be modified, the original activity information of the enzyme to be modified is obtained; Obtaining an active region of the enzyme to be modified, and obtaining at least one target loop structure in the active region; Based on the amino acid sequence information corresponding to each of the multiple basic protein loop structures in the protein sequence database, the amino acid sequence corresponding to at least one target loop structure in the active region is replaced multiple times to obtain multiple modified enzymes and modified amino acid sequence information corresponding to each of the multiple modified enzymes; Based on the protein activity prediction model and the modified amino acid sequence information corresponding to each of the multiple modified enzymes, the modified activity information corresponding to each of the modified amino acid sequence information corresponding to each of the multiple modified enzymes is obtained, and based on the modified activity information corresponding to each of the multiple modified enzymes, the final modification plan of the enzyme to be modified is determined.

[0007] Preferably, in the protein sequence database, clustering is performed according to the amino acid sequence information corresponding to each of the multiple basic proteins, and the structural representation of the amino acid sequence information corresponding to each of the multiple basic proteins is unified.

[0008] Preferably, the protein sequence database also includes information on the number of occurrences of each of the multiple basic protein loop structures, and distance information on secondary structures connected to each of the multiple basic protein loop structures.

[0009] Further preferably, according to the distance information of the secondary structures connected to each of the multiple basic protein loop structures, the amino acid sequence corresponding to at least one target loop structure in the active region is replaced multiple times based on the same distance.

[0010] Further preferably, according to the occurrence frequency information of each of the multiple basic protein loop structures, the amino acid sequence corresponding to at least one target loop structure in the active region is replaced multiple times based on the priority high frequency frequency.

[0011] Preferably, the method further comprises: The modified amino acid sequence information corresponding to each of the multiple modified enzymes and the modified activity information corresponding to each of the modified amino acid sequence information corresponding to each of the multiple modified enzymes are input into a protein activity prediction model, and the modified amino acid sequence information corresponding to each of the multiple modified enzymes is compared with the original amino acid sequence information, and the modified activity information corresponding to each of the modified amino acid sequence information corresponding to each of the multiple modified enzymes is compared with the original activity information, and the protein activity prediction model is optimized according to the comparison results to obtain the protein activity prediction model.

[0012] Preferably, the protein activity prediction model is a reproducible neural network model, and the training of the protein activity prediction model uses the difference between the transformed activity information corresponding to the transformed amino acid sequence information corresponding to each of the multiple transformed enzymes and the original activity information as a loss function.

[0013] In a second aspect, the present invention provides a computer device, comprising: processor; A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the above-mentioned method for obtaining an enzyme modification scheme based on a ring structure by executing the executable instructions.

[0014] In a third aspect, the present invention provides a computer-readable storage medium, specifically: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for obtaining the ring structure-based enzyme modification scheme is implemented.

[0015] Compared with the prior art, the present invention has the following technical effects: The present invention uses a pre-trained protein activity prediction model to reconstruct the protein structure of the enzyme to be modified, determine the potential area with higher enzyme functional activity, and use the protein activity prediction model to rationally design and modify the ring structure of the potential area with higher enzyme activity, thereby improving the efficiency of screening high-yield enzymes. Multiple rounds of optimization of the activity of the enzyme to be modified based on the amino acid sequence of the protein ring structure can more effectively improve the screening efficiency of enzyme modification. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flow chart of a neural network-based enzyme modification method shown in the disclosed embodiment of the present invention. DETAILED DESCRIPTION

[0017] Here, exemplary embodiments will be described in detail, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are only examples of devices and methods consistent with some aspects of the present invention.

[0018] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "the" used in the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0019] An embodiment of a neural network-based enzyme modification method of the present invention is described in detail below in conjunction with the accompanying drawings.

[0020] The present invention provides an enzyme modification method based on neural network, such as Figure 1 As shown, the following steps are included: S1. Obtain a protein sequence database, wherein the protein sequence database includes amino acid sequence information corresponding to each of a plurality of basic protein loop structures; S2. Based on the protein activity prediction model and the original amino acid sequence information of the enzyme to be modified, obtaining the original activity information of the enzyme to be modified; S3, obtaining the active region of the enzyme to be modified, and obtaining at least one target loop structure in the active region; Based on the amino acid sequence information corresponding to each of the multiple basic protein loop structures in the protein sequence database, the amino acid sequence corresponding to at least one target loop structure in the active region is replaced multiple times to obtain multiple modified enzymes and modified amino acid sequence information corresponding to each of the multiple modified enzymes; S4. Based on the protein activity prediction model and the modified amino acid sequence information corresponding to each of the multiple modified enzymes, obtain the modified activity information corresponding to each of the modified amino acid sequence information corresponding to each of the multiple modified enzymes, and based on the modified activity information corresponding to each of the multiple modified enzymes, determine the final modification plan of the enzyme to be modified.

[0021] This method uses a protein activity prediction model to reconstruct the protein structure of the enzyme to be modified, that is, by determining the potential area with high enzyme functional activity, and using the protein activity prediction model to rationally design and modify the ring structure of the potential area with high enzyme activity, thereby providing a method with high efficiency in screening high-yield enzymes. Multiple rounds of optimization of the activity of the enzyme to be modified based on the amino acid sequence of the protein ring structure can more effectively improve the screening efficiency of enzyme modification.

[0022] In one embodiment, the website of the protein sequence database may be http: / / loopfinder.zjut.edu.cn / .

[0023] In one embodiment, the protein sequence database is clustered according to the amino acid sequence information corresponding to each of the multiple basic proteins, and the structural representation of the amino acid sequence information corresponding to each of the multiple basic proteins is unified. Proteins in the existing protein loop structure data set are clustered according to sequence similarity to avoid oversampling of certain protein loop regions, for example, the PDB database. The retained amino acid sequence information is resampled and distributed to construct a protein structure data set that conforms to the natural distribution. Existing protein information can be extracted from the protein crystal structure database.

[0024] In another embodiment, the protein sequence database also includes information on the number of occurrences of each of the multiple basic protein loop structures, and distance information on secondary structures connected to each of the multiple basic protein loop structures.

[0025] Furthermore, according to the distance information of the secondary structures connected to each of the multiple basic protein loop structures, the amino acid sequence corresponding to at least one target loop structure in the active region is replaced multiple times based on the same distance.

[0026] Alternatively, further, according to the occurrence frequency information of each of the multiple basic protein loop structures, the amino acid sequence corresponding to at least one target loop structure in the active region is replaced multiple times based on the priority high frequency number.

[0027] In one embodiment, the method further comprises: The modified amino acid sequence information corresponding to each of the multiple modified enzymes and the modified activity information corresponding to each of the modified amino acid sequence information corresponding to each of the multiple modified enzymes are input into a protein activity prediction model, and the modified amino acid sequence information corresponding to each of the multiple modified enzymes is compared with the original amino acid sequence information, and the modified activity information corresponding to each of the modified amino acid sequence information corresponding to each of the multiple modified enzymes is compared with the original activity information, and the protein activity prediction model is optimized according to the comparison results to obtain the protein activity prediction model.

[0028] In one embodiment, the protein activity prediction model is a reproducible neural network model, and the training of the protein activity prediction model uses the difference between the transformed activity information corresponding to the transformed amino acid sequence information corresponding to each of the multiple transformed enzymes and the original activity information as a loss function.

[0029] Furthermore, the protein activity prediction model includes an encoding layer, a long short-term memory layer, a dense layer, and a decoding layer. The encoding layer is used to encode the input amino acid sequence information into matrix information, the long short-term memory layer is used to store the activity information corresponding to the amino acid sequence information, the dense layer is used to extract the activity information corresponding to the amino acid sequence information and map the extracted activity information to the decoding layer, and the decoding layer is used to output the prediction result as amino acid sequence information.

[0030] The prediction model is constructed based on the reproducible neural network model. The network model first has an encoding layer, whose main function is to encode the amino acid sequence information into a matrix table; then a long short-term memory layer, which uses the ReLU function as the activation function; then a dense layer, whose ReLU is used as the activation function; and finally a decoding layer, which passes the predicted data through the decoding layer and finally outputs the prediction result as an amino acid sequence.

[0031] The formula for reproducing the neural network model is as follows: Among them, L represents the Lth filter, F represents the size of the filter (here F = 3), C represents the input channel (C = 12), W is the weight, b is the bias, X is the input data, (i, j, k) is the position of the output data, (m, n, d) is the position of the input data, c represents the number of input channels C, ranging from 0 to C; x is the independent variable of the ReLU function.

[0032] In this embodiment, the original data is encoded into a matrix table, and the matrix table format is standardized. The encoding layer matrix length is the maximum sequence length plus 1. The encoding layer formula is shown in formula (3) and formula (4): h t is the hidden state at the current time step t, σ is the activation function, and W hh is the weight matrix from hidden state to hidden state, h t−1 is the hidden state of the previous time step, W xh is the weight matrix input to the hidden state, x t is the input of the current time step, b h is the bias term of the hidden layer.

[0033] y t is the output of the current time step t, and the softmax function is used to convert the linear output into a probability distribution, W hy is the weight matrix from hidden state to output, b y is the bias term of the output layer.

[0034] The training uses the Adam optimization algorithm, and finally trains and constructs a protein activity prediction model. In this embodiment, the crystal structure amino acid sequence of the natural enzyme is used as input, and the modified amino acid sequence information corresponding to each of the multiple modified enzymes is compared to obtain the modified activity information and the original activity information, and the amino acid sequence with increased activity is screened as a potential modification scheme for optimizing enzyme function. Based on this, the site to be modified can be determined.

[0035] Although the present invention includes many specific implementation details, these should not be interpreted as limiting the scope of any invention or the scope of protection claimed, but are mainly used to describe the features of the specific embodiments of specific inventions. Certain features described in multiple embodiments of the present invention may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may work in certain combinations as described above and even initially claimed as such, one or more features from the claimed combination may be removed from the combination in some cases, and the claimed combination may point to a sub-combination or a variation of a sub-combination.

[0036] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that these operations be performed in the particular order shown or performed sequentially, or requiring that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.

[0037] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In some implementations, multitasking and parallel processing may be advantageous.

[0038] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for enzyme modification based on neural network, characterized in that: The following steps are involved: Acquire a protein sequence database, wherein the protein sequence database includes amino acid sequence information corresponding to each of a plurality of basic protein loop structures; Based on the protein activity prediction model and the original amino acid sequence information of the enzyme to be modified, the original activity information of the enzyme to be modified is obtained; Obtaining an active region of the enzyme to be modified, and obtaining at least one target loop structure in the active region; Based on the amino acid sequence information corresponding to each of the multiple basic protein loop structures in the protein sequence database, the amino acid sequence corresponding to at least one target loop structure in the active region is replaced multiple times to obtain multiple modified enzymes and the modified amino acid sequence information corresponding to each of the multiple modified enzymes; Based on the protein activity prediction model and the modified amino acid sequence information corresponding to each of the multiple modified enzymes, the modified activity information corresponding to each of the modified amino acid sequence information corresponding to each of the multiple modified enzymes is obtained, and based on the modified activity information corresponding to each of the multiple modified enzymes, the final modification plan of the enzyme to be modified is determined.

2. The method for enzyme modification based on neural network according to claim 1, characterized in that: In the protein sequence database, clustering is performed according to the amino acid sequence information corresponding to each of the multiple basic proteins, and the structural representation of the amino acid sequence information corresponding to each of the multiple basic proteins is unified.

3. A neural network-based enzyme modification method according to claim 1 or 2, characterized in that: The protein sequence database also includes the number of occurrences of each of the multiple basic protein loop structures, and distance information of secondary structures connected to each of the multiple basic protein loop structures.

4. The method for enzyme modification based on neural network according to claim 3, characterized in that: According to the distance information of the secondary structures connected to each of the multiple basic protein loop structures, the amino acid sequence corresponding to at least one target loop structure in the active region is replaced multiple times based on the same distance.

5. The method for enzyme modification based on neural network according to claim 3, characterized in that: According to the occurrence frequency information of each of the multiple basic protein loop structures, the amino acid sequence corresponding to at least one target loop structure in the active region is replaced multiple times based on the priority high frequency frequency.

6. The method for enzyme modification based on neural network according to claim 1, characterized in that: The method further comprises: The modified amino acid sequence information corresponding to each of the multiple modified enzymes and the modified activity information corresponding to each of the modified amino acid sequence information corresponding to each of the multiple modified enzymes are input into a protein activity prediction model, and the modified amino acid sequence information corresponding to each of the multiple modified enzymes is compared with the original amino acid sequence information, and the modified activity information corresponding to each of the modified amino acid sequence information corresponding to each of the multiple modified enzymes is compared with the original activity information, and the protein activity prediction model is optimized according to the comparison results to obtain the protein activity prediction model.

7. A neural network-based enzyme modification method according to claim 1 or 6, characterized in that: The protein activity prediction model is a reproducible neural network model. The training of the protein activity prediction model uses the difference between the transformed activity information corresponding to the transformed amino acid sequence information corresponding to each of the multiple transformed enzymes and the original activity information as a loss function.

8. A computer device, characterized in that: include: processor; A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the method for obtaining an enzyme modification scheme based on a ring structure as described in any one of claims 1 to 7 by executing the executable instructions.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for obtaining an enzyme modification scheme based on a ring structure according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Asparaginase mutant and application thereof

    CN106434612A

  • Protein sequence design method, protein structure design method, device and electronic equipment

    CN114155912A

  • Deep learning-based rational design method for enzyme modification

    CN115798581A

  • Programmed design method of topological protein

    CN118155706A