An artificial intelligence-based enzyme engineering method, system, device, and medium

By using an AI-based enzyme engineering method, structurally compatible enzyme sequences are generated using protein reverse folding models and creation strategies. This solves the problems of high cost and low efficiency in enzyme modification, and enables efficient and low-cost enzyme variant creation and expansion of the base editing toolkit.

CN121237206BActive Publication Date: 2026-07-21INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
Filing Date
2025-09-16
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing enzyme modification methods are costly and inefficient, making it difficult to create novel engineered enzyme variants efficiently and at low cost. In particular, there are problems with insufficient editing efficiency and specificity in the base editing process.

Method used

An AI-based enzyme engineering method is employed to obtain the structural information of the enzyme to be engineered, generate multiple structurally compatible protein sequences using a protein reverse folding model, and combine these with a creation strategy to generate single mutant variants and mutant combination variants, thereby achieving efficient and low-cost enzyme modification.

Benefits of technology

It enables the efficient and low-cost generation of novel enzyme variants, expands the base editing toolbox, improves the accuracy and applicability of base editing, reduces training costs, is applicable to the optimization of multiple different families of engineered enzymes, and provides in-depth insights into enzyme sequence and functional generation space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237206B_ABST
    Figure CN121237206B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of enzyme engineering, and discloses an enzyme engineering method, system, device and medium based on artificial intelligence, which comprises the following steps: obtaining a plurality of protein sequences compatible in structure based on a given skeleton structure through a protein unfolding model according to structural information of an engineering enzyme to be modified; obtaining mutation amino acid frequency and wild-type amino acid frequency at each sequence position according to the plurality of protein sequences compatible in structure; and obtaining a single-mutation variant set and a mutation combination variant set of the engineering enzyme to be modified by combining a first creation strategy, a second creation strategy and a third creation strategy according to the mutation amino acid frequency and the wild-type amino acid frequency at each sequence position and the structural information of the engineering enzyme to be modified. The present application can efficiently and at low cost obtain the single-mutation variant set and the mutation combination variant set of the engineering enzyme to be modified, realize precise and efficient gene editing, and effectively expand the application range and safety of gene editing.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese patent application No. 202411307848.9, filed on September 19, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This invention relates to the field of enzyme engineering technology, and in particular to an enzyme engineering method, system, device and medium based on artificial intelligence. Background Technology

[0004] Proteins are essential components of living organisms, participating in physiological and biochemical activities. Protein engineering aims to modify and design proteins to advance the understanding and reconstruction of life processes, and has been driven by artificial intelligence technology in recent years. Proteins possess inherent flexibility, namely the ability to alter their structure and function by changing their amino acid sequence. Protein engineering utilizes this characteristic to achieve rapid protein evolution. Compared to natural evolution, this process has the potential to increase the rate of evolution by orders of magnitude, thereby rapidly generating new protein variants that meet specific requirements.

[0005] Ideal protein engineering strategies aim to achieve optimal engineering performance with minimal effort. However, the complex adaptive landscape of proteins presents significant challenges to protein engineering. Current strategies include structure-guided rational protein design, directed evolution, and the development of AI models for specific protein families, but these strategies often struggle to achieve both low cost and high efficiency. Specifically, structure-guided rational protein design relies on human experience and expertise to tailor protein mutations to achieve desired functional changes, but its success rate is low and it may get bogged down in local optimal fitness. Directed evolution strategies suffer from potential evolutionary bottlenecks, high iteration costs, and difficulty in tailoring mutations to different situations. Protein engineering methods using deep learning models are limited by high training costs and poor generalizability across different proteins.

[0006] Current research utilizes the protein backfolding model ESM-IF1 to backfold the structure into an amino acid sequence, and then defines fitness based on the logarithmic ratio of amino acid frequencies at each position, thereby identifying mutations that may have high fitness. However, this method does not consider the potential influence of the protein structure itself and is only used for antibody evolution, failing to adequately characterize its applicability to enzymes.

[0007] Genome editing technologies, exemplified by the CRISPR-Cas system and its derivatives, have provided unprecedented opportunities for studying complex biological processes and exploring the root causes of genetic diseases. For example, single mutations account for 43.65% of known disease-related variations in the human genome and are closely related to economic traits in crops and livestock. Base editing technologies with deaminases as their main functional components, such as ABE (adenine base editor) and CBE (cytosine base editor), have been widely used, but they suffer from problems such as insufficient editing efficiency and specificity. Specifically, although ABE and CBE can achieve A-to-G and C-to-T transitions, the bystander editing problem, i.e., unintended base conversions, still exists, and the available deaminases are insufficient to meet customized needs.

[0008] Therefore, to address the above problems, there is an urgent need for an enzyme engineering method based on artificial intelligence to create novel engineered enzyme variants efficiently and at low cost, thereby achieving precise and efficient base editing and ultimately developing a diverse base editing toolkit. Summary of the Invention

[0009] This invention provides an enzyme engineering method, system, device, and medium based on artificial intelligence, which addresses the shortcomings of existing enzyme modification methods that are costly and inefficient.

[0010] This invention provides an enzyme engineering method based on artificial intelligence, comprising:

[0011] Obtain the structural information of the engineered enzyme to be modified;

[0012] Based on the structural information of the engineered enzyme to be modified, multiple structurally compatible protein sequences are obtained using a protein reverse folding model, based on the given backbone structure of the engineered enzyme to be modified.

[0013] Based on multiple structurally compatible protein sequences, the frequencies of mutated amino acids and wild-type amino acids at each sequence position were obtained;

[0014] Based on the frequency of mutated amino acids and wild-type amino acids at each sequence position, as well as the structural information of the enzyme to be engineered, a set of single mutant variants of the enzyme to be engineered is obtained by combining the first creation strategy and the second creation strategy.

[0015] Based on multiple structurally compatible protein sequences, and combined with a third creation strategy, a set of mutant variants of the engineered enzyme to be modified is obtained.

[0016] This invention also provides an enzyme engineering system based on artificial intelligence, comprising:

[0017] The data acquisition module is used to: acquire the structural information of the engineered enzyme to be modified;

[0018] The reverse folding module is used to: obtain multiple structurally compatible protein sequences based on the structural information of the engineered enzyme to be modified, using a protein reverse folding model and a given backbone structure of the engineered enzyme to be modified;

[0019] The processing module is used to: obtain the frequency of mutant amino acids and the frequency of wild-type amino acids at each sequence position based on multiple structurally compatible protein sequences;

[0020] The single mutation module is used to: obtain a set of single mutant variants of the engineered enzyme to be modified based on the mutant amino acid frequency and wild-type amino acid frequency at each sequence position and the structural information of the enzyme to be modified, combined with the first creation strategy and the second creation strategy.

[0021] The mutation combinatorial module is used to: obtain a set of mutant combinatorial variants of the engineered enzyme to be modified, based on the structural information of the enzyme to be modified and multiple protein sequences, combined with a third creation strategy.

[0022] The present invention also provides an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the computer program to implement any of the above-described artificial intelligence-based enzyme engineering methods.

[0023] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described artificial intelligence-based enzyme engineering methods.

[0024] The present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute any of the above-described artificial intelligence-based enzyme engineering methods.

[0025] The present invention also provides a deaminase mutant having a mutation at at least one amino acid site compared to the wild-type deaminase, and having at least one of higher base editing efficiency, narrower editing window and lower off-target efficiency compared to the wild-type deaminase.

[0026] The deaminase mutant described above, wherein the wild-type deaminase is the adenine deaminase shown in SEQ ID NO:8, and the deaminase mutant includes amino acids selected from E1M, H10R, A11S, A11V, L12I, G27A, A28C, V29I, L32K, N33D, E39D, E39K, R43Q, H48N, L59I, G62A, Y77F, Y77W, T79N, A87C, A89L, I91L, and S93 in the amino acid sequence shown in SEQ ID NO:8. One of the following: A, R94K, I95V, V98I, F100Y, G101S, G108S, A110L, A110S, G111E, G111A, L141M, V116I, L117F, L117I, N118A, P120E, M122N, M122I, M122L, H124Y, E130S.

[0027] The deaminase mutant described above, wherein the amino acid sequence shown in SEQ ID NO:8 includes one selected from E1M, A11S, G27A, V29I, L32K, N33D, E39D, E39K, L59I, G62A, Y77F, Y77W, T79N, A87C, I91L, S93A, R94K, I95V, V98I, F100Y, G101S, A110L, A110S, G111E, G111A, V116I, L117F, L117I, N118A, P120E, M122N, M122L, and E130S.

[0028] The deaminase mutant described above, wherein the amino acid sequence shown in SEQ ID NO:8 includes one selected from A11S, G27A, E39K, E39D, L59I, Y77F, Y77W, T79N, A87C, R94K, V98I, A110S, A110L, G111A, G111E, L117I, L117F, M122L, and E130S.

[0029] The deaminase mutant described above includes, in the amino acid sequence shown in SEQ ID NO:8, one selected from G27A, E39K, T79N, A87C, R94K, V98I, A110S, A110L, L117I, L117F, and M122L.

[0030] The deaminase mutant described above includes, in the amino acid sequence shown in SEQ ID NO:8, one selected from G27A, E39K, A87C, R94K, V98I, A110S, and A110L.

[0031] The deaminase mutant described above includes A87C in the amino acid sequence shown in SEQ ID NO:8.

[0032] The deaminase mutant described above includes V29I and T79N in the amino acid sequence shown in SEQ ID NO:8.

[0033] The deaminase mutant described above includes L59I and V98I in the amino acid sequence shown in SEQ ID NO:8.

[0034] The deaminase mutant described above includes L59I and G111E in the amino acid sequence shown in SEQ ID NO:8.

[0035] The deaminase mutant described above includes L59I and G108S in the amino acid sequence shown in SEQ ID NO:8.

[0036] The deaminase mutant described above includes R94K, G111E, and L117F in the amino acid sequence shown in SEQ ID NO:8.

[0037] The deaminase mutant described above includes A11S and G27A in the amino acid sequence shown in SEQ ID NO:8.

[0038] The deaminase mutant described above includes V29I and T79S in the amino acid sequence shown in SEQ ID NO:8.

[0039] The deaminase mutant described above includes L59I and G111A in the amino acid sequence shown in SEQ ID NO:8.

[0040] The deaminase mutant described above includes L59I and G108A in the amino acid sequence shown in SEQ ID NO:8.

[0041] The deaminase mutant described above includes R94K, G111A, and L117I in the amino acid sequence shown in SEQ ID NO:8.

[0042] The deaminase mutant described above includes I56A and L59I in the amino acid sequence shown in SEQ ID NO:8.

[0043] The deaminase mutant described above includes A110L and L117F in the amino acid sequence shown in SEQ ID NO:8.

[0044] The deaminase mutant described above includes C83S and C86S in the amino acid sequence shown in SEQ ID NO:8.

[0045] The deaminase mutant described above includes A11V and G27A in the amino acid sequence shown in SEQ ID NO:8.

[0046] The deaminase mutant described above includes A11V, G27A, and E39K in the amino acid sequence shown in SEQ ID NO:8.

[0047] The deaminase mutant described above includes A11V and G27A and E39K and T79N in the amino acid sequence shown in SEQ ID NO:8.

[0048] The deaminase mutant described above includes V29I and L59I and Y77W and T79N in the amino acid sequence shown in SEQ ID NO:8.

[0049] The deaminase mutant described above includes A110S and L117I in the amino acid sequence shown in SEQ ID NO:8.

[0050] The deaminase mutant described above, wherein the wild-type deaminase is a single-stranded DNA cytosine deaminase as shown in SEQ ID NO:7, and the deaminase mutant includes amino acids selected from P1G, P1M, A2P, K4A, K4E, P5C, P5G, S6A, K9D, P10A, T13C, P15K, A16P, K27D, D28Y, R29C, A30C, W35Y, G39E, N40A, N40T, V42C, G44D, S47T, A48P, D49G, D51T, P53C, A55S, T56R, K60I, W63Y, Y66M, A77C, H82A in the amino acid sequence shown in SEQ ID NO:7. One of the following: H82R, D84K, T88V, V90T, M91L, M91V, K99D, K105A, K105C, L106R, K114D, K114E, G115D, S116A, W119V, W119Y, M120L, M120V, R122V, F124K, N126D, G128S, K130R, K130T, Y132I, Q133D, Q133V, T137D, R139W, Y141F, V142A.

[0051] The deaminase mutant described above, wherein the amino acid sequence shown in SEQ ID NO:7 includes amino acids selected from P1G, P1M, A2P, K4A, K4E, P5C, P5G, S6A, K9D, P10A, T13C, P15K, A16P, K27D, D28Y, R29C, A30C, W35Y, N40A, N40T, V42C, G44D, S47T, A48P, D49G, D51T, A55S, T56R, K60I, W63Y, and Y6. One of the following: 6M, H82R, H82A, D84K, T88V, V90T, M91V, M91L, K99D, K105C, K105A, L106R, K114D, K114E, G115D, W119V, W119Y, M120L, F124K, N126D, G128S, K130R, K130T, Y132I, Q133D, Q133V, T137D, V142A.

[0052] The deaminase mutant described above, wherein the amino acid sequence shown in SEQ ID NO:7 includes one selected from P1G, A2P, K4A, K4E, P5G, P5C, S6A, T13C, P15K, A16P, D28Y, A30C, N40A, N40T, V42C, A55S, T56R, K60I, Y66M, H82R, H82A, D84K, T88V, V90T, K99D, L106R, K114D, W119V, M120L, F124K, K130R, Y132I, Q133D, Q133V, and V142A.

[0053] The deaminase mutant described above includes, in the amino acid sequence shown in SEQ ID NO:7, one selected from T13C, D28Y, N40A, N40T, V42C, T56R, Y66M, H82R, T88V, L106R, M120L, Y132I, and V142A.

[0054] The deaminase mutant described above includes V42L and R80W in the amino acid sequence shown in SEQ ID NO:7.

[0055] The deaminase mutant described above includes P15K and D19A in the amino acid sequence shown in SEQ ID NO:7.

[0056] The deaminase mutant described above includes D49G and D50N in the amino acid sequence shown in SEQ ID NO:7.

[0057] The deaminase mutant described above includes M120L and F134Y in the amino acid sequence shown in SEQ ID NO:7.

[0058] The deaminase mutant described above includes F124K and K130T in the amino acid sequence shown in SEQ ID NO:7.

[0059] The deaminase mutant described above includes Y132F and Y141F in the amino acid sequence shown in SEQ ID NO:7.

[0060] The deaminase mutant described above includes W63Y, H76Y, W119Y, Y132F, and Y141F in the amino acid sequence shown in SEQ ID NO:7.

[0061] The deaminase mutant described above, wherein the wild-type deaminase is a double-stranded DNA deaminase whose amino acid sequence is shown in SEQ ID NO:2, and the deaminase mutant includes one of the following amino acids selected from S2K, S2T, T9C, T9V, G24P, G24S, R38A, R38P, S47A, N55G, S57K, S57Q, S57T, V61L, T66P, Y72A, Y72F, T77I, T77L, T77V, E93A, E93P, E93S, A102S, D104A, D104G, V106P, S113A, S113P, and N125P in the amino acid sequence shown in SEQ ID NO:2.

[0062] The deaminase mutant described above, wherein the amino acid sequence shown in SEQ ID NO:2 includes one selected from S2K, T9V, T9C, G24S, G24P, R38A, R38P, S47A, S57K, V61L, T66P, Y72A, T77L, T77I, E93S, E93A, E93P, D104A, D104G, V106P, S113A, and S113P.

[0063] The deaminase mutant described above includes, in the amino acid sequence shown in SEQ ID NO:2, one selected from R38A, R38P, S47A, V61L, T66P, T77L, and E93S.

[0064] The present invention also provides a nuclease mutant, wherein the nuclease mutant has a mutation at at least one amino acid site compared with the wild-type nuclease, and the nuclease mutant has a higher nucleic acid cleavage efficiency compared with the wild-type nuclease.

[0065] The nuclease mutant described above, wherein the wild-type nuclease is the nuclease AcCas12n shown in SEQ ID NO:1, and the AcCas12n mutant includes one of the following amino acid sequences in the amino acid sequence shown in SEQ ID NO:1: F35Y, H37Y, L84F, A86K, W124Y, R136P, G185A, V243T, A261R, R281A, T286Q, R298A, V376L, R438K, N456D, A460R, R483P, and S486P.

[0066] The nuclease mutant described above, wherein the wild-type nuclease is the nuclease LbCas12a shown in SEQ ID NO:3, and the LbCas12a mutant includes one of the following in the amino acid sequence shown in SEQ ID NO:3: S12P, L54I, D156K, K206E, V228L, S333P, K374G, I419L, M456I, N607K, Y653F, E659P, F682W, Y794F, E795I, K979P, M986L, I996V, and A1173S.

[0067] The nuclease mutant described above, wherein the wild-type nuclease is the nuclease SpRYCas9 shown in SEQ ID NO:6, and the SpRYCas9 mutant includes one of the following in the amino acid sequence shown in SEQ ID NO:6: K111R, R139I, M161L, H167N, D182Q, D288P, S675T, S685D, F846H, N869D, L908I, E923V, D947G, T957L, M1021L, Y1036L, G1067P, and H1262N.

[0068] The present invention also provides a nuclear localization sequence mutant, wherein the nuclear localization sequence mutant has a mutation at at least one amino acid site with the wild-type nuclear localization sequence, and the nuclear localization sequence mutant has higher nuclear localization efficiency compared with the wild-type nuclear localization sequence.

[0069] The nuclear localization sequence mutant as described above, wherein the wild-type nuclear localization sequence is the nuclear localization sequence of amino acids shown in SEQ ID NO:5, and the nuclear localization sequence mutant includes one selected from K1I, R2Y, T3D, E8L, E10I, K12D, K13D, and K14P in the amino acid sequence shown in SEQ ID NO:5.

[0070] The present invention also provides a fusion protein comprising any of the deaminase mutant domains and / or nuclease mutant domains described above.

[0071] The fusion protein as described above further includes a nuclear localization sequence domain or any of the aforementioned nuclear localization sequence mutant domains.

[0072] It is understandable that nuclear localization sequences or nuclear localization sequence mutants can be fused at any position in the deaminase mutant domain, or at any position in the nuclease mutant domain, or simultaneously at any position in both the deaminase mutant domain and the nuclease mutant domain.

[0073] The present invention also provides a complex comprising the above-described fusion protein and an sgRNA targeting a specific sequence.

[0074] The present invention also provides a polynucleotide encoding the above-described fusion protein or the above-described complex.

[0075] The present invention also provides a vector comprising the above-mentioned polynucleotides.

[0076] The present invention also provides a gene editing method, comprising contacting the nucleic acid to be edited with the above-mentioned complex.

[0077] The present invention also provides a reverse transcriptase mutant, wherein the reverse transcriptase mutant has a mutation at at least one amino acid site compared with the wild-type reverse transcriptase, and the reverse transcriptase mutant has higher reverse transcription efficiency compared with the wild-type reverse transcriptase.

[0078] The reverse transcriptase mutant described above, wherein the wild-type reverse transcriptase is the reverse transcriptase whose amino acid sequence is shown in SEQ ID NO:4, and the reverse transcriptase mutant includes one of the following in the amino acid sequence shown in SEQ ID NO:4: G22P, T55D, S56A, T197N, D209P, Q213A, Q221L, G239A, M289L, T332V, L333P, W406L, C409P, R411Q, V433T, and Y460I.

[0079] The present invention also provides a fusion protein comprising the above-mentioned reverse transcriptase mutant domain and nickase domain.

[0080] The fusion protein described above has a cleavage enzyme domain that is a Cas protein.

[0081] The fusion protein as described above, wherein the nicking enzyme is selected from at least one of the above-mentioned nuclease mutants.

[0082] The fusion protein as described above further includes a nuclear localization sequence domain or a mutant nuclear localization sequence domain.

[0083] The present invention also provides a complex characterized in that it comprises the above-mentioned fusion protein and sgRNA targeting a specific sequence.

[0084] The present invention also provides a polynucleotide encoding the above-described fusion protein or the above-described complex.

[0085] The present invention also provides a vector comprising the above-mentioned polynucleotides.

[0086] The present invention also provides a gene editing method, comprising contacting the nucleic acid to be edited with the above-mentioned complex.

[0087] This invention provides an artificial intelligence-based enzyme engineering method, system, device, and medium. Based on population evolution theory, and according to the structural information of the enzyme to be engineered, it obtains multiple structurally compatible protein sequences using a protein reverse folding model and a given backbone structure of the enzyme. By combining a first creation strategy, a second creation strategy, and a third creation strategy, it can efficiently and cost-effectively obtain sets of single-mutant variants and sets of combined mutant variants of the enzyme to be engineered. The engineered enzyme variants developed according to this invention can greatly expand the base editing toolbox, achieving a comprehensive improvement in base editing engineering properties, enabling precise and efficient base editing, and effectively expanding the applicability and safety of base editing, making base editing strategies more customized.

[0088] Furthermore, this invention requires no training of proprietary artificial intelligence models or guidance from human experts, generating highly efficient mutations in a single round, significantly improving the efficiency of enzyme engineering. Secondly, this invention possesses strong versatility, applicable to the optimization of multiple different families of engineered enzymes, and is compatible with traditional methods. Thirdly, based on structural information such as inverse folding, evolutionary coupling information, and structure-guided learning, this invention provides numerous counterintuitive yet effective single mutants and mutation combinations that potentially avoid epistatic effects, facilitating a broad exploration of enzyme sequences and functional generation space. This invention not only guides rapid enzyme optimization but also provides deep insights into protein distribution and generation space under natural conditions. Attached Figure Description

[0089] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0090] Figure 1 A diagram illustrating the advantages and disadvantages of AI-assisted, traditional structure-guided rational design, and directed evolution methods in protein engineering.

[0091] Figure 2 This is a flowchart illustrating an artificial intelligence-based enzyme engineering method provided by the present invention.

[0092] Figure 3 The technical details of an enzyme engineering method based on artificial intelligence provided by this invention are shown in the figure.

[0093] Figure 4 For AiCE singlePerformance comparison and benchmarking of strategies in different structural regions; where A represents the fitness difference of single amino acid substitutions between flexible and non-flexible regions (31 DMS libraries); B represents the comparison of fitness differences after screening for heterogeneity and homogeneity in different structural regions (31 DMS libraries); C represents the prediction accuracy and number distribution of high-fitness (HF) mutations at different mutation rates (31 DMS libraries); D represents the enrichment trend of relatively high-fitness (RHF) mutations at different mutation rates (31 DMS libraries); E represents the performance of HF mutation prediction accuracy in 60 DMS libraries; F represents the comparison of HF mutation prediction performance of different filtering strategies; G represents the comparison of HF mutation prediction performance of each method; H represents AiCE. single Prediction accuracy at different quantile thresholds (51 monomeric protein libraries); I represents AiCE. single HF mutation prediction in Cas9 protein; J represents AiCE single Predictive performance of HF mutations in protein-protein and protein-nucleic acid complexes.

[0094] Figure 5 For AiCE single Performance comparison of strategies in different structural regions and supplementary details of benchmark tests; A shows the prediction accuracy of 15th and 25th quantile mutations in 31 DMS libraries at different incidence rates; B is the same as (A), but based on 60 DMS libraries; C shows the logistic regression analysis of different inverse folding models predicting mutations under 60 DMS library data, with the likelihood ratio test used for statistical testing; D shows the comparison of different filtering strategies in HF mutation prediction: ProteinMPNN (left), ESM-IF1 (middle), and LigandMPNN (right).

[0095] Figure 6 For AiCE multi The strategy predicts performance using a high-fitness mutation combination of four antibodies and two additional protein mutation libraries; where AB represents AiCE. multi Performance of CR6261 and CR9114 antibodies in HF multiple mutation prediction: A shows the binding of CR6261 to HA-H1 (left) and HA-H9 (right): correlation between predicted values ​​and measured dissociation constants and mutant fitness ranking; B shows the binding of CR9114 to HA-H1 (left) and HA-H3 (right): same analysis as (A); the rectangles in the scatter plot represent candidate HF multiple mutations; the bands show the mutant fitness quantile distribution; the vertical axis represents the fitness quantile in the mutant library; the blue and red lines represent the cumulative distribution of the original antibody and the mature antibody, respectively; C shows AiCE. multi The performance of His3 and ppluGFP2 in HF multiple mutation prediction is shown. Bubble color represents the number of mutations, and size represents relative fitness; D represents AiCE. multi Schematic diagram.

[0096] Figure 7 For AiCE multi Performance testing and parallel comparison of the strategy under different inverse folding models; where AB represents AiCE. multi Performance of SaProt in HF multiple mutation prediction for CR6261(A) and CR9114(B) antibodies, using ProteinMPNN and ESM-IF1 models; C is a comparison of the mutant fitness distribution predicted by different methods with the global distribution; D is the performance of SaProt in HF multiple mutation prediction for His3 and ppluGFP2.

[0097] Figure 8 To illustrate the phylogenetic evolution of deaminases created using AiCE, the SCP1.201 family in the diagram includes the Ddd1 and Sdd6 proteins, and the dCMP_cyt family includes the TadA8e protein.

[0098] Figure 9 Evaluation of the design and editing of single-stranded DNA deaminases TadA8e and Sdd6 variants; where A represents the optimization of TadA8e mutants, the average relative editing efficiency of mutants generated by different methods at three sites in HEK293T cells, HF mutations are marked in red with their number (n), the total number generated by each method (N) is also shown, and the blue line represents wild-type efficiency; B represents AiCE multi The average relative editing efficiency was determined based on three sites in HEK293T cells, compared with the multiple mutations predicted by the BLOSUM62 matrix. Green indicates mutants shared by both methods, and the rest are labeled the same as in (A). (C)AiCE multi Comparison of the relative editing efficiency of multiple mutations predicted by BLOSUM62 with that of single mutations, with each point representing the efficiency of a single mutation at three sites and three independent repeats; D represents three AiCEs. multi Predicting multiple mutations, two randomized low EC score multiple mutations (Low) multi E represents the relative editing efficiency of wild-type TadA8e at three sites; E represents the average window editing efficiency of 18 TadA8e mutants (13 single mutations and 5 multiple mutations), TadA8e, and ABE9 (a previously reported high-fidelity variant) at six sites in HEK293T cells; F represents the window editing efficiency distribution of Tc1, TadA8e, and ABE9 at 24 target sites in HEK293T cells; G represents the optimized Sdd6 mutant, with the average relative editing efficiency at two endogenous sites, and the rest are labeled the same as (A); H represents AiCE. multi The Sdd6 multiple mutations predicted by BLOSUM62 are labeled as in (B); I represents AiCE. multiComparison with BLOSUM62 for Sdd6 mutation prediction, labeled as in (C); J represents the relative editing efficiency of the selected Sdd6 mutant at six sites in HEK293T cells; K represents the structural analysis of Sdd6 and Sc9; LM represents the specificity detection of Sc9 in HEK293T cells (L) and HeLa cells (M), using an orthogonal R-loop experiment, and comparing off-target effects with wild-type Sdd6, the highly active variant rAPOBEC1, and the high-fidelity variant rAPOBEC1-YE1.

[0099] Figure 10 Design and editing evaluation of TadA8e deaminase window reduction variants; where A is a schematic diagram of the nCas9-based ABE vector; B shows the relative editing efficiency of TadA8e multiple mutants and their constituent mutations at three sites in HEK293T cells, with orange, gray, and green representing AiCE respectively. multi BLOSUM62 and mutants shared by both, column height is average, each point is an independent replicate; C is the average window editing efficiency of the selected TadA8e mutant at six sites in HEK293T cells, with PAM proximal defined as position 1; D is the cryo-electron microscopy structure of the TadA8e and SpCas9 complex (PDB ID: 6VPC), with yellow markings indicating constituent mutations, showing their spatial location and potential interactions in Tc1; E is the distribution of window editing efficiency of Tc1, TadA8e, and ABE9 at 24 target sites of eight genes in each of the three cell lines.

[0100] Figure 11 Multi-cell line editing evaluation of the TadA8e deaminase window reduction variant; AD represents the editing efficiency and window distribution of Tcom1, TadA8e, and ABE9 at 24 targets in HEK293T cells (A), HeLa cells (B), K562 cells (C), and U2OS cells (D), with data representing the average of three independent biological replicates.

[0101] Figure 12This study aims to design and evaluate highly specific variants of the deaminase Sdd6. A shows a schematic diagram of the CBE vector, including an nSpCas9(D10A)-based CBE and an orthogonal R-loop detection system based on nSaCas9(D10A). B shows the relative editing efficiencies of multiple Sdd6 mutants and their constituent mutations at two sites in HEK293T cells. C shows the editing efficiencies of selected Sdd6 mutants, Sdd6, and rAPOBEC1 at six sites in HEK293T cells. D shows the editing efficiencies of HEK2... 93T cell specificity assessment, including eight target sites and two off-target sites for specificity detection; E represents the multi-cell line editing efficiency, the average editing efficiency of Sc9, Sdd6, rAPOBEC1, and rAPOBEC1-YE1 at 24 endogenous sites in HEK293T, HeLa, K562, and U2OS cells; F represents the system specificity assessment in HeLa cells, including 16 target sites and 6 off-target sites; G represents the editing ratio of each editor target / off-target site calculated based on the results of (E).

[0102] Figure 13 This study evaluates the design and editing of double-stranded DNA deaminase Ddd1 variants. A shows a schematic of the DdCBE detection system, employing the fusion of split Ddd1 with TALE for C-to-T editing. B represents the editing efficiency of Ddd1 and DddA at two nuclear and six mitochondrial target sites in HEK293T cells. C illustrates the computational strategy for generating Ddd1 mutants, with the total number of mutants tested by each method marked outside the column and the number of HF mutants marked inside the column; the black line indicates overlap in mutant identification by different methods. D shows the relative editing efficiency of Ddd1 mutants at the ND1.2 and ND6.2 mitochondrial sites in HEK293T cells; the red line represents a relative efficiency of 1.1. E represents AiCE. multi The relative editing efficiencies of the multiple mutants predicted by BLOSUM62 and their corresponding single mutations are shown above. The mutation types are displayed, and single mutations forming combined mutations are connected by black lines. F represents the editing efficiency of enDdd1, D-h11, wild-type Ddd1, PACE-M3, DddA11, and DddA at four mitochondrial sites in HEK293T cells. GH represents the editing efficiency of Ddd1 mutant, Ddd1, DddA, and DddA11 at mitochondria (G) and nuclear targets (H) in HeLa cells and HEK293T cells. D-hs represents AiCE. single -ProteinMPNN obtained a set of 20 mutants with parameter β = 0.8.

[0103] Figure 14Details of the design and editing evaluation of the Ddd1 deaminase variant; A is a schematic diagram of the DdCBE vector construction; B shows the editing efficiency of Ddd1 and DddA at nuclear and mitochondrial targets; C shows the editing efficiency of Ddd1 mutants obtained from different models at two mitochondrial sites in HEK293T cells, with colorless indicating lower efficiency than Ddd1 and DddA, transparent indicating intermediate efficiency, solid indicating higher efficiency than both, and dashed lines representing wild-type Ddd1 efficiency; D shows the relative editing efficiency of Ddd1 mutants at four mitochondrial sites in HEK293T cells; E shows the AiCE... multi Comparison with BLOSUM62 matrix for Ddd1 mutation prediction; F shows the editing efficiency of DC6, DC7, Ddd1, and DddA at two mitochondrial sites in HEK293T and HeLa cells; G shows the structural visualization of enDdd1 and DddA11 (DC7). The two substitution sites of enDdd1 and DC6 are marked in blue, and the six substitution sites of DddA11 are marked in yellow; H shows the comparison of enDdd1 with other mutants in mitochondrial editing; I shows the editing activity heatmap of Ddd1 mutant within the endogenous sequence window, corresponding to (F); J shows the editing efficiency of Ddd1 mutant, Ddd1, DddA, and DddA11 at the nuclear target SIRT6 in HEK293T cells.

[0104] Figure 15 For the design and editing evaluation of other engineered enzyme variants; where A represents AiCE. single A schematic diagram of the optimized detection strategy for mutants in multiple proteins; BD represents the activity of the optimized NLS (B), nuclease (C), and reverse transcriptase M-MLV RT (D) mutants; E represents the relative editing efficiency distribution of the five proteins; F represents the structural correlation analysis of the eight AiCE optimized proteins, and the data shows the RMSD correlation value, reflecting the structural similarity between mutants.

[0105] Figure 16Details of vector construction and editing evaluation for other engineered enzyme variants are provided below. A is a schematic diagram of the NLS activity detection system; in the CBE system, the N-terminus is the mutant BPNLS, and the C-terminus is the wild-type BPNLS. B shows the average editing efficiency of the NLS mutant and wild-type at three sites in HeLa cells. C is a schematic diagram of the nuclease activity detection system; the LbCas12a, AcCas12n, and SpRYCas9 constructs all contain their RNA expression vectors for detecting indel efficiency. DF shows the average indel efficiency of the LbCas12a(D), AcCas12n(E), and SpRYCas9(F) mutants and wild-type at multiple sites in HEK293T cells. G is a schematic diagram of the M-MLV RT-based guided editing system, which includes truncated M-MLV RT mutants and RNA expression vectors for targeted editing. H shows the average editing efficiency of the M-MLV RT mutant and wild-type at three sites in HEK293T cells; all proteins are truncated versions lacking the RNase H domain.

[0106] Figure 17 For AiCE single The distribution of high-fitness mutations obtained by the strategy is shown in Figure A. A represents the distribution of mutations generated by ProteinMPNN, with a ridge plot showing the mutation incidence rate distribution and a dashed line representing the median. A line graph shows the ratio of mutated amino acids to wild-type at each position. Figure B shows the amino acid substitution patterns of TadA8e, Sdd6, and Ddd1, with a bar chart showing the substitution distribution and colored portions representing mutations screened by AiCE. Figure C shows the structural distribution of single mutations, with dark areas representing hotspots identified by traditional rational design, light blue representing experimental detection sites, and dark blue representing high-fitness (HF) mutations. The table summarizes the number of HF mutations (high), other mutations (low), and the total number (total) in hotspot (R) and non-hotspot (N) regions of the seven proteins. The structure of TadA8e (PDB ID: 6VPC) was resolved by cryo-electron microscopy, and the rest were predicted by AlphaFold3.

[0107] Figure 18 This is a schematic diagram of the structure of an enzyme engineering system based on artificial intelligence, provided by the present invention.

[0108] Figure 19 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0109] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, embodiments of this invention, and should not be construed as limiting the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. In the description of this invention, it should be understood that the terminology used is for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0110] The following is combined Figures 1-19 This invention describes an artificial intelligence-based enzyme engineering method, system, device, and medium. It should be noted that the executing entity of the artificial intelligence-based enzyme engineering method provided by this invention can be any network-side device / terminal-side device that meets the technical requirements, such as an artificial intelligence-based enzyme engineering device.

[0111] This invention combines general artificial intelligence, evolutionary theory, and structural information to propose a novel enzyme engineering method. Proteins possess flexibility, a capability that protein engineering leverages to modify their structure and function by altering their amino acid sequences. Compared to natural processes, protein engineering has the potential to accelerate evolution by several orders of magnitude, enabling the rapid generation of novel protein variants adapted to specific requirements. However, the complex fitness scenarios of proteins pose significant challenges to protein engineering. Current strategies, including structure-guided rational protein design, directed evolution, and model customization for specific protein families, struggle to achieve low-cost, high-efficiency enzyme engineering. Figure 1 ).

[0112] like Figure 1 As shown, structure-guided rational design can achieve the customization of mutants, but it is inefficient and requires a strong professional knowledge background. Directed evolution strategies, by artificially simulating the evolutionary process in nature, can achieve rapid creation and iteration of mutations. However, this approach suffers from high costs, limited starting templates, and often has relatively singular selection pressures, making it difficult to customize high-fitness mutations across multiple scenarios. Applying artificial intelligence to enzyme engineering is a rapidly developing new direction. Researchers often customize specific enzyme families' models through the transfer of large models, often achieving good optimization results, but this also brings problems such as excessively high model training costs and difficulty in effective generalization.

[0113] Current widely used large-scale models, including language models and structural models, effectively learn the formation and distribution patterns of proteins in nature. Effectively utilizing generalized protein models can significantly improve the challenges of model transfer and generalization. Therefore, this invention proposes an artificial intelligence-based enzyme engineering method. Figure 2Here is a flowchart illustrating its process. (Refer to...) Figure 2 The present invention provides an enzyme engineering method based on artificial intelligence, which may include:

[0114] Step S110: Obtain the structural information of the enzyme to be modified. Specifically, the structural information of the enzyme to be modified can be obtained from open-source protein structure databases or predicted by tools such as AlphaFold2 and AlphaFold3. The structural information of the enzyme to be modified may include structural type information, active site information, metal ion or cofactor information, domain and functional domain information, three-dimensional structural information (such as the three-dimensional coordinates of each atom in the protein structure of the enzyme to be modified), etc.

[0115] Step S120: Based on the structural information of the engineered enzyme to be modified, multiple structurally compatible reversed amino acid sequences are obtained using a protein reversed model based on the three-dimensional protein structure of the engineered enzyme. Specifically, any existing protein reversed model suitable for engineered enzyme modification, such as ESM-IF1, LigandMPNN, ProteinMPNN, etc., can be selected.

[0116] Step S130: Based on multiple structurally compatible protein sequences, obtain the frequency of mutant amino acids and the frequency of wild-type amino acids at each sequence position.

[0117] In one embodiment, step S130 may include:

[0118] Based on multiple structurally compatible protein sequences, the frequency of each mutated amino acid at each sequence position is obtained through a first expression;

[0119] Based on multiple structurally compatible protein sequences, the frequency of wild-type amino acids at each sequence position is obtained through a second expression.

[0120] The first expression is:

[0121]

[0122] In the first expression, x represents 20 common amino acid types, f x (i) represents amino acid x i The frequency of sequence position i in multiple structurally compatible protein sequences, x i This represents the amino acid at the i-th sequence position of the j-th protein sequence among multiple structurally compatible protein sequences. M represents the total number of structurally compatible protein sequences. 1{·} represents an indicator function that takes the value 1 when the condition is true and 0 otherwise.

[0123] The second expression is:

[0124]

[0125] In the second expression, f wt (i) and f mut (i) represents the frequency of wild-type amino acid wt and mutant amino acid mut at sequence position i in multiple structurally compatible protein sequences, respectively, and x i This represents the amino acid at the i-th sequence position of the j-th protein sequence among multiple structurally compatible protein sequences. M represents the total number of structurally compatible protein sequences. 1{·} represents an indicator function that takes the value 1 when the condition is true and 0 otherwise.

[0126] Step S140: Based on the frequency of mutated amino acids and wild-type amino acids at each sequence position, as well as the structural information of the engineered enzyme to be modified, and combining the first creation strategy and the second creation strategy, a set of single mutant variants of the engineered enzyme to be modified is obtained.

[0127] In one embodiment, step S140 may include:

[0128] Based on the frequency of mutant amino acids and the frequency of wild-type amino acids at each sequence position, the highest frequency mutant amino acid at each sequence position with a frequency higher than that of wild-type amino acids is obtained.

[0129] The frequency of the most frequent mutated amino acid at each sequence position is compared with a first preset threshold to obtain the first set of single mutant variants of the engineered enzyme to be modified.

[0130] The expression for the frequency of the most frequently mutated amino acid is:

[0131] f max (i)=max{f mut (i)∣f mut (i)>f wt (i)},

[0132] In the expression for the frequency of the most frequently mutated amino acid, f max (i) represents the frequency of the most frequently mutated amino acid at sequence position i among multiple structurally compatible protein sequences, f mut (i) represents the frequency of the mutated amino acid mut at sequence position i in multiple structurally compatible protein sequences, f wt(i) represents the frequency of the wild-type amino acid wt at sequence position i in multiple structurally compatible protein sequences. Based on the structural information of the enzyme to be engineered, the flexible region of the enzyme is obtained using protein secondary structure prediction methods. Specifically, protein secondary structure prediction methods, such as the DSSP algorithm, can be used to calculate the energy of each possible hydrogen bond pair based on the structural information of the enzyme to be engineered. If the hydrogen bond energy is lower than a preset threshold (e.g., -0.5 kcal / mol), it is considered that a hydrogen bond exists. If a hydrogen bond exists, it does not belong to the flexible region. The energy of the hydrogen bond pair is obtained by the following formula:

[0133]

[0134] In the formula, q i and q j r represents the charge of the atoms involved in hydrogen bond formation. ON r CH r OH r CN Indicates the distance between corresponding atom pairs (in words) (in units), the constant 332 converts the energy units to Karls per mole;

[0135] Determine whether the most frequently mutated amino acid at each sequence position is located in the flexible region of the enzyme to be modified, and obtain the most frequently mutated amino acid located in the flexible region of the enzyme to be modified.

[0136] The frequency of the most frequently mutated amino acid located in the flexible region of the engineered enzyme to be modified is compared with a second preset threshold to obtain a second set of single mutant variants of the engineered enzyme to be modified.

[0137] By combining the first and second single mutant variant sets of the enzyme to be engineered, a single mutant variant set of the enzyme to be engineered is obtained.

[0138] In one embodiment, step S140 can compare the frequency of the most frequent mutated amino acid at each sequence position with a first preset threshold according to the expression of the first creation strategy to obtain a first set of single mutant variants of the engineered enzyme to be modified, wherein the expression of the first creation strategy is:

[0139]

[0140] In the expression for the first creation strategy, f max (i) represents the frequency of the most frequent mutated amino acid at sequence position i, and β represents the first preset threshold. This represents the first set of single mutant variants of the engineered enzyme to be modified. The first set of single mutant variants includes the highest frequency mutant amino acid with a frequency not lower than a first preset threshold and its sequence position.

[0141] Then, according to the expression of the second creation strategy, the frequency of the most frequently mutated amino acid located in the flexible region of the enzyme to be modified is compared with a second preset threshold to obtain the second set of single mutant variants of the enzyme to be modified. The expression of the second creation strategy is as follows:

[0142]

[0143] In the expression for the second creation strategy, f max (i) represents the frequency of the most frequent mutated amino acid at sequence position i, γ represents the second preset threshold, and is_flexible(i) = 1 indicates that the condition that the most frequent mutated amino acid at sequence position i is located in the flexible region of the enzyme to be modified is true. This represents the second set of single mutant variants of the engineered enzyme to be modified. The second set of single mutant variants includes the sequence position of the highest frequency mutant amino acid located in the flexible region of the engineered enzyme and with a frequency greater than a second preset threshold.

[0144] Then, by combining the first and second single-mutant variant sets of the enzyme to be modified, the single-mutant variant set of the enzyme to be modified is obtained, and its expression is:

[0145]

[0146] In the expression for the set of single mutant variants of the engineered enzyme to be modified, This represents the first set of single mutant variants of the engineered enzyme to be modified. This represents the second set of single mutant variants of the engineered enzyme to be modified. This represents the set of single mutant variants of the engineered enzyme to be modified.

[0147] Step S150: Based on multiple structurally compatible protein sequences of the engineered enzyme to be modified, and combined with the third creation strategy, a set of mutant variants of the engineered enzyme to be modified is obtained.

[0148] In one embodiment, step S150 may include:

[0149] Using codon correspondence, multiple structurally compatible protein sequences are reversed into multiple pseudoDNA sequences. Based on the first expression, the deviation between the observed haplotype frequency and the expected haplotype frequency, which are independently distributed, is obtained and denoted as D.

[0150] Based on the deviation between the observed haplotype frequencies and the expected haplotype frequencies, which are independently distributed, and according to the second expression, the proportion of variance in the engineered enzyme to be modified that can be explained by one allele by the other allele is denoted as r. 2 ;

[0151] By combining the third creation strategy, the third expression can be used to screen out mutant combinations in the engineered enzyme whose variance ratio explained by the corresponding allele of the overall allele is higher than the first preset value, thus obtaining the first set of mutant combinations of the engineered enzyme to be modified.

[0152] The first expression is:

[0153] D = P AB -P A P B ,

[0154] In the first expression, P A P B , and P AB The population frequencies of alleles A and B, and heterozygote AB, are respectively.

[0155] The second expression is:

[0156]

[0157] In the second expression, d represents the deviation between the haplotype frequency and the expected haplotype frequency, which are independently distributed, and P... A P B , and P AB The population frequencies of alleles A and B, and heterozygote AB, are respectively.

[0158] In one embodiment, step S150 may include:

[0159] Based on multiple structurally compatible protein sequences, the log probability of amino acid a at position i in the sequence is obtained using the fourth expression, relative to the overall probability φ of amino acid a in the entire sequence. i And the log probability of amino acid b at position j in the sequence, relative to the overall probability φ of amino acid b in the entire sequence. j ;

[0160] Based on multiple structurally compatible protein sequences, the fifth expression is used to obtain the joint mutation coupling score between pairs of amino acids.

[0161] Based on the joint mutation coupling score between pairs of amino acids, and φ i and φ j And according to the sixth expression, the weighted joint mutation coupling score EC between pairs of amino acids is obtained;

[0162] By combining the third creation strategy, the third expression can be used to select mutant combination variants with statistical coupling scores greater than or equal to the second preset value, thus obtaining the second mutant combination variant set of the engineered enzyme to be modified.

[0163] The fourth expression is:

[0164]

[0165] In the fourth expression, f ai and f bi These represent the frequencies of amino acids a and b at positions i and j in multiple sequence alignment, respectively, and q... a and q b These are the background frequencies of amino acids a and b.

[0166] The fifth expression is:

[0167]

[0168] In the fifth expression, and These are the frequencies of amino acids a and b at positions i and j, respectively. This represents the probability that the i-th position is amino acid a and the j-th position is amino acid b.

[0169] The sixth expression is:

[0170]

[0171] In the sixth expression, The score represents the joint mutational coupling score between amino acids a and b at positions i and j.

[0172] The third expression is:

[0173]

[0174] In the third expression, where Represents a set of mutations The total number of sites is calculated, and score(i,j) represents the score between position i and position j, including the variance proportion score and the joint mutation coupling score. This calculation provides a measure of the overall evolutionary coupling strength of the mutation combination.

[0175] Preferably, the expression for the first set of mutant variants of the engineered enzyme to be modified is:

[0176]

[0177] In the expression for the first set of mutant combinations of the engineered enzyme to be modified, This represents the first set of mutant combinations of the engineered enzyme to be modified. Indicates the combination of mutations in multiple pseudoDNA sequences The variance proportion score;

[0178] Preferably, the expression for the second mutant variant combination set of the engineered enzyme to be modified is:

[0179]

[0180] In the expression for the second set of mutant variants of the engineered enzyme to be modified, This represents the set of second mutant variants of the engineered enzyme to be modified. Table of mutation combinations Statistical coupling score percentile 90 This refers to the 90th percentile of the global coupling score of the amino acid sequence.

[0181] In one implementation, the final set of mutant variants of the engineered enzyme to be modified is expressed as:

[0182]

[0183] In the expression for the final set of mutant variants of the engineered enzyme to be modified, This represents the final set of mutant variants of the engineered enzyme to be modified. This represents the first set of mutant combinations of the engineered enzyme to be modified. This represents the second set of mutant variants of the engineered enzyme to be modified.

[0184] The set of single mutant variants and the set of mutant combination variants of the engineered enzyme to be modified obtained in this embodiment are shown in Table 1, and the wild-type sequence is shown in Table 2.

[0185] Table 1 Mutation information of the engineered enzyme to be modified

[0186]

[0187]

[0188]

[0189]

[0190]

[0191] Table 2 Protein sequences of the engineered enzymes to be modified

[0192]

[0193]

[0194] The AI-based enzyme engineering method provided by this invention can obtain high-fitness mutations and their combinations from most protein backfolding models without human expert guidance. The results of modifying three commonly used engineered enzymes demonstrate that this invention is simple, efficient, and versatile. It can design single and combined mutants with optimized engineering properties or altered deamination substrates, successfully developing a series of novel engineered enzymes. This invention can serve as a state-of-the-art (SOTA) method for engineered enzyme optimization and has the potential to be a universal method for enzyme modification. It can also provide profound insights into the spatial distribution and generation patterns of protein functional domains.

[0195] This invention, based on experimentally determined or computationally predicted protein structures, combines a reverse-folding deep learning model with evolutionary theory to create effective single mutations and combinations. This invention eliminates the need to train proprietary AI models, avoiding the costs associated with human expert guidance and model training, and offers advantages such as simplicity, efficiency, and versatility. Using this invention, several engineered enzyme variants with enhanced engineering properties have been successfully developed, covering double-stranded DNA deaminase Ddd1, single-stranded DNA cytosine deaminase Sdd6, and adenine deaminase TadA8e, nucleases SpRYCas9 and AcCas12n, and...

[0196] LbCas12a, the nuclear localization sequence NLS, and the reverse transcriptase M-MLV RT, have been developed into a series of products with more precise base editing performance. Importantly, this invention can effectively drive mutation generation using both lightweight neural network models and large protein language models, further enabling the design of interpretable combinations of substrate-altered mutants and high-fitness mutations. This invention demonstrates significant success in the engineering of enzymes and even proteins based on artificial intelligence combined with evolutionary theory, providing potential directions for the iteration of protein engineering and the understanding of protein spatial distribution patterns.

[0197] The verification process of the present invention will be described in detail below. The enzyme engineering method based on artificial intelligence provided by the present invention will be referred to as AiCE.

[0198] Current conventional approaches to enzyme engineering mainly consist of two methods: structure-assisted rational design and directed protein evolution. Figure 1 The former, guided by human expert knowledge, allows for customized mutation design, but is often inefficient, time-consuming, and requires highly specialized knowledge. The latter is an effective method for mutation creation, simulating the natural selection process to iteratively create high-fitness mutants. However, this method also has inherent drawbacks, mainly including the need for multiple iterations, high cost, and the requirement for a starting template. In recent years, with the widespread application of artificial intelligence technology in the biological field, transfer learning using pre-trained models has become an effective new method for enzyme engineering. However, this approach also faces challenges such as high training costs and weak generalization.

[0199] This embodiment utilizes a reported universal protein backfolding model to propose a novel, lightweight strategy for achieving efficient enzyme engineering, named AiCE( Figure 2 , Figure 3 This embodiment first proposes AiCE. single This is a strategy for creating single-mutant mutants. The strategy is based on a prior assumption: artificial intelligence learns the distribution and generation patterns of proteins in nature, and can screen and fix single mutations with high fitness by artificially defining selection intensity to simulate natural selection. Based on population genetics, this embodiment posits that the fitness of a genetic locus maps to its distribution frequency in a population; high-frequency mutation sites suggest widespread retention in nature, and these mutations often possess high fitness and functional stability. Using publicly available DMS (Deep Mutation Scan) data from 60 sites, this embodiment demonstrates…

[0200] AiCE can restore beneficial mutations in tested proteins, including but not limited to kinases, sequence-specific nucleases, signaling proteins, receptors, and viruses. Figure 4 , Figure 5 This example further demonstrates that mutations with high fitness (i.e., those in the top 5% of the set) from both the AiCE-predicted mutation set and the DMS database exhibit enrichment in structurally flexible regions, proving the tolerance of structurally flexible regions to high-fit mutations. Therefore, this example developed AiCE. single This module incorporates additional screening for structurally flexible regions, and its feasibility and effectiveness in creating single mutations have been demonstrated in 60 DMS libraries. Figure 4 , Figure 5 (and benchmarks with other models, AiCE) single It can achieve a high-fitness mutation prediction accuracy of up to 16%, which is nearly twice the performance of advanced large AI models such as ESM3-open. Figure 4 Logistic regression analysis indicates that the structure is restricted in AiCE. single Performance is absolutely essential. Figure 5 The effectiveness of this method is evident not only in the design of monomeric structures, but also in the design of protein complexes (such as AsCas12f) and large proteins.

[0201] This embodiment further iteratively develops AiCE. multiThe method employs two strategies to calculate amino acid residues in evolutionary coupling: In the first strategy, this embodiment uses the optimal codon to reverse transcribe the amino acid sequence into a pseudonucleotide sequence, and obtains the corresponding mutation combinations by calculating the linkage disequilibrium between nucleotides; the second strategy calculates the co-evolution score between amino acids through statistical coupling analysis, thereby obtaining amino acids that may co-evolve, and mutation combinations with a high degree of evolutionary coupling may have higher fitness. This embodiment demonstrates that compared to random combinations, AiCE... multi It has a very high probability of creating high-fitness combinatorial mutants. Figure 6 Regardless of the inverse folding model used, based on AiCE multi All created combinatorial mutants showed a significant increase in fitness. This example uses AiCE... multi This can be further extended to the modification of other proteins such as His3 and ppluGFP2. Figure 6 The results all demonstrated its effectiveness in overcoming epistatic effects between mutations. Furthermore, the screening approach based on evolutionary coupling also provides direction for the interpretable creation of combinatorial mutations, namely, utilizing AiCE. multi Select mutation combinations with higher evolutionary coupling. Comparison with the baseline model SaProt shows that AiCE... multi Its predictive ability for combined mutations is comparable, but its computational cost is less than 1% of the former.

[0202] Enzyme-based base editors have the potential to effectively correct pathogenic mutations, which account for approximately 43% of human disease-related variants (data from dbVar: https: / / www.ncbi.nlm.nih.gov / dbvar / , 08 / 2024), or introduce beneficial mutations. However, many engineered enzymes can be enhanced to improve deficiencies in deamination activity, unpredictable off-target effects, and the generation of unintended base transitions (also known as bystander editing). This embodiment focuses on improving these deficiencies through protein engineering, using engineered enzymes Ddd1 and Sdd6 from the SCP1.201 family and TadA8e from the dCMP_cyt family as chassis enzymes to conduct a conceptual exploration of precise and efficient enzyme modification. Figure 8 ).

[0203] In this embodiment, AiCE is utilized. single And other strategies were used to design 131 TadA8e single mutants (T1-T131) and 114 Sdd6 single mutants (S1-S114) respectively. Figure 9 , Figure 10 , Figure 12 Using AiCE multiThe same strategy was used to design 24 TadA8e combinatorial mutants and 21 Sdd6 combinatorial variants (Tc1-Tc24, Sc1-Sc21). Figure 9 , Figure 10 , Figure 12 Preliminary validation of three endogenous targets in HEK293T cells revealed 19 TadA8e single mutants in this embodiment (11 of which were derived from AiCE). single (Created) and 10 combined mutants (6 of which were created by AiCE) multi The created type can improve deammoniation efficiency by more than 10% compared to the wild type. Figure 9 , Figure 10 Randomly combined mutants showed no improvement in deamination efficiency. Preliminary validation of two endogenous targets in HEK293T cells revealed 48 Sdd6 single mutants in this embodiment (21 of which were generated by AiCE). single (created) and 6 combined mutants (all 6 were created by AiCE) multi The created type can improve deammoniation efficiency by more than 10% compared to the wild type. Figure 9 , Figure 12 ).

[0204] Table 3. Base editing efficiency of TadA8e single mutants and combined mutants

[0205]

[0206]

[0207]

[0208]

[0209]

[0210] Table 4. Verification of base editing efficiency of single-stranded DNA cytosine deaminase (Sdd6) mutants

[0211]

[0212]

[0213]

[0214]

[0215] This embodiment uses the highly efficient TadA8e mutant as a base and randomly selects 18 highly efficient mutants (13 single mutants and 5 multiple mutants). Their activity is evaluated within the editing windows of 6 additional target sites. Two variants, T1 (E1M) and Tc1 (A11V / G27A), with improved efficiency and precision, were identified. T1 showed an approximately 70% increase in editing efficiency within the target window. Figure 9 , Figure 10 This embodiment further utilizes Tc1 to construct enABE8e, and comprehensively evaluates it, along with ABE9, one of the most precise adenine base editors, at 24 target sites across 8 genes in HEK293T, HeLa, K562, and U2OS cells. Figure 9 , Figure 10 The enABE8e variant has an editing window that is nearly half the width of ABE8e, and maintains comparable or higher editing efficiency on more than half of the test sites; compared to the ABE9 window, which is about 1 bp wider, its editing efficiency is significantly better than ABE9. This embodiment considers enABE8e to have powerful editing efficiency and a significantly narrower editing window, making it one of the most accurate adenine base editors currently available.

[0216] This embodiment also evaluated the creation of Sdd6 variants, randomly selecting 13 high-efficiency mutants (9 single mutants and 4 multiple mutants) for further evaluation at 6 additional target sites in HEK293T cells. Eleven of these mutants showed at least a 10% increase in deamination activity at one or more target sites. Figure 9 , Figure 11 Sc9 (F124K / K130T) is the most stable mutant, with editing efficiency improved by approximately 12% to 35%. Based on the Sdd6 mutant-ssDNA complex structure predicted by AlphaFold3, this embodiment suggests that Sc9 possesses the functionality of a high-fidelity variant, potentially reducing the flexibility of positively charged surface regions and thus possibly minimizing off-target effects. Figure 9 To effectively evaluate the editing specificity of the Sdd6 mutant, this embodiment uses a previous test method based on orthogonal R-loop. The test shows that compared with the wild type, Sc9 exhibits a higher target / off-target ratio (hereinafter referred to as "fidelity") at 8 target sites and 2 off-target sites. Specifically, the specificity is improved by about 30%, and the fidelity is also improved by about 20%.

[0217] This embodiment further characterized its deamination activity in different cell types. Compared with Sdd6, it improved the editing efficiency by 19%, 84%, and 31% in HEK293T, HeLa, and U2OS cells, respectively. Compared with rAPOBEC1, its editing efficiency improved by 44%, 195%, and 127%, respectively. Figure 12Regarding editing specificity, analysis of 16 target / off-target pairs at 6 off-target sites in HeLa cells showed that enCBE (Sc9-CBE) significantly improved specificity by approximately 28% and fidelity by approximately 132%. Figure 12 Compared with the high-fidelity variant rAPOBEC1-YE1, the specificity was improved by 51%. Figure 12 The two mutations in Sc9 are based on AiCE. single The mutation discovered through screening lacked prior literature support and was located outside of well-characterized catalytic or binding regions, making it difficult to predict using traditional methods. This highlights the powerful performance of the AiCE method, which provides a new foundation for rational protein design.

[0218] Mitochondrial DNA mutations can lead to a variety of genetic diseases. DdCBEs, developed by fusing transcription activator-like protein with Ddds (double-stranded DNA engineered enzymes), can achieve C-to-T alterations in mitochondrial DNA to correct pathogenic mutations, such as Leber hereditary optic neuropathy and maternally inherited deafness. Compared to the commonly used double-stranded engineered enzyme DddA, its homolog Ddd1 can edit the 5'-GC sequence, which is difficult for DddA to edit, thus expanding the potential applications of DdCBEs. However, although Ddd1 exhibits efficient deamination activity in the nuclear environment, its editing efficiency in the mitochondrial environment is significantly reduced. Figure 13 Therefore, this embodiment first uses the DddA homolog complex (PDB ID 8E5E) predicted by AlphaFold2 and resolved by cryo-electron microscopy as the target framework, and uses various prediction methods to nominate 138 single mutants and 15 multiple mutants. Figure 13 These mutants were split at the N94 site and constructed into paired TALE vectors. Their deamination efficiency was assessed at two endogenous mitochondrial target sites in HEK293T cells. Target site deep sequencing revealed that 20 single mutants (12 derived from AiCE) exhibited deamination efficiency at two endogenous mitochondrial target sites. single (Created) and 8 multi-mutants (7 of which were created by AiCE) multi The predicted mutation was identified as a high-fitness mutation. Figure 13 , Figure 14 Using AiCE singleThe accuracy of predicting high-fitness mutants was 27-32%, while the accuracy of other methods ranged from 0-23%. The top four variants with the highest prediction accuracy were tested at two additional non-5'-GC targets, with 21-32% showing higher deamination activity than DddA, and the best variant, D71(V61L), showing a 0.9-6.5-fold increase in deamination activity. These high-fitness mutants retained strong deamination activity in different cell lines. In this example, seven single mutants and eight multiple mutants were randomly selected from the first round of screening, and their deamination activity against the ND1.2 target site in HeLa cells was evaluated. Figure 13 Among these mutants, the deamination activity was significantly enhanced in 4 single mutants and 7 multiple mutants.

[0219] Table 5. Verification of base editing efficiency of double-stranded DNA deaminase (Ddd1) mutants

[0220]

[0221]

[0222]

[0223]

[0224]

[0225] This embodiment further demonstrates its compatibility with other methods. This embodiment combines the most efficient point mutation D-h11 with a previously reported PACE-M3 mutation to create the most efficient mitochondrial editor currently available, enDdd1, achieving a 14-fold increase in base editing efficiency under mitochondrial conditions. Figure 13 , Figure 14 enDdd1 also effectively edited the 5'-GC sequence, achieving an efficiency improvement of approximately 40% compared to the six-mutant DddA11 obtained through several rounds of phage-assisted continuous evolution. This fully demonstrates the effectiveness and compatibility of this embodiment.

[0226] This example also evaluates AiCE's ability to generate mutations in Ddd1 adapted to the nuclear environment. This embodiment will...

[0227] AiCE single Twenty mutants generated by ProteinMPNN were applied to base editing in a nuclear environment, and seven new mutations were found to enhance editing capabilities in a nuclear environment, improving efficiency by approximately 1.6 times. Figure 13 This is something that cannot be solved by current solutions relying on human experts.

[0228] This embodiment applies AiCE to other complex tasks to further enhance genome editing capabilities. This embodiment selects one nuclear localization sequence and four proteins as engineering targets. Figure 15 This includes: nuclear localization sequences (NLS) that may improve the nuclear localization and efficiency of the editing enzyme; type V CRISPR nucleases LbCas12a and AcCas12n; type II CRISPR nuclease SpRYCas9; and Moloney murine leukemia virus (M-MLV) reverse transcriptase (RT), which is crucial for guiding the editing process. This example uses AiCE. single ProteinMPNN designed mutations for each target and evaluated their genome editing potential in HEK293T or HeLa cells. Figure 15 , Figure 16 The results showed that the average prediction accuracies were 21%, 28%, 60%, 17%, and 60%, respectively, exceeding those of traditional protein engineering methods. Figure 15 These five proteins vary greatly in size (from tens to thousands of residues) and exhibit considerable structural heterogeneity. Figure 15 However, AiCE still demonstrates excellent performance in these complex protein modification tasks.

[0229] Obtaining high-functionality mutants at low cost has always been a goal of protein engineering. This difficulty stems from the multidimensional nature of protein structure and the complex relationship between sequence and function. Proteins essential to biological systems, such as enzymes, receptors, and channel proteins, typically possess dynamic structures that balance stability and flexibility to ensure their continued functionality. Reflecting these conformational dynamics through a single static structure is very difficult, if not impossible, and limits the effectiveness of traditional protein engineering methods. This study demonstrates that AiCE, as a simple protein mutation design method, can identify high-fitness mutants by utilizing the distribution of backfolded sequences. This method requires no additional model transfer or training costs, is not limited by model size or architecture, and outperforms traditional protein engineering methods. The AiCE method is based on the assumption that generalized protein backfolding models inherently contain the natural dynamics of protein sequence and structure generation and distribution. These fundamental principles can be extracted and applied using contemporary genetic theory. Although AiCE uses a generalized protein model rather than a model specific to a particular protein family, the sequences and distribution of generated high-fitness mutants still vary depending on the target protein backbone. Figure 17 Furthermore, this example demonstrates that high-fitness and high-frequency mutations identified by AiCE often manifest themselves in unconventional and counterintuitive ways in terms of mutation location and type. Figure 17 This may explain AiCE's compatibility with other methods and its potential as a new option for structure-guided rational protein design.

[0230] The method for verifying the editing efficiency of the above-mentioned engineered enzyme mutants includes the following steps:

[0231] I. Carrier Construction:

[0232] A TALE-mediated base editor plasmid, comprising Ddd1, a TALE array, a UGI (uracil glycosylase inhibitor), and a mitochondrial targeting signal, was optimized for codon usage to promote expression in human cells. These were commercially synthesized by GenScript. The above components were then cloned into the vectors pCMV-TALE-L-JAK2-Ddd9-N (Addgene#204853) and pCMV-TALE-R-JAK2-Ddd9-C (Addgene#204854). A single mutation in Ddd1 predicted by the AiCE method was introduced into a PCR fragment using primer design and cloned into the corresponding backbone vector containing the TALE array using the Uniclone One Step Seamless Cloning Kit (Genesand).

[0233] Single or combined amino acid mutations in Sdd6 predicted by AiCE were introduced into PCR fragments using primer design and cloned into the p2T-CMV-miniSdd6-BE4max-BlastR (Addgene#204850) vector backbone using the Uniclone One Step Seamless Cloning Kit (Genesand). The 2×UGI sequence was removed from the pnCas9-miniSdd6-PBE vector to construct the p2T-CMV-miniSdd6-BE4max-BlastR-delUGI vector backbone. The TadA8e sequence was cloned from pTPH413 (Addgene#185728). Then, the AiCE variant with the wild-type sequence was constructed into the p2T-CMV-miniSdd6-BE4max-BlastR-delUGI vector backbone using the methods described above.

[0234] The phU6 vector (Addgene#53188) was used to express sgRNA. The sgRNA vector was constructed using circular PCR, in which a spacer sequence was integrated into the primers. The F and R primers had a 20 bp overlap at the 5' end. The complete vector sequence was amplified using PCR, and the new spacer sequence was integrated. After amplification, the template plasmid was digested with DpnI restriction enzyme, which exhibits specificity for DNA sequences with methylation modifications. The digested PCR product was then transformed into Fast-T1 competent cells (Vazyme). PCR amplification was performed using 2×PhantaMax Master Mix (Vazyme).

[0235] II. Human cell transfection and DNA extraction:

[0236] Transfection was performed 16–24 hours post-inoculation. In transfection experiments using the TALE-mediated base editing system, 0.4 μl of Lipofectamine 2000 (ThermoFisher Scientific), 150 ng of TALE-L vector, 150 ng of TALE-R vector, and 10 ng of green fluorescent protein were co-incubated and transfected into cells.

[0237] In a transfection experiment involving a CRISPR-Cas-mediated base editing system, 0.4 μl of Lipofectamine 2000, 300 ng of a vector containing engineered enzymes, 100 ng of sgRNA expression vector, and 10 ng of green fluorescent protein were co-incubated and transfected into cells.

[0238] To examine off-target effects using R-loop analysis, four vectors were co-incubated with 0.4 μl of Lipofectamine 2000 and transfected. A total of 150 ng of pCMV-deam-BE4max vector, 150 ng of pCMV-nSaCas9 vector, and two corresponding sgRNA vectors (50 ng each), along with 10 ng of green fluorescent protein, were used. After a 72-hour incubation period, HEK293T cells were washed with PBS, and genomic DNA was extracted using the Triumfi Mouse Tissue Direct PCR Kit (Genesand) with lysis buffer and proteinase K.

[0239] III. Amplicon sequencing analysis

[0240] A series of primers with 5' barcodes were designed to amplify the target sequence. Amplicons were purified using the Thermo Scientific GeneJET Kit (Thermo Fisher Scientific) and quantified using a NanoDrop 2000 spectrophotometer (Thermo Fisher Scientific). Equal volumes of PCR products were pooled and then commercially sequenced using NGS (GENEWIZ). For target sites that are difficult to amplify, a single round of primers was designed for 500 bp amplification. The resulting product was diluted 20-fold, and then nested PCR was performed using 1 μl of the diluted product as a template with primers containing the barcode sequence.

[0241] In amplicon sequencing analysis, the sequencing data was first cleaned and split according to the sequencing primers. The following methods were used: (1) Data splitting and merging. The collected sequencing data was split into separate processes according to the sample-specific adapter sequences. The forward and reverse reads of each process were merged into a single read using the FLASH tool (v1.2.11) with the parameter set to "-m 5-M150" to ensure that the minimum overlap of the paired reads was 5 bp and the maximum overlap was 150 bp. (2) Sequence alignment and SNP extraction. The merged reads were compared with the wild-type reference sequence. The alignment strategy included determining the sequence window for viewing edit events by obtaining 7 base pairs before and after the reads, using the reference sequence as the index sequence, aligning the index sequence with the merged reads, and if there were index sequences at both ends in the merged reads, the middle sequence was extracted and then aligned with the window sequence to identify and extract SNP information. (3) Edit event counting. Based on the extracted SNP information, the editing efficiency, type, and other related indicators of all base editing events were calculated.

[0242] The artificial intelligence-based enzyme engineering system provided by the present invention will be described below. The artificial intelligence-based enzyme engineering system described below can be referred to in correspondence with the artificial intelligence-based enzyme engineering method described above.

[0243] Reference Figure 18 The present invention provides an enzyme engineering system based on artificial intelligence, which may include:

[0244] The data acquisition module is used to: acquire the structural information of the engineered enzyme to be modified;

[0245] The reverse folding module is used to: obtain multiple structurally compatible protein sequences based on the structural information of the engineered enzyme to be modified, using a protein reverse folding model and a given backbone structure of the engineered enzyme to be modified;

[0246] The processing module is used to: obtain the frequency of mutant amino acids and the frequency of wild-type amino acids at each sequence position based on multiple structurally compatible protein sequences;

[0247] The single mutation module is used to: obtain a set of single mutant variants of the engineered enzyme to be modified based on the mutant amino acid frequency and wild-type amino acid frequency at each sequence position and the structural information of the enzyme to be modified, combined with the first creation strategy and the second creation strategy.

[0248] The mutation combinatorial module is used to: obtain a set of mutant combinatorial variants of the engineered enzyme to be modified based on multiple structurally compatible protein sequences and in conjunction with a third creation strategy.

[0249] Figure 19 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 19 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute an artificial intelligence-based enzyme engineering method, which includes:

[0250] Obtain the structural information of the engineered enzyme to be modified;

[0251] Based on the structural information of the engineered enzyme to be modified, multiple structurally compatible protein sequences are obtained using a protein reverse folding model, based on the given backbone structure of the engineered enzyme to be modified.

[0252] Based on multiple structurally compatible protein sequences, the frequencies of mutated amino acids and wild-type amino acids at each sequence position were obtained;

[0253] Based on the frequency of mutated amino acids and wild-type amino acids at each sequence position, as well as the structural information of the enzyme to be engineered, a set of single mutant variants of the enzyme to be engineered is obtained by combining the first creation strategy and the second creation strategy.

[0254] Based on multiple structurally compatible protein sequences, and combined with a third creation strategy, a set of mutant variants of the engineered enzyme to be modified is obtained.

[0255] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0256] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the artificial intelligence-based enzyme engineering method provided by the above methods, the method comprising:

[0257] Obtain the structural information of the engineered enzyme to be modified;

[0258] Based on the structural information of the engineered enzyme to be modified, multiple structurally compatible protein sequences are obtained using a protein reverse folding model, based on the given backbone structure of the engineered enzyme to be modified.

[0259] Based on multiple structurally compatible protein sequences, the frequencies of mutated amino acids and wild-type amino acids at each sequence position were obtained;

[0260] Based on the frequency of mutated amino acids and wild-type amino acids at each sequence position, as well as the structural information of the enzyme to be engineered, a set of single mutant variants of the enzyme to be engineered is obtained by combining the first creation strategy and the second creation strategy.

[0261] Based on multiple structurally compatible protein sequences, and combined with a third creation strategy, a set of mutant variants of the engineered enzyme to be modified is obtained.

[0262] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the artificial intelligence-based enzyme engineering method provided by the methods described above, the method comprising:

[0263] Obtain the structural information of the engineered enzyme to be modified;

[0264] Based on the structural information of the engineered enzyme to be modified, multiple structurally compatible protein sequences are obtained using a protein reverse folding model, based on the given backbone structure of the engineered enzyme to be modified.

[0265] Based on multiple structurally compatible protein sequences, the frequencies of mutated amino acids and wild-type amino acids at each sequence position were obtained;

[0266] Based on the frequency of mutated amino acids and wild-type amino acids at each sequence position, as well as the structural information of the enzyme to be engineered, a set of single mutant variants of the enzyme to be engineered is obtained by combining the first creation strategy and the second creation strategy.

[0267] Based on multiple structurally compatible protein sequences, and combined with a third creation strategy, a set of mutant variants of the engineered enzyme to be modified is obtained.

[0268] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0269] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0270] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An enzyme engineering method based on artificial intelligence, characterized in that, include: Obtain the structural information of the engineered enzyme to be modified; Based on the structural information of the engineered enzyme to be modified, multiple structurally compatible protein sequences are obtained using a protein reverse folding model, based on the given backbone structure of the engineered enzyme to be modified. Based on multiple structurally compatible protein sequences, the frequencies of mutated amino acids and wild-type amino acids at each sequence position were obtained; Based on the frequency of mutated amino acids and wild-type amino acids at each sequence position, as well as the structural information of the enzyme to be engineered, a set of single mutant variants of the enzyme to be engineered is obtained by combining the first creation strategy and the second creation strategy. Based on multiple protein sequences of the enzyme to be engineered, and combined with the third creation strategy, a set of mutant combinations of the enzyme to be engineered is obtained. The process of obtaining a set of single mutant variants of the engineered enzyme based on the mutant amino acid frequency and wild-type amino acid frequency at each sequence position, along with the structural information of the enzyme to be modified, and combining the first and second creation strategies, includes: Based on the frequency of mutant amino acids and the frequency of wild-type amino acids at each sequence position, the highest frequency mutant amino acid at each sequence position with a frequency higher than that of wild-type amino acids is obtained. The frequency of the most frequent mutated amino acid at each sequence position is compared with a first preset threshold to obtain the first set of single mutant variants of the engineered enzyme to be modified. Based on the structural information of the enzyme to be modified, the flexible region of the enzyme to be modified is obtained using the protein secondary structure prediction method. Determine whether the most frequently mutated amino acid at each sequence position is located in the flexible region of the enzyme to be modified, and obtain the most frequently mutated amino acid located in the flexible region of the enzyme to be modified. The frequency of the most frequently mutated amino acid located in the flexible region of the engineered enzyme to be modified is compared with a second preset threshold to obtain a second set of single mutant variants of the engineered enzyme to be modified. By combining the first and second single mutant variant sets of the enzyme to be modified, a single mutant variant set of the enzyme to be modified is obtained. The process of obtaining a set of mutant variants of the engineered enzyme based on multiple protein sequences of the enzyme to be modified, combined with a third creation strategy, includes: Using codon correspondences, multiple structurally compatible protein sequences are reversed into multiple pseudoDNA sequences. Based on the first expression, the deviation between the observed haplotype frequency and the expected haplotype frequency, which are independently distributed, is obtained and denoted as [missing information]. ; Based on the deviation between the observed haplotype frequencies and the expected haplotype frequencies, which are independently distributed, and according to the second expression, the proportion of variance in the engineered enzyme to be modified that can be explained by one allele by the other allele is denoted as . ; Combining the third creation strategy, based on the third expression, the mutant combination variants in the engineered enzyme to be modified whose variance ratio explained by the corresponding allele of the overall allele is higher than the first preset value are selected, and the first mutant combination variant set of the engineered enzyme to be modified is obtained. Combining the third creation strategy, based on the third expression, mutant combination variants with statistical coupling scores greater than or equal to the second preset value are selected to obtain the second mutant combination variant set of the engineered enzyme to be modified; The third expression is: , In the third expression, Represents a set of mutations The total number of sites in the middle, Indicates position i and location j The scores include variance proportion score and joint mutation coupling score.

2. The enzyme engineering method based on artificial intelligence according to claim 1, characterized in that, The process of obtaining the frequency of mutated amino acids and the frequency of wild-type amino acids at each sequence position based on multiple structurally compatible protein sequences includes: Based on multiple structurally compatible protein sequences, the frequency of each amino acid at each sequence position is obtained through a first expression; Based on multiple structurally compatible protein sequences, the wild-type amino acid frequency and mutant amino acid frequency at each sequence position are obtained through a second expression; The first expression is: , In the first expression, This indicates the preset common amino acid types. Indicates amino acids Sequence position in multiple structurally compatible protein sequences The frequency on, This indicates the first of several structurally compatible protein sequences. The first protein sequence Amino acids at each sequence position, The total number of sequences representing structurally compatible protein sequences. This indicates an indicator function that takes the value 1 when the condition is true and 0 otherwise. The second expression is: , , In the second expression, and These represent wild-type amino acids. and mutant amino acids Sequence position in multiple structurally compatible protein sequences The frequency on, This indicates the first of several structurally compatible protein sequences. The first protein sequence Amino acids at each sequence position, The total number of sequences representing structurally compatible protein sequences. This indicates an indicator function that takes the value 1 when the condition is true and 0 otherwise.

3. The enzyme engineering method based on artificial intelligence according to claim 2, characterized in that, The expression for the frequency of the most frequently mutated amino acid is: , In the expression for the frequency of the most frequently mutated amino acid, Indicates sequence positions in multiple structurally compatible protein sequences The frequency of the most frequently mutated amino acid, Indicates mutated amino acids Sequence position in multiple structurally compatible protein sequences The frequency on, Indicates wild-type amino acids Sequence position in multiple structurally compatible protein sequences The frequency on; The step of comparing the frequency of the most frequently mutated amino acid at each sequence position with a first preset threshold to obtain the first set of single mutant variants of the engineered enzyme to be modified includes: According to the expression of the first creation strategy, the frequency of the most frequently mutated amino acid at each sequence position of multiple structurally compatible protein sequences is compared with a first preset threshold to obtain the first set of single mutant variants of the engineered enzyme to be modified. The expression of the first creation strategy is as follows: , In the expression of the first creation strategy, Indicates sequence positions in multiple structurally compatible protein sequences The frequency of the most frequently mutated amino acid, This represents the first preset threshold. This represents the first set of single mutant variants of the engineered enzyme to be modified. The first set of single mutant variants includes the highest frequency mutant amino acid with a frequency not lower than a first preset threshold and its sequence position.

4. The enzyme engineering method based on artificial intelligence according to claim 3, characterized in that, The process of comparing the frequency of the most frequently mutated amino acid located in the flexible region of the engineered enzyme to be modified with a second preset threshold yields a second set of single mutant variants of the engineered enzyme to be modified, including: According to the expression of the second creation strategy, the frequency of the most frequently mutated amino acid located in the flexible region of the enzyme to be modified is compared with a second preset threshold to obtain the second set of single mutant variants of the enzyme to be modified. The expression of the second creation strategy is as follows: , In the expression for the second creation strategy, Indicates sequence position The frequency of the most frequently mutated amino acid, This indicates the second preset threshold. Indicates sequence position The condition that the most frequently mutated amino acid is located in the flexible region of the enzyme to be modified is true. The second set of single mutant variants of the engineered enzyme to be modified includes the sequence position and type of the highest frequency mutant amino acid located in the flexible region of the engineered enzyme and with a frequency greater than a second preset threshold. The expression for the set of single mutant variants of the engineered enzyme to be modified is: , In the expression for the set of single mutant variants of the engineered enzyme to be modified, This represents the first set of single mutant variants of the engineered enzyme to be modified. This represents the second set of single mutant variants of the engineered enzyme to be modified. This represents the set of single mutant variants of the engineered enzyme to be modified.

5. The enzyme engineering method based on artificial intelligence according to claim 4, characterized in that, The first expression is: , In the first expression, , ,and These represent the population frequencies of alleles A and B, and heterozygote AB, respectively. The second expression is: , In the second expression, This indicates the deviation between the observed haplotype frequencies and the expected haplotype frequencies, which are independently distributed. , ,and These represent the population frequencies of alleles A and B, and heterozygote AB, respectively. The second set of mutant variants of the engineered enzyme to be modified, obtained by combining multiple structurally compatible protein sequences with a third creation strategy, includes: Based on multiple structurally compatible protein sequences, the fourth expression is used to obtain the position of the protein in the sequence. amino acids Logarithmic probability of occurrence, relative to amino acids Overall probability of occurrence in the entire sequence and in the sequence at position amino acids Logarithmic probability of occurrence, relative to amino acids Overall probability of occurrence in the entire sequence ; Based on multiple structurally compatible protein sequences, the fifth expression is used to obtain the joint mutation coupling score between pairs of amino acids. ; Based on the joint mutation coupling score between pairs of amino acids, and relative to amino acids The overall probability of occurrence throughout the entire sequence and relative to amino acids The overall probability of occurrence throughout the entire sequence is calculated, and the weighted joint mutation coupling score between each pair of amino acids is obtained according to the sixth expression. ; Combining the third creation strategy, based on the third expression, mutant combination variants with statistical coupling scores greater than or equal to the second preset value are selected to obtain the second mutant combination variant set of the engineered enzyme to be modified; The fourth expression is: , , In the fourth expression, and These represent the positions in multiple sequence alignment. and amino acids and frequency, and They represent amino acids and Background frequency; The fifth expression is: , In the fifth expression, and Representing positions respectively and location amino acids and frequency, This represents the probability that the i-th position is amino acid a and the j-th position is amino acid b; The sixth expression is: , In the sixth expression, Indicates position and location amino acids and The score for joint mutational coupling between pairs of amino acids; The expression for the first set of mutant combinations of the engineered enzyme to be modified is: , In the expression for the first set of mutant combinations of the engineered enzyme to be modified, This represents the first set of mutant combinations of the engineered enzyme to be modified. Indicates the combination of mutations in multiple pseudoDNA sequences The variance proportion score; The expression for the second mutant variant combination set of the engineered enzyme to be modified is: , In the expression for the second set of mutant variants of the engineered enzyme to be modified, This represents the set of second mutant variants of the engineered enzyme to be modified. Indicates mutation combination Statistical coupling score, This represents the 90th percentile of the global coupling score for the amino acid sequence. The expression for the final set of mutant variants of the engineered enzyme to be modified is as follows: , In the expression for the final set of mutant variants of the engineered enzyme to be modified, This represents the final set of mutant variants of the engineered enzyme to be modified. This represents the first set of mutant combinations of the engineered enzyme to be modified. This represents the second set of mutant variants of the engineered enzyme to be modified.

6. An enzyme engineering system based on artificial intelligence, characterized in that, include: The data acquisition module is used to: acquire the structural information of the engineered enzyme to be modified; The reverse folding module is used to: obtain multiple structurally compatible protein sequences based on the structural information of the engineered enzyme to be modified, using a protein reverse folding model and a given backbone structure of the engineered enzyme to be modified; The processing module is used to: obtain the frequency of mutant amino acids and the frequency of wild-type amino acids at each sequence position based on multiple structurally compatible protein sequences; The single mutation module is used to: obtain a set of single mutant variants of the engineered enzyme to be modified based on the mutant amino acid frequency and wild-type amino acid frequency at each sequence position and the structural information of the enzyme to be modified, combined with the first creation strategy and the second creation strategy. The mutation combination module is used to: obtain a set of mutant combination variants of the engineered enzyme to be modified based on the structural information of the enzyme to be modified and multiple protein sequences, combined with a third creation strategy; The process of obtaining a set of single mutant variants of the engineered enzyme based on the mutant amino acid frequency and wild-type amino acid frequency at each sequence position, along with the structural information of the enzyme to be modified, and combining the first and second creation strategies, includes: Based on the frequency of mutant amino acids and the frequency of wild-type amino acids at each sequence position, the highest frequency mutant amino acid at each sequence position with a frequency higher than that of wild-type amino acids is obtained. The frequency of the most frequent mutated amino acid at each sequence position is compared with a first preset threshold to obtain the first set of single mutant variants of the engineered enzyme to be modified. Based on the structural information of the enzyme to be modified, the flexible region of the enzyme to be modified is obtained using the protein secondary structure prediction method. Determine whether the most frequently mutated amino acid at each sequence position is located in the flexible region of the enzyme to be modified, and obtain the most frequently mutated amino acid located in the flexible region of the enzyme to be modified. The frequency of the most frequently mutated amino acid located in the flexible region of the engineered enzyme to be modified is compared with a second preset threshold to obtain a second set of single mutant variants of the engineered enzyme to be modified. By combining the first and second single mutant variant sets of the enzyme to be modified, a single mutant variant set of the enzyme to be modified is obtained. The process of obtaining a set of mutant variants of the engineered enzyme based on multiple protein sequences of the enzyme to be modified, combined with a third creation strategy, includes: Using codon correspondences, multiple structurally compatible protein sequences are reversed into multiple pseudoDNA sequences. Based on the first expression, the deviation between the observed haplotype frequency and the expected haplotype frequency, which are independently distributed, is obtained and denoted as [missing information]. ; Based on the deviation between the observed haplotype frequencies and the expected haplotype frequencies, which are independently distributed, and according to the second expression, the proportion of variance in the engineered enzyme to be modified that can be explained by one allele by the other allele is denoted as . ; Combining the third creation strategy, based on the third expression, the mutant combination variants in the engineered enzyme to be modified whose variance ratio explained by the corresponding allele of the overall allele is higher than the first preset value are selected, and the first mutant combination variant set of the engineered enzyme to be modified is obtained. Combining the third creation strategy, based on the third expression, mutant combination variants with statistical coupling scores greater than or equal to the second preset value are selected to obtain the second mutant combination variant set of the engineered enzyme to be modified; The third expression is: , In the third expression, Represents a set of mutations The total number of sites in the middle, Indicates position i and location j The scores include variance proportion score and joint mutation coupling score.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the artificial intelligence-based enzyme engineering method as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the artificial intelligence-based enzyme engineering method as described in any one of claims 1 to 5.