Artificial intelligence-based enzyme engineering method and system, device, and medium

By using an AI-based enzyme engineering method, structurally compatible protein sequences are generated using protein reverse folding models. Combined with a creation strategy, single mutant variants and mutant combination variants of the enzyme are generated, solving the problems of high cost and low efficiency in enzyme modification and achieving efficient enzyme modification and improved precision in base editing.

WO2026061352A1PCT designated stage Publication Date: 2026-03-26INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing enzyme modification methods are costly and inefficient, making it difficult to create novel engineered enzyme variants efficiently and at low cost. In particular, there are problems with insufficient editing efficiency and specificity in the base editing process.

Method used

An AI-based enzyme engineering method is employed to obtain the structural information of the enzyme to be engineered, generate multiple structurally compatible protein sequences using a protein reverse folding model, and combine these with a creation strategy to generate single mutant variants and mutant combinatorial variants, thereby achieving efficient enzyme modification.

Benefits of technology

It enables the efficient and low-cost generation of single mutant variants and mutant combinatorial variants of enzymes, expands the base editing toolbox, improves the accuracy and applicability of base editing, reduces reliance on human expert knowledge, and is applicable to the optimization of multiple different families of engineered enzymes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025121473_26032026_PF_FP_ABST
    Figure CN2025121473_26032026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of enzyme engineering. Disclosed are an artificial intelligence-based enzyme engineering method and system, a device, and a medium. The method comprises: on the basis of structural information of an engineered enzyme to be modified, using an inverse protein folding model to obtain a plurality of structurally compatible protein sequences on the basis of a given scaffold structure; on the basis of the plurality of structurally compatible protein sequences, obtaining a mutant amino acid frequency and a wild-type amino acid frequency at each sequence position; and on the basis of the mutant amino acid frequency and the wild-type amino acid frequency at each sequence position and the structural information of said engineered enzyme, as well as a first creation strategy, a second creation strategy, and a third creation strategy, obtaining a single-mutation variant set and a combinatorial-mutation variant set of said engineered enzyme. The present invention enables efficient and low-cost generation of single-mutation and combinatorial-mutation variant sets of an engineered enzyme to be modified, thereby achieving precise and efficient gene editing and effectively expanding the applicability and safety of gene editing.
Need to check novelty before this filing date? Find Prior Art

Description

An artificial intelligence-based enzyme engineering method, system, device and medium TECHNICAL FIELD

[0001] The present application relates to the technical field of enzyme engineering, and particularly relates to an artificial intelligence-based enzyme engineering method, system, device and medium. BACKGROUND

[0002] Proteins are important components of living organisms and participate in physiological and biochemical activities. Protein engineering aims to modify and design proteins to promote the understanding and reconstruction of life processes, and has been driven by artificial intelligence technology in recent years. Proteins have inherent flexibility, i.e., the ability to change their structure and function by changing the amino acid sequence. Protein engineering takes advantage of this feature to achieve rapid evolution of proteins. Compared with the natural evolution process, this process can potentially improve the evolution speed by orders of magnitude, thereby rapidly generating new protein variants that meet specific requirements.

[0003] An ideal protein engineering strategy aims to achieve the best engineering performance with the least effort. However, the complex adaptive landscape of proteins poses a great challenge to protein engineering. Current strategies include structure-guided rational protein design, directed evolution, and development of artificial intelligence models for specific protein families, but these strategies often struggle to achieve low cost and high efficiency. Specifically, structure-guided rational protein design relies on human expertise to customize protein mutations to achieve the desired functional changes, but the success rate is low and it can fall into the trap of local optimal fitness of proteins. The main defects of the directed evolution strategy are the potential evolutionary bottleneck, high iteration cost, and difficulty in customizing mutations for different situations. Protein engineering methods using deep learning models are limited by high training costs and poor generality between different proteins.

[0004] Current research uses the protein unfolding model ESM-IF1 to unfold the structure into an amino acid sequence, and then defines the fitness according to the logarithmic ratio of the amino acid frequency at each position, thereby determining mutations that may have high fitness. However, this method does not consider the influence of the protein structure itself, and is only used for antibody evolution, without fully characterizing the applicability of the method to enzymes.

[0005] Genome editing technologies, represented by CRISPR-Cas system and its derivative technologies, provide unprecedented opportunities for studying complex biological processes and exploring the root causes of genetic diseases. For example, single mutations account for 43.65% of known disease-related variations in the human genome, and are closely related to the economic traits of crops and livestock. Base editing technologies, such as ABE (adenine base editor) and CBE (cytosine base editor), which use deaminases as the main functional component, have been widely used, but there are problems such as insufficient editing efficiency and specificity. Specifically, although ABE and CBE can achieve A-to-G and C-to-T conversion, there are still problems of bystander editing, i.e. unintended base conversion, and available deaminases cannot meet the customized needs.

[0006] Therefore, in view of the above problems, there is an urgent need for an artificial intelligence-based enzyme engineering method to efficiently and cost-effectively create new engineered enzyme variants, thereby achieving precise and efficient base editing, and ultimately developing a diverse base editing toolbox. TECHNICAL PROBLEM

[0007] The present application provides an artificial intelligence-based enzyme engineering method, system, device and medium to solve the defects of high cost and low efficiency of existing enzyme modification methods. TECHNICAL SOLUTION

[0008] The present application provides an artificial intelligence-based enzyme engineering method, comprising:

[0009] obtaining structural information of an engineered enzyme to be modified;

[0010] According to the structural information of the engineered enzyme to be modified, a plurality of protein sequences compatible with the structure are obtained based on the given backbone structure of the engineered enzyme to be modified through a protein unfolding model;

[0011] According to the plurality of protein sequences compatible with the structure, the mutation amino acid frequency and the wild type amino acid frequency at each sequence position are obtained;

[0012] According to the mutation amino acid frequency and the wild type amino acid frequency at each sequence position and the structural information of the engineered enzyme to be modified, a single mutation variant set of the engineered enzyme to be modified is obtained by combining a first creation strategy and a second creation strategy;

[0013] According to the plurality of protein sequences compatible with the structure, a mutation combination variant set of the engineered enzyme to be modified is obtained by combining a third creation strategy.

[0014] The present application also provides an artificial intelligence-based enzyme engineering system, comprising:

[0015] The data acquisition module is configured to obtain structural information of an engineered enzyme to be modified;

[0016] The reverse folding module is configured to: according to the structural information of the engineering enzyme to be reformed, obtain a plurality of protein sequences compatible in structure based on a given skeleton structure of the engineering enzyme to be reformed through a protein reverse folding model;

[0017] The processing module is configured to: according to the plurality of protein sequences compatible in structure, obtain a mutation amino acid frequency and a wild type amino acid frequency at each sequence position;

[0018] The single mutation obtaining module is configured to: according to the mutation amino acid frequency and the wild type amino acid frequency at each sequence position and the structural information of the engineering enzyme to be reformed, and in combination with the first creation strategy and the second creation strategy, obtain a single mutation variant set of the engineering enzyme to be reformed.

[0019] The mutation combination obtaining module is configured to: according to the structural information of the engineering enzyme to be reformed and the plurality of protein sequences, and in combination with the third creation strategy, obtain a mutation combination variant set of the engineering enzyme to be reformed.

[0020] The present application also provides an electronic device comprising a processor and a memory storing a computer program, wherein the processor implements any of the above-mentioned artificial intelligence-based enzyme engineering methods when executing the computer program.

[0021] The present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement any of the above-mentioned artificial intelligence-based enzyme engineering methods.

[0022] The present application also provides a computer program product comprising a computer program, wherein the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to implement any of the above-mentioned artificial intelligence-based enzyme engineering methods.

[0023] The present application also provides a deaminase mutant, wherein the deaminase mutant has at least one mutation in an amino acid site compared with a wild type deaminase, and the deaminase mutant has at least one of higher base editing efficiency, narrower editing window and lower off-target efficiency compared with the wild type deaminase.

[0024] The deaminase mutant as described above, wherein the wild-type deaminase is an adenine deaminase whose amino acid sequence is represented by SEQ ID NO: 8, and the deaminase mutant includes one selected from the group consisting of E1M, H10R, A11S, A11V, L12I, G27A, A28C, V29I, L32K, N33D, E39D, E39K, R43Q, H48N, L59I, G62A, Y77F, Y77W, T79N, A87C, A89L, I91L, S93A, R94K, I95V, V98I, F100Y, G101S, G108S, A110L, A110S, G111E, G111A, L141M, V116I, L117F, L117I, N118A, P120E, M122N, M122I, M122L, H124Y, E130S in the amino acid sequence represented by SEQ ID NO: 8.

[0025] The deaminase mutant as described above, wherein the deaminase mutant includes one selected from the group consisting of E1M, A11S, G27A, V29I, L32K, N33D, E39D, E39K, L59I, G62A, Y77F, Y77W, T79N, A87C, I91L, S93A, R94K, I95V, V98I, F100Y, G101S, A110L, A110S, G111E, G111A, V116I, L117F, L117I, N118A, P120E, M122N, M122L, E130S in the amino acid sequence represented by SEQ ID NO: 8.

[0026] The deaminase mutant as described above, wherein the deaminase mutant includes one selected from the group consisting of A11S, G27A, E39K, E39D, L59I, Y77F, Y77W, T79N, A87C, R94K, V98I, A110S, A110L, G111A, G111E, L117I, L117F, M122L, E130S in the amino acid sequence represented by SEQ ID NO: 8.

[0027] The deaminase mutant as described above, wherein the deaminase mutant includes one selected from the group consisting of G27A, E39K, T79N, A87C, R94K, V98I, A110S, A110L, L117I, L117F, M122L in the amino acid sequence represented by SEQ ID NO: 8.

[0028] The deaminase mutant as described above, wherein the deaminase mutant includes one selected from the group consisting of G27A, E39K, A87C, R94K, V98I, A110S, A110L in the amino acid sequence represented by SEQ ID NO: 8.

[0029] The deaminase mutant as described above, which includes A87C in the amino acid sequence set forth in SEQ ID NO: 8.

[0030] The deaminase mutant as described above, which includes V29I and T79N in the amino acid sequence set forth in SEQ ID NO: 8.

[0031] The deaminase mutant as described above, which includes L59I and V98I in the amino acid sequence set forth in SEQ ID NO: 8.

[0032] The deaminase mutant as described above, which includes L59I and G111E in the amino acid sequence set forth in SEQ ID NO: 8.

[0033] The deaminase mutant as described above, which includes L59I and G108S in the amino acid sequence set forth in SEQ ID NO: 8.

[0034] The deaminase mutant as described above, which includes R94K and G111E and L117F in the amino acid sequence set forth in SEQ ID NO: 8.

[0035] The deaminase mutant as described above, which includes A11S and G27A in the amino acid sequence set forth in SEQ ID NO: 8.

[0036] The deaminase mutant as described above, which includes V29I and T79S in the amino acid sequence set forth in SEQ ID NO: 8.

[0037] The deaminase mutant as described above, which includes L59I and G111A in the amino acid sequence set forth in SEQ ID NO: 8.

[0038] The deaminase mutant as described above, which includes L59I and G108A in the amino acid sequence set forth in SEQ ID NO: 8.

[0039] The deaminase mutant as described above, which includes R94K and G111A and L117I in the amino acid sequence set forth in SEQ ID NO: 8.

[0040] The deaminase mutant as described above, which includes I56A and L59I in the amino acid sequence set forth in SEQ ID NO: 8.

[0041] The deaminase mutant as described above, comprising A110L and L117F in the amino acid sequence set forth in SEQ ID NO: 8.

[0042] The deaminase mutant as described above, comprising C83S and C86S in the amino acid sequence set forth in SEQ ID NO: 8.

[0043] The deaminase mutant as described above, comprising A11V and G27A in the amino acid sequence set forth in SEQ ID NO: 8.

[0044] The deaminase mutant as described above, comprising A11V and G27A and E39K in the amino acid sequence set forth in SEQ ID NO: 8.

[0045] The deaminase mutant as described above, comprising A11V and G27A and E39K and T79N in the amino acid sequence set forth in SEQ ID NO: 8.

[0046] The deaminase mutant as described above, comprising V29I and L59I and Y77W and T79N in the amino acid sequence set forth in SEQ ID NO: 8.

[0047] The deaminase mutant as described above, comprising A110S and L117I in the amino acid sequence set forth in SEQ ID NO: 8.

[0048] The deaminase mutant as described above, wherein the wild-type deaminase is a single-stranded DNA cytosine deaminase having the amino acid sequence set forth in SEQ ID NO: 7, and the deaminase mutant comprises in the amino acid sequence set forth in SEQ ID NO: 7 one selected from the group consisting of P1G, P1M, A2P, K4A, K4E, P5C, P5G, S6A, K9D, P10A, T13C, P15K, A16P, K27D, D28Y, R29C, A30C, W35Y, G39E, N40A, N40T, V42C, G44D, S47T, A48P, D49G, D51T, P53C, A55S, T56R, K60I, W63Y, Y66M, A77C, H82A, H82R, D84K, T88V, V90T, M91L, M91V, K99D, K105A, K105C, L106R, K114D, K114E, G115D, S116A, W119V, W119Y, M120L, M120V, R122V, F124K, N126D, G128S, K130R, K130T, Y132I, Q133D, Q133V, T137D, R139W, Y141F, V142A.

[0049] The deaminase mutant as described above, wherein the deaminase mutant comprises in the amino acid sequence set forth in SEQ ID NO: 7 one selected from the group consisting of P1G, P1M, A2P, K4A, K4E, P5C, P5G, S6A, K9D, P10A, T13C, P15K, A16P, K27D, D28Y, R29C, A30C, W35Y, N40A, N40T, V42C, G44D, S47T, A48P, D49G, D51T, A55S, T56R, K60I, W63Y, Y66M, H82R, H82A, D84K, T88V, V90T, M91V, M91L, K99D, K105C, K105A, L106R, K114D, K114E, G115D, W119V, W119Y, M120L, F124K, N126D, G128S, K130R, K130T, Y132I, Q133D, Q133V, T137D, V142A.

[0050] The deaminase mutant as described above, which comprises one selected from the group consisting of P1G, A2P, K4A, K4E, P5G, P5C, S6A, T13C, P15K, A16P, D28Y, A30C, N40A, N40T, V42C, A55S, T56R, K60I, Y66M, H82R, H82A, D84K, T88V, V90T, K99D, L106R, K114D, W119V, M120L, F124K, K130R, Y132I, Q133D, Q133V, V142A in the amino acid sequence set forth in SEQ ID NO: 7.

[0051] The deaminase mutant as described above, which comprises one selected from the group consisting of T13C, D28Y, N40A, N40T, V42C, T56R, Y66M, H82R, T88V, L106R, M120L, Y132I, V142A in the amino acid sequence set forth in SEQ ID NO: 7.

[0052] The deaminase mutant as described above, which comprises V42L and R80W in the amino acid sequence set forth in SEQ ID NO: 7.

[0053] The deaminase mutant as described above, which comprises P15K and D19A in the amino acid sequence set forth in SEQ ID NO: 7.

[0054] The deaminase mutant as described above, which comprises D49G and D50N in the amino acid sequence set forth in SEQ ID NO: 7.

[0055] The deaminase mutant as described above, which comprises M120L and F134Y in the amino acid sequence set forth in SEQ ID NO: 7.

[0056] The deaminase mutant as described above, which comprises F124K and K130T in the amino acid sequence set forth in SEQ ID NO: 7.

[0057] The deaminase mutant as described above, which comprises Y132F and Y141F in the amino acid sequence set forth in SEQ ID NO: 7.

[0058] The deaminase mutant as described above, which comprises W63Y, H76Y, W119Y, Y132F, and Y141F in the amino acid sequence set forth in SEQ ID NO: 7.

[0059] The deaminase mutant as described above, wherein the wild-type deaminase is a double-stranded DNA deaminase having the amino acid sequence of SEQ ID NO: 2, and the deaminase mutant comprises one selected from S2K, S2T, T9C, T9V, G24P, G24S, R38A, R38P, S47A, N55G, S57K, S57Q, S57T, V61L, T66P, Y72A, Y72F, T77I, T77L, T77V, E93A, E93P, E93S, A102S, D104A, D104G, V106P, S113A, S113P in the amino acid sequence of SEQ ID NO: 2.

[0060] The deaminase mutant as described above, wherein the deaminase mutant comprises one selected from S2K, T9V, T9C, G24S, G24P, R38A, R38P, S47A, S57K, V61L, T66P, Y72A, T77L, T77I, E93S, E93A, E93P, D104A, D104G, V106P, S113A, S113P in the amino acid sequence of SEQ ID NO: 2.

[0061] The deaminase mutant as described above, wherein the deaminase mutant comprises one selected from R38A, R38P, S47A, V61L, T66P, T77L, E93S in the amino acid sequence of SEQ ID NO: 2.

[0062] The present application also provides a nuclease mutant having a mutation of at least one amino acid site of a wild-type nuclease, and the nuclease mutant has a higher nucleic acid cleavage efficiency than the wild-type nuclease.

[0063] The nuclease mutant as described above, wherein the wild-type nuclease is a nuclease AcCas12n having the amino acid sequence of SEQ ID NO: 1, and the nuclease AcCas12n mutant comprises one selected from F35Y, H37Y, L84F, A86K, W124Y, R136P, G185A, V243T, A261R, R281A, T286Q, R298A, V376L, R438K, N456D, A460R, R483P, S486P in the amino acid sequence of SEQ ID NO: 1.

[0064] The nuclease mutant as described above, wherein the wild-type nuclease is a nuclease LbCas12a having an amino acid sequence set forth in SEQ ID NO: 3, and the nuclease mutant of the nuclease LbCas12a comprises one selected from S12P, L54I, D156K, K206E, V228L, S333P, K374G, I419L, M456I, N607K, Y653F, E659P, F682W, Y794F, E795I, K979P, M986L, I996V, A1173S in the amino acid sequence set forth in SEQ ID NO: 3.

[0065] The nuclease mutant as described above, wherein the wild-type nuclease is a nuclease SpRYCas9 having an amino acid sequence set forth in SEQ ID NO: 6, and the nuclease mutant of the nuclease SpRYCas9 comprises one selected from K111R, R139I, M161L, H167N, D182Q, D288P, S675T, S685D, F846H, N869D, L908I, E923V, D947G, T957L, M1021L, Y1036L, G1067P, H1262N in the amino acid sequence set forth in SEQ ID NO: 6.

[0066] The present application also provides a nuclear localization sequence mutant, which has a mutation of at least one amino acid site compared with a wild-type nuclear localization sequence, and has a higher nuclear localization efficiency compared with the wild-type nuclear localization sequence.

[0067] The nuclear localization sequence mutant as described above, wherein the wild-type nuclear localization sequence is a nuclear localization sequence having an amino acid sequence set forth in SEQ ID NO: 5, and the nuclear localization sequence mutant comprises one selected from K1I, R2Y, T3D, E8L, E10I, K12D, K13D, K14P in the amino acid sequence set forth in SEQ ID NO: 5.

[0068] The present application also provides a fusion protein comprising the deaminase mutant domain and / or the nuclease mutant domain as described above.

[0069] The fusion protein as described above, wherein the fusion protein further comprises a nuclear localization sequence domain or the nuclear localization sequence mutant domain as described above.

[0070] It can be understood that the nuclear localization sequence or the nuclear localization sequence mutant can be fused at any position of the deaminase mutant domain, at any position of the nuclease mutant domain, or at any position of both the deaminase mutant domain and the nuclease mutant domain.

[0071] The application further provides a complex comprising the fusion protein and an sgRNA targeting a target sequence.

[0072] The application further provides a polynucleotide encoding the fusion protein or the complex.

[0073] The application further provides a vector comprising the polynucleotide.

[0074] The application further provides a gene editing method comprising contacting a nucleic acid to be edited with the complex.

[0075] The application further provides a reverse transcriptase mutant, wherein the reverse transcriptase mutant has at least one mutation in an amino acid site compared with a wild-type reverse transcriptase, and the reverse transcriptase mutant has a higher reverse transcription efficiency than the wild-type reverse transcriptase.

[0076] The reverse transcriptase mutant as described above, wherein the wild-type reverse transcriptase is a reverse transcriptase with an amino acid sequence as shown in SEQ ID NO: 4, and the reverse transcriptase mutant comprises one selected from G22P, T55D, S56A, T197N, D209P, Q213A, Q221L, G239A, M289L, T332V, L333P, W406L, C409P, R411Q, V433T, Y460I in the amino acid sequence as shown in SEQ ID NO: 4.

[0077] The application further provides a fusion protein comprising the reverse transcriptase mutant domain and a nickase domain.

[0078] The fusion protein as described above, wherein the nickase domain is a Cas protein.

[0079] The fusion protein as described above, wherein the nickase is at least one selected from the nuclease mutants.

[0080] The fusion protein as described above, wherein the fusion protein further comprises a nuclear localization sequence domain or the nuclear localization sequence mutant domain.

[0081] The application further provides a complex comprising the fusion protein and an sgRNA targeting a target sequence.

[0082] The application further provides a polynucleotide encoding the fusion protein or the complex.

[0083] The application further provides a vector comprising the polynucleotide.

[0084] The application further provides a gene editing method comprising contacting a nucleic acid to be edited with the complex. Advantages

[0085] The application provides an enzyme engineering method, system, device and medium based on artificial intelligence, which is based on population evolution theory, obtains a plurality of protein sequences compatible in structure according to structural information of an engineering enzyme to be modified based on a given skeleton structure of the engineering enzyme to be modified through a protein reverse folding model, and combines a first creation strategy, a second creation strategy and a third creation strategy to efficiently and at low cost obtain a single-mutation variant set and a mutation combination variant set of the engineering enzyme to be modified. The engineering enzyme variant developed according to the application can greatly expand the base editing toolbox, realize all-round improvement of the base editing engineering properties, realize accurate and efficient base editing, effectively expand the application range and safety of base editing, and make the base editing strategy tend to be customized.

[0086] In addition, the application does not need to train a special artificial intelligence model, does not need to be guided by human expert knowledge, and can generate efficient mutations through only one round, greatly improving the efficiency of enzyme engineering. Secondly, the application has strong universality, is suitable for optimization of a plurality of different family engineering enzymes, and is compatible with traditional methods. Thirdly, the application provides a plurality of counterintuitive but effective single-mutation variants and mutation combination guidance capable of potentially avoiding epistatic effects based on reverse folding of structural information, evolution coupling information and structure-guided learning, and is helpful for extensive exploration of enzyme sequence and function generation space. The application not only can guide rapid optimization of enzymes, but also provides deep insights into protein distribution and generation space in a natural state. BRIEF DESCRIPTION OF DRAWINGS

[0087] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.

[0088] Fig. 1 is a schematic diagram comparing advantages and disadvantages of artificial intelligence assisted, traditional structure guided rational design and directed evolution method in protein engineering.

[0089] Fig. 2 is a flowchart of an enzyme engineering method based on artificial intelligence provided by the application.

[0090] Fig. 3 is a technical detail diagram of an enzyme engineering method based on artificial intelligence provided by the application.

[0091] Fig. 4 is an AiCE singlePerformance comparison and benchmarking of different filtering strategies in different structural regions; A, fitness difference between flexible and non-flexible regions for single amino acid substitution (31 DMS libraries); B, fitness difference comparison based on heterogeneous and homogeneous screening in different structural regions (31 DMS libraries); C, prediction accuracy and number distribution of high fitness (HF) mutations at different mutation rates (31 DMS libraries); D, enrichment trend of relative high fitness (RHF) mutations at different mutation rates (31 DMS libraries); E, prediction accuracy of HF mutations in 60 DMS libraries; F, performance comparison of different filtering strategies for HF mutations; G, performance comparison of different methods for HF mutations; H, AiCE single Prediction accuracy at different quantile thresholds (51 monomeric protein libraries); I, AiCE single Performance of HF mutation prediction in Cas9 protein; J, AiCE single Performance of HF mutation prediction in protein-protein and protein-nucleic acid complexes.

[0092] Figure 5. AiCE single Details of performance comparison and benchmarking of different filtering strategies in different structural regions; A, prediction accuracy of the 15th and 25th quantile mutations at different mutation rates in 31 DMS libraries; B, same as (A), but based on 60 DMS libraries; C, logistic regression analysis of different inverse folding models for predicting mutations in 60 DMS libraries, statistical test using likelihood ratio test; D, comparison of different filtering strategies in HF mutation prediction: ProteinMPNN (left), ESM-IF1 (middle), LigandMPNN (right).

[0093] Figure 6. AiCE multi Performance of AiCE strategy in combination of high fitness mutations in 4 antibodies and additional 2 protein mutation libraries; A-B, AiCE multi Performance of AiCE strategy in combination of high fitness mutations in 4 antibodies and additional 2 protein mutation libraries; A-B, AiCE multi Performance of AiCE strategy in combination of high fitness mutations in 4 antibodies and additional 2 protein mutation libraries; A-B, AiCE multi Principle diagram.

[0094] Figure 7. AiCEmulti Performance test and parallel comparison of different inverse folding models; Wherein A-B is AiCE multi Performance of CR6261 (A) and CR9114 (B) antibody HF multi-mutation prediction, the model uses ProteinMPNN and ESM-IF1; C is the comparison of the adaptive distribution of the mutants predicted by different methods and the global distribution; D is the performance of SaProt in His3 and ppluGFP2 in HF multi-mutation prediction.

[0095] Figure 8 is the phylogenetic evolutionary relationship of deaminases created by AiCE, which contains Ddd1 and Sdd6 proteins in SCP1.201 family and TadA8e protein in dCMP_cyt family.

[0096] Figure 9 is the design and editing evaluation of single-stranded DNA deaminase TadA8e and Sdd6 variants; Wherein, A is the optimization of TadA8e mutants, the average relative editing efficiency of mutants generated by different methods at three sites in HEK293T cells, HF mutations are marked in black and the number (n) is indicated, the total number (N) generated by each method is also shown, 1 is the wild type efficiency; B is AiCE multi Multi-mutation predicted by BLOSUM62 matrix, based on the average relative editing efficiency determined at three sites in HEK293T cells, Shared is the mutant shared by the two methods.(C) AiCE multi Comparison of relative editing efficiency of multi-mutation and single mutation predicted by BLOSUM62, each point is the efficiency of a single mutation at three sites, three independent repeats; D is three AiCE multi Predicted multi-mutation, two random low EC score multi-mutation (Low multi ) and wild type TadA8e at three sites; E is the average window editing efficiency of 18 TadA8e mutants (13 single mutations, 5 multi-mutations), TadA8e and ABE9 (reported high-precision variant) at six sites in HEK293T cells; F is the window editing efficiency distribution of Tc1, TadA8e and ABE9 at 24 target sites in HEK293T cells; G is the optimization of Sdd6 mutants, the average relative editing efficiency of two endogenous sites, the rest are marked as in (A); H is AiCE multi Sdd6 multi-mutation predicted by BLOSUM62, marked as in (B); I is AiCE multiComparison of BLOSUM62 prediction of Sdd6 mutations, labeled as (C); J is the relative editing efficiency of selected Sdd6 mutants at six sites in HEK293T cells; K is the structural analysis of Sdd6 and Sc9; L-M are the specificity assays of Sc9 in HEK293T cells (L) and HeLa cells (M) using orthogonal R-loop assay, compared with wild-type Sdd6, hyperactive variant rAPOBEC1 and hyper- fidelity variant rAPOBEC1-YE1 for off-target effects.

[0097] Figure 10 is the design and editing evaluation of TadA8e window-reduced variants; wherein A is the schematic diagram of nCas9-based ABE vector; B is the average window editing efficiency of selected TadA8e mutants at six sites in HEK293T cells, with the PAM proximal defined as position 1; C is the cryo-EM structure of TadA8e and SpCas9 complex (PDB ID: 6VPC), with the yellow labels indicating the constituent mutations and their spatial positions and potential interactions in Tc1; D is the window editing efficiency distribution of Tc1, TadA8e and ABE9 at 24 target sites of eight genes in three cell lines.

[0098] Figure 11 is the design and editing evaluation of Sdd6 high-specificity variants; wherein A is the schematic diagram of CBE vector, including nSpCas9 (D10A)-based CBE and nSaCas9 (D10A)-based orthogonal R-loop detection system; B is the editing efficiency of selected Sdd6 mutants, Sdd6 and rAPOBEC1 at six sites in HEK293T cells; C is the HEK293T cell-specific evaluation, including eight target sites and two off-target sites for specificity detection; D is the multi-cell line editing efficiency, including the average editing efficiency of Sc9, Sdd6, rAPOBEC1 and rAPOBEC1-YE1 at 24 endogenous sites in HEK293T, HeLa, K562 and U2OS cells; E is the specificity evaluation of the system in HeLa cells, including 16 target sites and 6 off-target sites; F is the editing ratio of each editor target site / off-target site calculated based on the results of (D).

[0099] Figure 12 is the design and editing evaluation of deaminase Ddd1 variants; A is the schematic of DdCBE detection system, split Ddd1 and TALE fusion were used for C-to-T editing; B is the editing efficiency of Ddd1 and DddA at two nuclear targets and six mitochondrial targets in HEK293T cells; C is the computational strategy to generate Ddd1 mutants, the total number of mutants tested for each method is labeled outside the column, the number of HF mutants is labeled inside the column, and the black line indicates the overlap of mutants identified by different methods; D is the relative editing efficiency of Ddd1 mutants at ND1.2 and ND6.2 two mitochondrial sites in HEK293T cells, the black line is the relative efficiency 1.1; E is the editing efficiency of AiCE multi and the corresponding single mutation of the multi-mutant predicted by BLOSUM62 matrix, the mutation type is shown above, and the single mutation constituting the combined mutation is connected by a black line; F is the editing efficiency of enDdd1, D-h11, wild-type Ddd1, PACE-M3, DddA11 and DddA at four mitochondrial sites in HEK293T cells; G-H is the editing efficiency of Ddd1 mutants, Ddd1, DddA and DddA11 at mitochondrial (G) and nuclear targets (H) in HeLa cells and HEK293T cells, D-hs is AiCE single - The set of 20 mutants obtained by ProteinMPNN with the parameter β = 0.8.

[0100] Figure 13 is the detailed design and editing evaluation of deaminase Ddd1 variants; A is the schematic of DdCBE vector construction; B is the editing efficiency of Ddd1 and DddA at nuclear and mitochondrial targets; C is the editing efficiency of Ddd1 mutants obtained by different models at two mitochondrial sites in HEK293T cells, colorless indicates efficiency lower than Ddd1 and DddA, transparent color indicates between the two, solid color indicates higher than the two, and the dotted line is the efficiency of wild-type Ddd1; D is the relative editing efficiency of Ddd1 mutants at four mitochondrial sites in HEK293T cells; E is AiCE multi Comparison of Ddd1 mutant prediction with BLOSUM62 matrix; F is the editing efficiency of DC6, DC7, Ddd1 and DddA at two mitochondrial sites in HEK293T and HeLa cells; G is the structure visualization of enDdd1 and DddA11 (DC7); H is the comparison of enDdd1 and other mutants in mitochondrial editing; I is the editing activity heatmap of Ddd1 mutants within the endogenous sequence window, corresponding to (F); J is the editing efficiency of Ddd1 mutants, Ddd1, DddA and DddA11 at the nuclear target SIRT6 in HEK293T cells.

[0101] Figure 14 is the design and editing evaluation of other engineered enzyme variants; A is AiCE singleSchematic diagram of detection strategy of optimized mutants in multiple proteins; B-D are the activities of optimized NLS (B), nuclease (C) and reverse transcriptase M-MLV RT (D) mutants; E is the relative editing efficiency distribution of five proteins; F is the structure correlation analysis of eight AiCE optimized proteins, the data shows the RMSD correlation value, reflecting the structural similarity between mutants.

[0102] Figure 15 is a specific vector construction and editing evaluation details of other engineered enzyme variants; wherein, A is a schematic diagram of NLS activity detection system, in CBE system, N-terminal is mutant BPNLS, and C-terminal is wild type BPNLS; B is the average editing efficiency of NLS mutants and wild type at three sites in HeLa cells; C is a schematic diagram of nuclease activity detection system, LbCas12a, AcCas12n and SpRYCas9 constructs all contain RNA expression vectors for detecting indel efficiency; D-F are the average indel efficiency of LbCas12a (D), AcCas12n (E) and SpRYCas9 (F) mutants and wild type at multiple sites in HEK293T cells; G is a schematic diagram of M-MLV RT-based prime editing system, which contains truncated M-MLV RT mutants and RNA expression vectors for targeted editing; H is the average editing efficiency of M-MLV RT mutants and wild type at three sites in HEK293T cells, and all proteins are truncated type lacking RNase H domain.

[0103] Figure 16 is AiCE single Distribution of high fitness mutations obtained by the strategy; wherein, A is the distribution of mutations generated by ProteinMPNN, the ridge chart shows the mutation rate distribution, and the dotted line is the median; the line chart shows the proportion of each position mutation amino acid and wild type; B is the structural distribution of single mutation, the dark area is the hot spot identified by traditional rational design, the light blue is the experimental detection site, and the dark blue is the high fitness (HF) mutation, the table summarizes the number of HF mutations (high), other mutations (low) and total number of HF mutations (total) in the hot spot region (R) and non-hot spot region (N) of seven proteins, TadA8e structure (PDB ID: 6VPC) is analyzed by cryo-EM, and the rest is predicted by AlphaFold3.

[0104] Figure 17 is a structural schematic diagram of an enzyme engineering system based on artificial intelligence provided by the present application.

[0105] Figure 18 is a structural schematic diagram of an electronic device provided by the present application. Embodiments of the present application

[0106] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below in combination with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments, and they should not be understood as limiting the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application. In the description of the present application, it should be understood that the terms used are only for the purpose of description and should not be understood as indicating or implying relative importance.

[0107] The AI-based enzyme engineering method, system, device and medium provided by the present application will be described below in combination with Figures 1-19. It should be noted that the execution subject of the AI-based enzyme engineering method provided by the present application can be any network side device / terminal side device that meets the technical requirements, such as an AI-based enzyme engineering device, etc.

[0108] The present application combines general artificial intelligence, evolutionary theory and structural information to propose a new enzyme engineering method. Proteins have flexibility, and protein engineering takes advantage of this ability to modify their structure and function by changing their amino acid sequences. Compared with natural processes, protein engineering has the potential to increase the speed of evolution by several orders of magnitude, so that new variants of proteins that can adapt to specific requirements can be quickly generated. However, the complex fitness landscape of proteins poses a huge challenge to protein engineering. Current strategies, including structure-guided rational design of proteins, directed evolution and model customization for specific protein families, are difficult to achieve low-cost and efficient enzyme engineering (Figure 1).

[0109] As shown in Figure 1, structure-guided rational design can achieve customization of mutants, but is inefficient and requires strong professional knowledge background. The directed evolution strategy can achieve rapid creation and iteration of mutations by artificially simulating the evolutionary process in nature. However, this scheme has problems such as high cost and limited starting template, and the selection pressure is often single, making it difficult to achieve customization of mutations with high fitness in multiple scenarios. The application of artificial intelligence to enzyme engineering is a new direction of current development, and researchers often customize exclusive models for specific enzyme families through the transfer of large models, often achieving good optimization results, but also bringing the problems of high model training cost and difficulty in effective generalization.

[0110] Current general large models, including language models and structural models, effectively learn the generation and distribution rules of proteins in nature. Effective utilization of general protein models can effectively improve the problems of model transfer and generalization. The present application therefore proposes an AI-based enzyme engineering method, and Figure 2 is a flowchart thereof. Referring to Figure 2, the AI-based enzyme engineering method provided by the present application can comprise:

[0111] Step S110, obtaining structural information of the engineering enzyme to be modified. Specifically, the structural information of the engineering enzyme to be modified can be obtained from an open-source protein structure database, or predicted by AlphaFold2, AlphaFold3 or the like. The structural information of the engineering enzyme to be modified can include structural type information, active site information, metal ion or cofactor information, domain and functional domain information, three-dimensional structure information (for example, the three-dimensional coordinates of each atom in the protein structure of the engineering enzyme to be modified), and the like.

[0112] Step S120, obtaining a plurality of structure-compatible reverse folding amino acid sequences based on the protein three-dimensional structure of the engineering enzyme to be modified according to the structural information of the engineering enzyme to be modified by a protein reverse folding model. Specifically, the protein reverse folding model can be any existing protein reverse folding model suitable for engineering enzyme modification, such as ESM-IF1, LigandMPNN, ProteinMPNN, etc.

[0113] Step S130, obtaining the mutation amino acid frequency and wild-type amino acid frequency at each sequence position according to the plurality of structure-compatible protein sequences.

[0114] In an embodiment, step S130 can include:

[0115] obtaining the frequency of each mutation amino acid at each sequence position according to the plurality of structure-compatible protein sequences by a first expression;

[0116] obtaining the frequency of wild-type amino acid at each sequence position according to the plurality of structure-compatible protein sequences by a second expression.

[0117] wherein the first expression is:

[0118] In the first expression, x represents 20 common amino acid types, f x (i) represents the frequency of amino acid x i in the sequence position i in the plurality of structure-compatible protein sequences, x i represents the amino acid at the i-th sequence position of the j-th protein sequence in the plurality of structure-compatible protein sequences, M represents the total number of sequences of the plurality of structure-compatible protein sequences, and 1{·} represents an indicator function, which takes a value of 1 when the condition is true, and otherwise takes a value of 0.

[0119] The second expression is:

[0120] In the second expression, f wt (i) and f mut(i) represents the frequency of the wild-type amino acid wt and the mutant amino acid mut at the sequence position i in the structurally compatible plurality of protein sequences, respectively, x i represents the amino acid at the i-th sequence position of the j-th protein sequence in the structurally compatible plurality of protein sequences, M represents the total number of sequences of the structurally compatible plurality of protein sequences, 1{·} represents an indicator function that takes the value 1 when the condition is true and 0 otherwise.

[0121] Step S140, according to the mutant amino acid frequency and the wild-type amino acid frequency at each sequence position and the structure information of the engineering enzyme to be modified, combining the first creation strategy and the second creation strategy, obtaining a single-mutation variant set of the engineering enzyme to be modified.

[0122] In an embodiment, step S140 can include:

[0123] According to the mutant amino acid frequency and the wild-type amino acid frequency at each sequence position, obtaining the highest frequency mutant amino acid at each sequence position with a frequency higher than the wild-type amino acid frequency and the highest frequency;

[0124] Comparing the frequency of the highest frequency mutant amino acid at each sequence position with the first preset threshold value respectively, obtaining a first single-mutation variant set of the engineering enzyme to be modified;

[0125] Wherein, the expression of the frequency of the highest frequency mutant amino acid is:

[0126] f max (i) = max{f mut (i) | f mut (i) > f wt (i)},

[0127] In the expression of the frequency of the highest frequency mutant amino acid, f max (i) represents the frequency of the highest frequency mutant amino acid at the sequence position i in the structurally compatible plurality of protein sequences, f mut (i) represents the frequency of the mutant amino acid mut at the sequence position i in the structurally compatible plurality of protein sequences, f wt (i) represents the frequency of the wild-type amino acid wt at the sequence position i in the structurally compatible plurality of protein sequences. According to the structure information of the engineering enzyme to be modified, the flexible region of the engineering enzyme to be modified is obtained by using the protein secondary structure prediction method. Specifically, the energy of each possible hydrogen bond pair can be calculated by, for example, the protein secondary structure prediction method of DSSP algorithm according to the structure information of the engineering enzyme to be modified. If the hydrogen bond energy is lower than a preset threshold value (for example, -0.5 kcal / mol), it is considered that there is a hydrogen bond, and the hydrogen bond does not belong to the flexible region. The energy of the hydrogen bond pair is obtained by the following formula:

[0128] where q i and q j denotes the charge of the atom involved in hydrogen bond formation, r ON , r CH , r OH , r CN denotes the distance between the corresponding atom pairs (in units), the constant 332 converts the energy units into kcal / mol;

[0129] determine whether the highest frequency mutation amino acid at each sequence position is located in the flexible region of the engineered enzyme to be modified, to obtain the highest frequency mutation amino acid located in the flexible region of the engineered enzyme to be modified;

[0130] compare the frequency of the highest frequency mutation amino acid located in the flexible region of the engineered enzyme to be modified with a second preset threshold, to obtain a second single-mutation variant set of the engineered enzyme to be modified;

[0131] combine the first single-mutation variant set and the second single-mutation variant set of the engineered enzyme to be modified, to obtain a single-mutation variant set of the engineered enzyme to be modified.

[0132] In an embodiment, step S140 can compare the frequency of the highest frequency mutation amino acid at each sequence position with a first preset threshold according to the expression of the first creation strategy, to obtain a first single-mutation variant set of the engineered enzyme to be modified, wherein the expression of the first creation strategy is:

[0133] In the expression of the first creation strategy, f max (i) denotes the frequency of the highest frequency mutation amino acid at sequence position i, and β denotes the first preset threshold, denotes the first single-mutation variant set of the engineered enzyme to be modified, and the first single-mutation variant set includes the highest frequency mutation amino acid with a frequency not lower than the first preset threshold and the sequence position thereof.

[0134] Then, according to the expression of the second creation strategy, the frequency of the highest frequency mutation amino acid located in the flexible region of the engineered enzyme to be modified is compared with a second preset threshold, to obtain a second single-mutation variant set of the engineered enzyme to be modified, wherein the expression of the second creation strategy is:

[0135] In the expression of the second creation strategy, f max (i) denotes the frequency of the highest frequency mutation amino acid at sequence position i, γ denotes the second preset threshold, and is_flexible(i) = 1 indicates that the condition that the highest frequency mutation amino acid at sequence position i is located in the flexible region of the engineered enzyme to be modified is true, The second single-mutant variant set of the engineering enzyme to be reformed, the second single-mutant variant set comprising sequence positions at which the highest frequency mutation amino acids located in the flexible region of the engineering enzyme to be reformed and the frequency is greater than a second preset threshold.

[0136] Then, the first single-mutant variant set and the second single-mutant variant set of the engineering enzyme to be reformed are combined to obtain a single-mutant variant set of the engineering enzyme to be reformed, and the expression is:

[0137] In the expression of the single-mutant variant set of the engineering enzyme to be reformed, The first single-mutant variant set of the engineering enzyme to be reformed is represented by: The second single-mutant variant set of the engineering enzyme to be reformed is represented by: The single-mutant variant set of the engineering enzyme to be reformed is represented by.

[0138] Step S150, according to the structure-compatible multiple protein sequences of the engineering enzyme to be reformed, combining the third creation strategy, obtaining a mutant combination variant set of the engineering enzyme to be reformed.

[0139] In an embodiment, step S150 can include:

[0140] Using the codon correspondence relationship, the structure-compatible multiple protein sequences are reversed into multiple pseudo-DNA sequences, and according to the first expression, the deviation that the observed haplotype frequency and the expected haplotype frequency are independently distributed is obtained, denoted as D;

[0141] According to the deviation that the observed haplotype frequency and the expected haplotype frequency are independently distributed, and according to the second expression, the variance proportion that one allele in the engineering enzyme to be reformed can be explained by another allele is obtained, denoted as r 2 ;

[0142] Combining the third creation strategy, the mutant combination variant in which the overall allele in the engineering enzyme to be reformed corresponds to the allele explained variance proportion higher than the first preset value can be screened out by the third expression, obtaining a first mutant combination variant set of the engineering enzyme to be reformed.

[0143] The first expression is:

[0144] D=P AB -P A P B ,

[0145] In the first expression, P A , P B , and P AB are the population frequencies of alleles A, B, and heterozygote AB, respectively.

[0146] The second expression is:

[0147] In the second expression, D is the deviation of haplotype frequency from the expected haplotype frequency assuming independent distribution, P A , P B , and P AB are the population frequencies of alleles A, B, and heterozygote AB, respectively.

[0148] In an embodiment, step S150 can comprise:

[0149] According to the plurality of protein sequences compatible with the structure, the fourth expression is used to obtain the logarithmic probability of the occurrence of amino acid a at position i in the sequence, relative to the overall probability of the occurrence of amino acid a in the entire sequence φ i , and the logarithmic probability of the occurrence of amino acid b at position j in the sequence, relative to the overall probability of the occurrence of amino acid b in the entire sequence φ j .

[0150] According to the plurality of protein sequences compatible with the structure, the fifth expression is used to obtain the joint mutation coupling score between two amino acids

[0151] According to the joint mutation coupling score between two amino acids, and φ i and φ j , and according to the sixth expression, the weighted joint mutation coupling score EC between two amino acids is obtained.

[0152] In combination with the third creation strategy, the third expression can be used to screen mutation combination variants with a statistical coupling score greater than or equal to a second preset value, to obtain a second set of mutation combination variants of the engineered enzyme to be modified.

[0153] The fourth expression is:

[0154] In the fourth expression, f ai and f bi are the frequencies of amino acids a and b at positions i and j in the multiple sequence alignment, q a and q b are the background frequencies of amino acids a and b.

[0155] The fifth expression is:

[0156] In the fifth expression, and are the frequencies of amino acids a and b at positions i and j, represents the probability that the i-th position is amino acid a and the j-th position is amino acid b.

[0157] The sixth expression is:

[0158] In the sixth expression, represents the joint mutation coupling score between the two amino acids a and b at position i and position j.

[0159] wherein the third expression is:

[0160] In the third expression, wherein represents the total number of sites in the mutation set score(i,j) represents the score between position i and position j, including the variance proportion score and the joint mutation coupling score. This calculation provides a measure of the overall evolutionary coupling strength of the mutation combination.

[0161] Preferably, the expression of the first set of mutation combination variants of the engineered enzyme to be improved is:

[0162] In the expression of the first set of mutation combination variants of the engineered enzyme to be improved, represents the first set of mutation combination variants of the engineered enzyme to be improved, represents the variance proportion score of the mutation combination in the plurality of pseudo DNA sequences;

[0163] Preferably, the expression of the second set of mutation variants of the engineered enzyme to be improved is:

[0164] In the expression of the second set of mutation variants of the engineered enzyme to be improved, represents the second set of mutation variants of the engineered enzyme to be improved, represents the statistical coupling score of the mutation combination scorepercentile 90 refers to the 90th percentile of the global coupling score of the amino acid sequence.

[0165] In an embodiment, the expression of the final set of mutation combination variants of the engineered enzyme to be improved is:

[0166] In the expression of the final set of mutation combination variants of the engineered enzyme to be improved, represents the final set of mutation combination variants of the engineered enzyme to be improved, represents the first set of mutation combination variants of the engineered enzyme to be improved, represents the second set of mutation variants of the engineered enzyme to be improved.

[0167] The single mutant variant set and the mutant combination variant set of the to-be-reformed engineered enzyme obtained in this embodiment are shown in Table 1, and the wild-type sequence is shown in Table 2.

[0168] Table 1 Mutation information of the to-be-reformed engineered enzyme

[0169] Table 2 Protein sequence of the to-be-reformed engineered enzyme

[0170] The enzyme engineering method based on artificial intelligence provided by the present application can obtain high fitness mutations and combinations thereof from most protein unfolding models without human expert guidance. The results of the modification of three commonly used engineered enzymes show that the present application is simple, efficient and universal, and can design single and combined mutants with various engineering property optimizations or deamination substrate changes, successfully developing a series of new engineered enzymes. The present application can be used as a SOTA (state-of-the-art) method for engineering enzyme optimization, and has the potential to be used as a universal method for enzyme modification, which can also provide deep insights into the exploration of protein function space distribution and generation rules.

[0171] The present application combines an unfolding deep learning model with evolutionary theory based on experimentally determined or computationally predicted protein structures to create effective single mutations and combinations. The present application does not require training of a specialized artificial intelligence model, avoiding the cost of human expert guidance and model training, and has the advantages of simplicity, efficiency, universality, etc. Using the present application, multiple engineered enzymes with enhanced engineering properties have been successfully developed, including double-stranded DNA deaminase Ddd1, single-stranded DNA cytosine deaminase Sdd6 and adenine deaminase TadA8e, nuclease SpRYCas9, AcCas12n and LbCas12a, nuclear localization sequence NLS and reverse transcriptase M-MLV RT, and a series of products with more accurate base editing performance have been developed. Importantly, whether it is a lightweight neural network model or a large protein language model, the present application can effectively drive mutation generation, and further possibly design substrate-changing mutants and explainable combinations of high fitness mutations. The present application demonstrates the significant success of artificial intelligence combined with evolutionary theory in engineering enzyme modification and even protein modification, providing a potential direction for protein engineering iteration and understanding of protein space distribution rules.

[0172] The verification process of the present application will be described in detail below. The enzyme engineering method based on artificial intelligence provided by the present application is referred to as AiCE.

[0173] Current routine strategies for enzyme engineering mainly include two ways: structure-aided rational design and directed evolution (Fig. 1). The former is guided by human expert knowledge, which can customize the design of mutations, but it is often inefficient and time-consuming, and requires extremely professional knowledge background. The latter is an effective way of creating mutations, which simulates the natural selection process to achieve iterative creation of high-adaptability mutants. However, this method also has its inherent defects, mainly reflected in the following aspects: it needs several rounds of iteration, high cost, and requires a starting template. In recent years, with the wide application of artificial intelligence technology in the biological field, the use of pre-trained model transfer learning has become an effective new method for enzyme engineering. However, this approach also faces the problems of high training cost and weak generalization.

[0174] This embodiment uses the reported general protein unfolding model to propose a new strategy for lightweight and efficient enzyme engineering, named AiCE (Fig. 2, Fig. 3). This embodiment first proposes AiCE single , which is a strategy for creating single-mutation mutants. This strategy is based on a prior assumption that artificial intelligence learns the distribution and generation rules of proteins in nature, and can simulate natural selection by artificially defining selection intensity to screen and fix single mutations with high adaptability. According to the idea of population genetics, this embodiment believes that the adaptability of genetic sites is mapped to their distribution frequency in the population, and high-frequency mutation sites imply that they are widely preserved in nature, and these mutations often have high adaptability and functional stability. Using the publicly available 60 DMS (Deep Mutational Scanning) data, this embodiment proves that AiCE can restore beneficial mutations of test proteins, including but not limited to kinases, sequence-specific nucleases, signaling proteins, receptors, viruses, etc. (Fig. 4, Fig. 5). This example further proves that whether it is the mutation set predicted by AiCE or the mutation set of the DMS database, the mutations with high adaptability (i.e., adaptability in the top 5% of the set) show enrichment in the structural flexibility region, which proves that the structural flexibility region is tolerant to high-adaptability mutations. Therefore, this example develops the AiCE single module, which adds additional screening of the structural flexibility region, and the feasibility and effectiveness of its single-mutation creation are also proven in the 60 DMS libraries (Fig. 4, Fig. 5). Compared with the benchmark test of other models, AiCE single can achieve up to 16% of the prediction accuracy of high-adaptability mutations, which is nearly twice the performance of advanced artificial intelligence models such as ESM3-open (Fig. 4); logistic regression analysis shows that structural constraints are essential for the performance of AiCE single (Fig. 5). The effectiveness of this method is not only reflected in the design of monomeric structures, but also in the design of protein complexes (such as AsCas12f) and large proteins.

[0175] This embodiment further iteratively develops AiCEmulti The method calculates amino acid residues in evolutionary coupling by two strategies: in the first strategy, the present embodiment uses the optimal codon to reverse transcribe the amino acid sequence into a pseudo-nucleic acid sequence, and obtains the corresponding mutation combination by calculating the linkage disequilibrium between nucleotides; the second strategy calculates the co-evolution score between amino acids by statistical coupling analysis, so as to obtain amino acids that may exist in co-evolution, and the mutation combination with a high degree of evolutionary coupling may have a higher fitness. The present embodiment proves that, compared with random combination, AiCE multi has a very high probability of creating a high fitness combination mutant (Figure 6), regardless of which reverse folding model is used, based on AiCE multi The created combination mutant has a significant improvement in fitness. The present embodiment extends AiCE multi to the modification of other proteins such as His3 and ppluGFP2 proteins (Figure 6), and the results show that it effectively overcomes the epistatic effect between mutations. In addition, the screening idea based on evolutionary coupling also provides a direction for the explainable creation of combination mutations, that is, using AiCE multi to select the mutation combination with a higher degree of evolutionary coupling. Comparison with the baseline model SaProt shows that AiCE multi has a comparable prediction ability for combination mutations (Figure 7), but the calculation cost is only 1% of that of the former.

[0176] Base editors based on engineered enzymes have the potential to effectively correct pathogenic mutations, which account for about 43% of human disease-related variations (data from dbVar: https: / / www.ncbi.nlm.nih.gov / dbvar / , 08 / 2024), or introduce beneficial mutations. However, many engineered enzymes can be improved by enhancing the deamination activity, unpredictable off-target effects, and the generation of unintended base conversions (also known as bystander editing). The present embodiment focuses on improving these defects through protein engineering, taking the engineered enzyme proteins Ddd1 and Sdd6 of the SCP1.201 family, and the engineered enzyme protein TadA8e of the dCMP_cyt family as the chassis enzyme, and conducts conceptual exploration of precise and efficient enzyme modification (Figure 8).

[0177] In this embodiment, the present embodiment uses AiCE single and other strategies to design 131 TadA8e single mutants (T1-T131) and 114 Sdd6 single mutants (S1-S114) (Figures 9, 10, and 11), respectively, and uses AiCE multiTadA8e combination mutants and 21 Sdd6 combination variants (Tc1-Tc24, Sc1-Sc21) (Fig. 9, Fig. 10, Fig. 11) were designed. After the preliminary validation of three endogenous targets in HEK293T cells, this embodiment found that 19 TadA8e single mutants (11 created by AiCE single ) and 10 combination mutants (6 created by AiCE multi ) could improve the deamination efficiency by more than 10% compared to the wild type (Fig. 9, Fig. 10), and none of the randomly combined mutants improved the deamination efficiency. After the preliminary validation of two endogenous targets in HEK293T cells, this embodiment found that 48 Sdd6 single mutants (21 created by AiCE single ) and 6 combination mutants (all created by AiCE multi ) could improve the deamination efficiency by more than 10% compared to the wild type (Fig. 9, Fig. 11).

[0178] Table 3, Base editing efficiency of TadA8e single mutants and combination mutants

[0179] Table 4, Base editing efficiency validation of single-stranded DNA cytosine deaminase (Sdd6) mutants

[0180] This embodiment took the high-efficiency TadA8e mutants as the chassis, randomly selected 18 high-efficiency mutants (13 single mutants and 5 multi-mutants), and evaluated their activities in the editing window of another 6 target sites, and identified two variants T1 (E1M) and Tc1 (A11V / G27A) with improved efficiency and accuracy, of which T1 improved the editing efficiency in the target window by about 70% (Fig. 9, Fig. 10). This embodiment further constructed enABE8e using Tc1, and comprehensively evaluated enABE8e and one of the most accurate adenine base editors, ABE9, across 24 target sites of 8 genes in HEK293T, HeLa, K562 and U2OS cells (Fig. 9, Fig. 10). The editing window of the enABE8e variant is nearly half as wide as that of ABE8e, and maintains comparable or higher editing efficiency at more than half of the test sites; compared to ABE9, the window is about 1 bp wide, but its editing efficiency is significantly better than that of ABE9. This embodiment believes that enABE8e has strong editing efficiency and significantly narrower editing window, and is one of the most accurate adenine base editors.

[0181] The present embodiment also evaluates the creation of Sdd6 variants, and randomly selects 13 high-efficiency mutants (9 single mutants and 4 multiple mutants) for further evaluation of another 6 target sites in HEK293T cells. Among them, 11 mutants have at least a 10% increase in deamination activity at one or more target sites (Fig. 9, Fig. 11). Sc9 (F124K / K130T) is the most stable mutant, with an editing efficiency increase of about 12% to 35%. Combined with the predicted Sdd6 mutant-ssDNA complex structure by AlphaFold3, the present embodiment considers that Sc9 has the function of a high-fidelity variant, which may reduce the flexibility of the positively charged surface region, thereby possibly reducing off-target effects (Fig. 9). To effectively evaluate the editing specificity of Sdd6 mutants, the present embodiment uses the previous orthogonal R-loop-based test method, and the test shows that compared with the wild type, Sc9 shows higher target / de-target ratio (hereinafter referred to as "fidelity") at 8 target sites and 2 off-target sites. Specifically, the specificity is increased by about 30%, and the fidelity is increased by about 20%.

[0182] The present embodiment further characterizes its deamination activity in different cell types. Compared with Sdd6, its editing efficiency in HEK293T, HeLa and U2OS cells is increased by 19%, 84% and 31%, respectively. Compared with rAPOBEC1, its editing efficiency is increased by 44%, 195% and 127%, respectively (Fig. 11). In terms of editing specificity, analysis of 16 pairs of target / off-target pairs at 6 off-target sites in HeLa cells shows that the specificity of enCBE (Sc9-CBE) is significantly improved by about 28%, and the fidelity is improved by about 132% (Fig. 11), and the specificity is improved by 51% compared with the high-fidelity variant rAPOBEC1-YE1 (Fig. 11). The two mutations in Sc9 are based on AiCE single The screening found that this mutation lacks previous literature support and is located outside the well-characterized catalytic or binding region, and is difficult to predict by traditional methods, which highlights the powerful performance of the AiCE method, which provides a new basis for rational protein design.

[0183] Mitochondrial DNA mutations can cause a variety of genetic diseases, and DdCBEs developed by fusing transcription activator-like factor proteins with Ddds (double-stranded DNA engineering enzymes) can achieve C-to-T changes in mitochondrial DNA to correct pathogenic mutations, such as Leber's hereditary optic neuropathy and maternally inherited deafness. Compared with the commonly used double-stranded engineering enzyme DddA, its homolog Ddd1 can edit 5'-GC sequences, while DddA is difficult to edit this sequence, thereby expanding the potential applications of DdCBEs. However, although Ddd1 exhibits efficient deamination activity in the nuclear environment, its editing efficiency in the mitochondrial environment is significantly reduced (Figure 12). Therefore, the first embodiment first uses AlphaFold2 to predict the folding of Ddd1 and the DddA homolog complex (PDB ID 8E5E) resolved by cryo-electron microscopy as the target skeleton, and uses various prediction methods to nominate 138 single mutants and 15 multiple mutants (Figure 12), which are split at the N94 site and respectively constructed into paired TALE vectors, and their deamination efficiency is evaluated at two endogenous mitochondrial target sites in HEK293T cells. Target deep sequencing shows that 20 single mutants (12 predicted by AiCE single ) and 8 multiple mutants (7 predicted by AiCE multi ) are identified as high fitness mutants (Figure 12, Figure 13). The accuracy of AiCE single predictions for high fitness mutants is 27-32%, while the accuracy of other methods is between 0-23%. The top four groups of variants with the highest prediction accuracy were tested at two additional non-5'-GC target sites, and 21-32% of them showed higher deamination activity than DddA, with the best variant D71 (V61L) showing a 0.9-6.5-fold increase in deamination activity. These high fitness mutants retain strong deamination activity in different cell lines, and the first embodiment randomly selects 7 single mutants and 8 multiple mutants from the first round of screening and evaluates their deamination activity at the ND1.2 target site in HeLa cells (Figure 12). Among these mutants, the deamination activity of 4 single mutants and 7 multiple mutants is significantly improved.

[0184] Table 5, Base editing efficiency verification of double-stranded DNA deaminase (Ddd1) mutant

[0185] This example further demonstrates the compatibility of this example with other methods. This example combines the above most efficient point D-h11 with the reported PACE-M3 mutation group to create the current most efficient mitochondrial editor enDdd1, achieving a 14-fold improvement in base editing efficiency in the mitochondrial environment (Fig. 12, Fig. 13). enDdd1 also effectively edited 5'-GC sequences, with an efficiency improvement of about 40% compared to the six-mutant DddA11 obtained after several rounds of phage-assisted continuous evolution. This fully embodies the effectiveness and compatibility of this example.

[0186] At the same time, this example evaluates the mutation ability of AiCE to generate Ddd1 adapted to the nuclear environment. This example applies AiCE single -ProteinMPNN to base editing in the nuclear environment, and finds that 7 new mutations can improve editing ability in the nuclear environment, with an efficiency improvement of about 1.6 times (Fig. 12), which is currently impossible to solve by relying on human expert solutions.

[0187] This example applies AiCE to other complex tasks to further improve genome editing capabilities. This example selects a nuclear localization sequence and four proteins as engineering targets (Fig. 14), including: a nuclear localization sequence NLS that may improve the nuclear localization and efficiency of editing enzymes; V-type CRISPR nuclease LbCas12a and AcCas12n; type II CRISPR nuclease SpRYCas9; and the key reverse transcriptase in guided editing, Moloney murine leukemia virus (M-MLV) reverse transcriptase (RT). This example uses AiCE single -ProteinMPNN to design mutations for each target and evaluate their genome editing potential in HEK293T or HeLa cells (Fig. 14, Fig. 15), with an average prediction accuracy of 21%, 28%, 60%, 17%, and 60%, respectively, exceeding traditional protein engineering methods (Fig. 14). These five proteins differ greatly in size (from tens of residues to thousands of residues) and exhibit considerable structural heterogeneity (Fig. 14), but AiCE still performs excellently in these complex protein modification tasks.

[0188] Obtaining high-functioning mutants at low cost has been a goal of protein engineering. This difficulty stems from the multidimensional nature of protein structure and the complex relationship between sequence and function. Proteins essential to biological systems, such as enzymes, receptors, and channel proteins, often have dynamic structures that balance stability and flexibility to ensure their sustained functionality. This dynamic structure makes it very difficult, if not impossible, to reflect these conformational dynamics through a single static structure, and limits the effectiveness of traditional protein engineering methods. This study shows that AiCE, as a simple protein mutation design method, can utilize the distribution of reverse folding sequences to identify high fitness mutations. This method does not require additional model migration or training costs, is not limited by model size or architecture, and is superior to traditional protein engineering methods. The AiCE method is based on the assumption that a generalized protein reverse folding model inherently contains the natural dynamics of protein sequence and structure generation and distribution. These fundamental principles can be extracted and applied using contemporary genetic theory. Although a general protein model is used in AiCE rather than a model specific to a particular protein family, the generated high fitness mutations and their distribution still differ for the target protein backbone (Figure 16). Furthermore, this example demonstrates that the high fitness mutations identified by AiCE and high frequency mutations often exhibit unconventional and counterintuitive forms in terms of mutation position and type (Figure 16). This can explain the compatibility of AiCE with other methods and its potential as a new option for structure-guided protein rational design.

[0189] The method for verifying the editing efficiency of the above-mentioned engineered enzyme mutant comprises the following steps:

[0190] I. Vector construction:

[0191] TALE system-mediated base editor plasmids, including Ddd1, TALE array, UGI (uracil glycosylase inhibitor), and mitochondrial targeting signal, were optimized for codon usage to facilitate human cell expression. These were commercially synthesized by GenScript. Subsequently, the above components were cloned into vectors pCMV-TALE-L-JAK2-Ddd9-N (Addgene #204853) and pCMV-TALE-R-JAK2-Ddd9-C (Addgene #204854). Single mutations in Ddd1 predicted by the AiCE method were introduced into PCR fragments by primer design and cloned into the corresponding TALE array-containing backbone vectors using Uniclone One Step Seamless Cloning Kit (Genesand).

[0192] Single or combined amino acid mutations of Sdd6 predicted by AiCE were introduced into PCR fragments by primer design and cloned into p2T-CMV-miniSdd6-BE4max-BlastR (Addgene #204850) vector backbone using Uniclone One Step Seamless Cloning Kit (Genesand). The 2x UGI sequence was removed from within the pnCas9-miniSdd6-PBE vector to construct the p2T-CMV-miniSdd6-BE4max-BlastR-delUGI vector backbone. The TadA8e sequence was cloned from pTPH413 (Addgene #185728). AiCE variants with wild-type sequence were then constructed into the p2T-CMV-miniSdd6-BE4max-BlastR-delUGI vector backbone using the methods described above.

[0193] The phU6 vector (Addgene #53188) was used to express sgRNA. The construction of sgRNA vectors was achieved by using circular PCR, in which the spacer sequence was integrated into the primers. The F and R primers have a 20 bp overlap at the 5' end. The complete vector sequence was amplified using PCR and the new spacer sequence was integrated. After amplification, the template plasmid was digested with Dpnl restriction enzyme, which is specific for DNA sequences with methylation modification. The digested PCR product was then transformed into Fast-T1 competent cells (Vazyme). PCR amplification was performed using 2x Phanta Max Master Mix (Vazyme).

[0194] II. Human cell transfection and DNA extraction:

[0195] Transfection was performed 16-24 hours after seeding. In transfection experiments using TALE-mediated base editing systems, 0.4 μΐ of Lipofectamine 2000 (ThermoFisher Scientific), 150 ng of TALE-L vector, 150 ng of TALE-R vector, and 10 ng of green fluorescent protein were co-incubated and transfected into cells.

[0196] In transfection experiments involving CRISPR-Cas-mediated base editing systems, 0.4 μΐ of Lipofectamine 2000, 300 ng of enzyme-containing vector, 100 ng of sgRNA expression vector, and 10 ng of green fluorescent protein were co-incubated and transfected into cells.

[0197] For off-target effect check using R-loop analysis, four vectors were co-incubated with 0.4 μΐ of Lipofectamine 2000 and transfected. A total of 150 ng of pCMV-deam-BE4max vector, 150 ng of pCMV-nSaCas9 vector, and two corresponding sgRNA vectors (50 ng each), and 10 ng of green fluorescent protein were used. After a 72-hour incubation period, the HEK293T cells were washed with PBS, and then genomic DNA was extracted using the Triumfi Mouse Tissue Direct PCR Kit (Genesand) with lysis buffer and proteinase K.

[0198] III. Amplicon sequencing analysis

[0199] A series of primers with barcodes at the 5' end were designed to amplify the target sequences. The amplicons were purified using the Thermo Scientific GeneJET Kit (Thermo Fisher Scientific) and quantified using a NanoDrop 2000 spectrophotometer (Thermo Fisher Scientific). Equal amounts of PCR products were pooled and then subjected to commercial sequencing using NGS (GENEWIZ). For target sites that were difficult to amplify, a round of primers should be designed for 500 bp amplification. The resulting product was diluted 20-fold, and then 1 μΐ of the diluted product was used as a template for nested PCR using primers containing barcode sequences.

[0200] In the amplicon sequencing analysis, the cleaned sequencing data was first split according to the sequencing primers. The following methods were used: (1) Data splitting and merging. The pooled sequencing data was split into individual treatments according to the sample-specific adapter sequence. The forward and reverse reads of each treatment were merged into a single read using the FLASH tool (v1.2.11) with the parameter settings "-m 5 -M 150", ensuring that the minimum overlap of double-end reads was 5 bp and the maximum overlap was 150 bp. (2) Sequence alignment and SNP extraction. The merged reads were compared with the wild-type reference sequence. The alignment strategy included determining the sequence window for viewing editing events by taking the previous and subsequent 7 base pairs, using the reference sequence as the index sequence, aligning the index sequence with the merged read, if there were two end index sequences in the merged read, extracting the middle sequence, and then aligning it with the window sequence to identify and extract SNP information. (3) Editing event counting. According to the extracted SNP information, the editing efficiency, type, and other related indicators of all base editing events were calculated.

[0201] The AI-based enzyme engineering system provided by the present application is described below, and the AI-based enzyme engineering system described below can be referred to in correspondence with the AI-based enzyme engineering method described above.

[0202] Referring to FIG. 17, the AI-based enzyme engineering system provided by the present application can include:

[0203] The data acquisition module is configured to acquire structural information of the enzyme to be reformed.

[0204] The reverse folding module is configured to obtain a plurality of protein sequences compatible in structure based on a given skeleton structure of the enzyme to be reformed according to the structural information of the enzyme to be reformed through a protein reverse folding model.

[0205] The processing module is configured to obtain a mutation amino acid frequency and a wild-type amino acid frequency at each sequence position according to the plurality of protein sequences compatible in structure.

[0206] The single mutation obtaining module is configured to obtain a single mutation variant set of the enzyme to be reformed according to the mutation amino acid frequency and the wild-type amino acid frequency at each sequence position and the structural information of the enzyme to be reformed in combination with a first creation strategy and a second creation strategy.

[0207] The mutation combination obtaining module is configured to obtain a mutation combination variant set of the enzyme to be reformed according to the plurality of protein sequences compatible in structure in combination with a third creation strategy.

[0208] FIG. 18 illustrates an entity structure diagram of an electronic device, as shown in FIG. 18, which can include a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 complete mutual communication through the communications bus 840. The processor 810 can invoke a logical instruction in the memory 830 to execute an AI-based enzyme engineering method, which includes:

[0209] Acquiring structural information of the enzyme to be reformed.

[0210] Obtaining a plurality of protein sequences compatible in structure based on a given skeleton structure of the enzyme to be reformed according to the structural information of the enzyme to be reformed through a protein reverse folding model.

[0211] Obtaining a mutation amino acid frequency and a wild-type amino acid frequency at each sequence position according to the plurality of protein sequences compatible in structure.

[0212] According to the mutation amino acid frequency and the wild type amino acid frequency at each sequence position and the structure information of the engineering enzyme to be modified, combining the first creation strategy and the second creation strategy, a single mutation variant set of the engineering enzyme to be modified is obtained;

[0213] According to the structure compatible multiple protein sequences, combining the third creation strategy, a mutation combination variant set of the engineering enzyme to be modified is obtained.

[0214] In addition, the logical instructions in the memory 830 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the part of the prior art or the part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0215] On the other hand, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the artificial intelligence based enzyme engineering method provided by the above method, which comprises:

[0216] Obtaining the structure information of the engineering enzyme to be modified;

[0217] According to the structure information of the engineering enzyme to be modified, a plurality of structure compatible protein sequences are obtained based on the given skeleton structure of the engineering enzyme to be modified through a protein unfolding model;

[0218] According to the structure compatible multiple protein sequences, the mutation amino acid frequency and the wild type amino acid frequency at each sequence position are obtained;

[0219] According to the mutation amino acid frequency and the wild type amino acid frequency at each sequence position and the structure information of the engineering enzyme to be modified, combining the first creation strategy and the second creation strategy, a single mutation variant set of the engineering enzyme to be modified is obtained;

[0220] According to the structure compatible multiple protein sequences, combining the third creation strategy, a mutation combination variant set of the engineering enzyme to be modified is obtained.

[0221] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the artificial intelligence-based enzyme engineering method provided by each of the above methods, the method comprising:

[0222] obtaining structural information of the engineering enzyme to be reformed;

[0223] According to the structural information of the engineering enzyme to be reformed, a plurality of protein sequences compatible in structure are obtained based on a given backbone structure of the engineering enzyme to be reformed through a protein unfolding model;

[0224] According to the plurality of protein sequences compatible in structure, a mutation amino acid frequency and a wild-type amino acid frequency at each sequence position are obtained;

[0225] According to the mutation amino acid frequency and the wild-type amino acid frequency at each sequence position and the structural information of the engineering enzyme to be reformed, a single-mutation variant set of the engineering enzyme to be reformed is obtained in combination with the first creation strategy and the second creation strategy;

[0226] According to the plurality of protein sequences compatible in structure, a mutation combination variant set of the engineering enzyme to be reformed is obtained in combination with the third creation strategy.

[0227] The device embodiments described above are only schematic, wherein the units illustrated as separate components can or can not be physically separate, and the components illustrated as units can or can not be physical units, i.e., can be located in one place or distributed on a plurality of network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement without creative labor.

[0228] From the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software plus a necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0229] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the same; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. Industrial applicability

[0230] The application provides an enzyme engineering method, system, device and medium based on artificial intelligence. Based on the population evolution theory, according to the structure information of the to-be-transformed engineering enzyme, a plurality of protein sequences compatible in structure are obtained through a protein unfolding model based on the given skeleton structure of the to-be-transformed engineering enzyme, and a single-mutation variant set and a mutation combination variant set of the to-be-transformed engineering enzyme are obtained efficiently and at low cost by combining a first creation strategy, a second creation strategy and a third creation strategy. The engineering enzyme variant developed according to the application can greatly expand the base editing toolbox, realize all-round improvement of the base editing engineering properties, realize precise and efficient base editing, effectively expand the application scope and safety of base editing, and make the base editing strategy tend to be customized.

[0231] In addition, the application does not need to train a special artificial intelligence model, does not need to be guided by human expert knowledge, and can generate efficient mutations through only one round, greatly improving the efficiency of enzyme engineering. Secondly, the application has strong universality, is suitable for optimization of multiple different family engineering enzymes, and is compatible with traditional methods. Thirdly, the application provides many counterintuitive but effective single mutants and mutation combination guidance that can potentially avoid epistatic effects based on the unfolding of structure information, evolution coupling information and structure-guided learning, which helps to explore the enzyme sequence and function generation space. The application not only guides the rapid optimization of enzymes, but also provides deep insights into the protein distribution and generation space in the natural state.

[0232] Cross-reference to related applications

[0233] This application claims priority to Chinese Patent Application No. 202411307848.9, filed on September 19, 2024, and Chinese Patent Application No. 202411351036.4, filed on September 26, 2024, the contents of which are hereby incorporated by reference in their entirety.

Claims

1. An artificial intelligence-based enzyme engineering method, characterized by, The method comprises the following steps: obtaining structural information of an engineering enzyme to be reformed; obtaining a plurality of protein sequences compatible in structure based on a given skeleton structure of the engineering enzyme to be reformed through a protein reverse folding model according to the structural information of the engineering enzyme to be reformed; obtaining a mutation amino acid frequency and a wild-type amino acid frequency at each sequence position according to the plurality of protein sequences compatible in structure; obtaining a single-mutation variant set of the engineering enzyme to be reformed by combining a first creation strategy and a second creation strategy according to the mutation amino acid frequency and the wild-type amino acid frequency at each sequence position and the structural information of the engineering enzyme to be reformed; obtaining a mutation combination variant set of the engineering enzyme to be reformed by combining a third creation strategy according to the plurality of protein sequences of the engineering enzyme to be reformed.

2. The artificial intelligence-based enzyme engineering method of claim 1, wherein, The step of obtaining a mutation amino acid frequency and a wild-type amino acid frequency at each sequence position according to the plurality of protein sequences compatible in structure comprises the following steps: obtaining a frequency of each amino acid at each sequence position through a first expression according to the plurality of protein sequences compatible in structure; obtaining a wild-type amino acid frequency and a mutation amino acid frequency at each sequence position through a second expression according to the plurality of protein sequences compatible in structure. Preferably, the first expression is: In the first expression, x represents a preset common amino acid type, f x (i) represents an amino acid x i In the second expression, x represents a preset common amino acid type, f i In the third expression, x represents a preset common amino acid type, f In the fourth expression, x represents a preset common amino acid type, f In the fifth expression, x represents a preset common amino acid type, f In the sixth expression, x represents a preset common amino acid type, f In the seventh expression, x represents a preset common amino acid type, f In the eighth expression, x represents a preset common amino acid type, f In the ninth expression, x represents a preset common amino acid type, f In the tenth expression, x represents a preset common amino acid type, f In the eleventh expression, x represents a preset common amino acid type, f In the twelfth expression, x represents a preset common amino acid type, f In the thirteenth expression, x represents a preset common amino acid type, f In the fourteenth expression, x represents a preset common amino acid type, f In the fifteenth expression, x represents a preset common amino acid type, f In the sixteenth expression, x represents a preset common amino acid type, f In the seventeenth expression, x represents Preferably, the second expression is: In the second expression, f wt (i) and f mut (i) represent the frequency of the wild-type amino acid wt and the mutant amino acid mut, respectively, at sequence position i in the structurally compatible plurality of protein sequences, x i represents the amino acid at the i-th sequence position of the j-th protein sequence in the structurally compatible plurality of protein sequences, M represents the total number of sequences of the structurally compatible plurality of protein sequences, and 1{·} represents the indicator function, which takes the value 1 when the condition is true and 0 otherwise.

3. The artificial intelligence-based enzyme engineering method of claim 2, wherein, The step of obtaining a single-mutation variant set of the engineering enzyme to be reformed by combining a first creation strategy and a second creation strategy according to the mutation amino acid frequency and the wild-type amino acid frequency at each sequence position and the structural information of the engineering enzyme to be reformed comprises the following steps: obtaining a highest-frequency mutation amino acid with a frequency higher than the wild-type amino acid frequency and a highest frequency at each sequence position according to the mutation amino acid frequency and the wild-type amino acid frequency at each sequence position; comparing the frequency of the highest-frequency mutation amino acid at each sequence position with a first preset threshold value respectively to obtain a first single-mutation variant set of the engineering enzyme to be reformed.

4. The artificial intelligence-based enzyme engineering method of claim 3, wherein, The expression for the frequency of the most frequent mutated amino acid is: f max (i) = max{f mut (i) | f mut (i) > f wt (i)} f max (i) denotes the frequency of the most frequent mutant amino acid at sequence position i in the structurally compatible plurality of protein sequences, f mut (i) denotes the frequency of the mutant amino acid mut at sequence position i in the structurally compatible plurality of protein sequences, f wt (i) denotes the frequency of the wild-type amino acid wt at sequence position i in the structurally compatible plurality of protein sequences; Preferably, the step of comparing the frequency of the highest-frequency mutation amino acid at each sequence position with a first preset threshold value respectively to obtain a first single-mutation variant set of the engineering enzyme to be reformed comprises the following steps: According to the expression of the first creation strategy, the frequency of the highest frequency mutation amino acid at each sequence position of the plurality of structurally compatible protein sequences is compared with the first preset threshold respectively, to obtain a first single-mutation variant set of the to-be-transformed engineered enzyme, wherein the expression of the first creation strategy is: In the expression of the first creation strategy, f max (i) represents the frequency of the most frequent mutation amino acid at sequence position i in the plurality of protein sequences compatible with the structure, and β represents a first preset threshold value, representing the first single-mutation variant set of the engineering enzyme to be reformed, and the first single-mutation variant set comprises the highest-frequency mutation amino acid with a frequency not lower than the first preset threshold value and a sequence position where the highest-frequency mutation amino acid is located.

5. The artificial intelligence-based enzyme engineering method of claim 4, wherein the artificial intelligence-based enzyme engineering method is characterized by, The step of obtaining a single-mutation variant set of the engineering enzyme to be reformed by combining a first creation strategy and a second creation strategy according to the mutation amino acid frequency and the wild-type amino acid frequency at each sequence position and the structural information of the engineering enzyme to be reformed comprises the following steps: obtaining a flexible region of the engineering enzyme to be reformed by using a protein secondary structure prediction method according to the structural information of the engineering enzyme to be reformed; judging whether the highest-frequency mutation amino acid at each sequence position is located in the flexible region of the engineering enzyme to be reformed to obtain a highest-frequency mutation amino acid located in the flexible region of the engineering enzyme to be reformed; comparing the frequency of the highest-frequency mutation amino acid located in the flexible region of the engineering enzyme to be reformed with a second preset threshold value to obtain a second single-mutation variant set of the engineering enzyme to be reformed; combining the first single-mutation variant set and the second single-mutation variant set of the engineering enzyme to be reformed to obtain the single-mutation variant set of the engineering enzyme to be reformed.

6. The artificial intelligence-based enzyme engineering method of claim 5, wherein, The frequency of the highest frequency mutation amino acid located in the flexible region of the to-be-constructed engineered enzyme is compared with a second preset threshold, to obtain a second single-mutation variant set of the to-be-constructed engineered enzyme, including: According to the expression of the second creation strategy, the frequency of the highest frequency mutation amino acid located in the flexible region of the to-be- modified engineered enzyme is compared with a second preset threshold to obtain a second single-mutation variant set of the to-be-modified engineered enzyme, wherein the expression of the second creation strategy is: In the expression of the second creation strategy, f max (i) represents the frequency of the highest frequency mutation amino acid at sequence position i, γ represents a second preset threshold, and is_flexible(i) = 1 represents that the highest frequency mutation amino acid at sequence position i is located in the flexible region of the to-be-transformed engineered enzyme, and the condition is true, The second single-mutation variant set of the to-be-constructed engineered enzyme, the second single-mutation variant set including sequence positions and types of the highest frequency mutation amino acid located in the flexible region of the to-be-constructed engineered enzyme and having a frequency greater than the second preset threshold; Preferably, the expression of the set of single mutant variants of the engineered enzyme to be improved is: In the expression of the set of single mutant variants of the engineered enzyme to be modified, denotes a first set of single mutant variants of an engineered enzyme to be improved, a second set of single mutant variants representing engineered enzymes, The single-mutation variant set of the to-be-constructed engineered enzyme.

7. The artificial intelligence-based enzyme engineering method according to claim 6, characterized by, The mutation combination variant set of the to-be-constructed engineered enzyme is obtained according to the plurality of protein sequences of the to-be-constructed engineered enzyme and in combination with the third creation strategy, including: The plurality of protein sequences compatible in structure are reversed into a plurality of pseudo DNA sequences by using a codon correspondence relationship, and a deviation of an observed haplotype frequency from an expected haplotype frequency is obtained according to a first expression, denoted as D; According to the deviation of the observed haplotype frequency from the expected haplotype frequency being an independent distribution, and according to the second expression, the proportion of variance that can be explained by one allele of the to-be- modified engineered enzyme by another allele is obtained, denoted as r 2 ; In combination with the third creation strategy, a mutation combination variant in which a proportion of variance explained by an overall allele pair is higher than a first preset value is screened out from the to-be-constructed engineered enzyme according to a third expression, to obtain a first mutation combination variant set of the to-be-constructed engineered enzyme; Preferably, the first expression is: D = P AB - P A P B , In the first expression, P A , P B , and P AB represent the population frequencies of alleles A, B, and heterozygote AB, respectively. Preferably, the second expression is: In the second expression, D represents the deviation of the observed haplotype frequency from the expected haplotype frequency assuming independent distribution, P A , P B , and P AB represent the population frequencies of alleles A, B, and heterozygote AB, respectively. Preferably, the second mutation combination variant set of the to-be-constructed engineered enzyme is obtained according to the plurality of protein sequences compatible in structure and in combination with the third creation strategy, including: According to the structural compatibility of the plurality of protein sequences, using a fourth expression, the logarithmic probability of the occurrence of an amino acid a at position i in the sequence, relative to the overall probability of the occurrence of the amino acid a in the entire sequence φ i , and the logarithmic probability of the occurrence of an amino acid b at position j in the sequence, relative to the overall probability of the occurrence of the amino acid b in the entire sequence φ j ; According to the structure compatible multiple protein sequences, using the fifth expression, the combined mutation coupling score between each pair of amino acids is obtained A weighted joint mutation coupling score between two amino acids is obtained according to a joint mutation coupling score between two amino acids, a total probability of occurrence of the amino acid a in the entire sequence, a total probability of occurrence of the amino acid b in the entire sequence, and according to a sixth expression; In combination with the third creation strategy, a mutation combination variant in which a statistical coupling score is greater than or equal to a second preset value is screened out according to the third expression, to obtain a second mutation combination variant set of the to-be-constructed engineered enzyme; Preferably, the fourth expression is: In the fourth expression, f ai and f bi represent the frequencies of amino acids a and b at positions i and j in the multiple sequence alignment, respectively, q a and q b represent the background frequencies of amino acids a and b, respectively. Preferably, the fifth expression is: In the fifth expression, and respectively represent the frequency of amino acids a and b at position i and position j, respectively, A probability that the i-th position is the amino acid a and the j-th position is the amino acid b; Preferably, the sixth expression is: In the sixth expression, A joint mutation coupling score between two amino acids a and b at the i-th position and the j-th position; Preferably, the third expression is: In the third expression, representing a set of mutations A total number of sites in the matrix, score(i, j) represents a score between the i-th position and the j-th position, including a variance proportion score and a joint mutation coupling score; Preferably, the first mutant combination variant set of the engineered enzyme to be modified is expressed as: In the expression of the first mutant combinatorial variant set of engineered enzymes to be modified, a first set of mutant combinatorial variants representing engineered enzymes, Representing combinations of mutations in multiple pseudosequences The variance proportion score of the matrix; Preferably, the second mutant variant set of the engineered enzyme to be modified is expressed by the formula: the expression of the second mutant variant set of engineered enzymes to be modified, a second set of combinatorial collections of mutant variants representing engineered enzymes, indicates a combination of mutations The statistical coupling score of the matrix, scorepercentile90 represents a 90th percentile of a global coupling score of the amino acid sequence; Preferably, the expression of the final set of mutant combinatorial variants of the engineered enzyme to be improved is: The expression of the final set of mutant combinatorial variants of the engineered enzyme to be modified is, represents the final set of mutant combinatorial variants of the engineered enzyme to be improved, a first set of mutant combinatorial variants representing engineered enzymes, The second mutation variant combination set of the to-be-constructed engineered enzyme.

8. An artificial intelligence-based enzyme engineering system, characterized by, The data acquisition module is configured to acquire structural information of a to-be-constructed engineered enzyme; The inverse folding module is configured to obtain a plurality of protein sequences compatible in structure based on a given skeleton structure of the to-be-constructed engineered enzyme by a protein inverse folding model according to the structural information of the to-be-constructed engineered enzyme; The processing module is configured to obtain a mutation amino acid frequency and a wild-type amino acid frequency at each sequence position according to the plurality of protein sequences compatible in structure; The single-mutation obtaining module is configured to obtain a single-mutation variant set of the to-be-constructed engineered enzyme according to the mutation amino acid frequency and the wild-type amino acid frequency at each sequence position and the structural information of the to-be-constructed engineered enzyme in combination with a first creation strategy and a second creation strategy; ​ The mutation combination module is configured to obtain, according to structural information of the to-be-reformed engineered enzyme, a plurality of protein sequences, and in combination with a third creation strategy, a set of mutation combination variants of the to-be-reformed engineered enzyme.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the artificial intelligence-based enzyme engineering method according to any one of claims 1-7 when executing the program.

10. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program implements the artificial intelligence-based enzyme engineering method according to any one of claims 1-7 when executed by the processor.

Citation Information

Patent Citations

  • M2 group-based candidate causal mutation site gene localization method

    CN113130005A

  • Protein structure prediction method and device, electronic equipment and storage medium

    CN117558337A

  • Protein engineering workflow using a generative model of protein families

    US20240282404A1

Cited By

  • Monte Carlo optimized protein thermal stability screening method fusing evolutionary information

    CN122050499A