A method for de novo design of uracil-dna glycosylase based on physical priors and deep learning
By employing a method based on physical priors and deep learning, catalytic functions are precisely embedded into novel protein backbones, solving the problems of low enzyme catalytic efficiency and structural redundancy in existing technologies, and enabling the design of highly efficient gene editing tools.
Patent Information
- Application Number
- CN202610911977.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2046-06-24
AI Technical Summary
Existing gene editing tools rely on the reuse of natural enzymes, which leads to structural redundancy and substrate-specific immobilization issues, making it difficult to achieve efficient delivery and significant topological remodeling. Furthermore, existing computational design frameworks lack precision control and verification systems for complex enzyme catalysis processes.
Using a physical prior and deep learning approach, a set of functional motifs is defined by homology sequence retrieval and clustering. A new protein backbone is generated by combining an all-atom diffusion model. Furthermore, by employing rigid body transformation and multi-round iterative optimization, an orthogonal prediction model and a Pareto optimal strategy are introduced to screen candidate enzymes.
This technology enables precise embedding of catalytic functions into novel protein backbones, with sub-angstrom level precision control of active sites. This improves the catalytic efficiency and stability of enzymes, breaks through the path dependence on the natural backbone, and provides a novel programmable gene editing tool.
Smart Images

Figure CN122435986B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of protein engineering and computational biology, and in particular to a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning. Background Technology
[0002] Precision genome editing technology has become a core driving force for crop trait improvement and human disease treatment. Among these advancements, the emergence of glycosylation enzyme base editors (GBEs) has broken through the limitations of classical base editors, which can only catalyze conversion mutations, enabling transversion mutations and providing a crucial tool for diverse genetic manipulations. In this system, uracil-DNA glycosylation enzymes (UNGs) are the core components performing key catalytic functions. However, existing gene editing tools mainly rely on the reuse and modification of natural enzymes (such as human UNGs). These natural backbones are constrained by evolutionary path dependence, often exhibiting limitations such as structural redundancy and substrate-specific immobilization. Reprogramming natural enzymes using traditional engineering methods (such as directed evolution or rational mutation) not only faces bottlenecks of high trial-and-error costs and low success rates but also struggles to achieve significant topological reshaping of proteins, failing to meet the stringent requirements of in vivo efficient delivery systems for enzyme miniaturization. Therefore, overcoming the limitations of natural evolution and designing novel, compact, ultra-stable, and functionally programmable DNA glycosylation enzymes from scratch has become a critical problem urgently needing to be solved in this field.
[0003] To break free from dependence on natural protein backbones, de novo protein design technology emerged. Early de novo protein design was primarily driven by computational simulations based on physicochemical principles (such as software like Rosetta), which treated design as an energy optimization task. However, physics-based methods face significant challenges when dealing with highly complex tasks: first, the exponential explosion of computational complexity limits their application in designing large molecular complexes due to the need for precise conformational sampling at the atomic scale; second, the approximate error of the energy function makes it difficult to perfectly simulate complex biophysical environments, and simply relying on minimizing physical energy often fails to accurately fix sub-angstrom catalytic geometry, resulting in designed enzymes with activity far lower than natural enzymes, and typically requiring multiple rounds of expensive directed evolution experiments for subsequent optimization.
[0004] In recent years, deep learning and generative AI (as shown in figure neural networks) have made breakthroughs in protein design, enabling the direct generation of novel protein backbones and sequences that meet specific geometric conditions. However, when the task focuses on the complex enzyme catalysis level, the limitations of purely generative AI methods remain significant. Current deep learning networks typically focus more on enabling proteins to have more stable overall folding capabilities rather than higher local catalytic efficiency. This is especially true for DNA glycosylation enzymes, whose catalytic process is extremely complex. The active site needs to achieve sub-angstrom-level geometric matching with the flipped-outbase conformation, while simultaneously establishing precise electrostatic and shape complementarity with the large, twisted, and heavily charged DNA backbone. This requires the active pocket to have extremely high pre-organization within the novel protein backbone. At present, single AI models often struggle to accurately control the intricate hydrogen bond network and electrostatic arrangement of the active site, and also find it difficult to effectively stabilize the transition states during the catalytic process.
[0005] In summary, current computational design frameworks still lack mature systems for controlling and validating computational precision when dealing with enzymes that undergo complex substrate conformational changes (such as UNG). How to precisely embed complex and discontinuous catalytic motifs into novel protein topologies, while ensuring overall scaffold diversity and folding stability, and achieving sub-angstrom level precision control of the active pocket, remains an unsolved technical challenge in the interdisciplinary field of protein engineering and computational biology. Summary of the Invention
[0006] This invention provides a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning, in order to overcome the shortcomings of existing technologies.
[0007] This invention provides a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning, comprising the following steps: S1. Using the crystal structure of the complex of natural uracil-DNA glycosylation enzyme and substrate analog as a natural template, a set of functional motifs containing multiple non-continuous peptide segments is defined through homology sequence retrieval and clustering, multiple sequence alignment and conservation analysis, and fixed functional motif definition. The set of functional motifs is divided into a catalytic core layer, a DNA binding interface layer and a structural support layer according to their functional roles. The types of amino acid residues in the set of functional motifs and their main chain and side chain conformations are completely fixed in subsequent design steps as geometric constraints for generating new protein backbones. S2. Input the three-dimensional coordinates of the functional motif set and the substrate analogue as conditions into the all-atom diffusion model, set a variable-length region to be generated between fixed functional motif fragments, and introduce guiding potential energy for ligand contact during the denoising process of the all-atom diffusion model to drive the all-atom diffusion model to generate a protein topological backbone that is geometrically complementary to and tightly packed with the substrate analogue and functional motif, as a candidate protein backbone; S3. For each candidate protein backbone generated in S2, the DNA double helix in the natural template is transplanted to the binding site of the candidate protein backbone using a rigid body transformation algorithm to obtain the corresponding protein-DNA complex model. Based on the protein-DNA complex model, the side chain conformation of the preset key amino acid residues in the fixed functional motif is replaced and repaired. S4. For each protein-DNA complex model obtained in S3, perform multiple rounds of alternating iterative optimization based on deep learning sequence generation and physical force field structure refinement to obtain a set of candidate sequence-structure pairs. S5. Using an orthogonal prediction model and a Pareto optimal strategy, a de novo uracil-DNA glycosylation enzyme candidate set is selected from the candidate sequence-structure pair set.
[0008] According to the present invention, a de novo design method for uracil-DNA glycosylation enzyme based on physical priors and deep learning is provided, wherein the natural uracil-DNA glycosylation enzyme is a human uracil-DNA glycosylation enzyme, and the PDB number of the crystal structure of the complex is 1EMH.
[0009] The definition of the functional motif set must satisfy at least one of the following: (1) the consistency frequency is greater than 90% in multiple sequence alignment of homologous sequences; (2) it is annotated as an active site or binding site in the UniProt database; (3) it has a key functional role verified by mutation experiments; (4) the spatial distance between it and any heavy atom of the substrate analog in the crystal structure is within 8.0 Å.
[0010] According to the present invention, a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning is provided, wherein the all-atom diffusion model is RFdiffusionAA; The overall length of the candidate protein backbone is set to be between 215 and 235 amino acid residues.
[0011] According to the present invention, a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning is provided, wherein step S3 may include: S31. For each candidate protein backbone generated in S2, use a rigid body transformation algorithm (e.g., Kabsch algorithm) to calculate the transformation matrix that minimizes the root mean square deviation between the Cα atom of the fixed functional motif in the candidate protein backbone and the corresponding atom in the natural template. S32. Using the transformation matrix obtained in S31, the DNA double helix in the natural template is transplanted to the binding site of the candidate protein backbone, thereby constructing a complete protein-DNA complex model. S33. Based on the protein-DNA complex model, the side chain conformation of the key amino acid residues in the fixed functional motif is replaced with the corresponding conformation in the natural template.
[0012] According to the present invention, a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning is provided, wherein the key amino acid residues include tryptophan.
[0013] According to the present invention, a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning is provided, wherein step S4 may include: S41. In each iteration, a graph neural network model that is aware of the ligand environment is first used to generate a batch of candidate amino acid sequences for the current candidate protein backbone. During the generation of candidate amino acid sequences, the types of amino acids in the functional motif set remain fixed. S42. Using a force field containing protein physicochemical energy terms and custom catalytic geometry constraints, perform side chain optimization and backbone energy minimization on the complex structure corresponding to each candidate amino acid sequence. In this process, the degrees of freedom of all atoms in the functional motif set are frozen by setting an atom movement control file. S43. Based on the total energy and geometric constraint deviation of the complex structure, the Pareto strategy is used to select the best-performing sequence-structure pair as the input for the next iteration. S44. As the iteration rounds progress, adjust the noise level parameters used in the graph neural network model to transition from high noise to low noise, so as to achieve a self-consistent process of sequence and structure from coarse exploration to high-precision convergence.
[0014] According to the present invention, a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning is provided, wherein the graph neural network model is LigandMPNN; The force field, which includes protein physicochemical energy terms and custom catalytic geometry constraints, is a RosettaFastRelax program combined with an enzyme design constraint file.
[0015] According to the present invention, a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning is provided. The alternating iterative optimization in step S4 consists of three rounds. The LigandMPNN model used in the first round of iteration is trained under 0.20 Å Gaussian noise, and the LigandMPNN model used in the second and third rounds of iteration is trained under 0.10 Å Gaussian noise.
[0016] According to the present invention, a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning is provided, wherein step S5 may include: S51. For the set of candidate sequence-structure pairs, the first prediction model is used to predict the protein-DNA complex containing the natural substrate, and the second prediction model is used to independently predict the protein monomers. S52. Calculate the evaluation index data based on the two prediction results obtained in S51. The evaluation index covers the following four dimensions: complex binding confidence, monomer folding stability, structural fidelity, and physicochemical compatibility. S53. Based on the evaluation index data, a qualifying screening is performed on the candidate sequence-structure pair set to eliminate designs that do not meet the preset basic requirements. S54. For the candidate designs that have passed the qualifying round, select at least two mutually orthogonal evaluation indicators as optimization objectives, construct a multidimensional objective space, and identify a set of non-dominated solutions by calculating the Pareto front in the multidimensional objective space, which constitutes the candidate set of de novo uracil-DNA glycosylation enzymes.
[0017] According to the present invention, a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning is provided, wherein the first prediction model is AlphaFold3 (AF3) and the second prediction model is ESMFold; the natural substrate is 2'-deoxyuridine (dU).
[0018] This invention provides a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning. The evaluation metrics include any one or any combination of the following: AF3 interface prediction template modeling score (ipTM), minimum interaction prediction alignment error score (ipSAE). min Global Predicted Local Distance Difference Test Score (pLDDT), functional motif pLDDT, ESMFold monomer pLDDT, ESMFold monomer pTM, functional motif Cα RMSD, global Cα RMSD, and shape complementarity of the protein-DNA interface ( S c ); The default basic requirements include any one or any combination of the following: ipTM ≥ 0.96, ipSAEmin ≥ 0.74, global Cα RMSD ≤ 1.8 Å, functional motif Cα RMSD ≤ 1.0 Å.
[0019] The present invention also provides a uracil-DNA glycosylase obtained by the de novo design method of uracil-DNA glycosylase based on physical prior and deep learning as described above, characterized in that the uracil-DNA glycosylase has an active pocket that matches the catalytic core of natural human uracil-DNA glycosylase with sub-angstrom precision, and its amino acid sequence has a global sequence identity of less than 40% with the natural template.
[0020] According to the present invention, a uracil-DNA glycosylation enzyme is provided, wherein the amino acid sequence of the uracil-DNA glycosylation enzyme is selected from the group consisting of SEQ ID NO:1-23.
[0021] The present invention also provides an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the computer program to implement any of the above-described methods for de novo design of uracil-DNA glycosylation enzymes based on physical priors and deep learning.
[0022] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning as described above.
[0023] The present invention also provides a computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute any of the above-described methods for de novo design of uracil-DNA glycosylation enzymes based on physical priors and deep learning.
[0024] The present invention provides a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning, which can bring at least the following technical effects: By introducing evolutionarily conserved physical priors encompassing three levels of catalysis, binding, and support, this invention can precisely extract the complex and discontinuous functional core of UNG from its natural framework and seamlessly embed it as a rigid constraint into a novel protein framework generated by AI with high topological diversity, effectively overcoming the path dependence of natural evolution.
[0025] By employing a sequence-structure collaborative iterative optimization strategy, combined with the ability to finely model local side-chain stacking and electrostatic arrangement using physical force fields, this invention can control the geometric conformational deviation of active sites to the sub-angstrom level while ensuring the overall folding stability of the backbone. Orthogonal prediction results confirm that the catalytic pocket of the designed enzyme can accurately reproduce the key hydrogen bond network with the substrate.
[0026] This invention overcomes the limitations of single-index evaluation by constructing a rigorous validation funnel that combines orthogonal model (AlphaFold3 and ESMFold) prediction with multi-dimensional indicators (confidence, folding stability, and structural fidelity), and introduces a Pareto optimal strategy for rational decision-making. This system significantly improves the expected success rate from computational design to experimental validation, providing a systematic precision control and validation standard for highly complex enzymatic computational design.
[0027] The de novo designed UNG candidates ultimately selected showed significant sequence differences from the natural template (global identity less than 40%) and significant structural deviations from their closest natural homologs (average RMSD approximately 3.0 Å). This indicates that the present invention successfully breaks away from the fine-tuning paradigm of the natural backbone, enabling the creation of novel functional molecules located in an independent sequence-structure space, laying a solid technical foundation for the development of novel, programmable gene-editing enzyme tools with independent intellectual property rights. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0029] Figure 1 This is a schematic diagram summarizing the overall process of an embodiment of a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning provided by the present invention.
[0030] Figure 2 This is a flowchart of functional motif definition and protein backbone generation in an embodiment of the present invention; wherein, (a) is a schematic diagram of a catalytic core definition strategy based on evolutionary information, and (b) is a schematic diagram of backbone generation and complex post-processing assembly under geometric constraints.
[0031] Figure 3 This is a funnel diagram of sequence-structure co-evolution and multidimensional prediction screening in an embodiment of the present invention; wherein, (a) is a schematic diagram of sequence-structure co-iterative optimization protocol, and (b) is a schematic diagram of structure verification and Pareto optimal screening funnel.
[0032] Figure 4 These are protein backbone generation and structural characterization diagrams based on functional motifs in embodiments of the present invention; wherein, (a) is a schematic diagram of the conserved catalytic core defined in the crystal structure of natural human UNG, (b) is a joint distribution diagram of the structural diversity of 5,000 generated backbones, and (c) is a comparison diagram of the secondary structure content of generated backbones and natural templates.
[0033] Figure 5 This is a sequence space exploration and iterative optimization index analysis diagram in the embodiment of the present invention; wherein, (a) is the global sequence consistency distribution and alignment gap frequency diagram of the design sequence and the template sequence, and (b) is the collaborative optimization trajectory diagram of key indicators in the three rounds of iterative optimization.
[0034] Figure 6 This is a comparison diagram of the orthogonal structure verification and benchmark of the de novo design population in this embodiment of the invention.
[0035] Figure 7 This is a Pareto optimal screening strategy and candidate quality enrichment analysis diagram in an embodiment of the present invention; wherein, (a) is a Pareto optimal screening diagram in the three-dimensional target space, and (b) is a comparison diagram of the index enrichment of the candidate library at different screening stages.
[0036] Figure 8 This is a characterization diagram of the diversity, novelty and atomic-level fidelity of the final candidates in the embodiments of the present invention; wherein, (a) is a pairwise comparison matrix diagram of the 23 final candidates, (b) is a structural novelty evaluation distribution diagram based on Foldseek, and (c) is an atomic-level fidelity superposition diagram of the first-ranked candidate design. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, embodiments of this invention, and should not be construed as limiting the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. In the description of this invention, it should be understood that the terminology used is for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0038] Figure 1This is a flowchart illustrating a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning, provided by the present invention. The execution entity of this de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning can be any suitable terminal-side device or network-side device, such as a de novo design device for uracil-DNA glycosylation enzymes.
[0039] See Figure 1 The present invention provides a de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning, which may include the following steps: S1. Using the crystal structure of the complex of natural uracil-DNA glycosylation enzyme and substrate analog as a natural template, a set of functional motifs containing multiple discontinuous peptide segments is defined through homology sequence retrieval and clustering, multiple sequence alignment and conservation analysis, and fixed functional motif definition. The set of functional motifs is divided into a catalytic core layer, a DNA binding interface layer and a structural support layer according to their functional roles. The types of amino acid residues in the set of functional motifs and their main chain and side chain conformations are completely fixed in subsequent design steps as geometric constraints for generating new protein backbones.
[0040] In one embodiment, the natural uracil-DNA glycosylation enzyme is a human uracil-DNA glycosylation enzyme, and the PDB number of the complex crystal structure is 1EMH.
[0041] In one embodiment, the functional motif set is defined to satisfy at least one of the following: (1) having a consistency frequency greater than 90% in multiple sequence alignments of homologous sequences; (2) being annotated as an active site or binding site in the UniProt database; (3) having a key functional role as verified by mutation experiments; and (4) having a spatial distance of less than 8.0 Å from any heavy atom of the substrate analog in the crystal structure.
[0042] S2. Input the three-dimensional coordinates of the functional motif set and the substrate analogue as conditions into the all-atom diffusion model. By configuring the contigmap parameter, a variable-length region to be generated is inserted between fixed functional motif fragments to define the overall length range of the target protein. In the denoising process of the all-atom diffusion model, a guiding potential energy for ligand contact is introduced to drive the all-atom diffusion model to generate a protein topological backbone that is geometrically complementary to and tightly packed with the substrate analogue and functional motif, thereby obtaining a batch (e.g., thousands) of candidate backbones with unique structural features as candidate protein backbones.
[0043] In one embodiment, the all-atom diffusion model is RFdiffusionAA; the overall length of the candidate protein backbone is set to be between 215 and 235 amino acid residues.
[0044] S3. For each candidate protein backbone generated in S2, the DNA double helix in the natural template is transplanted to the binding site of the candidate protein backbone using a rigid body transformation algorithm to obtain the corresponding protein-DNA complex model. Based on the protein-DNA complex model, the side chain conformation of the preset key amino acid residues in the fixed functional motif is replaced and repaired.
[0045] In one embodiment, step S3 may include: S31. For each candidate protein backbone generated in S2, use a rigid body transformation algorithm (e.g., Kabsch algorithm) to calculate the transformation matrix that minimizes the root mean square deviation (RMSD) between the Cα atom of the fixed functional motif in the candidate protein backbone and the corresponding atom in the natural template. S32. Using the transformation matrix obtained in S31, the DNA double helix in the natural template is transplanted to the binding site of the candidate protein backbone, thereby constructing a complete protein-DNA complex model. S33. Based on the protein-DNA complex model, the side chain conformation of the key amino acid residues (including tryptophan) in the fixed functional motif is replaced with the corresponding conformation in the natural template to correct the local geometric distortions that may be introduced during the generation process.
[0046] S4. For each protein-DNA complex model obtained in S3, perform multiple rounds of alternating iterative optimization based on deep learning sequence generation and physical force field structure refinement to obtain a set of candidate sequence-structure pairs.
[0047] In one embodiment, step S4 may include: S41. In each iteration, a graph neural network model that is aware of the ligand environment (e.g., LigandMPNN) is first used to generate a batch of candidate amino acid sequences for the current candidate protein backbone. During the generation of candidate amino acid sequences, the types of amino acids in the functional motif set remain fixed. S42. Using a force field containing protein physicochemical energy terms and custom catalytic geometry constraints (e.g., Rosetta FastRelax program combined with enzyme design constraint file), perform side chain optimization and backbone energy minimization on the complex structure corresponding to each candidate amino acid sequence. In this process, the degrees of freedom of all atoms in the functional motif set are frozen by setting an atom movement control file (MoveMap). S43. Based on the total energy and geometric constraint deviation of the complex structure, the Pareto strategy is used to select the best-performing sequence-structure pair as the input for the next iteration. S44. As the iteration rounds progress, adjust the noise level parameters used in the graph neural network model to transition from high noise to low noise, so as to achieve a self-consistent process of sequence and structure from coarse exploration to high-precision convergence.
[0048] In one embodiment, the alternating iterative optimization is performed in three rounds. The LigandMPNN model used in the first round of iteration is trained with 0.20 Å Gaussian noise, while the LigandMPNN model used in the second and third rounds of iteration is trained with 0.10 Å Gaussian noise.
[0049] S5. Using an orthogonal prediction model and a Pareto optimal strategy, a de novo uracil-DNA glycosylation enzyme candidate set is selected from the candidate sequence-structure pair set.
[0050] In one embodiment, step S5 may include: S51. For the candidate sequence-structure pair set, the first prediction model (AlphaFold3 (AF3)) is used to predict the protein-DNA complex containing the natural substrate (2'-deoxyuridine (dU)), and the second prediction model (ESMFold) is used to independently predict the protein monomers. S52. Calculate the evaluation index data based on the two prediction results obtained in S51. The evaluation index covers the following four dimensions: complex binding confidence, monomer folding stability, structural fidelity, and physicochemical compatibility. S53. Based on the evaluation index data, a qualifying screening is performed on the candidate sequence-structure pair set to eliminate designs that do not meet the preset basic requirements. S54. For the candidate designs that have passed the qualifying round, select at least two mutually orthogonal evaluation indicators as optimization objectives, construct a multidimensional objective space, and identify a set of non-dominated solutions by calculating the Pareto front in the multidimensional objective space, which constitutes the candidate set of de novo uracil-DNA glycosylation enzymes.
[0051] In one embodiment, the evaluation metrics include any one or any combination of the following: AF3 interface prediction template modeling score (ipTM), minimum interaction prediction alignment error score (ipSAE). min Global Predicted Local Distance Difference Test Score (pLDDT), functional motif pLDDT, ESMFold monomer pLDDT, ESMFold monomer pTM, functional motif Cα RMSD, global Cα RMSD, and shape complementarity of the protein-DNA interface ( Sc ).
[0052] In one embodiment, the preset basic requirements include any one or any combination of the following: ipTM ≥ 0.96, ipSAE min ≥ 0.74, global Cα RMSD ≤ 1.8 Å, functional motif Cα RMSD ≤ 1.0 Å.
[0053] The following specific embodiment further illustrates the de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning provided by the present invention.
[0054] The specific steps are as follows: Step 1: Extraction and Definition of Fixed Functional Motifs Based on Evolutionary and Structural Information 1. Data preparation: The high-resolution crystal structure of the complex of human uracil-DNA glycosylase (UNG) and the non-hydrolyzable substrate analog 2'-deoxypseudouridine-5'-monophosphate (P2U) was selected as the natural structural template. Its PDB number is 1EMH, the resolution is 1.80 Å, and the corresponding UniProt entry is P13051.
[0055] 2. Homologous Sequence Retrieval and Clustering: For the UniProtKB database, the search query (ec:3.2.2.27) AND (family:UNG) AND (fragment:false) AND (database:alphafolddb OR pdb) AND (length:[180 TO 400]) AND (ft_act_site:*) was used to retrieve 21,908 initial sequences containing active site annotations from the UNG family. Using MMseqs2 software with a minimum sequence consistency threshold of 0.90 and a coverage threshold of 0.80, the initial sequences were clustered to remove redundancy. The command was `mmseqs easy-cluster all.fasta ung_cluster_90_80 tmp / --min-seq-id 0.90 -c 0.8 --cov-mode 0 -e 1E-5`, yielding 10,466 representative sequences.
[0056] 3. Multiple Sequence Alignment and Conservation Analysis: The representative sequence described above was merged with the template sequence P13051, and multiple sequence alignment (MSA) was performed using the Super5 algorithm in MUSCLE v5.3 software. The MSA results were quantitatively analyzed using Jalview software, and amino acid sites with a consensus frequency greater than 90% were identified as highly evolutionarily conserved residues.
[0057] 4. Definition of Fixed Functional Motifs: Based on evolutionary, structural, and functional data, a set of functional motif residues that remain fixed in subsequent design steps is strictly defined. Any residue is included in this set if it meets one of the following conditions: (1) its consistency frequency in the MSA is >90%; (2) it is annotated as an "active site" or "binding site" in the UniProt P13051 entry; (3) it has been experimentally verified to have a key functional role, specifically including Y147 and N204 which involve substrate specificity, and L272 which is crucial for uracil excision activity; (4) in the 1EMH crystal structure, the distance between any heavy atom and any heavy atom of the substrate analog P2U is within 8.0 Å. Through the above screening, a set of fixed functional motifs containing multiple discontinuous peptides is defined and divided into three levels according to their functional roles: catalytic core layer, DNA binding interface layer, and structural support layer, such as Figure 2 As shown in (a).
[0058] Step 2: Conditional generation of a novel protein topological backbone based on a diffusion model Model configuration: The skeleton was generated using the all-atom diffusion model RFdiffusionAA. The three-dimensional coordinates of the fixed functional motif and substrate analog P2U defined in step one were used as geometric constraints input into the model.
[0059] Topology definition and generation: By configuring fragment mapping parameters (e.g., the contigmap parameter in the RFdiffusionAA model), variable-length regions to be generated are interspersed between fixed functional motif fragments, with the total length of the target protein ranging from 215 to 235 amino acid residues. In the denoised trajectory of the diffusion model, a guiding potential for ligand contact is introduced to promote the tight packing of the generated backbone around the ligand. This embodiment generated a total of 5,000 protein backbones with unique topologies.
[0060] Characterization of Generation Results: Structural analysis was performed on the 5,000 generated frameworks. Comparison with the natural template 1EMH showed that the average template modeling score (TM-score) of the generated frameworks was 0.87, and the average Cα RMSD was 2.39 Å, ranging from 1.4 to 3.2 Å. Quantitative analysis of secondary structures indicated that the average helicity of the generated frameworks was 38.4%, higher than the 30.5% of the natural template. All generated frameworks maintained a compact spherical conformation with an average radius of gyration of approximately 16.4 Å. These results demonstrate that the generation model, while preserving the overall folding characteristics of the UNG superfamily, achieves extensive sampling of the conformational space and exhibits a preference for thermodynamically more stable secondary structure elements, such as... Figure 2 (b) Figure 4As shown in (a)-(c).
[0061] Step 3: Post-processing assembly and geometric correction of protein-substrate complexes To address the issue that the raw output generated by RFdiffusionAA does not contain DNA strand information and some side chains may have geometric deviations, the following customized post-processing is performed: Rigid body transplantation of DNA double helix: Using the Kabsch algorithm, a rigid body transformation matrix is calculated that minimizes the root mean square deviation (RMSD) between the generated backbone and the Cα atoms of the fixed functional motif in the reference structure (1EMH). This matrix is then applied to precisely transplant the DNA double helix (B and C strands) in the reference structure to each corresponding binding site of the generated backbone.
[0062] Key side chain conformation repair: The side chain coordinates of all tryptophan residues in the fixed functional motif are replaced with high-resolution crystallographic coordinates in the 1EMH reference structure to correct local geometric distortions that may be introduced during the generation process.
[0063] Document standardization: The assembled composite model is standardized by supplementing complete PDB link and connection records to ensure it can be correctly identified and processed by downstream physics simulation tools, such as... Figure 2 (b) As shown in the dashed box.
[0064] Step 4: Cooperative Iterative Optimization of Sequence and Structure For the 5,000 complex models obtained in step three, three rounds of alternating iterative optimization based on deep learning sequence generation and physical force field structure refinement were performed.
[0065] First iteration: Sequence generation: The LigandMPNN model was used, with a sampling temperature set to 0.1°C. Model weights trained under 0.20 Å Gaussian noise were selected. The amino acid types of functional motif residues were strictly fixed, and 1,000 candidate sequences were generated for each backbone. The model was configured to directly output a full-atom model.
[0066] Sequence selection: The generated sequences are sorted according to the composite confidence score (calculated as 0.6×ligand_confidence + 0.4×overall_confidence), and the 8 sequences with the highest scores for each skeleton are selected to enter the physical refinement stage.
[0067] Physical Refinement: A custom formatting engine was developed to restore standard PDB linker records and inject the REMARK 666 header file required for the Rosetta enzyme design protocol. Energy minimization was performed using the Rosetta FastRelax program under the Beta_nov16 energy function. By setting a strict MoveMap, the backbone and side-chain degrees of freedom of functional motif residues were completely frozen, allowing relaxation only in non-motif regions. Simultaneously, a custom enzyme design constraint file (.cst) was introduced to impose geometric constraints on the ideal hydrogen bonds and catalytic distances between key protein residues and ligand P2U.
[0068] Pareto screening: For the refined structure, calculate its Rosetta total energy and constraint deviation. Using a Pareto screening strategy, identify and retain designs that achieve the optimal trade-off between minimizing total energy and minimizing constraint deviation, as input for the next iteration.
[0069] Second and third iterations: The sequence generation model was switched to a version of LigandMPNN trained at a lower noise level (0.10 Å) to accommodate the higher-quality skeletons after physical refinement. The number of sequence samples per skeleton was reduced to 200, while the sampling temperature remained constant at 0.1. The sequence selection, physical refinement, and Pareto selection steps from the first iteration were repeated. The trajectory of this co-iterative optimization is as follows: Figure 3 As shown in (a).
[0070] Evaluation of Iterative Optimization Effect: Longitudinal tracking analysis of the three iteration trajectories shows that as the iterations progress, the total Rosetta energy gradually converges to a lower energy state, while the sequence confidence of LigandMPNN steadily increases. Throughout the entire iteration process, the median geometric constraint deviation of the catalytic core remains below 0.6 REU (Rosetta Energy Unit), indicating that the geometric conformation of the active pocket remains stable during the optimization process. Figure 5 As shown in (b).
[0071] Step 5: Multidimensional screening based on orthogonal prediction models and Pareto optimal strategies Orthogonal structure prediction verification: AlphaFold3 (AF3) complex prediction: To simulate physiological catalytic states, the substrate analog P2U calculations used in the design phase were replaced with the natural substrate 2'-deoxyuridine (dU), and a custom chemical composition dictionary defining the topology and covalent linkages of dU was constructed for use by AF3. The inference process used five independent random seeds, and template input was disabled. Evaluation metrics included the interface prediction template modeling score (ipTM) and the minimum interaction prediction alignment error score (ipSAE). min ) and global pLDDT, etc.
[0072] ESMFold Monomer Prediction: Using ESMFold under standard settings, monomer structure prediction is performed on all design sequences to obtain the pLDDT and pTM scores of the monomers, which serve as an orthogonal evaluation of their intrinsic folding stability.
[0073] Multidimensional indicator calculation and qualifying selection: Based on the orthogonal prediction results and input model, 12 key indicators covering four dimensions—complex binding confidence, monomer folding stability, structural fidelity, and physicochemical compatibility—were calculated. The physical meaning and specific thresholds of each indicator are detailed in Table 1.
[0074] Table 1. Multidimensional evaluation indicators and screening criteria for designing UNG from scratch
[0075] Exploratory data analysis was conducted on the metric distribution of 5,000 design samples, and a series of hard thresholds were set for qualifying the competition. For example, ipTM ≥ 0.96 and ipSAE were required. min ≥ 0.74, global Cα RMSD (relative to the generated skeleton) ≤ 1.8 Å, etc. After this screening, the candidate library was reduced from 5,000 to 418 high-quality designs.
[0076] Pareto optimal decision: Among the 418 designs that passed the qualifying round, select the one that maximizes AF3 ipSAE. min Three mutually orthogonal key metrics—maximizing ESMFold pLDDT and minimizing functional motif Cα RMSD—were used as optimization objectives to construct a three-dimensional objective space. By calculating the Pareto front in this three-dimensional space, 23 non-dominated solutions were identified. These 23 designs were unsurpassed by any other design across all dimensions, representing a mathematically optimal balance between confidence, fold stability, and structural fidelity. They were determined as the final candidate set for de novo design of uracil-DNA glycosylation enzymes, such as... Figure 3 (b) and Figure 7 As shown in (a).
[0077] Comprehensive evaluation of the final candidate set: The quantitative analysis results for the above 23 final candidate designs are as follows: Structure fidelity: The average Cα RMSD of the fixed functional sequence achieves sub-angstrom accuracy of 0.47 Å.
[0078] Combining confidence level and physicochemical properties: average ipSAE min The average shape complementarity is 0.77. S c The value is 0.67.
[0079] Sequence and structural novelty: The median sequence identity with the natural template was only 38.7%; a Foldseek global structural homology search with the PDB database showed that the best matching results were all glycosylation enzyme structures (average TM-score approximately 0.80), but the average Cα RMSD was approximately 3.0 Å. Internal pairwise comparisons showed an average sequence identity of 48.1% and an average RMSD of 2.14 Å among the 23 candidates.
[0080] Atomic-level microscopic examination: The top-ranked candidate design achieved atomically precise superposition with the natural template 1EMH in the catalytic pocket region (RMSD < 0.5 Å). In the structure predicted by AF3, a direct hydrogen bond interaction spontaneously re-emerged between the catalytic aspartic acid residue (D64, corresponding to D145 in the natural sequence) and the substrate dU, confirming that de novo scaffold design can construct microenvironments with catalytic potential, such as... Figure 6 , Figure 7 (b) and Figure 8 As shown in (a)-(c).
[0081] The above results demonstrate that the method provided by this invention successfully achieved novel de novo design of uracil-DNA glycosylation enzymes. While accurately reproducing the catalytic geometry at the atomic level, it also yielded candidate enzyme molecules with novel sequences, compact structures, and high folding confidence, providing a reliable computational design strategy for developing novel gene editing tools.
[0082] Exemplary disclosure of the final candidate sequence: To further demonstrate the feasibility of the method and the structural novelty of the design products, the complete amino acid sequences of 23 non-dominated candidate designs obtained through the Pareto optimal decision screening described in step five are disclosed as exemplary embodiments. It should be clearly understood that these 23 sequences are specific product instances of the de novo design method described in this invention, and their identity is defined by the multidimensional computational screening process described in this invention, rather than an exhaustive list of all sequences that the method of this invention can generate. Those skilled in the art, guided by the method of this invention, can generate and screen more candidate sequences with similar structural features and confidence levels.
[0083] To comply with relevant patent law requirements, this application also submits a computer-readable sequence listing conforming to WIPO standard ST.26, which contains the 23 candidate amino acid sequences (SEQ ID NO: 1-23). These sequences are not repeated in tabular form in this specification; the contents of this sequence listing are incorporated herein by reference in their entirety.
[0084] This invention addresses the evolutionary path dependence of natural protein backbones, the accuracy bottleneck of physical computation methods, and the control blind spot of pure generative AI in catalytic motif embedding in existing technologies. It provides a de novo design method for uracil-DNA glycosylation enzymes based on a deep fusion of physical priors and deep learning. Compared with existing technologies, this invention has the following outstanding advantages: (1) This invention, through the introduction of deep conservation analysis of cross-species homologous sequences, precisely defines a three-tiered set of functional motifs comprising a catalytic core layer, a DNA-binding interface layer, and a structural support layer, and uses this as a rigid "physical prior" constraint, fully embedding it into the generation process of the all-atom diffusion model. The resulting design has less than 40% sequence similarity to the natural template and significant structural deviations from the closest natural homolog, truly breaking away from the fine-tuning paradigm of the natural skeleton and creating a novel functional molecule located in an independent sequence-structure space. This lays a solid technical foundation for developing a new generation of gene-editing tool enzymes with independent intellectual property rights and free from the patent barriers of natural enzymes.
[0085] (2) In the framework generation stage of the all-atom diffusion model, this invention uses the three-dimensional coordinates of the three-level functional motifs and substrate analogs as input conditions, and specifically introduces ligand contact guiding potential energy to drive the growth of a geometrically complementary and tightly packed new topological framework around the substrate analogs, anchoring the sub-angstrom-level geometric physical quantities of the active pocket from the source of framework generation. In the sequence-structure alternating iterative optimization stage, this invention couples sequence generation based on graph neural networks with physical force field refinement containing custom catalytic geometric constraints to form a high-frequency iterative closed loop of "AI design - physical verification". In each iteration, the atomic degrees of freedom of the functional motifs are fixed, and only the variable framework region of the generated sequence is optimized for side chains and minimized for energy. A strategy of gradually transitioning from high noise to low noise is adopted to achieve a self-consistent process of sequence and structure from coarse-grained plasticity exploration to high-precision deterministic convergence. Orthogonal prediction confirms that the active site of the finally designed enzyme can accurately reproduce the key hydrogen bond network with the substrate dU, which is a fundamental overcoming of the problem of "zero activity" or "weak activity" under the condition that both the sequence and the framework are greatly changed.
[0086] (3) At the prediction and validation level, this invention introduces two completely orthogonal models, AlphaFold3 and ESMFold, for cross-validation to predict the structure of protein-DNA complexes and monomers, respectively. The evaluation indicators highly correspond to the actual functional realization probability of the design, covering four dimensions: complex binding confidence (ipTM, ipSAEmin), monomer folding stability (pLDDT, pTM), structure fidelity (Cα RMSD of global and functional motifs), and interfacial physicochemical compatibility (Sc), comprehensively covering the functional realization chain from folding and binding to catalysis. At the decision-making and screening level, this invention introduces a Pareto optimal strategy for multi-objective trade-offs, rather than simple weighted scoring. By finding non-dominated solutions that achieve optimal balance among multiple potentially competing indicators (such as binding confidence and folding stability), rational ranking and decision-making of candidate designs are achieved. This system significantly reduces the false positive rate from computational prediction to wet experimental validation, greatly improves the expected success rate, and establishes a transferable systematic validation and decision-making standard for computationally complex enzymatic designs.
[0087] (4) In the design of this invention, the overall length constraint of the protein is set to ensure that the generated candidate backbone has a compact and concise structural feature from the source. At the same time, ESMFold is introduced into the screening system to independently evaluate the folding stability of naked protein monomers, ensuring that each candidate enzyme still has high thermodynamic stability and folding robustness when separated from complex chaperone proteins or DNA substrates. This intrinsic characteristic enables it to better adapt to the limited loading capacity of vectors such as adeno-associated virus, and to maintain the correct conformation and function in the complex molecular crowding environment within cells, providing a new path to solve the bottleneck problems of in vivo delivery and long-term stability of gene editing tools.
[0088] This invention provides an electronic device that may include: a processor, a communication interface 820, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus. The processor can invoke logical instructions in the memory to execute the steps of any of the above-described methods for de novo design of uracil-DNA glycosylation enzymes based on physical priors and deep learning.
[0089] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0090] On the other hand, the present invention also provides a computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to perform the steps of any of the above-described methods for de novo design of uracil-DNA glycosylation enzymes based on physical priors and deep learning.
[0091] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described methods for de novo design of uracil-DNA glycosylation enzymes based on physical priors and deep learning.
[0092] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0093] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning, characterized in that, Includes the following steps: S1. Using the crystal structure of the complex of natural uracil-DNA glycosylation enzyme and substrate analog as a natural template, a set of functional motifs containing multiple non-continuous peptide segments is defined through homology sequence retrieval and clustering, multiple sequence alignment and conservation analysis, and fixed functional motif definition. The set of functional motifs is divided into a catalytic core layer, a DNA binding interface layer and a structural support layer according to their functional roles. The types of amino acid residues in the set of functional motifs and their main chain and side chain conformations are completely fixed in subsequent design steps as geometric constraints for generating new protein backbones. S2. Input the three-dimensional coordinates of the functional motif set and the substrate analogue as conditions into the all-atom diffusion model, set a variable-length region to be generated between fixed functional motif fragments, and introduce guiding potential energy for ligand contact during the denoising process of the all-atom diffusion model to drive the all-atom diffusion model to generate a protein topological backbone that is geometrically complementary to and tightly packed with the substrate analogue and functional motif, as a candidate protein backbone; S3. For each candidate protein backbone generated in S2, the DNA double helix in the natural template is transplanted to the binding site of the candidate protein backbone using a rigid body transformation algorithm to obtain the corresponding protein-DNA complex model. Based on the protein-DNA complex model, the side chain conformation of the preset key amino acid residues in the fixed functional motif is replaced and repaired. S4. For each protein-DNA complex model obtained in S3, perform multiple rounds of alternating iterative optimization based on deep learning sequence generation and physical force field structure refinement to obtain a set of candidate sequence-structure pairs. S5. Using an orthogonal prediction model and Pareto optimality strategy, a de novo uracil-DNA glycosylation enzyme candidate set was selected from the candidate sequence-structure pair set. Step S3 includes: S31. For each candidate protein backbone generated in S2, use the rigid body transformation algorithm to calculate the transformation matrix that minimizes the root mean square deviation between the Cα atom of the fixed functional motif in the candidate protein backbone and the corresponding atom in the natural template. S32. Using the transformation matrix obtained in S31, the DNA double helix in the natural template is transplanted to the binding site of the candidate protein backbone, thereby constructing a complete protein-DNA complex model. S33. Based on the protein-DNA complex model, the side chain conformation of the key amino acid residues in the fixed functional motif is replaced with the corresponding conformation in the natural template.
2. The de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning according to claim 1, characterized in that, The natural uracil-DNA glycosylation enzyme is a human uracil-DNA glycosylation enzyme, and the PDB number of the crystal structure of the complex is 1EMH; The definition of the functional motif set must satisfy at least one of the following: (1) the consistency frequency is greater than 90% in multiple sequence alignment of homologous sequences; (2) it is annotated as an active site or binding site in the UniProt database; (3) it has been verified by mutation experiments to have a key functional role; (4) the spatial distance between it and any heavy atom of the substrate analog in the crystal structure is within 8.0 Å; The all-atom diffusion model is RFdiffusionAA; and / or, The overall length of the candidate protein backbone is set to be between 215 and 235 amino acid residues.
3. The de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning according to claim 1, characterized in that, Step S4 includes: S41. In each iteration, a graph neural network model that is aware of the ligand environment is first used to generate a batch of candidate amino acid sequences for the current candidate protein backbone. During the generation of candidate amino acid sequences, the types of amino acids in the functional motif set remain fixed. S42. Using a force field containing protein physicochemical energy terms and custom catalytic geometry constraints, the side chain optimization and backbone energy minimization are performed on the complex structure corresponding to each candidate amino acid sequence. In S42, the degrees of freedom of all atoms in the functional motif set are frozen by setting an atom movement control file. S43. Based on the total energy and geometric constraint deviation of the complex structure, the Pareto strategy is used to select the best-performing sequence-structure pair as the input for the next iteration. S44. As the iteration rounds progress, adjust the noise level parameters used in the graph neural network model to transition from high noise to low noise, so as to achieve a self-consistent process of sequence and structure from coarse exploration to high-precision convergence.
4. The de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning according to claim 3, characterized in that, The graph neural network model is LigandMPNN; The force field containing protein physicochemical energy terms and custom catalytic geometry constraints is a RosettaFastRelax program combined with an enzyme design constraint file; and / or, The alternating iterative optimization in step S4 consists of three rounds. The LigandMPNN model used in the first round of iteration is trained under 0.20 Å Gaussian noise, while the LigandMPNN model used in the second and third rounds of iteration is trained under 0.10 Å Gaussian noise.
5. The de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning according to claim 3, characterized in that, Step S5 includes: S51. For the set of candidate sequence-structure pairs, the first prediction model is used to predict the protein-DNA complex containing the natural substrate, and the second prediction model is used to independently predict the protein monomers. S52. Calculate the evaluation index data based on the two prediction results obtained in S51. The evaluation index covers the following four dimensions: complex binding confidence, monomer folding stability, structural fidelity, and physicochemical compatibility. S53. Based on the evaluation index data, a qualifying screening is performed on the candidate sequence-structure pair set to eliminate designs that do not meet the preset basic requirements. S54. For the candidate designs that have passed the qualifying round, select at least two mutually orthogonal evaluation indicators as optimization objectives, construct a multidimensional objective space, and identify a set of non-dominated solutions by calculating the Pareto front in the multidimensional objective space, which constitutes the candidate set of de novo uracil-DNA glycosylation enzymes.
6. The de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning according to claim 5, characterized in that, The first prediction model is AlphaFold3, and the second prediction model is ESMFold; The natural substrate is 2'-deoxyuridine; The evaluation metrics include any one or any combination of the following: AF3 interface prediction template modeling score, minimum interaction prediction alignment error score, global prediction local distance difference test score, functional motif pLDDT, ESMFold monomer pLDDT, ESMFold monomer pTM, functional motif Cα RMSD, global Cα RMSD, shape complementarity of the protein-DNA interface; and / or, The default basic requirements include any one or any combination of the following: ipTM ≥ 0.96, ipSAE min ≥ 0.74, global Cα RMSD ≤ 1.8 Å, functional motif Cα RMSD ≤ 1.0 Å.
7. A uracil-DNA glycosylation enzyme obtained by the de novo design method for uracil-DNA glycosylation enzyme based on physical priors and deep learning as described in any one of claims 1-6, characterized in that, The uracil-DNA glycosylase has an active pocket that matches the catalytic core of the natural human uracil-DNA glycosylase with sub-angstrom precision, and its amino acid sequence has less than 40% global sequence identity with the natural template.
8. The uracil-DNA glycosylation enzyme according to claim 7, characterized in that, The amino acid sequence of the uracil-DNA glycosylation enzyme is selected from the group consisting of SEQ ID NO:1-23.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the de novo design method for uracil-DNA glycosylation enzymes based on physical priors and deep learning as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Protein optimization design and screening method and device based on artificial intelligence algorithm
CN120260679A
Generative protein design via noise diffusion
WO2024158466A2