Protein optimization design and screening method and device based on artificial intelligence algorithm

By optimizing protein design through artificial intelligence algorithms, combining deep learning and generative models, key sites are identified and optimized luciferase variants are generated, solving the problems of low efficiency and high cost of traditional methods and achieving efficient and accurate luciferase optimization.

CN120260679BActive Publication Date: 2025-09-12ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510756581.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-12
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Traditional protein design methods are time-consuming, labor-intensive, costly, and inefficient, and are particularly challenging when optimizing complex functions. Traditional luciferases rely on exogenous substrates, which increases costs and has limited luminescence intensity.

Method used

A protein optimization design method based on artificial intelligence algorithms is used, combined with deep learning and generative models, to identify key sites and generate optimized protein variants through protein-substrate docking simulation, structure prediction, evolutionary analysis and graph neural networks.

Benefits of technology

It improves the efficiency and accuracy of protein design, enhances the luminescence intensity and stability of luciferase, and provides an efficient solution for the development of autonomous luminescence systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260679B_ABST
    Figure CN120260679B_ABST
Patent Text Reader

Abstract

The present invention discloses a protein optimization design and screening method and device based on an artificial intelligence algorithm. Taking fungal luciferase as an example, the method integrates multi-dimensional bioinformatics analysis and deep learning technology to achieve efficient protein engineering and optimization. First, an AI-driven molecular docking simulation method is used to identify the binding domain of luciferase and luciferin, and the key conserved sites are determined by combining literature, evolutionary analysis and structure prediction; then, a diffusion model is used to generate a protein functional domain skeleton under the constraint of fixed sites, and a graph neural network model is used to predict the protein sequence; finally, the generated sequences are scored and screened to obtain high-stability candidate variants. The present invention innovatively constructs a dynamic fusion strategy for conserved sites, which satisfies the diversity and adaptability of protein design sequences by logically optimizing fixed regions through the intersection priority of structure prediction and evolutionary sites and the expansion of the union of literature sites, breaking through the efficiency bottleneck of traditional solutions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics and artificial intelligence-driven computational biology, and in particular relates to a protein optimization design and screening method and device based on artificial intelligence algorithms. Background Art

[0002] With the continuous advancement of biology and biotechnology, protein design and optimization have gradually become key links in promoting disease treatment, precision medicine, and biotechnology innovation. However, traditional protein design methods often rely on screening and mutation strategies in the laboratory. These methods are often time-consuming, labor-intensive, costly, and inefficient, especially when faced with the optimization of complex functions, which poses enormous challenges. The introduction of artificial intelligence (AI) technology has brought unprecedented opportunities for protein design. Through advanced algorithms such as deep learning, reinforcement learning, and generative models, AI can automatically identify and mine potential patterns and patterns in massive amounts of protein data, greatly improving the efficiency and accuracy of protein design. AI models, especially deep neural networks (DNNs) and graph neural networks (GNNs), can predict the functional properties of proteins without experimental data by learning the complex relationship between protein sequences and their three-dimensional structures, thereby providing more efficient and accurate design solutions. In addition, generative models such as the Protein Message Passing Neural Network (ProteinMPNN) and RoseTTAFold Diffusion (RFDiffusion) can automatically generate optimized candidate proteins with specific functions within the protein variant space. By simulating intermolecular interactions and stability, they can predict and design protein variants with optimal performance under specific circumstances. This AI-driven protein design approach breaks through the limitations of traditional reliance on experience, making protein design and optimization more systematic, quantitative, and efficient, and promoting the rapid transition from basic research to applied development. In many practical applications, AI provides new possibilities for the targeted optimization of protein function, especially in fields such as enzyme design, antibody engineering, and vaccine development. The addition of artificial intelligence technology is leading protein design towards a more intelligent and efficient direction.

[0003] Luciferases are an important class of bioluminescent enzymes widely used in fields such as bioimaging, disease diagnosis, and therapeutic efficacy monitoring. Luciferases produce photons by catalyzing the oxidation of luciferin. This process not only provides a visualization tool for scientific research but also has widespread applications in medical testing. Traditional luciferases typically rely on exogenously added substrates (luciferin), which not only increases experimental costs but can also be limited by insufficient luciferin supply in some cases. In recent years, with the discovery of a bioluminescent pathway from the fungus Neonothopanus nambi, scientists have been able to construct autonomous luminescence systems in various eukaryotic organisms. This system converts caffeic acid into luciferin through a series of enzymes, ultimately catalyzing the production of photons by luciferase. However, the poor performance of native fungal luminescence systems in heterologous hosts results in limited luminescence intensity, limiting their potential for practical applications. This invention, through the introduction of artificial intelligence technology, has achieved a breakthrough in the optimized design of luciferases. By combining deep learning and generative models, AI can uncover key sites within the complex relationship between protein structure and function, systematically optimizing the performance of luciferase. This not only improves the enzyme's luminescence intensity and stability, but also generates high-performance luciferase variants in a short period of time, breaking through the limitations of traditional design methods and providing a more efficient and accurate solution for the development of autonomous luminescence systems. This AI-driven protein optimization design method not only enhances the application potential of luciferase, but also provides strong support for technological advancements in fields such as precision medicine and bioimaging. Summary of the Invention

[0004] The purpose of the present invention is to address the deficiencies of the existing technology and propose a protein optimization design and screening method and device based on artificial intelligence algorithm.

[0005] To achieve the above objectives, the present invention provides a protein optimization design and screening method based on an artificial intelligence algorithm, comprising the following steps:

[0006] (1) Using artificial intelligence-driven protein-substrate docking simulation methods to predict the binding domains of proteins and substrates; combining literature analysis, structure prediction, and evolutionary analysis to screen key conserved sites; taking the intersection of the sites screened by structure prediction and evolutionary analysis, and then taking the union of the sites extracted from the literature to form a conserved region, which is used as a fixed site;

[0007] (2) Generate protein domain skeletons under fixed site constraints using a diffusion model;

[0008] (3) Use the graph neural network model to predict the sequence of the protein functional domain skeleton generated in step (2);

[0009] (4) Scoring and screening are performed based on the predicted sequence, including sequence-level scoring and structure-level screening, to obtain candidate protein variants.

[0010] Furthermore, in step (1), artificial intelligence-driven protein-substrate docking simulation methods, including RFAA and DiffDock, are used to identify potential binding sites, thereby predicting the binding domain between the protein and the substrate molecule; visualization tools are used to reproduce the docking situation of the protein and its substrate molecule, and to analyze and screen amino acid residues with specific types and strengths of non-covalent forces between the protein and its substrate molecule.

[0011] Furthermore, in step (1), the screening of key conserved sites includes: extracting reported key functional sites based on literature information; identifying and screening potential binding sites based on protein structure prediction using pocket analysis tools and visualization tools; screening highly conserved sites based on evolutionary analysis using multiple sequence alignment, information entropy calculation, and hidden Markov model; taking the intersection of the sites obtained from structure prediction and evolutionary analysis screening, and then taking the union of the sites extracted from the literature to form a conserved region.

[0012] Furthermore, when the protein is fungal luciferase, the intersection of the sites obtained by structure prediction and evolutionary analysis was taken, and then the union was taken with the sites extracted from the literature, and 95 key sites were screened out, accounting for 35.58% of the full length of the sequence; the key sites are: 46F, 49V, 50G, 51P, 52S, 53Y, 54A, 55P, 56Q, 60G, 61Y, 63I, 64V, 66V, 67L, 68S, 69L, 70F, 71R, 74Q, 98G, 99T, 103I, 104T, 105S, 106H, 107I, 108I, 109Q, 110R, 111Q, 127D, 130I, 136R, 144S, 146S, 147K, 148F, 149E, 150F, 151H, 52A,154A,155I,156F,157L,166P,169I,173D,176R,177R,178T,179K,181E,182 I,183A,184H,185M,186H,187D,188Y,189H,190D,191C,192T,193L,194H,195L, 196A,197L,205V,212Q,213R,214H,215P,216L,217A,218G,221V,222P,223G,22 4P, 225P, 228W, 229T, 230F, 231L, 233A, 234P, 235R, 237E, 238E, 241R, 242V, 243V.

[0013] Furthermore, step (2) is specifically as follows: based on fixed sites, a denoising diffusion probability model is introduced to abstract the protein structure into six-dimensional information of each amino acid translation vector and rotation matrix, and gradually add noise as training input; in the denoising process, a loss function is constructed to measure the gap between the predicted structure and the true structure at each step, and the model is driven to restore the true protein structure, and finally the relative positions of each amino acid are dynamically adjusted from the noise, the folding framework of the functional domain is constructed, and the protein functional domain skeleton is generated.

[0014] Furthermore, step (3) is specifically as follows: based on the generated protein functional domain backbone structure without side chain atoms and sequence information, a graph neural network model is introduced to model the relationship between the protein sequence and the three-dimensional structure, that is, a deep learning model based on a graph neural network is used to learn the interactions and spatial constraints between amino acids to achieve sequence-structure mapping; by optimizing each amino acid position of a given backbone structure, the amino acid sequence that best matches it is dynamically predicted and generated.

[0015] Furthermore, in step (4), the sequence-level scoring includes: scoring the predicted sequence at the sequence level with ProteinMPNN_score and sequence repetitiveness ARR_score; wherein, ProteinMPNN_score is an indicator for measuring the compatibility between the predicted sequence and the target structure, and a sequence with a ProteinMPNN_score less than 1 indicates that it has good folding stability and can correctly form the desired three-dimensional conformation; the sequence repetitiveness score is used to evaluate the degree of repeated amino acids in the sequence, and an ARR_score greater than -2 indicates that the sequence has high diversity in the local region; through these two scores, high-quality sequences are preliminarily screened out.

[0016] Furthermore, in step (4), the structural level screening includes: performing structural prediction on high-quality sequences that have passed the preliminary screening, setting a screening threshold, and screening based on pLDDT and pTM scores, with the pLDDT threshold being no less than 80 and the pTM threshold being no less than 0.85, to ensure the structural stability and functional potential of these sequences; wherein, pLDDT is an indicator for evaluating the confidence of the local structure prediction of each residue, and a higher pLDDT value indicates a higher confidence in the structure prediction of the region; pTM is an indicator for measuring the confidence of the overall structure prediction, and a higher pTM value indicates that the predicted overall folding structure is closer to the true structure; based on the weighted sum of the ranking percentages of pLDDT and pTM, the protein variant with the highest comprehensive score is screened.

[0017] To achieve the above objectives, the present invention also provides a protein optimization design and screening device based on artificial intelligence algorithm, including one or more processors and a GPU processor, for implementing the above protein optimization design and screening method based on artificial intelligence algorithm.

[0018] To achieve the above objectives, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned protein optimization design and screening method based on artificial intelligence algorithm.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] 1. This paper provides a systematic, new process and method based on artificial intelligence algorithms for protein optimization design, breaking through the limitations of manual screening and experience-based reliance in traditional protein design, and significantly improving design efficiency and accuracy;

[0021] 2. We established an innovative method for identifying key protein sites, integrating and concatenating information related to protein sequence, structure, and function from different sources. This provides a new strategy for analyzing and targeting key protein sites. While ensuring protein function, it effectively explores and designs non-conserved regions, enhancing the diversity and adaptability of protein sequences.

[0022] 3. An efficient, flexible and reliable protein design framework has been constructed, integrating the latest artificial intelligence technologies and bioinformatics tools to adapt to different types of protein design needs. It provides a new solution to the challenges in traditional protein engineering, promotes the rapid development of biomedicine, environmental engineering, food industry and other fields, and lays a solid foundation for future interdisciplinary innovation and industrial application. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the technical solutions and specific implementations of the present invention, the following briefly introduces the drawings required for the technical description and specific implementations of the present invention. Obviously, the drawings described below are some implementations of the present invention, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0024] Figure 1 This is a flowchart of protein optimization design using fungal luciferase as an example in the present invention;

[0025] Figure 2 The distribution map of key sites of fungal luciferase from literature information in the present invention;

[0026] Figure 3 Schematic diagram of the docking state of fungal luciferase-luciferin in the present invention; wherein, Figure 3 (A) is a schematic diagram of the helical structure of fungal luciferase. Figure 3 (B) is a schematic diagram of fluorescein (3-hydroxylactocetin). Figure 3(C) is a schematic diagram of the helical structure of the fungal luciferase-luciferin docking model. Figure 3 (D) is a schematic diagram of the 3D structure of the fungal luciferase-luciferin docking model;

[0027] Figure 4 Schematic diagram of the biochemical characteristics of fungal luciferase in the present invention; wherein, Figure 4 (A) is a schematic diagram of the hydrophobicity of fungal luciferase. Figure 4 (B) is a schematic diagram of the van der Waals force of the fungal luciferase-luciferin docking model. Figure 4 (C) is a hydrogen bond diagram of the fungal luciferase-luciferin docking model. Figure 4 (D) is a schematic diagram of the ionic bond of the fungal luciferase-luciferin docking model;

[0028] Figure 5 This is the distribution map of key sites of fungal luciferase predicted by structure in the present invention;

[0029] Figure 6 This is the distribution map of key sites of fungal luciferase through evolutionary analysis in the present invention;

[0030] Figure 7 This is a heat map of potential fixed site correlations for fungal luciferase in the present invention;

[0031] Figure 8 A logic diagram is generated for the fixation sites of the fungal luciferase of the present invention; wherein the shaded portion represents the final fixation site, Figure 8 (a) is a schematic diagram of the key site set of fungal luciferase predicted by structure. Figure 8 (b) is a schematic diagram of the key site set of fungal luciferase through evolutionary analysis. Figure 8 (c) is a schematic diagram of the fungal luciferase key site set from literature information;

[0032] Figure 9 Schematic diagram of the final fixation site of the fungal luciferase in the present invention; wherein, Figure 9 (A) is a schematic diagram of the location of the fungal luciferase fixed site (white bold) in the sequence. Figure 9 (B) is a schematic diagram of the location of the fungal luciferase anchor site (white) in the secondary structure;

[0033] Figure 10 Generate an example skeleton diagram for the fungal luciferase of the present invention;

[0034] Figure 11 This is an example diagram of the scoring and screening output of the fungal luciferase prediction sequence in the present invention;

[0035] Figure 12is an example diagram of the predicted sequence and corresponding predicted structure of the fungal luciferase in the present invention; wherein, Figure 12 (A) is an example diagram of the predicted sequence of fungal luciferase. Figure 12 (B) is a schematic diagram of the predicted structure of fungal luciferase based on the example predicted sequence;

[0036] Figure 13 This is an example diagram of the final candidate sequence of the fungal luciferase in the present invention;

[0037] Figure 14 Schematic diagram of the fungal luciferase scoring and screening process based on predicted sequences in the present invention;

[0038] Figure 15 It is a structural schematic diagram of the device of the present invention. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to illustrate the present invention, rather than to be exhaustive. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.

[0040] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.

[0041] Example description: To verify the practical effect of this scheme, the luciferase from the fungus Neonothopanus nambi was used for optimization design testing. Its amino acid sequence can be directly obtained through the official website of the National Center for Biotechnology Information (NCBI) under the National Institutes of Health.

[0042] See also Figure 1The present invention provides a protein optimization design and screening method based on artificial intelligence algorithm. Taking the optimization of fungal luciferase as an example, a complete optimization framework from multi-dimensional key site identification of fungal luciferase to skeleton structure design to sequence prediction and screening is constructed. This method innovatively integrates protein sequence, structure and function related information, providing a new strategy for protein key site analysis. At the same time, it combines artificial intelligence-driven generative models to realize batch design of protein variants, solving the limitations of traditional protein design. It not only provides new possibilities for the development and application of luciferase autonomous luminescence systems, but also provides a new paradigm for the research of traditional protein engineering, which has far-reaching significance for promoting the rapid development of biomedicine, environmental engineering, food industry and other fields. Specifically comprising the following steps:

[0043] 1. The key sites are fixed.

[0044] Protein design is a complex undertaking, typically starting with a reference protein structure, modifying its sequence through mutation algorithms, and experimentally verifying the structure and functional activity of the mutants. Traditional protein design often relies on this cycle of mutation and experimental verification, identifying and fixing conserved regions in the protein and iteratively optimizing non-conserved regions to increase the success rate of design. However, this process is not only time-consuming and labor-intensive, but also limited by experimental resources and the mutation space to explore. In the design of certain specific proteins, such as fungal luciferase, due to the lack of detailed mechanistic studies and protein crystal structure analysis, directly using existing design methods may not be able to effectively address the complex luminescence mechanism. Especially when homologous sequences of the target protein are scarce, relying on existing large language models without specific fine-tuning may result in designed proteins that fail to meet experimental requirements. Therefore, accurately identifying and fixing key functional sites in a protein is a core step in the protein design process.

[0045] Key functional sites in proteins generally fall into three categories: active sites, binding sites, and structural sites. First, active sites are specific regions within a protein molecule responsible for catalyzing chemical reactions. They are typically composed of several key amino acid residues, whose precise arrangement and interactions provide an ideal environment for substrate binding and conversion. For enzymes, active sites are the core component of catalysis, effectively lowering the activation energy of the reaction and promoting its progress. An enzyme may contain multiple active sites, each with varying catalytic efficiency and specificity, collectively determining the enzyme's overall catalytic efficiency and selectivity. Second, binding sites refer to all locations on a protein that can bind to other molecules (such as substrates, inhibitors, and metal ions). While these sites are not necessarily directly involved in catalysis, they play a crucial role in the protein's interactions with other molecules. The precise design of binding sites has a significant impact on regulating protein functions, such as enzyme activity regulation and substrate affinity. Finally, structural sites are regions responsible for maintaining the stability of a protein's three-dimensional structure. These sites, through molecular mechanical interactions such as hydrogen bonds, hydrophobic interactions, and van der Waals forces, ensure that the protein maintains its stable spatial configuration. Although structural sites are not directly involved in enzymatic reactions, they play crucial roles in protein folding, stability, and function.

[0046] In protein design, the fixation of key functional sites is crucial for ensuring the core function and structural stability of the protein, while also increasing design diversity and the success rate of experimental verification. Therefore, the present invention employs a multi-strategy protein key functional site analysis method that combines literature analysis, evolutionary information, and structural prediction to determine the conditions for optimal protein design. This method allows for more flexible modification of non-conserved regions while maintaining the protein's core function, increasing the diversity of the designed sequences and thereby increasing the probability of ultimately obtaining effective mutants.

[0047] 1.1 Key sites from literature information.

[0048] Based on literature information, reported key functional sites were extracted. The key functional site information derived from the literature and verified experimentally provides a strong theoretical basis for the protein design process. Although the specific molecular mechanism of fungal luciferase luminescence has not yet been fully elucidated, it is generally believed that this process is related to the redox reaction, and the key amino acid residues that catalyze this reaction are usually located in the active center of the enzyme, that is, the catalytic pocket. These key residues are mostly polar amino acids, such as glutamic acid and histidine, which play a vital role in the expression of enzyme activity. In the catalytic mechanism of fungal luciferase, these polar amino acids promote the oxidation reaction by interacting with the substrate, thereby triggering bioluminescence. By summarizing a series of recent studies, it was found that when researchers compared the functions of 23 luciferase mutants, mutations at 13 sites (49V, 61Y, 104T, 109Q, 127D, 136R, 144S, 173D, 176R, 191C, 228W, 237E, 238E) significantly reduced the activity of luciferase. The site distribution is as follows: Figure 2 Although the results of the mutation experiment did not provide a clear relationship between polar and non-polar amino acid substitutions and changes in luciferase function, the 13 amino acid positions that lead to loss of luciferase activity were still regarded as key residues and were ensured to be unaffected during the design process.

[0049] 1.2 Key site analysis based on structure prediction: Based on protein structure prediction and protein-substrate docking simulation, potential binding sites are identified through pocket analysis tools (such as P2Rank).

[0050] The function of a protein is determined by its structure. In the protein design process, for proteins with unknown crystal structures, such as fungal luciferase, structure prediction based on deep learning methods can provide a good understanding of the protein's folding and functional regions, identify potential functional sites, binding sites, and their interactions with substrates or ligands, and thus optimize the design strategy. In enzymatic reactions, substrate binding often relies on the enzyme's catalytic pocket, which is the key functional region for efficient catalysis. Figure 3 (A) and Figure 3 (B) The fungal luciferase-luciferin (3-hydroxymilk echinopsine) docking model predicted by the artificial intelligence (AI)-driven docking simulation method found that luciferin is located in a specific cavity of the fungal luciferase, such as Figure 3 (C) and Figure 3 As shown in (D) in the figure, the non-covalent bonds in the complex predicted by the fungal luciferase-luciferin docking model were further visualized in the visualization tool PyMOL, as shown in Figure 4 As shown, Figure 4 (A) is a schematic diagram of the hydrophobicity of fungal luciferase. Figure 4 (B) is a schematic diagram of the van der Waals force of the fungal luciferase-luciferin docking model. Figure 4 (C) is a hydrogen bond diagram of the fungal luciferase-luciferin docking model. Figure 4 (D) is a schematic diagram of the ionic bonds in the fungal luciferase-luciferin docking model; these biochemical features help to better understand the binding mode between fungal luciferase and luciferin and help to determine the scope of functional residues.

[0051] It should be noted that the docking simulation operation of proteins and substrate molecules involves using AI-driven docking simulation methods, such as RoseTTAFold All-Atom (RFAA) and DiffDock, to deeply learn the patterns of protein-substrate interactions and identify potential binding sites in a short period of time, thereby achieving efficient prediction of the binding domain between proteins and substrate molecules. The docking situation of proteins and their substrate molecules is then reproduced using visualization tools, such as PyMOL.

[0052] In order to systematically identify the key structural sites between luciferase and its substrate, the present invention comprehensively adopts strategies such as structure prediction, molecular docking and binding pocket identification. First, the binding conformation of luciferin (3-hydroxylactocetin) and luciferase was predicted using two structural docking methods, RFAA and DiffDock. The results showed that the substrate was stably bound to the cavity region of the enzyme. Furthermore, the pocket recognition tool P2Rank was used to analyze the static structure of the enzyme, predict possible binding sites and record key residues. At the same time, taking into account the conformational changes of the enzyme during the catalytic process and the dynamics of potential binding sites under the "induced fit" mechanism, a spherical space with a radius of 10Å was constructed with the substrate as the center, and amino acid residues adjacent to the substrate were supplemented and screened as key sites that may be involved in binding and catalysis.

[0053] Furthermore, an in-depth analysis of the atomic interactions between the enzyme and substrate in the docked structure was conducted to more accurately identify key structural residues. Hydrogen bonds (distance <3.5 Å, angle tolerance <30°) were identified using the PyMOL visualization tool. Combined with interaction criteria such as van der Waals forces (center distance <4 Å, overlap ≥0.4–1.0 Å) and ionic bonds / salt bridges (distance between positively and negatively charged residues <4 Å), residues potentially playing a key role in binding stability and catalytic mechanism were screened.

[0054] Finally, by integrating the P2Rank prediction results, docking model analysis, spatial neighborhood screening and atomic-level interaction characteristics, 92 key structural sites were identified (46F, 49V, 50G, 51P, 52S, 53Y, 54A, 55P, 56Q, 60G, 63I, 64V, 66V, 67L, 68S, 69L, 70F, 71R, 74Q, 97E, 98G, 99T, 103I, 104T, 105S, 106H, 107I, 108I, 109Q, 110R, 111Q, 130I, 146S, 147K, 148F, 149E, 150F, 151H, 152A, 153K, 154A, 155I, 156F, 157 7L,166P,167L,168N,169I,176R,177R,178T,179K,181E,182I,183A,184H ,185M,186H,187D,188Y,189H,190D,191C,192T,193L,194H,195L,196A,1 97L,205V,212Q,213R,214H,215P,216L,217A,218G,221V,222P,223G,224P,225P,228W,229T,230F,231L,233A,234P,235R,241R,242V,243V). Its distribution is as follows Figure 5 shown.

[0055] 1.3 Key sites based on evolutionary analysis: Multiple sequence alignment (MSA), information entropy calculation, hidden Markov model (HMM) and other methods were used to screen highly conserved sites (information entropy < 0.3 and conservation score > -1.5).

[0056] In order to more comprehensively cover all key site information of luciferase, based on the large framework of deep learning and bioinformatics, this paper adopts multiple evolutionary analysis methods to extract potential patterns and information from large-scale sequence and structural data.

[0057] First, the multiple sequence alignment (MSA) method can systematically compare multiple homologous sequences, revealing the conservation and variability between sequences, and then inferring key functional sites, evolutionary relationships, and structural characteristics. By performing a simple analysis of the sequence files of luciferases from 43 fluorescent species, the frequency of each amino acid at each position is counted, and the information entropy is calculated. , 134 sites with information entropy less than 0.3 were screened as conserved sites; among them, Indicates the information entropy of the site. Indicates the number of different amino acid types appearing at this site, An index representing the amino acid type, Indicates the The frequency of occurrence of an amino acid at this position.

[0058] Secondly, to more comprehensively cover conserved sites, the present invention also uses DEEPMSA2, a deep learning-based multiple sequence alignment method. It can automatically extract features from sequence data using convolutional neural networks (CNNs). By learning the deep structural relationships and evolutionary information between sequences, it provides more accurate alignment results when dealing with complex sequence variation and diversity. DEEPMSA2 was used to analyze the target fungal luciferase and obtain a sequence logo. This efficient visualization tool is widely used to display conservation and variation patterns in DNA, RNA, or amino acid sequences. By displaying the relative heights of different bases or amino acids at each position, it intuitively reflects the frequency of occurrence and degree of conservation of that position, and 101 candidate fixed sites were selected based on this.

[0059] In addition, based on the above-mentioned multiple sequence alignment file, the present invention also uses the HMMER software to construct a probabilistic model of the protein family. HMMER adopts the Profile Hidden Markov Models (Profile HMMs) method, which is an extension of the hidden Markov model (HMM). It combines multiple sequence alignment information with the probabilistic framework of HMM and is widely used in the analysis of protein and DNA sequences. This method reveals the possibility of observing a specific observed state under a specific hidden state by calculating the emission probability from the hidden state to the observed state. In order to analyze the probabilistic model of the protein family, the present invention selected sites with an emission probability greater than 0.3 and a conservation score higher than -1.5 as conserved sites, analyzed the MSA results of the luciferase amino acid sequence, and determined the information entropy of the site by statistically calculating the score values ​​of different amino acids at each site. , Indicates the The score of the amino acid at the site is calculated and expressed by Get the conservation score of each site, and then check whether the conservation score of each site is greater than -1.5; the maximum emission probability of each site is , by checking whether the maximum emission probability of each site is greater than 0.3, 111 sites that meet the above two conditions are finally screened as potential fixed sites.

[0060] By combining the above methods and the evolution-related sites reported in the literature, the repeated sites were extracted, and a total of 186 fixed sites (30L, 36F, 37P, 39I, 40R, 41R, 42D, 43Y, 45T, 46F, 47L, 48E, 49V, 50G, 51P, 52S, 53Y, 54A, 55P, 56Q, 57N, 59R, 60G, 61Y, 62I, 63I, 64V, 66V, 67L, 68S, 69L, 70F, 71R, 73E, 74Q, 77L, 79I, 80Y, 81D, 83L, 84P, 85E, 86K, 87R, 89W) were obtained based on evolutionary analysis. ,90L,92D,93L,94P,96R,98G,99T,100R,101P,102S,103I,104T,105S ,106H,107I,108I,109Q,110R,111Q,112R,113T,114Q,117D,120F,125 L,126I,127D,128K,129V,130I,132R,133V,134Q,135A,136R,137H,13 8T,141T,143L,144S,145T,146S,147K,148F,149E,150F,151H,152A,1 54A,155I,156F,157L,159P,161I,163I,165D,166P,169I,170P,171S ,172H,173D,174T,175V,176R,177R,178T,179K,180R,181E,182I,183 A,184H,185M,186H,187D,188Y,189H,190D,192T,193L,194H,195L,1 96A,197L,198A,199A,200Q,201D,203K,204E,205V,206L,207K,208K, 209G,210W,211G,212Q,213R,214H,215P,216L,217A,218G,219P,220G,221V,222P,223G,224P,225P,226T,227E,228W,229T,230F,231L,232Y,233A,234P,235R,236N,237E,238E,239E,240A,241R,242V,243V,244E,246I,247V,248E,249A,250S,251I,253Y,254M,255T,256N). Figure 6 shown.

[0061] 1.4 Final determination of fixed sites: The intersection of the structure prediction and evolutionary analysis sites was taken, and then the union was taken with the experimental sites in the literature to form a conserved region accounting for 35%-38% of the total sequence length, which was used as the fixed site.

[0062] The present invention adopts a multi-dimensional integration strategy to identify and verify the key amino acid residues of luciferase, and constructs an analysis framework covering literature sites, structure-related sites and evolutionary sites. Among them, the literature sites include 13 experimentally verified functional key residues; there are 92 structural sites, which are identified through three-dimensional structure analysis and are responsible for maintaining protein stability and function; there are 186 evolutionary sites, which are highly conserved among species, reflecting their functional importance. Figure 7 As shown in the figure, through correlation heat map analysis, it was found that the correlation between the fixed sites obtained based on structure prediction and evolutionary analysis screening was high, suggesting that the structure and evolutionary information were highly consistent, while the correlation between the fixed sites determined by literature information and the above two was significantly reduced, suggesting that although the structure and evolutionary information were highly consistent, there were still deviations from the currently known experimental results.

[0063] In order to further optimize the fixed site design method, the present invention innovatively designed a set of fixed site determination methods, such as Figure 8 As shown, Figure 8 (a) is a schematic diagram of the key site set of fungal luciferase predicted by structure. Figure 8 (b) is a schematic diagram of the key site set of fungal luciferase through evolutionary analysis. Figure 8(c) is a schematic diagram of the key site set of fungal luciferase from literature information; specifically, the intersection of the fixed sites (i.e., key conserved sites) obtained based on structure prediction and evolutionary analysis was taken, and then the union was taken with the fixed sites determined by literature information. Finally, 95 key sites were screened out, accounting for 35.58% of the total sequence length (46F, 49V, 50G, 51P, 52S, 53Y, 54A, 55P, 5 6Q,60G,61Y,63I,64V,66V,67L,68S,69L,70F,71R,74Q,98G,99T,103I,104T,105S,10 6H,107I,108I,109Q,110R,111Q,127D,130I,136R,144S,146S,147K,148F,149E,150F, 151H,152A,154A,155I,156F,157L,166P,169I,173D,176R,177R,178T,179K,181E,18 2I,183A,184H,185M,186H,187D,188Y,189H,190D,191C,192T,193L,194H,195L,196A, 197L,205V,212Q,213R,214H,215P,216L,217A,218G,221V,222P,223G,224P,225P,228W,229T,230F,231L,233A,234P,235R,237E,238E,241R,242V,243V). The specific distribution of these 95 key sites is as follows: Figure 9 (A) and Figure 9 As shown in (B).

[0064] This innovative method for identifying fixed sites effectively balances the varying predictions from AI and bioinformatics tools, identifying core potential sites that, along with real-world experimental results, serve as key sites for fixation. These sites will be retained during the design and engineering of fungal luciferases, supporting the stable realization of protein function. This approach addresses the diversity and adaptability of protein design sequences, breaking through the efficiency bottlenecks of traditional approaches.

[0065] In general, the strategy for identifying conserved regions as fixed sites combines a comprehensive approach of literature research, evolutionary analysis, and structure prediction. Specifically, the following steps are taken: First, by combining existing literature data, known key functional sites can be identified and confirmed. These functional sites are often core regions of protein function and stability and need to remain unchanged during the design process. Second, the function of a protein is determined by its structure. Generally speaking, in an enzymatic reaction, the substrate binds to the active region of the enzyme, which is the catalytic pocket. For protein docking results predicted by AI-driven docking simulation methods (such as RFAA and DiffDock), the P2Rank program predicts the pocket region of the protein and records the key residue positions of the pocket. In addition, through evolutionary analysis, such as performing multiple sequence alignments (MSA) of similar proteins from similar species, the frequency of each amino acid at each position is counted. Determining conserved sites by calculating information entropy is a key step in analyzing protein sequences from an evolutionary perspective. To achieve the most comprehensive coverage of conserved sites, we employed graphical methods to visualize conservation and patterns within DNA, RNA, or amino acid sequences. For example, we used DEEPMSA2 to analyze target proteins and generate sequence logos. This display of the height of different bases or amino acids at each position in the sequence reflects the frequency of occurrence at that position and the degree of conservation at that position. Furthermore, based on the aforementioned multiple sequence alignment files, we constructed a hidden Markov model to establish a probabilistic model of protein families to explore evolutionary relationships and further improve the accuracy of identifying conserved regions.

[0066] 2. Design protein functional domain skeleton based on fixed sites.

[0067] The specific operation of protein functional domain backbone design is as follows: Based on fixed conserved sites, the denoising diffusion probabilistic model (DDPM) is introduced. The protein structure is abstracted into six-dimensional information of the translation vector and rotation matrix of each amino acid, and noise is gradually added as training input. During the denoising process, a loss function is constructed to measure the gap between the predicted structure and the true structure at each step, driving the model to restore the true protein structure. Ultimately, the relative positions of each amino acid are dynamically adjusted from the noise to accurately construct the folding framework of the functional domain. Among them, diffusion models include but are not limited to the denoising diffusion probabilistic model (DDPM) and RFDiffusion. The model optimizes the protein backbone structure through gradual noise addition and denoising to meet specific functional requirements. The generated backbone structure contains more than 500 conformational variants.

[0068] Protein scaffold design is a key technology in protein engineering, aiming to construct stable and functionally specific protein structures to optimize natural proteins or create novel proteins. Protein scaffold design generally employs two strategies: optimizing existing scaffolds and de novo design. Optimizing existing scaffolds involves sequence adjustments, key site mutations, and structural fine-tuning to improve stability or enhance function. Computer-aided design (CAD) often utilizes tools such as Rosetta, FoldX, and PyRosetta, which screen for optimal mutations through energy calculations. De novo design, however, utilizes basic secondary structures (such as α-helices, β-sheets, and loops) to assemble specific structural topologies or functional requirements. In recent years, deep learning models such as AlphaFold3 and RoseTTAFold have significantly improved the success rate of de novo design. Furthermore, optimization methods based on evolutionary information are also an important area of ​​focus in protein scaffold design. By analyzing large numbers of homologous sequences, key co-evolving amino acid pairs can be identified and used to optimize scaffold structures. Diffusion models and generative adversarial networks (GANs) have demonstrated strong capabilities in protein structure generation.

[0069] In the present invention, a scaffold enzyme functional site design strategy was adopted to construct a stable scaffold structure around the key motifs that support and stabilize its biological functions, and an AI-based diffusion model (such as RFDiffusion) was used to generate the skeleton of the fungal luciferase with fixed sites. RFDiffusion is a protein design framework based on a diffusion model proposed by David Baker's laboratory. It simulates the inverse process of diffusion from random noise to the target structure, and gradually generates a high-precision protein skeleton that meets the design constraints. Compared with traditional design methods based on sampling or energy optimization, RFDiffusion can not only globally capture the geometric properties of the protein skeleton, but also finely reconstruct key sites based on functional constraints, thereby realizing the de novo design of complex structures such as binders and symmetric polymers. For protein skeleton design that requires fixed sites (Anchor Sites), RFDiffusion provides an efficient method that can optimize the topological structure of the overall skeleton while retaining functional amino acids, making it more stable and meeting specific functional requirements.

[0070] Combined with previous analysis, 95 fixed sites of fungal luciferase were locked, and masks were set using AI-based diffusion models (such as RFDiffusion) to ensure that key sites remain unchanged, while allowing the skeletons of other areas to be freely generated. This defined areas that required high fidelity and free exploration, and guided the structure to converge towards positions that met the constraints during the generation process. In this process, starting from the wild-type fungal luciferase, we gradually converged to a physically reasonable protein backbone, allowing the model to be diffusely generated under restricted conditions, and gradually optimized the protein backbone to satisfy both global stability and fixed sites (diffusion model input format: 45-45 / A46-46 / 2-2 / A49-56 / 3-3 / A60-61 / 1-1 / A63-64 / 1-1 / A66-71 / 2-2 / A74-74 / 23-23 / A98-99 / 3-3 / A103-111 / 15-15 / A127-127 / 2-2 / A130-130 / 5-5 / Based on the above constraints, 500 new protein backbone structures were finally obtained. An example of the structure of one of the generated protein backbones is shown below. Figure 10 shown.

[0071] 3. Protein sequence prediction based on backbone structure.

[0072] Based on the designed protein functional domain backbone structure without side chain atoms and sequence information, a graph neural network model is further introduced to accurately model the complex relationship between protein sequence and three-dimensional structure. For example, a deep learning model based on a graph neural network can learn the interactions and spatial constraints between amino acids to achieve accurate sequence-structure mapping. By optimizing each amino acid position in a given backbone structure, the amino acid sequence that best matches it is dynamically predicted and generated. Sequence prediction uses a deep learning-based protein sequence-structure mapping method, including but not limited to ProteinMPNN, with a temperature parameter of 0.1-0.2 controlling sequence sampling. 24 candidate sequences are generated for each backbone, and the number of predicted candidate sequences exceeds 36,000.

[0073] Protein sequence prediction based on the protein backbone plays a crucial role in protein design and engineering, particularly in enzyme optimization, antibody design, and vaccine development. Its core goal is to infer the corresponding amino acid sequence based on the known or predicted three-dimensional protein structure (i.e., the backbone) so that it can correctly fold into this structure and possess the desired function. This process requires not only ensuring structural stability but also ensuring that the sequence can perform the desired function (such as catalysis, binding, and stability). Therefore, the following principles are crucial when considering design approaches. First, sequence-structure compatibility: Each amino acid residue is constrained in three-dimensional space by its neighbors and the folding of the entire molecule. The designed sequence must be compatible with the backbone geometry to fold into a stable three-dimensional structure. Second, functional residue retention: For functional proteins (such as enzymes and receptors), certain specific amino acid residues (such as catalytic sites and substrate binding sites) must maintain their function. Therefore, special attention must be paid to the amino acid type and position of these key sites. Finally, sequence stability: The designed sequence must be sufficiently stable to allow the protein to maintain its three-dimensional structure under varying conditions (such as temperature, pH, and ionic strength).

[0074] Sequence prediction methods based on energy functions (e.g., Rosetta) or deep learning-based sequence-structure mapping (e.g., AlphaFold), as well as deep learning models and graph neural networks (GNNs) (e.g., ProteinMPNN), are currently the most mainstream approaches for protein sequence prediction. In the present invention, based on the protein backbone generated by the diffusion model, a deep learning-driven protein sequence design model (e.g., ProteinMPNN) is used for protein sequence prediction. ProteinMPNN is a deep learning-driven protein sequence design tool developed by David Baker's laboratory that optimizes amino acid sequences to ensure the stability of the target protein structure, enabling it to successfully fold into a predetermined three-dimensional structure. Compared to traditional design methods that rely on physical models (e.g., Rosetta), ProteinMPNN significantly improves sequence recovery, enabling more accurate protein sequence recovery and significantly increasing design efficiency. The optimized fungal luciferase backbone structure generated by the diffusion model undergoes target function-based design optimization to ensure stability and functional requirements. Using it as input for sequence design can effectively establish a mapping between structure and sequence, thereby generating a sequence that correctly folds and possesses the desired function.

[0075] In practice, the present invention fixed functional sites within the backbone using the same design constraints as the diffusion model. These functional sites may include catalytic sites, substrate binding sites, or other regions directly related to fungal luciferase activity. During the design process, considering the close correlation between sampling temperature and the model's selection of amino acids from a probability distribution, relatively low temperatures (0.1, 0.15, and 0.2) were used to guide model sampling. Lower temperatures tend to favor the most likely amino acids in the probability distribution, resulting in more concentrated and similar sequences. This design choice ensures greater determinism and conserved sequences, better meeting the design requirements of the present invention. A lower temperature sampling strategy is particularly effective when functional sites need to remain stable. To ensure diversity and cover a wider range of sequence space, 24 different sequences were generated for each target backbone. This strategy ensures a diverse set of sequences even at lower sampling temperatures, providing a greater number of potential candidate sequences. Throughout the design process, a total of 36,000 candidate luciferase sequences were generated for 500 fungal luciferase backbones.

[0076] 4. Scoring screening based on prediction sequence.

[0077] Although the above design method has generated a huge number of sequences, in order to ensure the quality and functional potential of the protein sequences designed by AI, the predicted sequences need to be screened through a strict scoring mechanism. First, the ProteinMPNN_score and sequence repetitiveness (ARR_score) were scored for the 36,000 predicted sequences at the sequence level. ProteinMPNN_score is an important indicator to measure the compatibility of the predicted sequence with the target structure. A sequence with ProteinMPNN_score <1 indicates that it has good folding stability and can correctly form the desired three-dimensional conformation; the sequence repetitiveness score is used to evaluate the degree of repeated amino acids in the sequence, and ARR_score>-2 indicates that the sequence has high diversity in the local region, avoiding irrational sequences of continuous repetition of multiple amino acids, which is crucial to ensuring the functional activity of the protein. Through these two preliminary scoring threshold screenings at the sequence level, 1,640 high-quality fungal luciferase sequences were finally screened out from 36,000 candidate sequences, such as Figure 11 As shown, it accounts for 4.55% of the total number of designed sequences. Figure 11Each row represents a predicted protein, displaying the name automatically generated during the RFDiffusion and ProteinMPNN processes, the predicted amino acid sequence, the sampling temperature, and two ProteinMPNN scores: ProteinMPNN_score and ARR_score. These preliminary screening sequences effectively eliminate irrational sequences that lack reasonable structure or function. They not only meet the requirements for structural stability but also have good functional potential, laying the foundation for subsequent sequence screening and applied research.

[0078] In addition, the present invention also introduced deep learning-based structure prediction models (such as ESMFold and AlphaFold3) to perform structure prediction on 1,640 sequences that passed the initial screening. These sequences were screened based on pLDDT (Predicted Local Distance Difference Test) and pTM (Predicted Template Modeling score) scores to further ensure the structural stability and functional potential of these sequences. pLDDT is a metric that assesses the confidence of the local structure prediction for each residue, ranging from 0 to 100. A high pLDDT value indicates high confidence in the structure prediction for that region. Specifically, pLDDT ≥ 90 indicates very high confidence, 70 ≤ pLDDT < 90 indicates high confidence, 50 ≤ pLDDT < 70 indicates low confidence, and pLDDT < 50 indicates very low confidence. pTM is a metric that measures the confidence of the overall structure prediction, ranging from 0 to 1. A higher pTM value indicates that the predicted overall folded structure is likely to be closer to the true structure. Typically, a pTM score above 0.5 indicates that the overall predicted fold of the protein is likely similar to the true structure. Based on the prediction and evaluation results, sequences with pLDDT values ​​above 80 and pTM values ​​above 0.85 were prioritized to ensure high confidence in both local and global structures.

[0079] To comprehensively evaluate the confidence of the local and global structures of a sequence and ensure the scientificity and accuracy of the screening process, the present invention calculates the ranking percentages of the pLDDT and pTM scores of each sequence obtained from the above screening. A comprehensive score, S, is introduced, which weights the ranking percentages of the pLDDT and pTM scores by adding them together with specific weights. The specific formula is as follows: . Among them, Rank_pLDDT and Rank_pTM represent the ranking percentages of pLDDT and pTM, respectively, and ω1 and ω2 are the weight coefficients of pLDDT and pTM, respectively. Taking into account that pLDDT reflects the local structure confidence, pTM reflects the overall folding accuracy, and the overall structure has a more critical impact on protein function, ω1=0.4 and ω2=0.6 are set. The sequences are sorted in ascending order according to the comprehensive score S. The lower the score, the higher the comprehensive ranking. Finally, the top 20 sequences with the comprehensive score were selected as the final candidate sequences of luciferase to ensure that they have high confidence in both local and overall structures. The final predicted sequence examples and the predicted structures based on the example predicted sequences are shown as follows. Figure 12 (A) and Figure 12 As shown in (B), the final screening results are as follows Figure 13 shown. Figure 13 In the figure, each row represents a predicted protein, showing the name of the protein automatically generated in the ESMFold structure prediction after pMPNN (i.e. ProteinMPNN), the predicted sequence, two scores of ESMfold (pLDDT and pTM), sampling temperature, two scores of ProteinMPNN (pMPNN_score and ARR_score, pMPNN_score is ProteinMPNN_score), and the weighted score.

[0080] Based on the above information, if Figure 14 As shown, the present invention has formed a specific screening scheme for fusion sequence and structure evaluation for the screening of fungal luciferase in the embodiments of the present invention, ensuring that the selected luciferase variants have the best performance in terms of structural stability and functional potential. It should be noted that the specific scoring threshold and ranking weight settings in this screening scheme can be dynamically adjusted according to the designed target protein and the protein performance to be optimized to meet different design requirements.

[0081] It should be noted that the types of proteins to which the present invention can be applied also include: proteins whose ligands are small molecules, such as transport proteins (such as glucose transporters), receptor proteins (such as certain G protein-coupled receptors), structural proteins (such as rhodopsin), regulatory proteins (such as certain protein kinases), etc.

[0082] Corresponding to the aforementioned embodiments of the protein optimization design and screening method based on artificial intelligence algorithms, the present invention also provides embodiments of the protein optimization design and screening device based on artificial intelligence algorithms.

[0083] See also Figure 15The protein optimization design and screening device based on artificial intelligence algorithm provided in an embodiment of the present invention includes one or more processors and a GPU processor, which is used to implement the protein optimization design and screening method based on artificial intelligence algorithm in the above embodiment.

[0084] The embodiment of the protein optimization design and screening device based on artificial intelligence algorithm of the present invention can be applied to any device with data processing capability, and the device with data processing capability can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capability in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for execution. From the hardware level, if Figure 15 As shown, it is a hardware structure diagram of any device with data processing capability where the protein optimization design and screening device based on artificial intelligence algorithm of the present invention is located, except Figure 15 In addition to the central processing unit, memory, network interface, non-volatile memory, GPU processor, and I / O device shown, any device with data processing capabilities in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.

[0085] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0086] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present invention. A person of ordinary skill in the art can understand and implement the present invention without inventive work.

[0087] Corresponding to the aforementioned embodiments of the protein optimization design and screening method based on artificial intelligence algorithm, an embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the protein optimization design and screening method based on artificial intelligence algorithm in the aforementioned embodiments is implemented.

[0088] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The computer-readable storage medium may also be any device with data processing capabilities, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.

[0089] The present invention can be widely applied to different computing environments such as stand-alone high-performance computing, cluster computing, and cloud computing, ensuring that the protein optimization and screening process can be efficiently completed under various computing power configurations.

[0090] The above contents are only preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. For those skilled in the art, various changes, combinations, simplifications, modifications, substitutions, and readjustments should all be considered equivalent replacements and will not depart from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and within the scope of protection of the present invention, many other equivalent embodiments are also included.

Claims

1. A protein optimization design and screening method based on artificial intelligence algorithm, characterized in that: The following steps are involved: (1) Using artificial intelligence-driven protein-substrate docking simulation methods to predict the binding domains of proteins and substrates; combining literature analysis, structure prediction, and evolutionary analysis to screen key conserved sites; taking the intersection of the sites screened by structure prediction and evolutionary analysis, and then taking the union of the sites extracted from the literature to form a conserved region, which is used as a fixed site; (2) Generate protein domain skeletons under fixed site constraints using a diffusion model; (3) Use the graph neural network model to predict the sequence of the protein functional domain skeleton generated in step (2); (4) Scoring and screening are performed based on the predicted sequence, including sequence-level scoring and structure-level screening, to obtain candidate protein variants.

2. The protein optimization design and screening method based on artificial intelligence algorithm according to claim 1, characterized in that: In step (1), artificial intelligence-driven protein-substrate docking simulation methods, including RFAA and DiffDock, are used to identify potential binding sites, thereby predicting the binding domain between the protein and the substrate molecule; visualization tools are used to reproduce the docking of the protein and its substrate molecule, and to analyze and screen amino acid residues with specific types and strengths of non-covalent forces between the protein and its substrate molecule.

3. The protein optimization design and screening method based on artificial intelligence algorithm according to claim 1, characterized in that: In step (1), the screening of key conserved sites includes: extracting reported key functional sites based on literature information; identifying and screening potential binding sites based on protein structure prediction using pocket analysis tools and visualization tools; screening highly conserved sites based on evolutionary analysis using multiple sequence alignment, information entropy calculation, and hidden Markov model; taking the intersection of the sites obtained from structure prediction and evolutionary analysis screening, and then taking the union of the sites extracted from the literature to form a conserved region.

4. The protein optimization design and screening method based on artificial intelligence algorithm according to claim 3, characterized in that: When the protein is fungal luciferase, the intersection of the sites obtained by structure prediction and evolutionary analysis was taken, and then the union was taken with the sites extracted from the literature, and 95 key sites were screened out, accounting for 35.58% of the full length of the sequence; the key sites are: 46F, 49V, 50G, 51P, 52S, 53Y, 54A, 55P, 56Q, 60G, 61Y, 63I, 64V, 66V, 67L, 68S, 69L, 70F, 71R, 74Q, 98G, 99T, 103I, 104T, 105S, 106H, 107I, 108I, 109Q, 110R, 111Q, 127D, 130I, 136R, 144S, 146S, 147K, 148F, 149E, 150F, 151H, 152A ,154A,155I,156F,157L,166P,169I,173D,176R,177R,178T,179K,181E,182I, 183A,184H,185M,186H,187D,188Y,189H,190D,191C,192T,193L,194H,195L,19 6A,197L,205V,212Q,213R,214H,215P,216L,217A,218G,221V,222P,223G,224 P,225P,228W,229T,230F,231L,233A,234P,235R,237E,238E,241R,242V,243V.

5. The protein optimization design and screening method based on artificial intelligence algorithm according to claim 1, characterized in that: Step (2) is as follows: based on fixed sites, a denoising diffusion probability model is introduced to abstract the protein structure into six-dimensional information of each amino acid translation vector and rotation matrix, and noise is gradually added as training input; During the denoising process, a loss function is constructed to measure the gap between the predicted structure and the true structure at each step, and to drive the model to restore the true protein structure. Ultimately, the relative positions of each amino acid are dynamically adjusted from the noise, the folding framework of the functional domain is constructed, and the protein functional domain skeleton is generated.

6. The protein optimization design and screening method based on artificial intelligence algorithm according to claim 1, characterized in that: Step (3) is specifically as follows: based on the generated protein functional domain backbone structure without side chain atoms and sequence information, a graph neural network model is introduced to model the relationship between the protein sequence and the three-dimensional structure, that is, a deep learning model based on a graph neural network is used to learn the interactions and spatial constraints between amino acids to achieve sequence-structure mapping; by optimizing each amino acid position of a given backbone structure, the amino acid sequence that best matches it is dynamically predicted and generated.

7. The protein optimization design and screening method based on artificial intelligence algorithm according to claim 1, characterized in that: In step (4), the structural level screening includes: performing structural prediction on high-quality sequences that have passed the preliminary screening, setting a screening threshold, and screening based on pLDDT and pTM scores, with the pLDDT threshold being no less than 80 and the pTM threshold being no less than 0.85, to ensure the structural stability and functional potential of these sequences; wherein, pLDDT is an indicator for evaluating the confidence of the local structure prediction of each residue, and a higher pLDDT value indicates a higher confidence in the local structure prediction; pTM is an indicator for measuring the confidence of the overall structure prediction, and a higher pTM value indicates that the predicted overall folding structure is closer to the true structure; based on the weighted sum of the ranking percentages of pLDDT and pTM, the protein variant with the highest comprehensive score is screened.

8. A protein optimization design and screening device based on artificial intelligence algorithm, characterized in that: The method comprises one or more central processing units and GPU processors, and is used to implement the protein optimization design and screening method based on artificial intelligence algorithm according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the protein optimization design and screening method based on an artificial intelligence algorithm as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Methods and systems for aligning sequences

    CN105637098A

  • Molecular modification method for glycosyltransferase based on artificial intelligence

    CN119724342A