Protein optimization design and screening method and device based on artificial intelligence algorithm

Through artificial intelligence algorithms, the multi-dimensional bioinformatics analysis is integrated and the luciferase design is optimized, which solves the problem of inefficiency of traditional methods, and achieves efficient luciferase performance improvement and the development of autonomous luminescence systems, which promotes the development of biomedicine and bioimaging fields.

CN120260679AActive Publication Date: 2025-07-04ZHEJIANG LAB

Patent Information

Application Number
CN202510756581.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Traditional protein design methods are time-consuming, laborious, and inefficient, and are difficult to effectively solve especially when complex functions are optimized. The existing luciferase has limited luminescence intensity in heterologous hosts, limiting its practical application.

Method used

Artificial intelligence algorithm is used to integrate multi-dimensional bioinformatics analysis, and through protein-substrate docking simulation, graph neural network model and diffusion model, key conserved sites are identified, protein functional domain backbone is generated, sequence prediction and scoring screening is performed, and luciferase performance is optimized.

Benefits of technology

It improves the efficiency and accuracy of protein design, improves the luminescence intensity and stability of luciferase, provides an efficient development solution for autonomous luminescence system, and promotes technological progress in the fields of biomedicine and bioimaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260679A_ABST
    Figure CN120260679A_ABST
Patent Text Reader

Abstract

The invention discloses a protein optimization design and screening method and device based on an artificial intelligence algorithm, and the method takes fungal luciferase as an example, and integrates multi-dimensional bioinformatics analysis and deep learning technology to realize efficient protein engineering transformation and optimization. The method comprises the following steps: firstly, identifying a binding domain of luciferase and fluorescein by utilizing an AI-driven molecular docking simulation method, and determining a key conservative site by combining literature, evolutionary analysis and structural prediction; then, generating a protein functional domain skeleton under the constraint of a fixed site by adopting a diffusion model, and performing protein sequence prediction by utilizing a graph neural network model; and finally, scoring and screening the generated sequences to obtain high-stability candidate variants. According to the method, a conservative site dynamic fusion strategy is innovatively constructed, a fixed region is optimized through logic of structure prediction, evolution site intersection priority and literature site union set expansion, the diversity and adaptability of a protein design sequence are met, and the efficiency bottleneck of a traditional scheme is broken through.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics and artificial intelligence-driven computational biology, and particularly relates to a method and device for protein optimization design and screening based on artificial intelligence algorithms. Background Art

[0002] With the continuous progress in the fields of biology and biotechnology, protein design and optimization have gradually become a key link in promoting disease treatment, precision medicine, and biotechnological innovation. However, traditional protein design methods often rely on screening and mutation strategies in the laboratory, which are usually time-consuming, laborious, costly, and inefficient. Especially when facing complex function optimization, they face huge challenges. The introduction of artificial intelligence (AI) technology has brought unprecedented opportunities to protein design. Through advanced algorithms such as deep learning, reinforcement learning, and generative models, AI can automatically identify and mine potential laws and patterns in a vast amount of protein data, greatly improving the efficiency and accuracy of protein design. AI models, especially deep neural networks (DNNs) and graph neural networks (GNNs), can predict the functional characteristics of proteins without experimental data by learning the complex relationship between protein sequences and their three-dimensional structures, thus providing more efficient and accurate design solutions. In addition, generative models such as Protein Message Passing Neural Network (ProteinMPNN) and RoseTTAFold Diffusion (RFDiffusion) can automatically generate optimized candidate proteins with specific functions in the protein variant space. By simulating intermolecular interactions and stability, they can predict and design protein variants with optimal performance in a specific environment. This AI-driven protein design method breaks through the limitations of traditional experience dependence, making protein design and optimization more systematic, quantitative, and efficient, and promoting the rapid transformation from basic research to application development. In many practical applications, AI provides new possibilities for the directional optimization of protein functions. Especially in fields such as enzyme design, antibody engineering, and vaccine development, the addition of artificial intelligence technology is leading protein design towards a more intelligent and efficient direction.

[0003] Luciferase is an important class of bioluminescent enzymes that are widely used in multiple fields such as bioimaging, disease diagnosis, and monitoring of treatment effects. Luciferase generates photons through the catalytic oxidation reaction of luciferin. This process not only provides a visualization tool for scientific research but also has extensive applications in medical detection. Traditional luciferases usually rely on exogenously added substrates (luciferin), which not only increases the experimental cost but may also be limited by insufficient luciferin supply in some cases. In recent years, with the discovery of the bioluminescence pathway from the fungus Neonothopanus nambi, scientists have been able to construct self-luminescent systems in various eukaryotic organisms. This system converts caffeic acid into luciferin through the action of a series of enzymes, and finally, luciferase catalyzes the generation of photons. However, due to the poor performance of the native fungal luminescence system in heterologous hosts, the luminescence intensity is limited, restricting its potential in practical applications. The present invention brings a breakthrough in the optimized design of luciferase by introducing artificial intelligence technology. By combining deep learning and generative models, AI can mine key sites from the complex relationship between protein structure and function, systematically optimize the performance of luciferase, not only improve the luminescence intensity and stability of the enzyme but also generate luciferase variants with excellent performance in a short time, breaking through the limitations of traditional design methods and providing a more efficient and accurate solution for the development of self-luminescent systems. This AI-driven protein optimization design method not only enhances the application potential of luciferase but also provides strong support for technological progress in fields such as precision medicine and bioimaging. Summary of the Invention

[0004] The object of the present invention is to propose a method and device for protein optimization design and screening based on artificial intelligence algorithms in view of the deficiencies of the prior art.

[0005] To achieve the above object, the present invention provides a method for protein optimization design and screening based on artificial intelligence algorithms, including the following steps: (1) Using an artificial intelligence-driven protein-substrate docking simulation method to predict the binding domain of the protein and the substrate; combining literature analysis, structure prediction, and evolutionary analysis to screen key conserved sites; taking the intersection of the sites screened by structure prediction and evolutionary analysis, and then taking the union with the sites extracted from the literature to form a conserved region, and using this conserved region as the fixed sites; (2) Using a diffusion model to generate the backbone of the protein functional domain under the constraint of the fixed sites; (3) Using a graph neural network model to predict the sequence of the protein functional domain backbone generated in step (2); (4) Conducting scoring and screening based on the predicted sequence, including sequence-level scoring and structure-level screening, to obtain candidate protein variants.

[0006] Further, in step (1), an artificial intelligence-driven protein-substrate docking simulation method, including RFAA and DiffDock, is used to identify potential binding sites, thereby predicting the binding domain between a protein and a substrate molecule; a visualization tool is used to reproduce the docking of the protein and its substrate molecule, and amino acid residues with specific types and intensities of non-covalent interactions between the protein and its substrate molecule are analyzed and screened.

[0007] Further, in step (1), the screening of key conserved sites includes: extracting reported key functional sites based on literature information; identifying and screening potential binding sites based on protein structure prediction through a pocket analysis tool and a visualization tool; screening highly conserved sites based on evolutionary analysis using multiple sequence alignment, information entropy calculation, and hidden Markov model; taking the intersection of the sites screened by structure prediction and evolutionary analysis, and then taking the union with the sites extracted from the literature to form a conserved region.

[0008] Further, when the protein is fungal luciferase, the intersection of the sites screened by structure prediction and evolutionary analysis is taken, and then the union with the sites extracted from the literature is taken to screen out 95 key sites, accounting for 35.58% of the full length of the sequence; the key sites are: 46F, 49V, 50G, 51P, 52S, 53Y, 54A, 55P, 56Q, 60G, 61Y, 63I, 64V, 66V, 67L, 68S, 69L, 70F, 71R, 74Q, 98G, 99T, 103I, 104T, 105S, 106H, 107I, 108I, 109Q, 110R, 111Q, 127D, 130I, 136R, 144S, 146S, 147K, 148F, 149E, 150F, 151H, 152A, 154A, 155I, 156F, 157L, 166P, 169I, 173D, 176R, 177R, 178T, 179K, 181E, 182I, 183A, 184H, 185M, 186H, 187D, 188Y, 189H, 190D, 191C, 192T, 193L, 194H, 195L, 196A, 197L, 205V, 212Q, 213R, 214H, 215P, 216L, 217A, 218G, 221V, 222P, 223G, 224P, 225P, 228W, 229T, 230F, 231L, 233A, 234P, 235R, 237E, 238E, 241R, 242V, 243V.

[0009] Furthermore, step (2) is specifically as follows: Based on the fixed sites, a denoising diffusion probability model is introduced. The protein structure is abstracted into six-dimensional information of the translation vector and rotation matrix of each amino acid, and noise is gradually added as the training input. During the denoising process, a loss function is constructed to measure the gap between the predicted structure and the true structure at each step, and the model is driven to restore the true protein structure. Finally, the relative positions of each amino acid are dynamically adjusted from the noise to construct the folding framework of the functional domain and generate the protein functional domain backbone.

[0010] Furthermore, step (3) is specifically as follows: Based on the generated protein functional domain backbone structure without side-chain atoms and sequence information, a graph neural network model is introduced to model the relationship between the protein sequence and the three-dimensional structure. That is, a deep learning model based on the graph neural network is used to learn the interactions and spatial constraints between amino acids to achieve sequence-structure mapping. By optimizing each amino acid position of the given backbone structure, the amino acid sequence that best matches it is dynamically predicted and generated.

[0011] Furthermore, in step (4), the sequence-level scoring includes: scoring ProteinMPNN_score and sequence repeatability ARR_score for the predicted and generated sequence at the sequence level. Among them, ProteinMPNN_score is an index to measure the fitness of the predicted sequence to the target structure. A sequence with a ProteinMPNN_score less than 1 indicates better folding stability and the ability to correctly form the desired three-dimensional conformation. The sequence repeatability score is used to evaluate the degree of repeated amino acids in the sequence. An ARR_score greater than -2 indicates a higher diversity in the local region of the sequence. Through these two scores, high-quality sequences are initially screened.

[0012] Furthermore, in step (4), the structure-level screening includes: performing structure prediction on the high-quality sequences that have passed the initial screening, and setting a screening threshold. Screening is based on pLDDT and pTM scores. The pLDDT threshold is not lower than 80, and the pTM threshold is not lower than 0.85 to ensure the structural stability and functional potential of these sequences. Among them, pLDDT is an index to evaluate the confidence of the local structure prediction of each residue. The higher the pLDDT value, the higher the confidence in the structure prediction of this region. pTM is an index to measure the confidence of the overall structure prediction. The higher the pTM value, the closer the predicted overall folding structure is to the true structure. Based on the weighted sum of the ranking percentages of pLDDT and pTM, the protein variant with the highest comprehensive score is screened.

[0013] To achieve the above object, the present invention also provides a protein optimization design and screening device based on an artificial intelligence algorithm, including one or more processors and a GPU processor, for implementing the above-mentioned protein optimization design and screening method based on an artificial intelligence algorithm.

[0014] To achieve the above object, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned protein optimization design and screening method based on artificial intelligence algorithms is implemented.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. It provides a systematic and new process and method for protein optimization design based on artificial intelligence algorithms, breaking through the limitations of manual screening and experience dependence in traditional protein design, and significantly improving the design efficiency and accuracy. 2. It establishes an innovative method for identifying key sites of proteins, integrating and concatenating protein sequence, structure, and function-related information from different sources, providing a new strategy for analyzing and locking key sites of proteins, effectively exploring and designing non-conserved regions while ensuring protein function, and enhancing the diversity and adaptability of protein sequences. 3. It constructs an efficient, flexible, and reliable protein design framework, integrating the latest artificial intelligence technologies and bioinformatics tools to meet different types of protein design requirements, providing a new solution to the challenges in traditional protein engineering, promoting the rapid development of fields such as biomedicine, environmental engineering, and food industry, and laying a solid foundation for future interdisciplinary innovation and industrial application. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions or specific implementation manners in the present invention, the following will briefly introduce the drawings required for the technical description and specific implementation manners of the present invention. Obviously, the drawings in the following description are some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0017] Figure 1 It is a flow chart of protein optimization design taking fungal luciferase as an example in the present invention; Figure 2 It is a distribution map of key sites of fungal luciferase from literature information in the present invention; Figure 3 It is a schematic diagram of the docking state of fungal luciferase - luciferin in the present invention; among them, Figure 3 in (A) is a schematic diagram of the helical structure of fungal luciferase, Figure 3 in (B) is a schematic diagram of luciferin (3-hydroxyhyptis suaveolens alkaloid), Figure 3 in (C) is a schematic diagram of the helical structure of the fungal luciferase - luciferin docking model, Figure 3 in (D) is a schematic diagram of the 3D structure of the fungal luciferase - luciferin docking model; Figure 4 Schematic diagram of the biochemical properties of the fungal luciferase in the present invention; wherein, Figure 4 In (A) is the schematic diagram of the hydrophobicity of the fungal luciferase, Figure 4 In (B) is the schematic diagram of the van der Waals force of the fungal luciferase-luciferin docking model, Figure 4 In (C) is the schematic diagram of the hydrogen bond of the fungal luciferase-luciferin docking model, Figure 4 In (D) is the schematic diagram of the ionic bond of the fungal luciferase-luciferin docking model; Figure 5 Distribution map of the key sites of the fungal luciferase predicted by structure in the present invention; Figure 6 Distribution map of the key sites of the fungal luciferase analyzed by evolution in the present invention; Figure 7 Heat map of the correlation of the potential fixation sites of the fungal luciferase in the present invention; Figure 8 Schematic diagram of the generation logic of the fixation sites of the fungal luciferase in the present invention; wherein, the shaded part represents the final fixation sites, Figure 8 In (a) is the schematic diagram of the set of key sites of the fungal luciferase predicted by structure, Figure 8 In (b) is the schematic diagram of the set of key sites of the fungal luciferase analyzed by evolution, Figure 8 In (c) is the schematic diagram of the set of key sites of the fungal luciferase from literature information; Figure 9 Schematic diagram of the final fixation sites of the fungal luciferase in the present invention; wherein, Figure 9 In (A) is the schematic diagram of the position of the fixation site (white and bold) of the fungal luciferase in the sequence, Figure 9 In (B) is the schematic diagram of the position of the fixation site (white) of the fungal luciferase in the secondary structure; Figure 10 Example diagram of the generated backbone of the fungal luciferase in the present invention; Figure 11 Example diagram of the output of the screening of the predicted sequence scores of the fungal luciferase in the present invention; Figure 12 Example diagram of the predicted sequence and corresponding predicted structure of the fungal luciferase in the present invention; wherein, Figure 12 In (A) is the example diagram of the predicted sequence of the fungal luciferase, Figure 12 In (B) is the schematic diagram of the predicted structure of the fungal luciferase based on the example predicted sequence; Figure 13 Example diagram of the final candidate sequence of the fungal luciferase in the present invention; Figure 14Schematic diagram of the screening process for fungal luciferase scoring based on the predicted sequence in the present invention; Figure 15 It is a schematic structural diagram of the device of the present invention. Detailed implementation manners

[0018] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts are within the protection scope of the present invention.

[0019] Here, exemplary embodiments will be described in detail, and examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are only examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0020] Example illustration: To verify the actual effect of this solution, luciferase from the fungus Neonothopanus nambi was used for optimization design testing, and its amino acid sequence can be directly obtained through the official website of the National Center for Biotechnology Information (NCBI) under the National Institutes of Health (NIH) of the United States.

[0021] See Figure 1 , a method for protein optimization design and screening based on an artificial intelligence algorithm provided by the present invention, taking the optimization of fungal luciferase as an example, constructs a complete optimization framework from the recognition of multi-dimensional key sites of fungal luciferase to the backbone structure design to sequence prediction and screening. This method innovatively integrates information related to protein sequence, structure and function, provides a new strategy for protein key site analysis, and at the same time combines an artificial intelligence-driven generation model to realize the batch design of protein variants, solving the limitations of traditional protein design. It not only provides new possibilities for the development and application of the luciferase self-luminescence system, but also provides a new paradigm for the research of traditional protein engineering, and has far-reaching significance for promoting the rapid development of fields such as biomedicine, environmental engineering, and food industry. Specifically, it includes the following steps: 1. Fixing key sites.

[0022] Protein design is a complex engineering task that typically begins with a reference protein structure. The sequence is modified through mutation algorithms, and the structure and functional activity of the mutants are verified through experiments. Traditional protein design often relies on this cycle of mutation and experimental verification, identifying and fixing conserved regions in the protein and iteratively optimizing non-conserved regions to improve the success rate of design. However, this process is not only time-consuming and laborious but also limited by experimental resources and the explored mutation space. In the design of certain specific proteins, such as fungal luciferase, due to the lack of detailed mechanism studies and protein crystal structure analysis, directly using existing design methods may not effectively solve the problem of its complex luminescence mechanism. Especially when the homologous sequences of the target protein are scarce, relying on existing large language models without specific fine-tuning, the designed proteins may not meet the experimental requirements. Therefore, accurately identifying and fixing the key functional sites of proteins is a core step in the protein design process.

[0023] Key functional sites of proteins generally include three categories: active sites, binding sites, and structural sites. First, the active site is a specific region in the protein molecule responsible for catalyzing chemical reactions. It is usually composed of several key amino acid residues, which provide an ideal environment for the binding and transformation of substrates through precise arrangement and interaction. For enzyme proteins, the active site is the core part that executes the catalytic function, capable of effectively reducing the activation energy of the reaction and promoting the progress of the reaction. An enzyme may contain multiple active sites, and the catalytic efficiency and specificity of these sites may vary, jointly determining the overall catalytic efficiency and selectivity of the enzyme. Second, the binding site refers to all the positions on the protein that can bind to other molecules (such as substrates, inhibitors, metal ions, etc.). Although these sites do not necessarily directly participate in the catalytic reaction, they play an important role in the interaction between the protein and other molecules. The precise design of the binding site has an undeniable impact on regulating the function of the protein, such as the activity regulation of enzymes and substrate affinity. Finally, the structural site is the region responsible for maintaining the stability of the three-dimensional structure of the protein. These sites ensure that the protein can maintain its stable spatial configuration through molecular mechanics interactions such as hydrogen bonds, hydrophobic interactions, and van der Waals forces. Although the structural site does not directly participate in the enzymatic reaction, it plays a crucial role in protein folding, stability, and function exertion.

[0024] In protein design, the fixation of key functional sites is crucial for ensuring the core function and structural stability of proteins, while enhancing the diversity of design and the success rate of experimental verification. Therefore, the present invention adopts a multi-strategy method for analyzing key functional sites of proteins that combines literature analysis, evolutionary information, and structure prediction to determine the conditions for protein optimization design. Through this method, on the basis of ensuring the core function of proteins, the non-conserved regions can be changed more flexibly, increasing the diversity of the designed generated sequences, thereby improving the probability of finally obtaining effective mutants.

[0025] 1.1 Key sites from literature information.

[0026] Extract reported key functional sites based on literature information. The experimentally verified key functional site information from the literature provides a strong theoretical basis for the protein design process. Although the specific molecular mechanism of fungal luciferase luminescence has not been fully elucidated, it is generally believed that this process is related to redox reactions, and the key amino acid residues catalyzing this reaction are usually located in the active center of the enzyme, that is, the catalytic pocket. These key residues are mostly polar amino acids such as glutamic acid and histidine, etc., which play a crucial role in the active expression of the enzyme. In the catalytic mechanism of fungal luciferase, these polar amino acids promote the oxidation reaction through their interaction with the substrate, thereby triggering bioluminescence. By summarizing a series of recent research contents, it is found that researchers compared the functions of 23 luciferase mutants, and the mutations at 13 sites (49V, 61Y, 104T, 109Q, 127D, 136R, 144S, 173D, 176R, 191C, 228W, 237E, 238E) significantly reduced the activity of luciferase, and their site distributions are as Figure 2 shown. Although the results of the mutation experiments did not give a clear relationship between polar-nonpolar amino acid substitutions and the changes in luciferase function, these 13 amino acid sites that led to the loss of luciferase activity were still regarded as key residues and ensured that they were not affected in the design process.

[0027] 1.2 Key site analysis based on structure prediction: Based on protein structure prediction and docking simulation of proteins with substrates, potential binding sites are identified through pocket analysis tools (such as P2Rank).

[0028] The function of a protein is determined by its structure. During the protein design process, for proteins with unknown crystal structures such as fungal luciferase, structure prediction using deep learning-based methods can provide a good understanding of the protein's folding and functional regions, identify potential functional sites, binding sites, and their interactions with substrates or ligands, thereby optimizing the design strategy. In an enzymatic reaction, the binding of the substrate usually depends on the catalytic pocket of the enzyme, which is a key functional region for achieving efficient catalysis. See Figure 3 in (A) of Figure 3 and the docking model of fungal luciferase - luciferin (3-hydroxycorynanthine) predicted by an artificial intelligence (AI)-driven docking simulation method in (B) of Figure 3 Figure 3 Figure 4 Figure 4 in (A) of Figure 4 Figure 4 in (B) of Figure 4

[0029] It should be noted that the docking simulation operation of protein and substrate molecules is as follows: Using AI-driven docking simulation methods such as RoseTTAFold All-Atom (RFAA) and DiffDock to deeply learn the mode of protein-substrate interaction, identify potential binding sites in a short time, so as to achieve efficient prediction of the binding domain between protein and substrate molecules. Using visualization tools such as PyMOL to reproduce the docking of protein and its substrate molecules.

[0030] ​​​​​To identify the key structural sites between luciferase and substrate in the system, the present invention comprehensively employs strategies such as structure prediction, molecular docking, and binding pocket recognition. First, two structure docking methods, RFAA and DiffDock, were used to predict the binding conformation of luciferin (3-hydroxyhydroxymorine) and luciferase. The results showed that the substrate was stably bound in the cavity region of the enzyme. Further, the pocket recognition tool P2Rank was used to analyze the static structure of the enzyme, predict possible binding sites, and record key residues. At the same time, considering the conformational changes of the enzyme during catalysis and the dynamics of potential binding sites under the "induced fit" mechanism, a spherical space with a radius of 10 Å was constructed centered on the substrate to supplement the screening of amino acid residues adjacent to the substrate as key sites that may be involved in binding and catalysis.

[0031] On this basis, an in-depth analysis was carried out on the atomic interactions between the enzyme and the substrate in the docking structure to more precisely identify the key structural residues. Hydrogen bonds (distance < 3.5 Å, angle tolerance < 30°) were identified through the PyMOL visualization tool, and residues that may play a key role in binding stability and catalytic mechanism were screened in combination with interaction criteria such as van der Waals forces (center distance < 4 Å, overlap ≥ 0.4–1.0 Å) and ionic bonds / salt bridges (distance between positive and negative charge residues < 4 Å).

[0032] Finally, by integrating the prediction results of P2Rank, the analysis of the docking model, the screening of spatial neighborhoods, and the interaction characteristics at the atomic level, 92 key structural sites (46F, 49V, 50G, 51P, 52S, 53Y, 54A, 55P, 56Q, 60G, 63I, 64V, 66V, 67L, 68S, 69L, 70F, 71R, 74Q, 97E, 98G, 99T, 103I, 104T, 105S, 106H, 107I, 108I, 109Q, 110R, 111Q, 130I, 146S, 147K, 148F, 149E, 150F, 151H, 152A, 153K, 154A, 155I, 156F, 157L, 166P, 167L, 168N, 169I, 176R, 177R, 178T, 179K, 181E, 182I, 183A, 184H, 185M, 186H, 187D, 188Y, 189H, 190D, 191C, 192T, 193L, 194H, 195L, 196A, 197L, 205V, 212Q, 213R, 214H, 215P, 216L, 217A, 218G, 221V, 222P, 223G, 224P, 225P, 228W, 229T, 230F, 231L, 233A, 234P, 235R, 241R, 242V, 243V) were determined, and their distribution is as Figure 5 shown.

[0033] 1.3 Key sites based on evolutionary analysis: Methods such as Multiple Sequence Alignment (MSA), information entropy calculation, and Hidden Markov Model (HMM) are used to screen for highly conserved sites (information entropy < 0.3 and conservation score > -1.5).

[0034] To more comprehensively cover all key site information of luciferase, based on the general framework of deep learning and bioinformatics, the present invention uses a variety of evolutionary analysis methods to extract potential rules and information from large-scale sequence and structure data.

[0035] First, the method of Multiple Sequence Alignment (MSA) can systematically align multiple homologous sequences, reveal the conservation and variability among sequences, and then infer key functional sites, evolutionary relationships, and structural features. By simply analyzing the sequence files of luciferases from 43 fluorescent species, the frequency of each amino acid at each position is counted, and the information entropy is calculated , and 134 sites with information entropy less than 0.3 are selected as conserved sites; among them, represents the information entropy of this site (Information Entropy), represents the number of different amino acid types appearing at this site, represents the index of the amino acid type, represents the th amino acid's frequency of appearance at this site.

[0036] Secondly, to more comprehensively cover conserved sites, the present invention also uses DEEPMSA2, a deep learning-based multiple sequence alignment method, which can automatically extract features from sequence data using a convolutional neural network (CNN). By learning the deep structural relationships and evolutionary information among sequences, it provides more accurate alignment results when dealing with complex sequence variations and diversities. By analyzing the target fungal luciferase with DEEPMSA2, a sequence Logo is obtained. This highly efficient visualization tool is widely used to display the conservation and variation patterns in DNA, RNA, or amino acid sequences. By showing the relative heights of different bases or amino acids at each position, it intuitively reflects the frequency of appearance and the degree of conservation at that position, and 101 candidate fixed sites are selected based on this.

[0037] In addition, based on the above-mentioned multiple sequence alignment file, the present invention also uses the HMMER software to construct a probability model of the protein family. HMMER adopts the Profile Hidden Markov Models (Profile HMMs) method, which is an extension of the Hidden Markov Model (HMM), combining multiple sequence alignment information with the probability framework of HMM, and is widely used in the analysis of protein and DNA sequences. This method reveals the likelihood of observing a specific observation state under a specific hidden state by calculating the emission probability from the hidden state to the observation state. To analyze the protein family probability model, the present invention selects sites with an emission probability greater than 0.3 and a conservation score higher than -1.5 as conserved sites, analyzes the MSA results of the luciferase amino acid sequence, and determines the information entropy of each site by statistically calculating the scoring values of different amino acids at each site. , represents the scoring value of the th amino acid at this site, and through the conservation score of each site is obtained, and then it is checked whether the conservation score of each site is greater than -1.5; the maximum emission probability of each site is , by checking whether the maximum emission probability of each site is greater than 0.3, finally 111 sites that meet the above two conditions are screened as potential fixed sites.

[0038] Combining the above methods and the evolution-related sites reported in the literature, the repeated sites were extracted, and a total of 186 fixed sites were obtained based on evolutionary analysis (30L, 36F, 37P, 39I, 40R, 41R, 42D, 43Y, 45T, 46F, 47L, 48E, 49V, 50G, 51P, 52S, 53Y, 54A, 55P, 56Q, 57N, 59R, 60G, 61Y, 62I, 63I, 64V, 66V, 67L, 68S, 69L, 70F, 71R, 73E, 74Q, 77L, 79I, 80Y, 81D, 83L, 84P, 85E, 86K, 87R, 89W, 90L, 92D, 93L, 94P, 96R, 98G, 99T, 100R, 101P, 102S, 103I, 104T, 105S, 106H, 107I, 108I, 109Q, 110R, 111Q, 112R, 113T, 114Q, 117D, 120F, 125L, 126I, 127D, 128K, 129V, 130I, 132R, 133V, 134Q, 135A, 136R, 137H, 138T, 141T, 143L, 144S, 145T, 146S, 147K, 148F, 149E, 150F, 151H, 152A, 154A, 155I, 156F, 157L, 159P, 161I, 163I, 165D, 166P, 169I, 170P, 171S, 172H, 173D, 174T, 175V, 176R, 177R, 178T, 179K, 180R, 181E, 182I, 183A, 184H, 185M, 186H, 187D, 188Y, 189H, 190D, 192T, 193L, 194H, 195L, 196A, 197L, 198A, 199A, 200Q, 201D, 203K, 204E, 205V, 206L, 207K, 208K, 209G, 210W, 211G, 212Q, 213R, 214H, 215P, 216L, 217A, 218G, 219P, 220G, 221V, 222P, 223G, 224P, 225P, 226T, 227E, 228W, 229T, 230F, 231L, 232Y, 233A, 234P, 235R, 236N, 237E, 238E, 239E, 240A, 241R, 242V, 243V, 244E, 246I, 247V, 248E, 249A, 250S, 251I, 253Y, 254M, 255T, 256N), and their distribution is as Figure 6 shown.

[0039] 1.4 Final determination of fixed sites: Take the intersection of the sites predicted by structure and those analyzed by evolution, and take the union with the experimental sites in the literature to form a conserved region accounting for 35%-38% of the full-length sequence. This conserved region is regarded as the fixed site.

[0040] The present invention adopts a multi-dimensional integration strategy to identify and verify the key amino acid residues of luciferase, and constructs an analysis framework covering literature sites, structure-related sites and evolutionary sites. Among them, the literature sites include 13 experimentally verified functionally critical residues; there are 92 structure sites, which are identified through three-dimensional structure analysis and are responsible for maintaining protein stability and function; there are 186 evolutionary sites, which are highly conserved among species and reflect their functional importance. As Figure 7 shown, through correlation heat map analysis, it is found that the correlation of the fixed sites screened based on structure prediction and evolutionary analysis is relatively high, indicating that the structural and evolutionary information is highly consistent. However, the correlation of the fixed sites determined by literature information with the above two is significantly reduced, suggesting that although the structural and evolutionary information is highly consistent, there are still deviations from the currently known experimental results.

[0041] To further optimize the design method of fixed sites, the present invention innovatively designs a set of methods for determining fixed sites, as Figure 8 shown, Figure 8 in which (a) is a schematic diagram of the set of key sites of fungal luciferase predicted by structure, Figure 8 and (b) in which is a schematic diagram of the set of key sites of fungal luciferase analyzed by evolution, Figure 8Among them, (c) is a schematic diagram of the key sites of fungal luciferase from literature information; specifically: take the intersection of the fixed sites (i.e., key conserved sites) screened by structure prediction and evolutionary analysis, and then take the union with the fixed sites determined by literature information. Finally, 95 key sites are screened out, accounting for 35.58% of the full length of the sequence (46F, 49V, 50G, 51P, 52S, 53Y, 54A, 55P, 56Q, 60G, 61Y, 63I, 64V, 66V, 67L, 68S, 69L, 70F, 71R, 74Q, 98G, 99T, 103I, 104T, 105S, 106H, 107I, 108I, 109Q, 110R, 111Q, 127D, 130I, 136R, 144S, 146S, 147K, 148F, 149E, 150F, 151H, 152A, 154A, 155I, 156F, 157L, 166P, 169I, 173D, 176R, 177R, 178T, 179K, 181E, 182I, 183A, 184H, 185M, 186H, 187D, 188Y, 189H, 190D, 191C, 192T, 193L, 194H, 195L, 196A, 197L, 205V, 212Q, 213R, 214H, 215P, 216L, 217A, 218G, 221V, 222P, 223G, 224P, 225P, 228W, 229T, 230F, 231L, 233A, 234P, 235R, 237E, 238E, 241R, 242V, 243V). The specific distribution of these 95 key sites is as shown in Figure 9 in (A) and Figure 9 shown in (B) of

[0042] Through this innovative method for determining fixed sites, it well balances different prediction results based on AI calculations and bioinformatics tools, and extracts the core potential sites. Finally, together with the experimental results in the real world, they are used as the key sites to be fixed. These sites will be retained in the design and engineering transformation of fungal luciferase to support the stable realization of protein functions. It meets the diversity and adaptability of the protein design sequence and breaks through the efficiency bottleneck of traditional schemes.

[0043] In general, the strategy of determining conserved regions as fixed sites combines a comprehensive approach of literature research, evolutionary analysis, and structure prediction. The specific operations are as follows: First, combined with existing literature data, known key functional sites can be identified and confirmed. These functional sites are usually the core areas of protein function and stability and need to remain unchanged during the design process; second, the function of a protein is determined by its structure. Generally speaking, in an enzymatic reaction, the substrate binds to the enzyme's action area, which is the enzyme's catalytic pocket. For the protein docking results predicted by AI-driven docking simulation methods (such as RFAA and DiffDock), the pocket area of ​​the protein is predicted by the P2Rank program, and the key residue sites of the pocket are recorded; in addition, through evolutionary analysis, such as multiple sequence alignment (MSA) of similar proteins from similar species, the frequency of occurrence of each amino acid at each position is counted, and determining the conserved sites by calculating the information entropy is a key step in analyzing protein sequences from an evolutionary perspective. At the same time, in order to cover the conserved sites as comprehensively as possible, a graphical method for visualizing the conservation and patterns in DNA, RNA or amino acid sequences was used. For example, the target protein was analyzed by DEEPMSA2 to obtain the sequence logo, which reflects the frequency of occurrence at that position and the degree of conservation at that position by displaying the height of different bases or amino acids at each position in the sequence. In addition, based on the above multiple sequence alignment files, a protein family probability model was established by constructing a hidden Markov model to mine evolutionary relationships and further improve the recognition accuracy of conserved regions.

[0044] 2. Design protein functional domain skeleton based on fixed sites.

[0045] The specific operation of protein functional domain skeleton design is as follows: based on fixed conserved sites, the denoising diffusion probability model (DDPM) is introduced to abstract the protein structure into six-dimensional information of each amino acid translation vector and rotation matrix, and gradually add noise as training input. In the denoising process, the gap between the predicted structure and the true structure at each step is measured by constructing a loss function, and the model is driven to restore the true protein structure. Finally, the relative positions of each amino acid are dynamically adjusted from the noise to accurately construct the folding framework of the functional domain. Among them, the diffusion model includes but is not limited to the denoising diffusion probability model (DDPM) and RFDiffusion. The model optimizes the protein skeleton structure by gradually adding noise and denoising to meet specific functional requirements. The generated skeleton structure contains more than 500 conformational variants.

[0046] The design of protein functional domain scaffolds is an important technique in protein engineering, aiming to construct stable protein structures with specific functions to optimize natural proteins or create novel proteins. Protein scaffold design usually employs two types of strategies: optimizing existing scaffolds and de novo design. Methods for optimizing existing scaffolds include sequence adjustment, mutation of key sites, and structural fine-tuning to enhance stability or functionality. Computer-aided design often uses tools such as Rosetta, FoldX, and PyRosetta to screen for optimal mutations through energy calculations. De novo design involves piecing together basic secondary structures (α-helices, β-sheets, loops, etc.) according to specific structural topologies or functional requirements. In recent years, deep learning models such as AlphaFold3 and RoseTTAFold have greatly improved the success rate of de novo design. In addition, optimization methods based on evolutionary information are also an important direction in protein scaffold design. By analyzing a large number of homologous sequences, key co-evolving amino acid pairs can be identified and the scaffold structure can be optimized accordingly. In particular, diffusion models and generative adversarial networks (GANs) have also demonstrated powerful capabilities in protein structure generation.

[0047] In the present invention, a scaffolding enzyme functional site design strategy is adopted. A stable scaffold structure is constructed around the key motifs that support and stabilize its biological function. An AI-based diffusion model (such as RFDiffusion) is used to generate the scaffold for the fungal luciferase at fixed sites. RFDiffusion is a protein design framework based on diffusion models proposed by the David Baker laboratory. It gradually generates a high-precision protein scaffold that meets the design constraints by simulating the reverse diffusion process from random noise to the target structure. Compared with traditional design methods based on sampling or energy optimization, RFDiffusion can not only globally capture the geometric properties of the protein scaffold but also finely reconstruct key sites based on functional constraints, thus enabling de novo design of complex structures such as binders and symmetric polymers. For the design of protein scaffolds that require anchor sites, RFDiffusion provides an efficient method to optimize the overall scaffold topology while retaining functional amino acids, making it more stable and meeting specific functional requirements.

[0048] Combined with the previous analysis, 95 fixed sites of fungal luciferase were locked. A mask was set using an AI-based diffusion model (such as RFDiffusion) to ensure that the key sites remained unchanged, while allowing the backbone of other regions to be freely generated. Thus, the regions requiring high-fidelity and free exploration were defined, guiding the structure to converge to positions that meet the constraints during the generation process. In this process, starting from the wild-type fungal luciferase, it gradually converges to a physically reasonable protein backbone, enabling the model to perform diffusion generation under restricted conditions and gradually optimizing the protein backbone to meet both global stability and the functional requirements of the fixed sites (diffusion model input format: 45-45 / A46-46 / 2-2 / A49-56 / 3-3 / A60-61 / 1-1 / A63-64 / 1-1 / A66-71 / 2-2 / A74-74 / 23-23 / A98-99 / 3-3 / A103-111 / 15-15 / A127-127 / 2-2 / A130-130 / 5-5 / A136-136 / 7-7 / A144-144 / 1-1 / A146-152 / 1-1 / A154-157 / 8-8 / A166-166 / 2-2 / A169-169 / 3-3 / A173-173 / 2-2 / A176-179 / 1-1 / A181-197 / 7-7 / A205-205 / 6-6 / A212-218 / 2-2 / A221-225 / 2-2 / A228-231 / 1-1 / A233-235 / 1-1 / A237-238 / 2-2 / A241-243 / 24-24 / 0). Based on the above constraints, 500 new protein backbone structures were finally obtained. An example of the structure of one of the generated protein backbones is shown as Figure 10 shown.

[0049] 3. Predict the protein sequence based on the backbone structure.

[0050] Based on the designed protein functional domain backbone structure without side-chain atoms and sequence information, a graph neural network model was further introduced to accurately model the complex relationship between the protein sequence and the three-dimensional structure. For example, using a deep learning model based on a graph neural network can learn the interactions and spatial constraints between amino acids to achieve an accurate sequence-structure mapping. By optimizing each amino acid position of the given backbone structure, the amino acid sequence that best matches it was dynamically predicted and generated. Among them, sequence prediction uses a deep learning-based protein sequence-structure mapping method, including but not limited to ProteinMPNN, controlling sequence sampling with a temperature parameter of 0.1-0.2. 24 candidate sequences were generated for each backbone, and the number of predicted candidate sequences exceeded 36,000.

[0051] Protein sequence prediction based on the protein backbone plays a crucial role in protein design and protein engineering, especially in fields such as enzyme optimization, antibody design, and vaccine development. Its core goal is to infer the corresponding amino acid sequence based on the known or predicted three-dimensional structure of the protein (i.e., the backbone), enabling it to correctly fold into this structure and possess the expected function. This process not only requires ensuring the stability of the structure but also ensuring that the sequence can perform the required functions (such as catalysis, binding, stability, etc.). Therefore, when considering design methods, the following principles are mainly followed. First is the compatibility between the sequence and the structure: Each amino acid residue is constrained by adjacent residues and the folding of the entire molecule in three-dimensional space. The designed sequence must be able to match the geometry of the backbone to fold into a stable three-dimensional structure. Second is the retention of functional residues: For functional proteins (such as enzymes, receptors, etc.), certain specific amino acid residues (such as catalytic sites, substrate-binding sites, etc.) must maintain their functions. Therefore, special attention needs to be paid to the amino acid types and positions of these key sites. Finally is the stability of the sequence: The designed sequence needs to have sufficient stability so that the protein can maintain its three-dimensional structure under different environments (such as temperature, pH, ionic strength, etc.).

[0052] Sequence prediction methods based on energy functions (such as Rosetta) or deep learning-based sequence-structure mapping (such as AlphaFold), as well as those based on deep learning models and graph neural networks GNN (such as ProteinMPNN), are currently relatively mainstream protein sequence prediction means. In the present invention, based on the protein backbone generated by the above diffusion model, a deep learning-driven protein sequence design model (such as ProteinMPNN) is adopted for protein sequence prediction. ProteinMPNN is a deep learning-driven protein sequence design tool developed by the David Baker laboratory. It can optimize the amino acid sequence while ensuring the stability of the target protein structure, enabling it to successfully fold into a predetermined three-dimensional structure. Compared with traditional design methods that rely on physical models (such as Rosetta), ProteinMPNN has a significant improvement in sequence recovery, can more accurately recover the protein sequence, and significantly improves the design efficiency. The optimized fungal luciferase backbone structure generated by the diffusion model has undergone design optimization based on the target function, ensuring stability and functional requirements. Using it as the input for sequence design can well establish the mapping relationship between the structure and the sequence, thereby generating a sequence that can correctly fold and has a predetermined function.

[0053] In the specific implementation process, according to the same design constraints as the diffusion model, the functional sites in the backbone are fixed. These functional sites may include catalytic sites, substrate binding sites, or other regions directly related to the activity of fungal luciferase. During the design process, considering that the sampling temperature is closely related to the process of the model selecting amino acids from the probability distribution, a relatively low temperature (0.1, 0.15, 0.2) is adopted to guide the model sampling. Because a lower temperature will make the model more inclined to select the most probable amino acids in the probability distribution when selecting amino acids, the generated sequences are more concentrated and similar. This design choice ensures that the generated sequences are more deterministic and conservative, and at the same time more meet the design requirements of the present invention. Especially in the case where the functional sites need to be stable, the sampling strategy at a lower temperature is very effective. When generating sequences, in order to ensure diversity and cover more sequence space, 24 different sequences are set for each target backbone. This strategy ensures that diverse sequences can still be obtained at a lower sampling temperature, thus providing more potential candidate sequences. During the entire design process, a total of 36,000 candidate luciferase sequences corresponding to 500 fungal luciferase backbones are generated.

[0054] 4. Scoring and screening based on the predicted sequences.

[0055] Although the above design method generates a large number of sequences, in order to ensure the quality and functional potential of the protein sequences designed by AI, it is also necessary to screen the predicted sequences through a strict scoring mechanism. First, at the sequence level, the ProteinMPNN_score and sequence repeatability (ARR_score) are scored for the 36,000 sequences generated by prediction. ProteinMPNN_score is an important indicator to measure the fitness of the predicted sequence to the target structure. Sequences with ProteinMPNN_score < 1 indicate better folding stability and can correctly form the expected three-dimensional conformation; the sequence repeatability score is used to evaluate the degree of repeated amino acids in the sequence. ARR_score > -2 indicates that the sequence has a high diversity in the local region, avoiding irrational sequences with continuous repetition of multiple amino acids, which is crucial for ensuring the functional activity of the protein. Through the preliminary scoring threshold screening at these two sequence levels, finally 1,640 high-quality fungal luciferase sequences are screened out from the 36,000 candidate sequences, as Figure 11 shown, accounting for 4.55% of the total designed sequences. Figure 11Among them, each row represents a predicted protein, showing the name automatically generated for the protein during the RFDiffusion and ProteinMPNN processes, the predicted amino acid sequence, the sampling temperature, and two scores of ProteinMPNN, namely ProteinMPNN_score and ARR_score. These preliminarily screened sequences effectively exclude those irrational sequences that do not possess reasonable structures or functions, not only meeting the requirements of structural stability but also having good functional potential, laying a foundation for subsequent sequence screening and application research.

[0056] In addition, the present invention also introduces deep learning-based structure prediction models (such as ESMFold, AlphaFold3), performs structure prediction on the 1640 sequences preliminarily screened, and screens based on pLDDT (Predicted Local Distance Difference Test) and pTM (Predicted Template Modeling score) scores to further ensure the structural stability and functional potential of these sequences. pLDDT is an index for evaluating the confidence of local structure prediction for each residue, ranging from 0 to 100. A high pLDDT value indicates a high confidence in the structure prediction of this region. Specifically, pLDDT≥90 indicates very high confidence, 70≤pLDDT<90 indicates relatively high confidence, 50≤pLDDT<70 indicates relatively low confidence, and pLDDT<50 indicates very low confidence. pTM, on the other hand, is an index for measuring the confidence of overall structure prediction, ranging from 0 to 1. The higher the pTM value, the closer the predicted overall folded structure may be to the true structure. Generally, a pTM score higher than 0.5 means that the overall predicted fold of the protein may be similar to the true structure. Based on the prediction and evaluation results, sequences with a pLDDT value higher than 80 and a pTM value higher than 0.85 are preferentially selected to ensure that the selected sequences have high confidence in both local and overall structures.

[0057] Meanwhile, in order to comprehensively evaluate the local and overall structure confidence of the sequences and ensure the scientificity and precision of the screening process, the present invention calculates the ranking percentages of the pLDDT and pTM scores for each of the sequences obtained through the above screening respectively, and introduces a comprehensive score S, which weights and sums the ranking percentages of pLDDT and pTM with specific weights. The specific formula is as follows: Among them, Rank_pLDDT and Rank_pTM represent the ranking percentages of pLDDT and pTM respectively, and ω1 and ω2 are the weight coefficients of pLDDT and pTM respectively. Considering that pLDDT reflects the local structure confidence, pTM reflects the overall folding accuracy, and the overall structure has a more critical impact on protein function, so ω1 = 0.4 and ω2 = 0.6 are set. The sequences are sorted in ascending order according to the comprehensive score S. The lower the score, the higher the comprehensive ranking. Finally, the top 20 sequences with the highest comprehensive scores are selected as the final candidate sequences of luciferase to ensure that they have high confidence in both local and overall structures. The example of the final predicted sequence and the predicted structure based on the example predicted sequence are as Figure 12 shown in (A) of Figure 12 and (B) of Figure 13 as shown. Figure 13 In it, each row represents a predicted protein, showing the name automatically generated during the ESMFold structure prediction after pMPNN (i.e., ProteinMPNN), the predicted sequence, two scores of ESMfold (pLDDT and pTM), the sampling temperature, two scores of ProteinMPNN (pMPNN_score and ARR_score, pMPNN_score is ProteinMPNN_score), and the weighted score.

[0058] Based on the above information, as Figure 14 shown, the present invention has formed a specific screening scheme integrating sequence and structure evaluation for the screening of fungal luciferase in the embodiments of the present invention, ensuring that the screened luciferase variants have the best performance in terms of structural stability and functional potential. It should be noted that the specific scoring thresholds and sorting weight settings in this screening scheme can be dynamically adjusted according to the designed target protein and the protein performance to be optimized to meet different design requirements.

[0059] It should be noted that the types of proteins to which the present invention can be applied also include: proteins with small molecule ligands, such as transport proteins (such as glucose transporters), receptor proteins (such as certain G protein-coupled receptors), structural proteins (such as rhodopsin), regulatory proteins (such as certain protein kinases), etc.

[0060] Corresponding to the embodiments of the protein optimization design and screening method based on the artificial intelligence algorithm described above, the present invention also provides embodiments of a protein optimization design and screening device based on the artificial intelligence algorithm.

[0061] See Figure 15, the protein optimization design and screening device based on the artificial intelligence algorithm provided by the embodiment of the present invention includes one or more processors and a GPU processor, which are used to implement the protein optimization design and screening method based on the artificial intelligence algorithm in the above embodiment.

[0062] The embodiment of the protein optimization design and screening device based on the artificial intelligence algorithm of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 15 shown, it is a hardware structure diagram of any device with data processing capabilities where the protein optimization design and screening device based on the artificial intelligence algorithm of the present invention is located. In addition to Figure 15 the central processing unit, memory, network interface, non-volatile memory, GPU processor, and I / O device shown, the any device with data processing capabilities where the device in the embodiment is located usually also includes other hardware according to the actual functions of the any device with data processing capabilities, which will not be elaborated here.

[0063] For the implementation processes of the functions and roles of each unit in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.

[0064] For the device embodiment, since it basically corresponds to the method embodiment, please refer to the partial description of the method embodiment for the relevant parts. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0065] Corresponding to the embodiment of the protein optimization design and screening method based on the artificial intelligence algorithm described above, the embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the protein optimization design and screening method based on the artificial intelligence algorithm in the above embodiment.

[0066] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc., equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or will be output.

[0067] The present invention can be widely applied to different computing environments such as single-machine high-performance computing, cluster computing, cloud computing, etc., ensuring that the protein optimization and screening processes can be efficiently completed under various computing power configurations.

[0068] The above content is only a preferred embodiment of the present invention, and should not be construed as a limitation on the scope of the present invention's patent. For those skilled in the art, various changes, combinations, simplifications, modifications, substitutions, and re-adjustments, etc., should all be equivalent replacement methods and will not depart from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above-mentioned implementation examples, the present invention is not limited to the above implementation examples only. Within the protection scope of the present invention, there are also more other equivalent implementation examples.

Claims

1. A method for optimizing the design and screening of proteins based on artificial intelligence algorithms, characterized in that, It includes the following steps: (1) Using an artificial intelligence-driven protein-substrate docking simulation method to predict the binding domain between a protein and a substrate; screening key conserved sites by combining literature analysis, structure prediction, and evolutionary analysis; taking the intersection of the sites obtained by structure prediction and evolutionary analysis, and then taking the union with the sites extracted from the literature to form a conserved region, and using this conserved region as the fixed sites; (2) Using a diffusion model to generate the backbone of a protein functional domain under the constraint of the fixed sites; (3) Using a graph neural network model to predict the sequence of the protein functional domain backbone generated in step (2); (4) Performing scoring and screening based on the predicted sequence, including sequence-level scoring and structure-level screening, to obtain candidate protein variants.

2. The method for protein optimization design and screening based on artificial intelligence algorithm according to claim 1, wherein In step (1), using an artificial intelligence-driven protein-substrate docking simulation method, including RFAA and DiffDock, to identify potential binding sites, thereby predicting the binding domain between a protein and a substrate molecule; using a visualization tool to reproduce the docking of the protein and its substrate molecule, and analyzing and screening the amino acid residues with specific types and intensities of non-covalent interactions between the protein and its substrate molecule.

3. The method for protein optimization design and screening based on artificial intelligence algorithm according to claim 1, characterized in that, In step (1), the screening of key conserved sites includes: extracting reported key functional sites based on literature information; identifying and screening potential binding sites through a pocket analysis tool and a visualization tool based on protein structure prediction; screening highly conserved sites using multiple sequence alignment, information entropy calculation, and hidden Markov model based on evolutionary analysis; taking the intersection of the sites obtained by structure prediction and evolutionary analysis, and then taking the union with the sites extracted from the literature to form a conserved region.

4. The method for protein optimization design and screening based on artificial intelligence algorithm according to claim 3, wherein When the protein is fungal luciferase, take the intersection of the sites obtained by structure prediction and evolutionary analysis, and then take the union with the sites extracted from the literature to screen out 95 key sites, accounting for 35.58% of the full length of the sequence; the key sites are: 46F, 49V, 50G, 51P, 52S, 53Y, 54A, 55P, 56Q, 60G, 61Y, 63I, 64V, 66V, 67L, 68S, 69L, 70F, 71R, 74Q, 98G, 99T, 103I, 104T, 105S, 106H, 107I, 108I, 109Q, 110R, 111Q, 127D, 130I, 136R, 144S, 146S, 147K, 148F, 149E, 150F, 151H, 152A, 154A, 155I, 156F, 157L, 166P, 169I, 173D, 176R, 177R, 178T, 179K, 181E, 182I, 183A, 184H, 185M, 186H, 187D, 188Y, 189H, 190D, 191C, 192T, 193L, 194H, 195L, 196A, 197L, 205V, 212Q, 213R, 214H, 215P, 216L, 217A, 218G, 221V, 222P, 223G, 224P, 225P, 228W, 229T, 230F, 231L, 233A, 234P, 235R, 237E, 238E, 241R, 242V, 243V.

5. The method for protein optimization design and screening based on artificial intelligence algorithm according to claim 1, characterized in that, Step (2) is specifically: Based on the fixed sites, introduce the denoising diffusion probability model, abstract the protein structure into six-dimensional information of the translation vector and rotation matrix of each amino acid, and gradually add noise as the training input; During the denoising process, construct a loss function to measure the gap between the predicted structure and the real structure at each step, and drive the model to restore the real protein structure. Finally, dynamically adjust the relative positions of each amino acid from the noise, construct the folding framework of the functional domain, and generate the protein functional domain backbone.

6. The method for optimizing design and screening of proteins based on artificial intelligence algorithms according to claim 1, characterized in that Step (3) is specifically: Based on the generated protein functional domain backbone structure without side-chain atoms and sequence information, introduce a graph neural network model to model the relationship between the protein sequence and the three-dimensional structure, that is, use a deep learning model based on the graph neural network to learn the interactions and spatial constraints between amino acids, and realize sequence-structure mapping; by optimizing each amino acid position of the given backbone structure, dynamically predict and generate the amino acid sequence that best matches it.

7. The method for protein optimization design and screening based on artificial intelligence algorithm according to claim 1, wherein, In step (4), the sequence-level scoring includes: scoring ProteinMPNN_score and sequence repeatability ARR_score for the sequences generated by prediction at the sequence level; among them, ProteinMPNN_score is an indicator to measure the fitness of the predicted sequence with the target structure, and a sequence with a ProteinMPNN_score less than 1 indicates better folding stability and can correctly form the expected three-dimensional conformation; the sequence repeatability score is used to evaluate the degree of repeated amino acids in the sequence, and an ARR_score greater than -2 indicates a higher diversity in the local region of the sequence; through these two scores, high-quality sequences are preliminarily screened out.

8. The method for protein optimization design and screening based on artificial intelligence algorithm according to claim 7, characterized in that, In step (4), the structure-level screening includes: performing structure prediction on the high-quality sequences preliminarily screened, and setting a screening threshold, and screening based on pLDDT and pTM scores. The pLDDT threshold is not lower than 80, and the pTM threshold is not lower than 0.85 to ensure the structural stability and functional potential of these sequences; among them, pLDDT is an indicator to evaluate the confidence of local structure prediction for each residue, and the higher the pLDDT value, the higher the confidence in the structure prediction of this region; pTM is an indicator to measure the confidence of overall structure prediction, and the higher the pTM value, the closer the predicted overall folding structure is to the true structure; based on the weighted sum of the ranking percentages of pLDDT and pTM, the protein variant with the highest comprehensive score is screened.

9. A protein optimization design and screening device based on artificial intelligence algorithms, characterized in that, It includes one or more processors and a GPU processor for implementing the protein optimization design and screening method based on the artificial intelligence algorithm according to any one of claims 1-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the protein optimization design and screening method based on the artificial intelligence algorithm according to any one of claims 1-8.

Citation Information

Patent Citations

  • Methods and systems for aligning sequences

    CN105637098A

  • Protein sequence design method based on graph neural network

    CN118197395A

  • Molecular modification method for glycosyltransferase based on artificial intelligence

    CN119724342A

  • Protein design automation for protein libraries

    EP1482434A2

  • Method and system for classifying biological species based on codon information of gene

    JP2012003737A

Cited By

  • Artificial intelligence-based luciferase high-throughput optimization screening method and system

    CN120452547A

  • A high-throughput optimization screening method and system for luciferase based on artificial intelligence

    CN120452547B

  • Antigen-antibody docking and antibody generation method based on coarse-grained structure model

    CN121034391A

  • Physical prior and deep learning-based uracil-DNA glycosylase de novo design method

    CN122435986A

  • A method for de novo design of uracil-dna glycosylase based on physical priors and deep learning

    CN122435986B