A high-throughput optimization screening method and system for luciferase based on artificial intelligence
Through artificial intelligence-driven high-throughput screening methods, combined with protein engineering and deep learning models, the problem of poor expression of luciferase in heterologous hosts was solved, and functionally stable luciferase mutants were quickly screened, thereby enhancing its application potential in the fields of biomedicine and environmental protection.
Patent Information
- Application Number
- CN202510948494.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Traditional luciferase systems express poorly in heterologous hosts, with limited luminescence efficiency and stability. Traditional protein optimization methods are time-consuming and difficult to predict the impact of mutations on protein folding and function, which limits the large-scale application of luciferase.
A high-throughput screening method based on artificial intelligence is used, combined with protein engineering and deep learning models, to screen luciferase mutants through multi-module comprehensive evaluation, including multiple analyses at the sequence and structural levels, to screen out protein variants with optimal physicochemical properties.
It has achieved rapid and efficient screening of luciferase mutants with optimized functionality and structural stability, breaking through the limitations of traditional protein design and improving the application reliability of luciferase in complex biological environments.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics and artificial intelligence-driven computational biotechnology, specifically to the technical field of luciferase multi-module screening methods based on artificial intelligence algorithms, and in particular to an artificial intelligence-based high-throughput optimization screening method and system for luciferase. Background Art
[0002] Luciferase, a bioluminescent enzyme, plays a vital role in biomedicine, gene expression research, disease diagnosis, and drug screening. It releases photons by catalyzing the oxidation of luciferin, providing a powerful visualization tool for scientific research. This bioluminescence phenomenon is not only used for dynamic observation at the cellular and molecular levels in basic scientific research, but is also clinically applied to disease monitoring, treatment efficacy assessment, and precision diagnosis. Although traditional luciferase systems are widely used, they generally suffer from defects such as the need for the addition of exogenous substrate luciferin and poor expression in heterologous hosts, further limiting their potential for large-scale application.
[0003] In recent years, scientists have tried to improve the luminescence intensity, substrate affinity and system stability of luciferases from different sources by engineering and optimizing them. Neonothopanus nambi The autonomous luminescence system discovered in provides a new approach for realizing luciferase without the need for exogenous substrates. However, the performance of this luminescence system in heterologous hosts is still relatively limited, and both the luminescence efficiency and stability need to be improved. Modification through traditional protein optimization and evolution methods often requires a lot of time and manpower costs, and it is difficult to model complex structure-function relationships and effectively predict the impact of mutations on protein folding and function, resulting in accidental results. With the introduction of artificial intelligence, especially the application of deep learning and generative models, new opportunities have been brought about for the evolutionary modification of luciferase, especially providing strong support in structure prediction, identification of key functional sites, skeleton generation, sequence generation, etc.
[0004] With the rapid generation of a large number of protein sequences driven by artificial intelligence, how to screen out a small range of sequence libraries with expected functionality and structural stability from a massive amount of designed sequences and reduce the cost of the experimental process has become a new problem. Based on the above background, the present invention has designed a system for multi-module comprehensive evaluation and screening of large-scale protein sequences designed by AI. Combining multiple analyses at the sequence level and the structural level, it can quickly screen out protein variants with optimal physical and chemical properties in a large-scale protein sequence library, while ensuring the functional activity and stability of the protein, and ensuring its application reliability in complex biological environments. This multi-module screening system not only provides new possibilities for the development and application of the luciferase autonomous luminescence system, but also breaks through the limitations of traditional protein design, giving AI-generated proteins more accurate screening and optimization capabilities, and promoting the innovative application of functional protein design in biomedicine, environmental protection and other fields. Summary of the Invention
[0005] The present invention aims to address the shortcomings of existing technologies by providing an artificial intelligence-based high-throughput luciferase optimization screening method and system. By combining protein engineering, computational biology, and deep learning models, the present invention enables rapid and efficient screening of luciferase mutants with optimized functional properties.
[0006] The object of the present invention is achieved through the following technical solutions: In a first aspect, an embodiment of the present invention provides a high-throughput optimization screening method for luciferase based on artificial intelligence, comprising the following steps:
[0007] (1) Construct a high-throughput protein mutation sequence library based on evolutionary analysis and mutation strategy simulation;
[0008] (2) Use a variety of sequence-level evaluation models and tools to predict and evaluate the physicochemical properties, enzyme function indicators, and mutation adaptability of protein sequences in the protein high-throughput mutation sequence library to obtain sequence-level evaluation results;
[0009] (3) Use a variety of structure prediction models and tools to predict protein structures of protein sequences in the protein high-throughput mutation sequence library to perform prediction and evaluation of various structural prediction indicators, physicochemical information, protein-substrate binding ability and protein-substrate affinity of the protein, and obtain structural level evaluation results;
[0010] (4) Based on the sequence-level evaluation results and the structure-level evaluation results, a customized screening scheme is adopted by combining threshold filtering with dynamic weighted scoring to optimize the screening of protein high-throughput mutation sequence libraries according to different screening objectives.
[0011] Furthermore, the step (1) specifically includes:
[0012] Obtain the complete amino acid sequence of the target protein, determine the fixed sites through evolutionary analysis and relevant literature information to identify the conserved regions; for non-conserved regions outside the conserved regions, use random mutagenesis or directed mutagenesis strategies to generate multiple mutant sequences to construct a high-throughput protein mutation sequence library.
[0013] Furthermore, in step (2), the physicochemical properties include aromaticity, instability index, isoelectric point, flexibility and thermal stability; the enzyme function index includes enzyme catalytic efficiency; and the mutation adaptability includes protein pseudo-likelihood and pseudo-perplexity, mutation marginal probability, wild-type marginal probability, mask marginal probability, site index cross entropy and mean entropy of each site.
[0014] Furthermore, the step (2) specifically includes:
[0015] For protein sequences in the high-throughput protein mutation sequence library, the BioPython tool is used to obtain the aromaticity, instability index, isoelectric point and flexibility of the protein, and the TemStaPro model is used to obtain the thermal stability of the protein, including the thermal stability classification results, stable temperature zones and conflict markers clash at different temperatures; the DLKcat tool is used to predict the enzyme catalytic efficiency of the protein; the ESM-2 model is used to predict the mutation effect of the protein, and obtain the pseudo-likelihood and pseudo-perplexity of the protein, mutation marginal probability, wild-type marginal probability, mask marginal probability, site index cross entropy and the mean entropy of each site.
[0016] Furthermore, in step (3), the structure prediction indicators include predicted local distance difference, predicted global structure quality, predicted average error, root mean square deviation, template modeling score and mutant-wild type protein structure similarity index; the physicochemical information includes solvent accessible surface area; the protein-substrate binding force includes hydrogen bonds, hydrophobic interactions, salt bridges, water bridges, halogen bonds and π-π stacking non-covalent interaction modes present in the three-dimensional structure of the protein-substrate complex.
[0017] Furthermore, the step (3) specifically includes:
[0018] For protein sequences in the protein high-throughput mutation sequence library, the ESMfold model is used to predict protein structure, and the predicted local distance difference, predicted global structure quality and predicted average error are output; the CEAlign tool and TM-align tool are used to calculate the root mean square deviation and template modeling score between the mutant and wild-type protein structures, respectively; the Foldseek tool is used to obtain the mutant-wild-type protein structure similarity indicators, including structural consistency percentage, alignment length, number of mismatches and alignment score; the FreeSASA tool is used to obtain the solvent accessible surface area, including the total solvent accessible surface area, polar region solvent accessible surface area, non-polar region solvent accessible surface area and relative solvent accessible surface area; the RFAA model is used to generate the three-dimensional structure of the protein-substrate complex, and the PLIP tool is then used to detect the hydrogen bonds, hydrophobic interactions, salt bridges, water bridges, halogen bonds and π-π stacking non-covalent interaction modes therein; based on the predicted protein structure, the PLANET model is used to obtain the protein-substrate affinity.
[0019] Furthermore, the customized screening scheme specifically includes:
[0020] ① An enzyme function-oriented dual-track activity-adaptability screening scheme, including: selecting enzyme catalytic efficiency, protein-substrate affinity, template modeling score, and site index cross entropy to set thresholds for threshold initial screening of mutant sequences in the protein high-throughput mutant sequence library; calculating the percentage ranking of enzyme catalytic efficiency, protein-substrate affinity, and predicted local distance difference index for each mutant sequence obtained after the initial screening and filtration, and assigning weights, and obtaining a comprehensive score for each mutant sequence by weighted summation; calculating the percentage ranking of template modeling scores for each mutant sequence obtained after the initial screening and filtration, and selecting and outputting the multidimensional evaluation results corresponding to the N mutant sequences with the highest and N lowest template modeling score percentage rankings;
[0021] ② A stability-prioritized structure-function progressive screening scheme, including: selecting predicted local distance differences, predicted global structural quality, the ratio of polar region solvent accessible surface area to total solvent accessible surface area, and conflict markers clash to set thresholds for preliminary structural stability screening of mutant sequences in the protein high-throughput mutant sequence library; assigning weights to the predicted local distance differences, predicted global structural quality, and the ratio of polar region solvent accessible surface area to total solvent accessible surface area of each mutant sequence after the preliminary screening, and obtaining a comprehensive stability score for each mutant sequence by weighted summation, and selecting the top P mutant sequences with the largest comprehensive stability scores to enter the next round of screening; calculating the percentage ranking of enzyme catalytic efficiency and alignment score for each mutant sequence obtained after the first two rounds of screening and assigning weights, and obtaining a comprehensive score for each mutant sequence by weighted summation, and selecting and outputting the multidimensional evaluation results corresponding to the top Q mutant sequences with the largest comprehensive scores;
[0022] ③ Mutational adaptability-driven evolution-combination collaborative optimization screening scheme, including: selecting mask marginal probability, mean entropy of each site and pseudo-perplexity to set thresholds for them to perform evolutionary adaptability initial screening and filtration of mutant sequences in the protein high-throughput mutation sequence library; calculating the binding ability score of each mutant sequence after the initial screening based on hydrophobic interaction distance and hydrogen bond distance; calculating the percentage ranking of each mutant sequence obtained after the first two rounds of screening based on the mask marginal probability, mean entropy of each site and pseudo-perplexity and performing weight allocation, and obtaining the evolutionary adaptability score of each mutant sequence by weighted summation; weighting the binding ability score and evolutionary adaptability score and performing weighted summation to obtain the comprehensive score of each mutant sequence, and selecting and outputting the multidimensional evaluation results corresponding to the top M mutant sequences with the largest comprehensive scores.
[0023] A second aspect of an embodiment of the present invention provides a system for implementing the above-mentioned artificial intelligence-based high-throughput optimization screening method for luciferase, comprising:
[0024] Sequence mutation generation module, used to construct a high-throughput protein mutation sequence library based on evolutionary analysis and mutation strategy simulation;
[0025] The sequence-level AI prediction and evaluation module is used to predict and evaluate the physicochemical properties, enzyme function indicators, and mutation adaptability of protein sequences in the protein high-throughput mutation sequence library through a variety of sequence-level evaluation models and tools to obtain sequence-level evaluation results;
[0026] The structure-level prediction and evaluation module is used to predict protein structures of protein sequences in the protein high-throughput mutation sequence library using a variety of structure prediction models and tools, so as to perform prediction and evaluation of various structural prediction indicators, physicochemical information, protein-substrate binding ability and protein-substrate affinity of the protein, and obtain structure-level evaluation results;
[0027] The customized screening module is used to customize the screening scheme based on the sequence-level evaluation results and the structure-level evaluation results, combining threshold filtering with dynamic weighted scoring, so as to optimize the screening of protein high-throughput mutation sequence libraries according to different screening objectives.
[0028] A third aspect of the present invention provides a screening device based on the above-mentioned artificial intelligence-based high-throughput optimization screening method for luciferase, comprising:
[0029] A storage module is used to store the constructed protein high-throughput mutation sequence library and the sequence-level evaluation results and structure-level evaluation results of each mutation sequence therein;
[0030] The computational module integrates multiple sequence-level evaluation models and tools as well as multiple structure prediction models and tools to predict and evaluate protein sequences in the high-throughput protein mutation sequence library at the sequence and structure levels in terms of physicochemical properties, enzyme function indicators, mutation adaptability, structure prediction indicators, physicochemical information, protein-substrate binding ability, and protein-substrate affinity, and obtain sequence-level evaluation results and structure-level evaluation results;
[0031] A dynamic weighting processor is used to perform threshold filtering and dynamic weight allocation, and to optimize the screening of mutant sequences in the protein high-throughput mutant sequence library using customized screening schemes.
[0032] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, is used to implement the above-mentioned artificial intelligence-based high-throughput optimization screening method for luciferase.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] (1) This invention provides a systematic, new process and method based on artificial intelligence algorithms for large-scale protein evolutionary screening, breaking through the limitations of manual screening and experience-based reliance in traditional protein design. It integrates bioinformatics tools and AI models, significantly improving the efficiency and accuracy of protein optimization screening, and realizing the automation and high-throughput quantification of the protein screening process.
[0035] (2) This invention innovatively integrates a multi-module protein screening and evaluation system, breaking through the limitations of single-index screening and constructing a three-level evaluation system of "sequence-structure-function". Through comprehensive evaluation of multiple dimensions such as structure, function, affinity, and stability, the performance of batch-generated mutant proteins is optimized. While improving the prediction accuracy, it also achieves accurate quantitative evaluation of protein performance, ensuring the superiority of the designed proteins in various biological properties.
[0036] (3) The present invention provides multiple efficient and flexible "threshold filtering-dynamic weighting" multi-stage evaluation systems, which promote the integration and balance of multi-objective optimization; through the comprehensive optimization of multiple performance indicators of proteins, it can maximize its stability and adaptability while ensuring functional realization, and at the same time flexibly adjust the screening weights of various indicators at any time to meet the needs of different application scenarios, providing an intelligent solution for enzyme engineering modification. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is the overall flow chart of the artificial intelligence-based high-throughput optimization screening method for luciferase in the present invention;
[0038] Figure 2 This is a distribution diagram of simulated fixed sites of luciferase in the present invention (wherein the light-colored and bold ones are fixed sites);
[0039] Figure 3 It is a flow chart of the luciferase sequence level screening and evaluation system of the present invention;
[0040] Figure 4 This is an example diagram of the BioPython prediction results of the luciferase mutation library in the present invention;
[0041] Figure 5 This is an example diagram of the thermal stability prediction results of the luciferase mutant library in the present invention;
[0042] Figure 6 This is an example diagram of the prediction results of the enzyme catalytic efficiency of the luciferase mutant library in the present invention;
[0043] Figure 7 This is an example diagram of the mutation adaptability prediction results of the luciferase mutation library in the present invention;
[0044] Figure 8 This is a flow chart of the luciferase structure-level screening and evaluation system of the present invention;
[0045] Figure 9 is an example diagram of the ESMfold predicted structure of the variant and wild-type fungal luciferase in the present invention; wherein, Figure 9 A in the figure is the ESMfold predicted structure of the variant fungal luciferase; Figure 9 Figure B is the ESMfold predicted structure of wild-type fungal luciferase;
[0046] Figure 10 This is an example diagram of the reliability index of the structure prediction of the luciferase mutant library in the present invention;
[0047] Figure 11 This is an example diagram of the structural comparison results of the luciferase mutant library in the present invention;
[0048] Figure 12This is an example diagram of the SASA calculation results of the luciferase mutant library in the present invention;
[0049] Figure 13 : is an example diagram of the predicted structure of the variant and wild-type fungal luciferase-substrate RFAA of the present invention; wherein, Figure 13 A in the figure is the predicted structure of variant fungal luciferase-substrate RFAA; Figure 13 B in the figure is the predicted structure of wild-type fungal luciferase-substrate RFAA;
[0050] Figure 14 is an example diagram of the prediction results of luciferase variant-substrate binding ability in the present invention;
[0051] Figure 15 This is an example diagram of the substrate affinity prediction results of the luciferase mutant library in the present invention;
[0052] Figure 16 This is an example diagram summarizing the main evaluation results of the luciferase mutant library in the present invention;
[0053] Figure 17 This is a schematic diagram of the results of the enzyme function-oriented "activity-adaptability" dual-track screening scheme in the luciferase screening scheme of the present invention;
[0054] Figure 18 This is a schematic diagram of the results of the stability-prioritized "structure-function" progressive screening scheme in the luciferase screening scheme of the present invention;
[0055] Figure 19 Schematic diagram of the results of the "evolution-combination" synergistic optimization screening scheme driven by mutation adaptability in the luciferase screening scheme of the present invention. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following, in conjunction with the accompanying drawings, takes the evolutionary screening of luciferase from fungi as an example to further describe the artificial intelligence-based high-throughput optimization screening method for luciferase of the present invention. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0057] The present invention will be described in detail below with reference to the accompanying drawings. Unless there is any conflict, the features of the following embodiments and implementations may be combined with each other.
[0058] like Figure 1As shown, the artificial intelligence (AI)-based high-throughput luciferase optimization screening method of the present invention specifically comprises the following steps:
[0059] (1) Construct a high-throughput protein mutation sequence library based on evolutionary analysis and mutation strategy simulation.
[0060] (1.1) Obtain the complete amino acid sequence of the target protein.
[0061] In this example, the target protein is based on fungi Neonothopanus nambi Wild-type luciferase from fungi can be obtained from the UniProt database. Neonothopanus nambi The complete amino acid sequence of the wild-type luciferase from the source is shown in SEQ ID No. 1, and its complete amino acid sequence is specifically MRINISLSSLFERLSKLSSRSIAITCGVVLASAIAFPIIRRDYQTFLEVGPSYAPQNFRGYIIVCVLSLFRQEQKGLAIYDRLPEKRRWLADLPFREGTRPSITSHIIQRQRTQLVDQEFATRELIDKVIPRVQARHTDKTFLSTSKFEFHAKAIFLLPSIPINDPLNIPSHDTVRRTKREIAHMHDYHDCTLHLALAAQDGKEVLKKGWGQRHPLAGPGVPGPPTEWTFLYAPRNEEEARVVEMIVEASIGYMTNDPAGKIVENAK. The target protein consists of 267 amino acids in total.
[0062] (1.2) For the complete amino acid sequence of the target protein, fixed sites are determined through evolutionary analysis and relevant literature information to identify conserved regions.
[0063] In this example, 134 fixed sites were identified through evolutionary analysis for the complete amino acid sequence of the target protein, i.e., fungal luciferase. Specifically, a multiple sequence alignment (MSA) was performed on the complete amino acid sequence of the target protein with homologous sequences from various similar species. This is to identify similarities and differences between the amino acid sequences of fungal luciferases from different species, reveal evolutionary relationships, and identify conserved regions. The MSA was then used to analyze the conserved regions. entropy < 0.3) identified 134 fixed sites (36F,37P,40R,42D,43Y,46F,47L,50G,51P,52S,53Y,54A,55P,56Q,57N,60G,61Y,63I,64V,67L,70F,71R,73E,79I,80Y,84P,85E,86K,87R,89W,90L,93L,96R,98G,100R, 101P,104T,105S,106H,107I,108I,109Q,110R,111Q,114Q,120F,125L,129V,130I,132R,134Q,13 6R,137H,141T,143L,146S,147K,148F,149E,151H,152A,154A,155I,156F,159P,163I,166P,170P, 171S,172H,173D,174T,175V,176R,177R,178T,179K,180R,181E,182I,183A,184H,185M,186H,18 7D,188Y,189H,190D,192T,194H,195L,196A,197L,198A,199A,200Q,201D,203K,205V,208K,209G, ,214H,215P,216L,217A,218G,219P,220G,222P,223G,224P,225P,226T,227E,228W,229T,230F,232Y,233A,234P,235R,239E,242V,243V,244E,246I,248E,249A,253Y,254M). In the context of multiple sequence alignment, entropy is used to measure the diversity or uncertainty of amino acid residues at a certain position; for a given amino acid position, the entropy is calculated based on the frequency of occurrence of different amino acids at that position. If only one amino acid appears at a site, the entropy value of the site is 0, indicating complete conservation; if multiple amino acids appear at similar frequencies, the entropy value is higher, indicating that the site has greater variability.Based on the above principles, after performing multiple sequence alignment and entropy analysis, sites within the sequence that meet the MSA entropy threshold can be selected as fixed sites. An MSA entropy <0.3 generally indicates that the corresponding amino acid site or region is highly conserved. The MSA entropy threshold is not absolute and may be adjusted based on specific circumstances and experience in different studies. Based on relevant literature information, 13 fixed sites (49V, 61Y, 104T, 109Q, 127D, 136R, 144S, 173D, 176R, 191C, 228W, 237E, and 238E) were identified. Based on published literature, mutations in these 13 sites affect fungal luciferase activity, and therefore these 13 sites were selected as fixed sites.
[0064] Furthermore, the fixed sites determined by evolutionary analysis and relevant literature information were integrated, and finally 140 fixed sites were identified (36F, 37P, 40R, 42D, 43Y, 46F, 47L, 49V, 50G, 51P, 52S, 53Y, 54A, 55P, 56Q, 57N, 60G, 61Y, 63I, 64V, 67L, 70F, 71R, 73E, 79I, 80Y, 84P, 85E, 86K, 87R, 89W, 90L, 93L ,96R,98G,100R,101P,104T,105S,106H,107I,108I,109Q,110R,111Q,114Q,120F,125L,127D,129V,130I,1 32R,134Q,136R,137H,141T,143L,144S,146S,147K,148F,149E,151H,152A,154A,155I,156F,159P,163I,1 66P,170P,171S,172H,173D,174T,175V,176R,177R,178T,179K,180R,181E,182I,183A,184H,185M,186H,1 87D,188Y,189H,190D,191C,192T,194H,195L,196A,197L,198A,199A,200Q,201D,203K,205V,208K,209G,2 10W,211G,212Q,213R,214H,215P,216L,217A,218G,219P,220G,222P,223G,224P,225P,226T,227E,228W,229T,230F,232Y,233A,234P,235R,237E,238E,239E,242V,243V,244E,246I,248E,249A,253Y,254M), specific sites such as Figure 2 As shown, the light-colored bold sites represent fixed sites.
[0065] The above-mentioned fixed sites are believed to be critical for the catalytic activity and luminescence properties of fungal luciferase and are identified as conserved regions that need to be protected from mutation. Regions outside the conserved regions in the complete amino acid sequence of fungal luciferase, such as surface-exposed hydrophobic regions and linker regions, may affect the stability, conformation, or affinity of fungal luciferase but do not directly affect its core catalytic function. Therefore, they are considered non-conserved regions for high-throughput mutagenesis.
[0066] It should be noted that conserved regions refer to amino acid segments in protein sequences that remain highly similar or unchanged during evolution. Therefore, the fixed sites can be determined by combining sites reported in relevant literature information that have an important impact on luciferase function and the entropy threshold (MSA entropy <0.3) established through evolutionary analysis; fixed sites refer to sites that remain unchanged during the construction of a high-throughput protein mutation sequence library, which also includes conserved sites determined based on conservation analysis.
[0067] (1.3) For non-conserved regions outside the conserved regions, random mutagenesis or directed mutagenesis is used to generate multiple mutant sequences to simulate the construction of a high-throughput mutant sequence library of the target protein.
[0068] Specifically, high-throughput mutagenesis is required for non-conserved regions outside of conserved regions. Mutational design can be either random or directed mutagenesis. Random mutagenesis targets amino acid residues in non-conserved regions, using the NNS / NNK random mutagenesis strategy (e.g., N represents any base, and S / K represents a specific base) to generate multiple possible amino acid substitutions; other random mutagenesis strategies can also be employed. Directed mutagenesis allows for targeted changes in specific regions, such as replacing hydrophobic amino acids with more hydrophobic ones or reversing polar sites to improve stability. In this example, a fully open mutagenesis rule function was constructed for random mutagenesis at 127 non-fixed sites in the non-conserved region based on a random mutagenesis strategy (including but not limited to the NNS / NNK random mutagenesis strategy) and a specific purpose (targeting the 127 non-fixed sites in the non-conserved region). Ultimately, 742 mutant sequences were generated, forming a high-throughput mutant sequence library for the fungal luciferase protein. This fully open mutation rule function is used to generate a library of protein mutants and save it as a FASTA file. The specific logic is as follows: first, the wild-type sequence and fixed sites are defined, and the fixed site numbers are converted to zero-based indices; then, a specified number of mutants are generated through a loop. In each loop, the wild-type sequence is converted into a list, and amino acid substitutions are randomly selected at non-fixed sites. The list is then converted into a string and stored in the mutant list; finally, each mutant is written to a file in FASTA format.
[0069] (2) Using a variety of sequence-level evaluation models and tools, we predict and evaluate the physicochemical properties, enzyme function indicators, and mutation adaptability of protein sequences in the high-throughput protein mutation sequence library at the sequence level to obtain sequence-level evaluation results. Among them, sequence-level evaluation models and tools include BioPython, TemStaPro, DLKcat, ESM-2, etc. Physicochemical properties include aromaticity, instability index, isoelectric point, flexibility, and thermal stability; enzyme function indicators include enzyme catalytic efficiency; mutation adaptability includes protein pseudo-likelihood (PLL) and pseudo-perplexity (PPX), marginal probability of mutation (Marginal Probability of Wild-Type), masked marginal probability (masked marginal probability), site-wise index cross-entropy (SICE), and per-position entropy mean.
[0070] It should be noted that in the AI-driven protein optimization screening and evolution process, sequence-level screening is to ensure that the sequence has the required physicochemical properties, functionality, and mutation adaptability. In order to comprehensively evaluate the high-throughput protein mutation sequence library of luciferase, a variety of tools and models were combined to evaluate the physicochemical properties, functional potential, and mutation adaptability of the protein sequences in the high-throughput protein mutation sequence library from multiple dimensions, such as Figure 3 As shown, it specifically includes the following sub-steps:
[0071] (2.1) For protein sequences in the high-throughput protein mutation sequence library, the bioinformatics analysis tool BioPython is used to obtain the aromaticity, instability index, isoelectric point, and flexibility of the protein. The deep learning-based protein language model TemStaPro is used to obtain the thermal stability of the protein, including thermal stability classification results at different temperatures, stable temperature zones, and conflict markers, to comprehensively evaluate the potential functionality and stability of the protein.
[0072] Specifically, BioPython is a widely used bioinformatics tool library that provides a rich set of modules and functions for processing and analyzing biological data, particularly tasks such as sequence analysis, biological computations, interaction with online databases, and analysis of structural data. BioPython was used to analyze the physicochemical properties of a high-throughput protein mutation sequence library, including aromaticity, instability index, isoelectric point, and flexibility.
[0073] Aromaticity analysis uses the aromaticity method in the BioPython ProteinAnalysis class to quantitatively assess the distribution and quantity of aromatic amino acids (such as phenylalanine, tryptophan, and tyrosine) in a protein sequence. The aromaticity method measures the aromaticity of a protein sequence by calculating the proportion of these aromatic amino acids in the entire protein sequence. Protein aromaticity is closely related to the aromatic amino acid content in the protein sequence. These aromatic residues, due to their aromatic ring structure, can contribute to protein structural stability through hydrophobic interactions and π-π stacking effects, playing a key role in protein folding. Aromatic amino acids are commonly found in the core regions of proteins. They not only enhance thermal stability by strengthening internal intermolecular interactions but also maintain structural integrity under extreme environmental conditions. Furthermore, aromatic residues play a crucial role in protein molecular recognition and substrate binding. In enzyme-catalyzed reactions, they can influence protein binding affinity and catalytic efficiency through interactions with substrates or ligands. Therefore, aromaticity not only plays a crucial role in protein stability and folding but is also directly linked to protein functional optimization and performance enhancement.
[0074] The instability index (instability_index) is an important tool provided by the ProteinAnalysis class in BioPython's SeqUtils.ProtParam module for preliminary assessment of protein stability. This index predicts protein stability by analyzing the types, order, and physicochemical properties of amino acids in the protein sequence. Specifically, the index considers the effects of specific amino acid residues (such as proline, aspartic acid, and glutamic acid) on protein folding, which are often associated with structural instability, misfolding, or hydrophobicity. A higher instability index generally indicates that the protein is prone to unfolding, structural loosening, or degradation due to a lack of a stable three-dimensional conformation, thereby losing biological function. Therefore, the instability index can help identify regions that may lead to protein misfolding or loss of function, and can be used to screen out design variants that are difficult to form a stable fold or are prone to degradation, thereby improving the stability and functional potential of proteins in the mutation library.
[0075] Flexibility is assessed using the flexibility method provided by the ProteinAnalysis class in BioPython's SeqUtils.ProtParam module, which calculates a flexibility score for each amino acid residue and quantifies the overall flexibility of the protein based on the average of these scores. The flexibility score reflects the flexibility of the protein in its three-dimensional structure and can reveal the structural stability and functional adaptability of the protein. Regions with higher flexibility are better able to adapt to diverse substrates and reaction conditions, which are crucial for catalytic reactions or molecular recognition processes. Proteins with insufficient flexibility typically exhibit lower adaptability. This inflexibility may cause the protein to become structurally rigid and unable to effectively regulate the interface or folding state that interacts with molecules, thereby limiting its functional performance. Therefore, by analyzing the flexibility score of a protein, it can provide an important reference for protein functional optimization and substrate adaptability, and help identify protein design variants with higher adaptability.
[0076] The isoelectric point is calculated using the isoelectric_point method in the ProteinAnalysis class in the SeqUtils.ProtParam module of BioPython. This method predicts the charge state of a protein at different pH values based on the charge properties and distribution of its amino acid residues. The isoelectric point can be used to infer the charge distribution of a protein at different pH values and its behavior in the cellular environment. It should be understood that proteins have different charge states at different pH values. When the pH is above the isoelectric point, the protein is negatively charged; when the pH is below the isoelectric point, the protein is positively charged. The isoelectric point is the point at which the net charge of a protein is zero at a specific pH and is directly related to its solubility, stability, and mobility in an electric field. Therefore, accurately predicting a protein's isoelectric point is a key factor in designing proteins with good solubility and stability, especially maintaining stability at a specific pH. This helps to better manipulate the protein's physicochemical properties and optimize its functional performance in different environments.
[0077] The prediction and analysis of the above key physical and chemical properties provide multi-dimensional support for the comprehensive evaluation of the overall stability of the designed sequence, which helps to understand the stability, functionality and adaptability of the protein from different perspectives. In the specific implementation process, the bioinformatics analysis tool library BioPython was used to predict the aromaticity, instability, flexibility and isoelectric point of the protein constructed by the fungal luciferase protein high-throughput mutation sequence library. Some of the output results are as follows Figure 4 shown.
[0078] In addition, the stability of proteins under physiological conditions, especially thermal stability, is also a key factor to consider in the evaluation and screening of high-throughput mutant sequence libraries of fungal luciferase proteins. Luciferases usually function at relatively low working temperatures, while high temperatures may destroy their three-dimensional structure, leading to loss of function. The thermal stability of proteins mainly stems from the synergistic effects of multiple interactions within them, including hydrogen bonds, salt bridges, aromatic stacking effects, and non-polar van der Waals forces. These interactions work together to maintain the folded state of the protein and enhance its resistance to high temperature conditions.
[0079] In order to accurately evaluate the thermal stability of protein design, a protein language model based on deep learning, TemStaPro, was introduced. This model can predict the thermal stability of proteins under different temperature conditions and provide their maximum stable temperature. The model extracts key features from protein sequences through a deep learning network, evaluates the stability of molecules at different ambient temperatures, and generates a thermal stability score. Specifically, TemStaPro can not only predict the stability of proteins at ideal temperatures, but also simulate their behavior at different temperatures. In specific applications, since TemStaPro itself may produce conflicting results in the process of predicting stable temperature zones, such as being stable in high temperature zones but unstable in low temperature zones, in a specific embodiment, such conflicts are marked as clash (i.e., the clash value is *). For those thermal stability prediction results without conflicts, they can be given priority, and it is believed that these sequences can maintain better structural stability under high temperature conditions.
[0080] Thermal stability analysis helps to evaluate the ability of proteins to maintain activity and structure at high temperatures, and is used to evaluate the maximum stable temperature of proteins. Thermal stability prediction analysis based on TemStaPro helps to screen out protein variants that are stable under high temperature conditions, and provides accurate theoretical support for further optimizing protein stability and function. In the specific implementation process, TemStaPro was used to predict the thermal stability of proteins in the constructed fungal luciferase protein high-throughput mutation sequence library, that is, the protein amino acid sequence file in FASTA format or FA format was input into TemStaPro, and the thermal stability data of the protein can be output, including the binary classification results of the thermal stability of the protein at different temperatures, stable temperature zones and conflict markers clash, etc. Some of the output results are as follows: Figure 5 As shown, a conflict mark clash of “*” represents a conflict, and a conflict mark clash of “-” represents no conflict.
[0081] (2.2) For protein sequences in the protein high-throughput mutation sequence library, the enzyme catalytic efficiency of the protein is predicted using the deep learning-based enzyme catalytic constant prediction tool DLKcat.
[0082] It should be noted that enzyme catalytic efficiency is the core parameter that characterizes enzyme function, among which Kcat is particularly important. Kcat represents the catalytic constant or catalytic efficiency. The two are essentially related concepts, but they describe the catalytic properties of enzymes from different perspectives. It measures the number of substrate molecules that each enzyme molecule can catalyze in unit time, and the unit is s. -1 In enzyme reactions that follow Michaelis-Menten kinetics, there is a clear relationship between Kcat, the maximum reaction rate (Vmax), and the enzyme concentration (Et), specifically Kcat = Vmax / Et.
[0083] In this embodiment, in order to accurately predict and optimize the catalytic efficiency of the enzyme, the deep learning-based enzyme catalytic constant prediction tool DLKcat is used to predict the enzyme catalytic efficiency of the protein with high precision. DLKcat is a Matlab / Python software package that aims to solve the problem of slow experimental measurement of enzyme activity parameters. By combining the interpretability of the mechanism model with the predictability of the deep learning model, the enzyme Kcat value can be predicted quickly and accurately. The input of DLKcat is a protein amino acid sequence file in FASTA format or FA format and the SMILES structure of the substrate, where the protein sequence is used to characterize the structural information of the enzyme, and the SMILES structure of the substrate is used to accurately describe the molecular structure of the substrate involved in the reaction; the output of DLKcat is the predicted enzyme catalytic efficiency (Kcat value) of the protein, and its unit is s -1 , used to measure the catalytic activity of proteins.
[0084] In DLKcat, a graph neural network (GNN) is combined with an attention-based convolutional neural network (Attention-CNN) to extract features from substrate and protein sequences. The feature vectors of the substrate and protein are then concatenated, and their interaction is modeled through a fully connected network. The final output is the predicted Kcat value. The Kcat value is a key metric for measuring the rate of protein-catalyzed reactions. It directly reflects the catalytic efficiency of an enzyme and is crucial for rapidly identifying efficient catalysts, guiding enzyme engineering, and optimizing bioprocesses.
[0085] DLKcat combines a graph neural network (for modeling substrate structure) with an attention-based convolutional neural network (for parsing protein sequences). By capturing the complex interactions between substrates and protein sequences, it can capture the impact of these interactions on enzyme catalytic activity and efficiently predict the catalytic efficiency of enzymes. DLKcat's computational architecture, through detailed modeling of substrate structure and protein sequence, can identify subtle changes in enzyme catalytic efficiency caused by amino acid mutations and provide highly accurate predictions of catalytic constants. The graph neural network simulates the spatial relationship between substrate molecules and enzyme active sites, while the attention-based convolutional neural network deeply explores potential functional sites in protein sequences. Combining the advantages of both can significantly improve prediction accuracy and generalization capabilities.
[0086] In specific applications, the Kcat prediction value provided by DLKcat provides an important quantitative basis for the catalytic efficiency of protein sequences. A higher Kcat value indicates that the protein sequence can convert the substrate more efficiently during the catalytic process, which indicates that the enzyme activity is enhanced. Compared with the wild-type protein (predicted Kcat value is 0.2037), the predicted higher Kcat value indicates that the designed sequence may have better catalytic performance and enzymatic reaction ability. Therefore, the prediction based on DLKcat not only provides an accurate quantitative evaluation for enzyme optimization, but also can guide the innovative design of protein engineering in improving enzyme efficiency and improving function. In the specific implementation process, DLKcat was used to predict the enzyme catalytic efficiency of the constructed fungal luciferase protein high-throughput mutant sequence library, and some of its output results are shown below. Figure 6 shown.
[0087] (2.3) For protein sequences in the high-throughput protein mutation sequence library, the deep learning-based protein language model ESM-2 is used to perform high-precision predictions on the mutation effects of proteins. The pseudo-likelihood and pseudo-perplexity of the protein, the marginal probability of mutation, the marginal probability of wild type, the marginal probability of masking, the site exponential cross-entropy and the mean entropy of each site are obtained to quantify the impact of amino acid substitutions on function.
[0088] It should be noted that ESM-2 is a deep learning-based protein language model used to make high-precision predictions of protein mutation effects. It can quantitatively evaluate amino acid mutations in protein sequences, helping researchers predict the potential impact of mutations on protein stability and function. Trained on ultra-large-scale protein sequence datasets, ESM-2 provides a quantitative assessment of the functional impact of each mutation, accurately identifies amino acid sites that are critical to protein function during the mutation process, captures long-term dependencies in protein sequences, and predicts the protein fitness landscape after mutation, outputting multiple important metrics that characterize the effects of mutations. It not only reflects the impact of specific mutations on protein stability and function, but also reveals differences in fitness between mutants, providing theoretical support for the rational design of high-throughput mutation sequence libraries for functionally optimized proteins.
[0089] In this example, ESM-2 was used to analyze sequence variants of fungal luciferase, focusing on evaluating the effects of amino acid mutations on protein stability and function. Specifically, the protein amino acid sequence in FASTA format or FA format and the wild-type protein sequence in FASTA format were used as input to ESM-2. By calculating the fitness difference between the mutant protein sequence and the wild-type protein sequence, the mutation effect of the protein was predicted, the adaptability of the protein mutation was evaluated, and mutants with better adaptability were selected. The parameter and threshold selection here only lists a more appropriate parameter combination, and there are many other options. Finally, ESM-2 outputs protein pseudo-likelihood and pseudo-perplexity, mutation marginal probability, wild-type marginal probability, mask marginal probability, site index cross entropy, and the mean entropy of each site.
[0090] Among them, protein pseudo-likelihood is used to measure the rationality of a specific amino acid in the current sequence context. The higher its value, the rarer the amino acid mutation is under the amino acid distribution learned by ESM-2, which means that the mutation may have a greater impact on protein function or stability. Protein pseudo-perplexity is an exponential expression of protein pseudo-likelihood and is used to measure the confidence of the entire protein sequence. The higher its value, the more unstable the protein sequence may be or the more likely it is to be impaired in protein function.
[0091] The marginal probability of protein mutation reflects the probability that a certain site mutates to the target amino acid. The lower the value, the rarer the mutation is in natural evolution, and the more likely it is to affect the function or stability of the protein.
[0092] The wild-type marginal probability of a protein is used to quantify the model's confidence in the wild-type amino acid at that site. A higher value indicates that the amino acid at that site is more conserved throughout evolutionary history. Therefore, mutations at that site are likely to disrupt protein function and stability.
[0093] The masked marginal probability of a protein refers to the probability distribution predicted by the model for the amino acids at the masked positions in the input sequence.
[0094] The site-specific exponential cross entropy of a protein refers to the exponential cross entropy at each site in the protein. It reflects the difference in information entropy between the amino acid distribution at the mutation site and the wild-type distribution. It is often used to measure the degree of evolutionary constraint at that site. A higher value indicates a more sensitive site to mutation. The mean entropy for each site is then calculated by averaging these values.
[0095] Through the above mutation adaptability index, the potential impact of different mutations on the function of fungal luciferase can be quantified, thereby screening candidate sequences that may have excellent performance or higher stability. This ESM-2-based method is more efficient than traditional structure-based or experimental screening methods, can greatly reduce the workload of experimental screening, and provide reliable theoretical support for the optimal design of fungal luciferase. In the specific implementation process, ESM-2 was used to perform mutation adaptability-related predictions on the constructed fungal luciferase protein high-throughput mutation sequence library, and some of its output results are as follows: Figure 7 shown.
[0096] (3) Using a variety of structure prediction models and tools, we predict protein structures from the protein high-throughput mutation sequence library at the structural level, and then conduct prediction and evaluation of various structural prediction indicators, physicochemical information, protein-substrate binding ability, and protein-substrate affinity to obtain structural level evaluation results. Among them, structure prediction models and tools include ESMfold, CEAlign, TM-align, Foldseek, FreeSASA, RFAA (RoseTTAFoldAll-Atom), PLIP, and PLANET (Protein-Ligand Affinity Prediction with NeuralNetworks). Structural prediction indicators include predicted local distance difference (pLDDT), predicted global structure quality (pTM), predicted average error (pAE), root mean square deviation (RMSD), template modeling score (TM-score) and mutant-wild-type protein structure similarity index; physical and chemical information includes solvent accessible surface area (SASA); protein-substrate binding force includes non-covalent interaction modes such as hydrogen bonds, hydrophobic interactions, salt bridges, water bridges, halogen bonds and π-π stacking in the three-dimensional structure of the protein-substrate complex.
[0097] It should be noted that the function of a protein is determined by its three-dimensional structure. The protein's folding state and spatial conformation directly determine its ability to interact with other molecules, catalytic efficiency, and biological activity. Therefore, when evaluating and screening high-throughput protein mutation sequence libraries, it is necessary not only to consider changes in the amino acid sequence but also to conduct in-depth analysis at the structural level to ensure that the mutated protein can maintain or improve its functionality.
[0098] In this example, a variety of models and tools, including deep learning-based structure prediction models, solvent accessible surface area prediction, protein-substrate complex structure prediction, protein-substrate interaction force analysis, and affinity prediction, were combined to evaluate and screen a high-throughput protein mutation sequence library at the structural level, aiming to more accurately identify protein sequences with potential functionality and stability, laying a solid foundation for subsequent protein engineering and functional research. Figure 8 As shown, it specifically includes the following sub-steps:
[0099] (3.1) For protein sequences in the high-throughput protein mutation sequence library, protein structure prediction is performed using the deep learning-based protein structure prediction model ESMfold, and pLDDT, pTM, and pAE are output; the root mean square deviation between the mutant and wild-type protein structures is calculated using the protein structure alignment tool CEAlign; the TM-score between the mutant and wild-type protein structures is calculated using the protein structure alignment tool TM-align; the structural similarity between the mutant and wild-type proteins is evaluated using the protein structure search tool Foldseek, and the mutant-wild-type protein structure similarity indicators are obtained, including structural identity percentage (fident), alignment length (alnlen), number of mismatches (mismatch), and alignment score (bits).
[0100] It should be noted that in order to ensure that the fungal luciferase variants have a stable folded structure, the deep learning-based protein structure prediction model ESMfold was introduced to predict the protein structure of the constructed fungal luciferase protein high-throughput mutation sequence library. ESMfold is a high-performance protein language model that uses the Transformer architecture and has 15 billion parameters. It can predict the three-dimensional structure of a protein based on a single amino acid sequence within a few minutes. Compared with the traditional AlphaFold2, ESMfold's prediction speed is increased by an order of magnitude and the structure prediction process is simplified, giving it a significant advantage in large-scale protein design and screening.
[0101] In this example, ESMfold was used to predict the three-dimensional structure of fungal luciferase variants. Figure 9As shown in A in , it also outputs key structural prediction indicators, including pLDDT, pTM, and pAE. ESMfold predicts the potential three-dimensional structure of a protein based on its amino acid sequence, captures the complex mapping relationship between sequence and structure, and accurately predicts the conformation of the protein in its natural state, such as the spatial conformation of the protein backbone, secondary structure elements, loop regions, and possible spatial distribution of active sites. In the ESMfold prediction process, the ESMfold prediction results of wild-type fungal luciferase are used as a comparative reference group to calculate the pLDDT and pTM values of the target protein to evaluate its structural reliability and global similarity with the reference structure. The ESMfold prediction results of wild-type fungal luciferase are shown in Figure 2. Figure 9 As shown in Figure 1B. Specifically, pLDDT is mainly used to measure the model's confidence in the local structure, and its numerical value ranges from 0 to 1. Predicted structures with a pLDDT higher than 0.8 are usually considered to be more reliable predictions, that is, pLDDT>0.8 indicates that the local structure is reliable. pTM is mainly used to evaluate the global structural quality of the prediction results, which is associated with the topological similarity of the reference structure. Its numerical value range is also between 0 and 1. The closer pTM is to 1, the more reliable the global folding of the prediction results. Predicted structures with a pTM higher than 0.7 are usually considered to be more reliable predictions, that is, pTM>0.7 indicates that the global folding is reliable. pAE reflects the model's error in the distance between residues. The lower the pAE value, the more accurate the spatial arrangement of the prediction results.
[0102] To better visualize the differences between the predicted protein structure and the wild-type, protein structure alignment tools such as CEAlign and TM-align were used to align and score the predicted three-dimensional structure with the wild-type protein structure. Specifically, the CEAlign tool was used to calculate the root mean square deviation (RMSD) between the mutant and wild-type protein structures, while the TM-align tool was used to calculate the TM-score (TM-score) between the mutant and wild-type protein structures to verify the prediction accuracy. The CEAlign tool aligns protein structures using the CE algorithm. Its predicted RMSD (root mean squared deviation) is a standard measure of the difference between two three-dimensional structures. It assesses similarity by calculating the root mean square of the sum of squared distances between corresponding atoms in the structures. Lower RMSD values indicate smaller differences between the predicted and target structures, reflecting higher prediction accuracy. TM-align is a powerful tool for protein structure alignment and is widely used to assess the overall topological similarity between two proteins. The TM-score offers significant advantages over traditional RMSD metrics. It adjusts the weighting of local structural differences and large errors, placing greater emphasis on global fold similarity. It also introduces a length-dependent normalization factor to eliminate the influence of structure length on the score. The TM-score ranges from 0 to 1, with 1 representing a perfect match. Scores of 0.5 and above typically indicate significant similarity between two structures, while scores below 0.17 suggest that the two structures may be completely unrelated. A TM-score > 0.5 is considered significant structural similarity.
[0103] In a specific embodiment, ESMfold was used to perform high-quality structure prediction on a high-throughput mutant sequence library of fungal luciferase, providing high-quality candidate variants for subsequent protein structure-based evaluation and analysis. At the same time, the pLDDT, pTM, pAE, RMSD, and TM-score values were output in the form of header information for each variant structure file, serving as an important basis for evaluating the reliability of protein structure prediction. Some of the output results are shown below. Figure 10 shown.
[0104] Furthermore, Foldseek is an open-source protein structure search tool designed for fast and highly sensitive protein structure alignment. It achieves efficient alignment of large-scale protein structure data by encoding protein tertiary structure information into sequences and using a pre-trained 3Di substitution matrix and MMseqs2 to perform sequence-like searches. It not only identifies similarities by calculating the structural distance matrix between proteins and matching them with known databases, but also applies cluster analysis to group structurally similar proteins to reveal protein families that may share similar functions.
[0105] In this example, a library of fungal luciferase variant structures predicted by ESMfold was constructed, and the ESMfold-predicted structure of the wild-type fungal luciferase was used as a reference for comparison to systematically evaluate the similarity and reliability of the predicted structures of the fungal luciferase variants. This topology-based structural alignment algorithm, by matching geometric features of protein structures rather than relying on specific amino acid sequences, can identify proteins with significant sequence differences but highly similar structures. In the specific implementation, Foldseek was used to assess the similarity between the mutant protein structures and the wild-type protein structures based on the ESMfold-predicted mutant protein structures. The resulting similarity index, specifically comprising four key metrics: structural identity percentage, alignment length, number of mismatches, and alignment score, was output. fident reflects the similarity between the variant and wild-type structures; a higher value indicates a more similar structure. alnlen and mismatch represent the amino acid sequence length of the variant and the number of mismatches between the variant and the wild-type sequence, respectively. bits reflects the confidence level of the comparison; a higher value indicates a closer functional folding state between the variant and the wild-type structure.
[0106] In a specific embodiment, the mutant-wild-type protein structure similarity index obtained by Foldseek can be used as an important supplementary index for protein screening, especially when some mutants are difficult to distinguish through other screening methods. This method provides a means of sorting based on highly sensitive structural similarity. In addition, it can also help to more comprehensively evaluate the reliability and potential functions of predicted protein structures, providing an important basis for further optimizing mutant screening strategies. Some of its output results are as follows: Figure 11 shown.
[0107] The above-mentioned scoring indicators such as pLDDT, pTM, pAE, RMSD, TM-score and mutant-wild-type protein structure similarity can well evaluate the overall quality of the predicted protein structure and help to quickly screen out high-quality candidate protein structures.
[0108] (3.2) For protein sequences in the protein high-throughput mutation sequence library, the bioinformatics analysis tool FreeSASA was used to obtain the solvent accessible surface area, including total SASA, polar region SASA, non-polar region SASA, and relative SASA.
[0109] It should be noted that SASA analysis is crucial in screening and optimizing protein sequences. It can provide information on protein surface exposure, help evaluate protein stability, folding properties, and potential for interaction with other molecules in solution, reflect the structural stability and functional adaptability of proteins, and provide a basis for subsequent functional studies. In this invention, two methods for calculating SASA are prepared:
[0110] ① SASA calculations were performed using the Shrake-Rupley algorithm in BioPython with a default sampling point of 200. The solvent accessible surface area of the entire protein structure was evaluated using the compute method.
[0111] ② SASA is calculated using the Lee-Richards algorithm in FreeSASA, with the default calculation accuracy set to 100. Through the calc function and classifyResults function, not only can the overall total SASA value and relative SASA value be analyzed, but also different types of regions can be classified and calculated, such as the SASA value of the polar region (Polar) and the SASA value of the non-polar region (Apolar).
[0112] Compared to the Shrake-Rupley algorithm in BioPython, FreeSASA provides a more refined analysis of surface properties, more accurately revealing the solvent accessibility of different protein sites and their potential impact. Therefore, FreeSASA is preferred for obtaining solvent accessible surface area.
[0113] SASA calculations performed by FreeSASA can more accurately analyze the surface properties and stability of proteins. SASA values not only reflect the surface exposure of proteins, but also reveal the exposure of different amino acid residues. A larger SASA value means that more amino acid residues are exposed to the solvent, which is usually closely related to the biological activity, substrate binding ability, and potential for intermolecular interactions of proteins. Conversely, a smaller SASA value indicates that the surface of the protein is relatively hidden, which may enhance the intrinsic stability of the protein, but may also limit its effective interaction with external molecules. Therefore, SASA analysis can not only help identify protein sequences with good surface exposure and easier participation in molecular recognition and function, but also evaluate the adaptability and stability of proteins under different environmental conditions, thereby providing a reliable theoretical basis for sequence optimization screening. In the specific implementation process, FreeSASA was used to perform SASA calculations on the constructed fungal luciferase protein high-throughput mutation sequence library, and some of its output results are shown below. Figure 12 shown.
[0114] (3.3) For protein sequences in the high-throughput protein mutation sequence library, the protein-substrate complex is modeled with high precision using the deep learning-based protein-ligand complex structure prediction model RFAA to generate the three-dimensional structure of the protein-substrate complex. The molecular interaction analysis tool PLIP is then used to detect non-covalent interaction patterns such as hydrogen bonds, hydrophobic interactions, salt bridges, water bridges, halogen bonds, and π-π stacking in the three-dimensional structure of the protein-substrate complex.
[0115] In this embodiment, in order to deeply analyze the interaction between the screened proteins and their substrates, RFAA was used to perform docking prediction on each protein-substrate pair and generate the corresponding three-dimensional structure of the protein-substrate complex. RFAA is an advanced deep learning model that can perform all-atomic-level modeling of biological molecules including proteins, nucleic acids, small molecules, metals, and covalently modified components. By combining the multimodal attention mechanism with the diffusion generation modeling strategy, RFAA combines the residue-based representation of amino acids and DNA bases with the atomic representation of other chemical entities, and successfully predicts the three-dimensional structure of complex biological molecular components; RFAA can accurately predict the three-dimensional structure of the protein-substrate complex based on the sequence information of the mutant protein in the absence of a known complex template, and provide a structural basis for subsequent interaction analysis. In the specific implementation process, the constructed high-throughput mutant sequence library of fungal luciferase protein and the corresponding substrate molecule information are used as input data. RFAA simulates the interaction between protein and substrate through its deep neural network architecture to generate a stable three-dimensional structure of the protein-substrate complex. The results are shown in the figure. Figure 13 The three-dimensional structure of the mutant fungal luciferase-substrate complex is shown in Figure 13 As shown in A, the three-dimensional structure of the wild-type fungal luciferase-substrate complex is shown in Figure 13 As shown in B.
[0116] Furthermore, after obtaining the predicted three-dimensional structure of the protein-substrate complex, the team integrated PLIP, an analytical tool for predicting non-covalent interactions between protein-ligand complexes, to analyze the non-covalent interaction patterns between each screened protein and ligand. PLIP is an open-source tool specifically designed to identify and analyze non-covalent interactions between proteins and ligands. It can automatically detect and quantify various interaction types in protein-ligand complexes, including hydrogen bonds, water bridges, salt bridges, halogen bonds, hydrophobic interactions, π-π stacking, π-cation interactions, and metal coordination. Its detection mechanism is primarily based on the spatial position and geometric relationships between atoms, ensuring accurate and reliable analysis, clarifying the effects of mutations on binding patterns, identifying potential functional mutation sites, and providing structural guidance for optimizing substrate binding ability and catalytic efficiency.
[0117] During the specific implementation process, it was found that there are four main non-covalent interaction modes between the protein-luciferase complex, namely hydrophobic interaction, salt bridge, π-π stacking and hydrogen bond. Among these non-covalent interaction modes, the distance between the protein chain and the ligand chain / ligand aromatic group center is the main factor in measuring the strength of the non-covalent interaction. Therefore, by outputting and storing the type and number of amino acid residues involved in each non-covalent interaction of each complex and their distance to the ligand, it serves as an important criterion for the subsequent evaluation of the binding ability between the screening protein and luciferase, such as Figure 14 shown.
[0118] (3.4) For protein sequences in the protein high-throughput mutation sequence library, based on the predicted corresponding protein structure, the protein-substrate affinity is obtained using the affinity prediction model PLANET based on a multi-objective graph neural network.
[0119] In this embodiment, in order to more comprehensively evaluate the interaction between fungal luciferase mutants and substrates, PLANET is used to obtain protein-substrate affinity to quantify its binding effect. PLANET is an affinity prediction model of a multi-objective graph neural network that can efficiently analyze the complex interactions between proteins and ligands. Therefore, PLANET can be used to efficiently predict the protein-substrate affinity of proteins. PLANET uses a two-dimensional molecular graph of the predicted protein structure and substrate as input, and adopts a multi-objective graph neural network to comprehensively learn key features such as the geometry, charge distribution and hydrophobicity of the binding site, thereby providing high-precision protein-substrate affinity prediction results. Protein-substrate affinity is a key indicator for measuring the tightness of binding between a protein and its substrate. It directly affects the efficiency and specificity of enzyme-catalyzed reactions and plays a vital role in screening high-affinity variants, optimizing substrate recognition ability and improving catalytic performance.
[0120] In the specific implementation process, during the optimization screening of the high-throughput mutant sequence library of fungal luciferase proteins, evaluating the binding affinity between mutants and substrates is crucial for screening efficient mutants. Therefore, the predicted protein structure of fungal luciferase mutants and the chemical structure of the substrate were input into PLANET to predict their binding affinity scores and obtain protein-substrate affinity. Some of the results are shown below. Figure 15 shown.
[0121] Based on these predictions, we can prioritize candidate luciferase proteins that bind more tightly to luciferin and have potentially higher catalytic efficiency by setting a reasonable affinity threshold and incorporating a weighted scoring mechanism into the subsequent screening process. This approach not only enhances the scientific validity of the screening strategy but also provides more targeted candidate mutation sequences for subsequent experimental verification.
[0122] (4) Based on the sequence-level evaluation results obtained in step (2) and the structure-level evaluation results obtained in step (3), a customized screening scheme is developed by combining threshold filtering with dynamic weighted scoring to optimize the screening of mutant sequences in the protein high-throughput mutant sequence library according to different screening objectives to obtain high-quality mutant sequences.
[0123] In this embodiment, based on the sequence-level evaluation results obtained in step (2) and the structure-level evaluation results obtained in step (3), the prediction evaluation results of different evaluation tools are comprehensively analyzed according to specific needs, and different evaluation tools in steps (2) and (3) are freely selected to construct a customized screening scheme with flexibility. Different customized screening schemes are proposed, including an enzyme function-oriented activity-adaptability dual-track screening scheme, a stability-prioritized structure-function progressive screening scheme, and a mutation-adaptability-driven evolution-combination synergistic optimization screening scheme, to adapt to different screening targets. In different customized screening schemes, the evaluation results of different evaluation tools can be combined in parallel or serially, so that the screening process can take into account multiple characteristics, such as catalytic efficiency, structural stability, substrate binding ability, etc., or focus on optimizing a specific function to meet application requirements. The core of this customized screening strategy lies in flexibility. Parameters can be adjusted according to target attributes, thereby effectively narrowing the candidate range, improving screening efficiency, and ensuring that the screened proteins have superior performance in multiple key indicators.
[0124] Specifically, in order to ensure the systematic and efficient screening and thus comprehensively evaluate the comprehensive performance of each candidate protein in the fungal luciferase mutant library, the main evaluation results of each candidate protein were summarized based on the above sequence level and structure level evaluation results, as follows: Figure 16 As shown, the results are exported as CSV files for final conditional screening. In practice, different screening strategies are used to selectively combine different evaluation results to perform customized screening of candidate proteins using threshold filtering or dynamic weighted scoring. Below are demonstrations of three customized screening schemes developed based on different screening strategies.
[0125] ① Enzyme function-oriented activity-adaptability dual-track screening program:
[0126] The strategic design basis of this screening program is: taking the core function of fungal luciferase as the first requirement, prioritizing the optimization of enzyme catalytic efficiency and protein-substrate affinity, while taking into account the TM-score in the structural prediction index to maintain the basic functional framework. Furthermore, the site index cross entropy in the mutation adaptability is introduced as a negative screening indicator to eliminate high-risk mutants that may cause functional disorders. Finally, through the pLDDT and TM-score involved in the activity-stability balance analysis, two types of candidate protein sequences, structurally conservative and radical, are screened to provide a diverse selection space for subsequent experiments. The specific implementation process of this screening program is as follows:
[0127] (a1) Initial screening threshold filtering: First, select enzyme catalytic efficiency, protein-substrate affinity, TM-score and site index cross entropy as the initial screening indicators and set the threshold respectively: Kcat value > 0.2037s -1 , protein-substrate affinity threshold>6.3, TM-score≥0.5, site index cross entropy threshold<50, where Kcat value represents the enzyme catalytic efficiency threshold, 0.2037s -1 is the wild-type prediction value, and TM-score represents the TM-score threshold; then, according to the threshold of the initial screening index, the mutant sequences (i.e., mutants) in the protein high-throughput mutation sequence library are subjected to threshold initial screening and filtering.
[0128] (a2) Dynamic weighted comprehensive score: First, the percentage ranking of the three indicators of enzyme catalytic efficiency, protein-substrate affinity and pLDDT is calculated for each mutant sequence obtained after the threshold initial screening and filtration; then the percentage rankings of these three indicators are weighted: the weight of the percentage ranking of enzyme catalytic efficiency is 40%, the weight of the percentage ranking of protein-substrate affinity is 30%, and the weight of the percentage ranking of pLDDT is 30%, that is, Kcat value_rank—40%, protein-substrate affinity_rank—30%, pLDDT_rank—30%; finally, the comprehensive score of each mutant sequence is obtained by weighted summation according to the weight corresponding to the percentage ranking of each indicator.
[0129] It should be understood that the above thresholds and weights can be dynamically adjusted according to actual needs.
[0130] (a3) Diversity output: First, the percentage ranking of the TM-score index is calculated for each mutant sequence obtained after the threshold initial screening and filtration; then, the five mutant sequences with the highest TM-score percentage ranking and the five mutant sequences with the lowest TM-score percentage ranking are selected, representing two types of candidate mutant sequences, namely, structural conservative and structural radical, and the enzyme catalytic efficiency, protein-substrate affinity, TM-score, site index cross entropy, pLDDT, percentage ranking of enzyme catalytic efficiency, percentage ranking of protein-substrate affinity, percentage ranking of pLDDT, and comprehensive score corresponding to the 10 selected mutant sequences are output to provide a diversity selection space for subsequent experiments. The final output of this screening scheme is as follows: Figure 17 shown.
[0131] ② Stability-first structure-function progressive screening scheme:
[0132] The strategic design of this screening program is based on the premise that the heterologous expression environment is highly stringent, prioritizing the optimization of the fungal luciferase mutant protein folding stability (involving the pLDDT and pTM indicators) and solvent-accessible surface area characteristics (involving the proportion of polar region SASA compared to total SASA). At the same time, secondary optimization is carried out by introducing thermal stability conflict detection (involving conflict markers clash) and structural similarity (involving alignment scores). Ultimately, based on the optimization of structural stability, mutants with excellent function, specifically manifested in higher enzyme catalytic efficiency, are screened to ensure that the mutants have high functional activity while maintaining stable folding. The specific implementation process of this screening program is as follows:
[0133] (b1) Initial screening of structural stability: First, four indicators, including pLDDT, pTM, the ratio of polar region SASA to total SASA, and the conflict marker clash, were selected as initial screening indicators for structural stability and thresholds were set respectively: pLDDT ≥ 70, pTM ≥ 0.7, polar region SASA / total SASA > 35%, clash = "-", where clash represents the conflict marker clash used in thermal stability conflict detection, which is used to screen conflict-free mutants, and its value "-" represents that the mutant has no conflict; then, according to the threshold of the initial screening indicator for structural stability, the mutant sequences in the high-throughput protein mutation sequence library were initially screened for structural stability.
[0134] (b2) Comprehensive stability score: First, weights are assigned to the three indicators of pLDDT, pTM, and the proportion of polar region SASA to total SASA for each mutant sequence after the initial structural stability screening: the weight of pLDDT is 50%, the weight of pTM is 30%, and the weight of the proportion of polar region SASA to total SASA is 20%, that is, pLDDT - 50%, pTM - 30%, polar region SASA / total SASA - 20%; then, the comprehensive stability score of each mutant sequence is obtained by weighted summation based on the weights corresponding to each indicator, and then the top 30 mutant sequences with the largest comprehensive stability scores are selected to enter the next round of screening.
[0135] It should be understood that the above thresholds and weights can be dynamically adjusted according to actual needs.
[0136] (b3) Functional optimization screening: First, the percentage ranking of the two indicators, enzyme catalytic efficiency (Kcat value) and alignment score, of each mutant sequence obtained after the first two rounds of screening was calculated; then the percentage rankings of these two indicators were weighted: the weight of the percentage ranking of enzyme catalytic efficiency was 60%, and the weight of the percentage ranking of alignment score was 40%, that is, Kcat_rank—60%, alignment score_rank—40%; finally, the comprehensive score of each mutant sequence was obtained by weighted summation based on the weights corresponding to the percentage rankings of each indicator, and the pLDDT, pTM, polar region SASA, total SASA, enzyme catalytic efficiency (Kcat value), alignment score, Kcat_rank, alignment score_rank, and comprehensive score corresponding to the top 10 mutant sequences with the largest comprehensive score were selected as the final output. The final output of this screening scheme is as follows: Figure 18 shown.
[0137] ③ Mutational adaptability-driven evolution combined with collaborative optimization screening scheme:
[0138] The strategic design of this screening program is based on the following: targeting the potential needs of long-term evolution of fungal luciferase, focusing on optimizing the evolutionary rationality of the mutation sites of fungal luciferase mutants (involving mask marginal probability, mean entropy of each site, and pseudo-perplexity) and non-covalent binding ability (involving hydrophobic interaction distance and hydrogen bond distance), thereby screening out new variants that can maintain evolutionary adaptability while also possessing strong substrate binding ability. In addition, by introducing multi-indicator dynamic weights, this strategy can effectively balance mutation conservation and functional innovation. The specific implementation process of this screening program is as follows:
[0139] (c1) Initial screening of evolutionary adaptability: First, the three indicators of mask marginal probability, mean entropy of each site and pseudo perplexity are selected as the initial screening indicators of evolutionary adaptability and the thresholds are set respectively: mask marginal probability > 40, mean entropy of each site < 1.6, pseudo perplexity < 100; then, according to the set thresholds of the initial screening indicators, the mutant sequences in the protein high-throughput mutation sequence library are filtered for evolutionary adaptability.
[0140] (c2) Binding ability score: First, the binding ability score of each mutant sequence after the initial evolutionary adaptability screening was calculated based on the hydrophobic interaction distance (Hydrophobic Interactions) and hydrogen bond distance (Hydrogen Bonds): For each mutant sequence's hydrophobic interaction distance, each interaction satisfying Hydrophobic Interactions (dist) < 3Å (strong hydrophobic interaction) was scored 1 point, each interaction satisfying 3.0Å ≤ Hydrophobic Interactions (dist) < 4.0Å (moderate interaction) was scored 0.5 points, and those satisfying Hydrophobic Interactions (dist) > 4Å were not scored. For each mutant sequence's hydrogen bond distance, each hydrogen bond satisfying Hydrogen Bonds (dist) < 3Å (strong hydrogen bond) was scored 1 point. The cumulative score of the hydrophobic interaction distance and hydrogen bond distance for each mutant sequence was then calculated as the binding ability score, which is used to reflect the tightness of binding between the mutant sequence and the substrate.
[0141] (c3) Dynamic weighted comprehensive ranking: First, the percentage ranking of the three indicators of mask marginal probability, mean entropy of each site and pseudo perplexity is calculated for each mutation sequence obtained after the first two rounds of screening; then the percentage rankings of these three indicators are weighted: the weight of the percentage ranking of mask marginal probability is 40%, the weight of the percentage ranking of mean entropy of each site is 30%, and the weight of the percentage ranking of pseudo perplexity is 30%, that is, mask marginal probability_rank-40%, 1-mean entropy of each site_rank-30%, 1-pseudo perplexity_rank-30%; then the evolutionary adaptability score of each mutation sequence is obtained by weighted summation according to the weight corresponding to the percentage ranking of each indicator. Furthermore, the weights of the binding ability score and the evolutionary adaptability score are assigned: the weight of the binding ability score is 40%, and the weight of the evolutionary adaptability score is 60%, that is, the binding ability score -40%, and the evolutionary adaptability score -60%; then the weighted sum of the weights of the binding ability score and the evolutionary adaptability score is obtained to obtain the comprehensive score of each mutation sequence. This comprehensive score can ensure the balance between evolutionary rationality and functional innovation in the screening results to a certain extent. Finally, the mask marginal probability, mean entropy of each site, pseudo perplexity, binding ability score, evolutionary adaptability score, and comprehensive score corresponding to the top 10 mutation sequences with the largest comprehensive scores are selected as the final output. The final output of this screening scheme is as follows: Figure 19 shown.
[0142] In summary, this paper proposes an AI-based high-throughput, multi-dimensional optimization and screening method for luciferase. Using the large-scale, high-throughput screening of a fungal luciferase protein-mimicking mutant library as an example, a systematic evaluation and screening system is established, spanning sequence, structure, and function. At the sequence level, the present invention conducts preliminary screening of mutants based on physicochemical properties (aromaticity, flexibility, etc.), enzyme kinetic parameters (Kcat prediction), and evolutionary adaptability (ESM-2 prediction analysis). Subsequently, at the structural level, three-dimensional conformations are predicted using ESMfold, and folding quality is assessed using pLDDT and pTM. Simultaneously, Foldseek is used for structural alignment, SASA is used to calculate surface properties, and PLIP is used to analyze key molecular interactions, enabling a comprehensive structure-function evaluation of the mutants.
[0143] In terms of screening strategy, a three-level evaluation system of "sequence-structure-function" was constructed, and an efficient and accurate screening framework was constructed by combining threshold filtering with dynamic weighted scoring. The specific process diagram is shown below. Figure 1As shown, a three-tiered "sequence-structure-function" evaluation system enables precise optimization and screening of protein performance. This innovative approach integrates multiple sequence-level evaluation models and tools with multiple structure prediction models and tools. By integrating high-precision, multi-dimensional information, it achieves automation and high-throughput protein screening, effectively overcoming the limitations of traditional experience-based screening and significantly improving screening efficiency and accuracy.
[0144] It is worth mentioning that the embodiments of the present invention also provide an artificial intelligence-based high-throughput luciferase optimization screening system for implementing the artificial intelligence-based high-throughput luciferase optimization screening method of the above embodiment. The system includes a sequence mutation generation module, a sequence-level AI prediction and evaluation module, a structure-level prediction and evaluation module, and a customized screening module.
[0145] In this embodiment, the sequence mutation generation module is used to construct a protein high-throughput mutation sequence library based on evolutionary analysis and mutation strategy simulation.
[0146] In this embodiment, the sequence-level AI prediction and evaluation module is used to perform predictive evaluation of the physicochemical properties, enzyme function indicators, and mutation adaptability of protein sequences in the protein high-throughput mutation sequence library at the sequence level through a variety of sequence-level evaluation models and tools to obtain sequence-level evaluation results.
[0147] In this embodiment, the structure-level prediction and evaluation module is used to predict the protein structure of protein sequences in the protein high-throughput mutation sequence library from the structural level through a variety of structure prediction models and tools, and then perform prediction and evaluation of various structural prediction indicators, physicochemical information, protein-substrate binding ability and protein-substrate affinity of the protein to obtain structure-level evaluation results.
[0148] In this embodiment, a customized screening module is used to customize the screening scheme based on the sequence-level evaluation results obtained by the sequence-level AI prediction and evaluation module and the structure-level prediction and evaluation results obtained by the structure-level prediction and evaluation module, by combining threshold filtering with dynamic weighted scoring, so as to optimize the screening of mutant sequences in the protein high-throughput mutation sequence library according to different screening targets to obtain high-quality mutant sequences.
[0149] In summary, the present invention takes the "sequence-structure-function" joint evaluation system as its core, integrating multiple sequence-level evaluation models and tools as well as multiple structure prediction models and tools for intelligent protein screening. It aims to achieve efficient optimization and application scenario customization of functional enzymes such as luciferase through a multi-module collaborative screening strategy.
[0150] Corresponding to the aforementioned embodiment of the artificial intelligence-based high-throughput optimization screening method for luciferase, the present invention also provides an embodiment of an artificial intelligence-based high-throughput optimization screening device for luciferase.
[0151] An embodiment of the present invention provides an artificial intelligence-based luciferase high-throughput optimization screening device, comprising a storage module, a computing module, and a dynamic weighting processor. The storage module is used to store the constructed protein high-throughput mutation sequence library and the sequence-level evaluation results and structure-level evaluation results of each mutation sequence therein. The computing module integrates a variety of sequence-level evaluation models and tools as well as a variety of structure prediction models and tools to perform predictive evaluations on the protein sequences in the protein high-throughput mutation sequence library in terms of physicochemical properties, enzyme function indicators, mutation adaptability, structure prediction indicators, physicochemical information, protein-substrate binding force, and protein-substrate affinity from the sequence level and structure level, and obtain sequence-level evaluation results and structure-level evaluation results. The dynamic weighting processor is used to perform threshold filtering and dynamic weight allocation, and a customized screening scheme is used to optimize the screening of mutation sequences in the protein high-throughput mutation sequence library to obtain high-quality mutation sequences.
[0152] This device is designed to optimize protein design and screening tasks and can be implemented on a variety of high-performance computing platforms, including but not limited to servers, workstations, and cloud computing environments equipped with graphics processing units (GPUs). It combines GPU parallel computing with an efficient processor architecture. The device can read and load computer programs from storage media to efficiently execute multi-target protein evolutionary screening tasks. Utilizing various implementation methods, including software, hardware, or a combination of both, it can be flexibly configured to maximize computational performance and task processing efficiency. From a hardware architecture perspective, the device's core computing units may include but are not limited to GPU-accelerated computing units, central processing units (CPUs), high-speed memory, network interfaces, and non-volatile storage devices. To meet diverse computing needs, the device supports expansion of hardware components, such as high-speed storage devices and dedicated computing accelerators, to further optimize computational efficiency and increase task throughput. In different application scenarios, the device dynamically allocates computing resources based on actual needs. It can be deployed centrally on a single device or distributed across multiple computing nodes to meet large-scale, efficient multi-module protein evolutionary screening tasks.
[0153] In addition, the present invention also provides a computer-readable storage medium storing the computer program of the present invention. When executed on a GPU-accelerated computing device, the program can implement the artificial intelligence-based high-throughput luciferase optimization screening method. The storage medium can include an internal storage unit (such as an SSD, high-bandwidth memory HBM) and an external storage device (such as an NVMe storage array or a distributed storage system). It not only supports program storage but also saves intermediate data, calculation results, and final screening output generated during execution. Through a flexibly configured storage and computing system, it can ensure efficient operation in a variety of computing environments, such as stand-alone high-performance computing (HPC), cluster computing, and cloud computing, meeting the protein design requirements under different computing power configurations and having excellent scalability and adaptability.
[0154] The method described in this paper is not only applicable to the optimization and screening of fungal luciferases but can also be extended to the rational design and screening of other functional proteins. Its standardized process is widely applicable in fields such as biomedicine, environmental engineering, and the food industry, providing a novel screening strategy for the efficient development of functional proteins with significant application value and industrial prospects.
[0155] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only.
[0156] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
[0157] The above embodiments are intended only to illustrate the design concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. The scope of protection of the present invention is not limited to the above embodiments. Therefore, any equivalent changes or modifications made based on the principles and design concepts disclosed in the present invention are within the scope of protection of the present invention.
Claims
1. A high-throughput optimization screening method for luciferase based on artificial intelligence, characterized in that: The following steps are involved: (1) Construct a high-throughput protein mutation sequence library based on evolutionary analysis and mutation strategy simulation; (2) Use a variety of sequence-level evaluation models and tools to predict and evaluate the physicochemical properties, enzyme function indicators, and mutation adaptability of protein sequences in the protein high-throughput mutation sequence library to obtain sequence-level evaluation results; (3) Use a variety of structure prediction models and tools to predict protein structures of protein sequences in the protein high-throughput mutation sequence library to perform prediction and evaluation of various structural prediction indicators, physicochemical information, protein-substrate binding ability and protein-substrate affinity of the protein, and obtain structural level evaluation results; (4) Based on the sequence-level and structure-level evaluation results, a customized screening scheme is developed by combining threshold filtering with dynamic weighted scoring to optimize the screening of high-throughput protein mutation sequence libraries according to different screening objectives; The customized screening program specifically includes: ① An enzyme function-oriented dual-track activity-adaptability screening scheme, including: selecting enzyme catalytic efficiency, protein-substrate affinity, template modeling score, and site index cross entropy to set thresholds for threshold initial screening of mutant sequences in the protein high-throughput mutant sequence library; calculating the percentage ranking of enzyme catalytic efficiency, protein-substrate affinity, and predicted local distance difference index for each mutant sequence obtained after the initial screening and filtration, and assigning weights, and obtaining a comprehensive score for each mutant sequence by weighted summation; calculating the percentage ranking of template modeling scores for each mutant sequence obtained after the initial screening and filtration, and selecting and outputting the multidimensional evaluation results corresponding to the N mutant sequences with the highest and N lowest template modeling score percentage rankings; ② A stability-prioritized structure-function progressive screening scheme, including: selecting predicted local distance differences, predicted global structural quality, the ratio of polar region solvent accessible surface area to total solvent accessible surface area, and conflict markers clash to set thresholds for preliminary structural stability screening of mutant sequences in the protein high-throughput mutant sequence library; assigning weights to the predicted local distance differences, predicted global structural quality, and the ratio of polar region solvent accessible surface area to total solvent accessible surface area of each mutant sequence after the preliminary screening, and obtaining a comprehensive stability score for each mutant sequence by weighted summation, and selecting the top P mutant sequences with the largest comprehensive stability scores to enter the next round of screening; calculating the percentage ranking of enzyme catalytic efficiency and alignment score for each mutant sequence obtained after the first two rounds of screening and assigning weights, and obtaining a comprehensive score for each mutant sequence by weighted summation, and selecting and outputting the multidimensional evaluation results corresponding to the top Q mutant sequences with the largest comprehensive scores; ③ Mutational adaptability-driven evolution-combination collaborative optimization screening scheme, including: selecting mask marginal probability, mean entropy of each site and pseudo-perplexity to set thresholds for them to perform evolutionary adaptability initial screening and filtration of mutant sequences in the protein high-throughput mutation sequence library; calculating the binding ability score of each mutant sequence after the initial screening based on hydrophobic interaction distance and hydrogen bond distance; calculating the percentage ranking of each mutant sequence obtained after the first two rounds of screening based on the mask marginal probability, mean entropy of each site and pseudo-perplexity and performing weight allocation, and obtaining the evolutionary adaptability score of each mutant sequence by weighted summation; weighting the binding ability score and evolutionary adaptability score and performing weighted summation to obtain the comprehensive score of each mutant sequence, and selecting and outputting the multidimensional evaluation results corresponding to the top M mutant sequences with the largest comprehensive scores.
2. The artificial intelligence-based high-throughput luciferase optimization screening method according to claim 1, characterized in that: The step (1) specifically includes: Obtain the complete amino acid sequence of the target protein, determine the fixed sites through evolutionary analysis and relevant literature information to identify the conserved regions; for non-conserved regions outside the conserved regions, use random mutagenesis or directed mutagenesis strategies to generate multiple mutant sequences to construct a high-throughput protein mutation sequence library.
3. The artificial intelligence-based high-throughput luciferase optimization screening method according to claim 1, characterized in that: In the step (2), the physicochemical properties include aromaticity, instability index, isoelectric point, flexibility and thermal stability; the enzyme function index includes enzyme catalytic efficiency; the mutation adaptability includes protein pseudo-likelihood and pseudo-perplexity, mutation marginal probability, wild-type marginal probability, mask marginal probability, site index cross entropy and mean entropy of each site.
4. The artificial intelligence-based high-throughput luciferase optimization screening method according to claim 3, characterized in that: The step (2) specifically includes: For protein sequences in the high-throughput protein mutation sequence library, the BioPython tool is used to obtain the aromaticity, instability index, isoelectric point and flexibility of the protein, and the TemStaPro model is used to obtain the thermal stability of the protein, including the thermal stability classification results, stable temperature zones and conflict markers clash at different temperatures; the DLKcat tool is used to predict the enzyme catalytic efficiency of the protein; the ESM-2 model is used to predict the mutation effect of the protein, and obtain the pseudo-likelihood and pseudo-perplexity of the protein, mutation marginal probability, wild-type marginal probability, mask marginal probability, site index cross entropy and the mean entropy of each site.
5. The artificial intelligence-based high-throughput luciferase optimization screening method according to claim 1, characterized in that: In the step (3), the structure prediction indicators include predicted local distance difference, predicted global structure quality, predicted average error, root mean square deviation, template modeling score and mutant-wild type protein structure similarity index; the physicochemical information includes solvent accessible surface area; the protein-substrate binding force includes hydrogen bonds, hydrophobic interactions, salt bridges, water bridges, halogen bonds and π-π stacking non-covalent interaction modes present in the three-dimensional structure of the protein-substrate complex.
6. The artificial intelligence-based high-throughput optimization screening method for luciferase according to claim 5, characterized in that: The step (3) specifically includes: For protein sequences in the protein high-throughput mutation sequence library, the ESMfold model is used to predict protein structure, and the predicted local distance difference, predicted global structure quality and predicted average error are output; the CEAlign tool and TM-align tool are used to calculate the root mean square deviation and template modeling score between the mutant and wild-type protein structures, respectively; the Foldseek tool is used to obtain the mutant-wild-type protein structure similarity indicators, including structural consistency percentage, alignment length, number of mismatches and alignment score; the FreeSASA tool is used to obtain the solvent accessible surface area, including the total solvent accessible surface area, polar region solvent accessible surface area, non-polar region solvent accessible surface area and relative solvent accessible surface area; the RFAA model is used to generate the three-dimensional structure of the protein-substrate complex, and the PLIP tool is then used to detect the hydrogen bonds, hydrophobic interactions, salt bridges, water bridges, halogen bonds and π-π stacking non-covalent interaction modes therein; based on the predicted protein structure, the PLANET model is used to obtain the protein-substrate affinity.
7. A system for implementing the artificial intelligence-based high-throughput optimization screening method for luciferase according to any one of claims 1 to 6, characterized in that: include: Sequence mutation generation module, used to construct a high-throughput protein mutation sequence library based on evolutionary analysis and mutation strategy simulation; The sequence-level AI prediction and evaluation module is used to predict and evaluate the physicochemical properties, enzyme function indicators, and mutation adaptability of protein sequences in the protein high-throughput mutation sequence library through a variety of sequence-level evaluation models and tools to obtain sequence-level evaluation results; The structure-level prediction and evaluation module is used to predict protein structures of protein sequences in the protein high-throughput mutation sequence library using a variety of structure prediction models and tools, so as to perform prediction and evaluation of various structural prediction indicators, physicochemical information, protein-substrate binding ability and protein-substrate affinity of the protein, and obtain structure-level evaluation results; The customized screening module is used to customize the screening scheme based on the sequence-level evaluation results and the structure-level evaluation results, combining threshold filtering with dynamic weighted scoring, so as to optimize the screening of protein high-throughput mutation sequence libraries according to different screening objectives.
8. A screening device based on the artificial intelligence-based high-throughput optimization screening method for luciferase according to any one of claims 1 to 6, characterized in that: include: A storage module is used to store the constructed protein high-throughput mutation sequence library and the sequence-level evaluation results and structure-level evaluation results of each mutation sequence therein; The computational module integrates multiple sequence-level evaluation models and tools as well as multiple structure prediction models and tools to predict and evaluate protein sequences in the high-throughput protein mutation sequence library at the sequence and structure levels in terms of physicochemical properties, enzyme function indicators, mutation adaptability, structure prediction indicators, physicochemical information, protein-substrate binding ability, and protein-substrate affinity, and obtain sequence-level evaluation results and structure-level evaluation results; A dynamic weighting processor is used to perform threshold filtering and dynamic weight allocation, and to optimize the screening of mutant sequences in the protein high-throughput mutant sequence library using customized screening schemes.
9. A computer-readable storage medium, characterized in that A program is stored thereon, which, when executed by a processor, is used to implement the artificial intelligence-based high-throughput optimization screening method for luciferase according to any one of claims 1 to 6.
Citation Information
Patent Citations
Protein optimization design and screening method and device based on artificial intelligence algorithm
CN120260679A
Method for designing optimized mutant protein sequence using amino acid coevolutionary information
US20210202039A1