Evolutionary biology and computational biology-based screening of rna evolution precursors

By combining methods from evolutionary biology and computational biology, nucleoside analogs were screened as candidate small molecules for RNA precursors, addressing the lack of systematicity and integration in existing studies and achieving efficient and accurate RNA precursor screening.

CN119517174BActive Publication Date: 2025-11-21NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411508319.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-28
Publication Date
2025-11-21
Estimated Expiration
2044-10-28

AI Technical Summary

Technical Problem

Existing research lacks systematicness and integration in exploring RNA nucleoside evolutionary precursors. Traditional methods mainly focus on specific nucleoside analogs and small molecules, neglecting a wider range of potential precursor molecules. Furthermore, evolutionary biology research lacks experimental data to support its findings.

Method used

Based on evolutionary biology and computational biology methods, we constructed a phylogenetic tree and three-dimensional structure of bacterial transcriptases to screen for transcriptases with broad-spectrum characteristics. We then used molecular docking and interaction fingerprinting techniques to screen for nucleoside analogs with similar interactions to the original nucleotides and predicted their double-strand structures.

Benefits of technology

The system screened nucleoside analogs as candidate small molecules for RNA precursors, improving the systematicness and accuracy of the screening, providing theoretical support for computational biology, and discovering RNA precursor molecules with broad binding capacity and stable double-stranded structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119517174B_ABST
    Figure CN119517174B_ABST
Patent Text Reader

Abstract

The application relates to the field of biotechnology and discloses an RNA evolutionary precursor screening method based on evolutionary biology and computational biology, which comprises the following steps: constructing a phylogenetic tree, screening an amino acid sequence, and predicting the three-dimensional structure of a bacterial transcription enzyme dimer; constructing a nucleotide analogue small molecule database; docking the predicted transcription enzyme dimer with nucleoside analogues, screening small molecules with excellent binding modes, screening out molecules similar to original nucleotide-protein interactions as potential candidate molecules; and performing double-stranded structure prediction on the screened candidate molecules to obtain RNA small molecule evolutionary precursors. The application combines evolutionary biology and computational biology methods to systematically screen nucleoside analogues as precursor candidate small molecules of RNA in the pre-RNA world, compared with traditional experimental research on RNA precursors, the method is more systematic, and compared with pure evolutionary biology-based life origin exploration, the method provides theoretical result support of computational biology.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biotechnology, in particular, to an RNA evolutionary precursor screening method based on evolutionary biology and computational biology. BACKGROUND

[0002] At the end of last century, with the discovery of RNA ribozyme catalytic activity, Walter Gilbert, the Nobel Prize winner in chemistry, first proposed the "RNA world" hypothesis. The hypothesis believes that before the emergence of modern life, proteins and DNA may not exist, and RNA, as a key molecule, both transmits genetic information and has catalytic function. Further assumptions suggest that RNA itself is the product of evolution, and there is an earlier "pre-RNA world" in which nucleoside analogs play a role as carriers of genetic information. If both hypotheses are true, the earliest transcriptional proteins may retain the ability to interact with RNA evolutionary precursors, and traces of these interactions still exist in active sites. Therefore, by studying the interactions between widely sampled RNA polymerases and RNA nucleotide analogs in the bacterial life tree, potential nucleoside small molecules can be screened to provide evidence for exploring molecules that perform information transmission functions in the pre-RNA world.

[0003] Although existing research has made some progress in exploring the evolutionary precursors of RNA nucleosides, there are certain limitations. Some research focuses on nucleoside analog small molecules that have been synthesized in the laboratory, and verifies the interaction between these artificially synthesized molecules and related proteins through experiments, and further infers their possibility as RNA nucleoside precursors based on these experimental results. This type of research usually has solid experimental data as support and can provide more direct evidence to support its argument. However, it is worth noting that this research method often uses a reverse reasoning approach. That is, it starts from known, artificially synthesized molecules and reversely deduces that they may be RNA nucleoside precursors. This reasoning approach, although to some extent, can reveal some potential relationships, but lacks other sources of molecular evidence to demonstrate the possibility of RNA nucleoside precursors. In addition, this type of research also lacks a certain degree of systematicness. Since the research focus is mainly on specific nucleoside analog small molecules, it may overlook potential precursor molecules in a broader range. At the same time, different studies often lack sufficient connection and integration, making it difficult to form a comprehensive, systematic theoretical framework to explain the evolutionary precursors of RNA nucleosides.

[0004] Another part of the research adopts the method of evolutionary biology when exploring the structure of proteins related to the origin of life. This kind of research usually focuses on constructing the evolutionary tree of proteins, revealing the genetic relationship and evolutionary history between them through the alignment and analysis of homologous protein sequences in different species. On this basis, researchers will further compare the three-dimensional structures of these proteins to explore their conservation and differences in structure and function. However, such research often stops at constructing the evolutionary tree and comparing the three-dimensional structure, lacking further experimental data support. This means that although these studies can provide valuable information about the evolutionary history of proteins, they may have relatively limited understanding of the specific protein function, active site, and interaction with other molecules.

[0005] Currently, there is no effective solution to the problems in the related art. SUMMARY

[0006] To solve the problems in the related art, the present application proposes an RNA evolutionary precursor screening method based on evolutionary biology and computational biology, to overcome the above technical problems existing in the prior art.

[0007] To this end, the specific technical solutions adopted by the present application are as follows:

[0008] The RNA evolutionary precursor screening method based on evolutionary biology and computational biology comprises the following steps:

[0009] S1. Construct a phylogenetic tree according to the amino acid sequences of bacterial transcription enzymes, screen the amino acid sequences of bacterial transcription enzymes with broad-spectrum characteristics, and predict the three-dimensional structure of the bacterial transcription enzyme dimer;

[0010] S2. Select ribose isomers and modern nucleotide analogs, and construct a complete nucleotide analog small molecule database through base phosphate replacement and spatial isomerization change method;

[0011] S3. Use molecular docking software to dock the predicted transcription enzyme dimer with nucleoside analogs, predict their binding mode, and screen small molecules with excellent binding mode;

[0012] S4. Screen out molecules similar to the original nucleotide-protein interaction through molecular interaction relationship fingerprint analysis technology, and use them as potential candidate molecules;

[0013] S5. Use a double-stranded prediction model to predict the double-stranded structure of the screened candidate molecules, and analyze the RNA small molecule evolutionary precursor based on the prediction results.

[0014] Further, the method comprises the following steps of:

[0015] S11, collecting a plurality of amino acid sequences of related bacterial transcription enzymes, and screening a preset number of amino acid sequences of bacterial transcription enzymes according to sequence length and labeled sequences;

[0016] S12, aligning the screened protein amino acid sequences by using multi-sequence alignment software, and constructing a phylogenetic tree based on maximum likelihood method and evolutionary model;

[0017] S13, screening amino acid sequences of bacterial transcription enzymes with broad spectrum characteristics, and predicting the protein three-dimensional structure model of the screened amino acid sequences by Alphafold2.

[0018] Further, the related bacterial transcription enzyme amino acid sequence is 20000, the preset number of bacterial transcription enzyme amino acid sequences is 2300, and the amino acid sequence of the bacterial transcription enzyme with broad spectrum characteristics is 80.

[0019] Further, the method comprises the following steps of:

[0020] S121, aligning the screened protein amino acid sequences by using multi-sequence alignment software, and inputting the sequence comparison results into bioinformatics software;

[0021] S122, selecting JTT evolutionary model, setting parameters, and running maximum likelihood algorithm according to the set parameters to obtain the phylogenetic tree.

[0022] Further, the method comprises the following steps of:

[0023] S131, evaluating the reliability of the phylogenetic tree and the evolutionary relationship between different sequences based on the topological structure, branch length and bootstrap support rate information of the phylogenetic tree;

[0024] S132, screening the amino acid sequences of bacterial transcription enzymes with broad spectrum characteristics according to the evaluation results, and saving the screened amino acid sequences in a preset format;

[0025] S133, input the amino acid sequence file in the preset format into the locally running Alphafold2 environment, and use the Alphafold2 model to predict the protein three-dimensional structure model of the screened amino acid sequence.

[0026] Further, the selected ribose isomer and modern nucleotide analog are constructed into a complete nucleotide analog small molecule database by base phosphate replacement and spatial isomerization change method, including the following steps:

[0027] S21, select the isomer of ribose, and replace the base and phosphate group to generate a partial small molecule structure space;

[0028] S22, based on the chemical informatics toolkit, combine the selected modern nucleotide analog to construct a complete nucleotide analog small molecule database.

[0029] Further, the predicted transcription enzyme dimer and nucleoside analog are docked by using molecular docking software, the binding mode is predicted, and small molecules with excellent binding mode are screened, including the following steps:

[0030] S31, the predicted transcription enzyme dimer and nucleoside analog are docked by using molecular docking software, and the binding mode is predicted;

[0031] S32, score the docking changes by the scoring function, and select small molecules with excellent binding mode according to the scoring results.

[0032] Further, the number of search spaces in each docking process is 24.

[0033] Further, the molecules similar to the original nucleotide-protein interaction are screened by molecular interaction relationship fingerprint analysis technology, and are used as potential candidate molecules, including the following steps:

[0034] S41, identify the docking conformation by using the screening tool, and screen the ligand conformation with reasonable hydrogen bond interaction with the template DNA according to the filter;

[0035] S42, based on the protein-ligand interaction fingerprint toolkit, generate a binary interaction fingerprint to identify the interaction between the ligand and the amino acid residues of the protein;

[0036] S43, compare the molecular interaction fingerprints generated for each small molecule with the interaction fingerprints of the original nucleotide, calculate the similarity degree of the two, and select a preset number of molecules as potential candidate molecules according to the similarity result.

[0037] Further, the double-stranded structure of the screened candidate molecules is predicted by using the double-stranded prediction model, and the RNA small molecule evolution precursor is obtained based on the prediction result analysis, including the following steps:

[0038] S51, parameters of RNA A-type helix model are used as guidance, wherein the parameters include helix rise height 2.55 angstrom, helix twist angle 32.69, and tilt angle 22.62;

[0039] S52, the RNA small molecule evolution precursor is obtained according to the test result of whether the nucleic acid analog can adopt the RNA-like helix conformation.

[0040] The beneficial effects of the present application are:

[0041] 1) The present application combines evolutionary biology and computational biology methods to systematically screen nucleoside analogs as candidate small molecules of RNA precursors in the pre-RNA world. Compared with the traditional experimental method of researching RNA precursors, this method is more systematic, and compared with the simple evolutionary biology-based exploration of the origin of life, it provides theoretical results support of computational biology.

[0042] 2) The present application first calculates the binding energy of the candidate nucleotide molecules and the broad-spectrum screened bacterial transcription enzyme by molecular docking, screens the candidate molecules with lower binding energy, and analyzes the reasonable binding mode. Then, the candidate molecules similar to the binding mode of the original nucleoside small molecules are screened by molecular fingerprint comparison. Finally, the energy and possibility of these candidate molecules to construct double-stranded structures are calculated and predicted, so as to screen the most suitable small molecules as RNA nucleoside precursors. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0044] Fig. 1 is a flow chart of the RNA evolution precursor screening method based on evolutionary biology and computational biology according to the embodiment of the present application;

[0045] Fig. 2 is a principle schematic diagram of the RNA evolution precursor screening method based on evolutionary biology and computational biology according to the embodiment of the present application;

[0046] Fig. 3 is a technical roadmap of the RNA evolution precursor screening method based on evolutionary biology and computational biology according to the embodiment of the present application. Detailed Implementation

[0047] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention. The components in the drawings are not drawn to scale, and similar component symbols are generally used to represent similar components.

[0048] According to embodiments of the present invention, a method for screening RNA evolution precursors based on evolutionary biology and computational biology is provided.

[0049] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figs. 1-3 As shown, the RNA evolution precursor screening method based on evolutionary biology and computational biology according to an embodiment of the present invention includes the following steps:

[0050] S1. Construct a phylogenetic tree based on the amino acid sequences of bacterial transcriptases, screen for amino acid sequences of bacterial transcriptases with broad-spectrum characteristics, and predict the three-dimensional structure of bacterial transcriptase dimers.

[0051] Specifically, suitable protein sequences are screened and a three-dimensional model is constructed. First, relevant bacterial transcription enzyme amino acid sequences are retrieved on a public protein amino acid sequence website UniProt, obtaining 20,000 transcription enzyme amino acid sequences. Then, 2,300 amino acid sequences are screened according to sequence length (the length of the beta chain is not less than 600 amino acids and not more than 1,600 amino acids) and sequence annotation. Next, the screened protein amino acid sequences are aligned using multi-sequence alignment software, and a powerful bioinformatics software (such as MEGA, Molecular Evolutionary Genetics Analysis) is used to further construct a phylogenetic tree using the Maximum Likelihood (ML) algorithm and a specific evolutionary model. The specific steps include: inputting the sequence alignment results into the bioinformatics software, selecting the JTT (Jones-Taylor-Thornton) evolutionary model, setting parameters, running the maximum likelihood algorithm, and allowing the bioinformatics software to calculate based on the provided sequence data and evolutionary model. Next, after the algorithm is run, the results of the phylogenetic tree are obtained. The topology of the tree, branch length, and bootstrap support rate are carefully checked to evaluate the reliability of the tree and the evolutionary relationship between different sequences. Considering the practical limitations of computational cost, it is not possible to perform detailed three-dimensional structure prediction and molecular docking analysis on all 2,300 bacterial transcription enzyme amino acid sequences. Therefore, in the broad spectrum, this embodiment carefully selects 80 representative bacterial transcription enzyme amino acid sequences. These sequences not only cover all major branches and types in the bacterial transcription enzyme evolutionary map, but also fully reflect the evolutionary relationship and structural characteristics between different bacterial transcription enzymes. Finally, Alphafold2 is further used to predict the three-dimensional structure model of the 80 protein amino acid sequences selected. The specific steps include: saving the selected protein amino acid sequences in FASTA format, submitting the prepared FASTA file to the locally running Alphafold2 environment, selecting the multimer module to model rpoβ and rpoβ', selecting the maximum number of prediction models as 5 times, and the prediction TM score is not less than 0.85. Alphafold2 model completes the prediction of the three-dimensional structure of the selected protein through multiple modules such as multi-sequence alignment, convolutional neural network, and spatial conformation prediction.

[0052] S2, selecting ribose isomers and modern nucleotide analogs, constructing a complete nucleotide analog small molecule database by base phosphate replacement and spatial isomerization change method;

[0053] Specifically, a suitable nucleotide analog small molecule database is screened. A nucleotide is a three-part structure containing a ribose + phosphate + base. In the process of constructing the nucleotide analog database, 227 ribose isomers + 13 modern nucleotide analogs are selected in this embodiment, and further replaced by base phosphate and changed in space isomerism (generated by the EmbedMultipleConfs function under RDKit, which is a very common and powerful chemical information toolkit), forming a nucleotide analog small molecule database of 5964.

[0054] S3, using molecular docking software to dock the predicted transcription enzyme dimer with nucleotide analogs, predicting their binding mode, and screening small molecules with excellent binding mode;

[0055] Specifically, molecular docking. In this key step of molecular docking, the embodiment adopts docking software (such as SMINA software) for processing. The essence of molecular docking is to combine protein models with small molecules, and then the small molecules dynamically change on the active site through rotation, structure adjustment, etc. The docking software will score these changes using a specific empirical function algorithm, and generally the lower the binding energy, the higher the score, which means that the ligand binds more closely to the target. In order to further improve the accuracy of docking, a self-defined scoring function is specially designed and optimized in this embodiment. The core purpose of this function is to accurately capture the interaction between the hydroxyl group and the magnesium ion, which is crucial for screening nucleotide candidates that can interact more accurately with the magnesium ion active site. In this docking process, the number of search spaces in each docking process is set to 24, that is, for each small molecule, 24 different conformations will be generated to dock with the protein. Considering that there are 5964 small molecules and 80 protein models, therefore, finally this embodiment will obtain a large data set, containing 5964 (number of small molecules) x 24 (number of conformations of each small molecule) x 80 (number of protein models) = million-level data points. In the face of such a large data set, screening work is particularly important. In this embodiment, the first step is to preliminarily screen according to the score, and select the top 20% of molecules as the candidate molecule set after the first step of screening. This step aims to screen out the molecules that are most likely to effectively bind to the target from the vast amount of data, providing strong support for subsequent research.

[0056] The scoring function in this embodiment is an empirical scoring function formula provided by SMINA software: Docking_score = -0.035579 * gauss1 - 0.005156 * gauss2 + 0.840245 * repulsion - 0.035069 * hydrophobic - 0.587439 * non_dir_h_bond, two Gaussian terms, one repulsion term, one hydrophobic term, and one non-directional hydrogen bond term, wherein Docking_score is the docking score, representing the output result of the entire formula, representing the comprehensive evaluation score of the ligand and receptor after docking, gauss1 and gauss2 are two quantities related to the Gaussian function, used to simulate the potential energy of interaction between atoms, repulsion is a repulsion term, representing the repulsion between the ligand and the receptor, when the atoms of the two are too close in space, the electron clouds will overlap and produce repulsive force, hydrophobic is a hydrophobic term, representing the hydrophobic interaction between the ligand and the receptor, if the hydrophobic term is small, it means that the hydrophobic interaction between the ligand and the receptor is good, which is conducive to binding, and non_dir_h_bond is a non-directional hydrogen bond term, used to quantify the contribution of hydrogen bond to ligand-receptor binding.

[0057] S4, screening similar molecules to the original nucleotide-protein interaction by molecular interaction relationship fingerprint analysis technology, and taking them as potential candidate molecules;

[0058] Specifically, the molecular action mode is screened. Molecules with high docking scores often have relatively reasonable binding conformations, but sometimes the structure with lower energy is not the conformation when the chemical reaction occurs. Only small molecules with a protein binding mode similar to the original molecule (NTP) are more valuable. In this embodiment, LigGrep is used as a screening tool (a free and open source program that is specifically used to screen and identify docking poses that match the user-specified receptor / ligand interaction) to identify the correct docking conformation, i.e. nucleotides need to form base complementary pairing (hydrogen bond connection) with corresponding DNA to transmit genetic information, and the selected conformation requires hydrogen bonding with the corresponding DNA base. A user-specified filter (written in JSON format) is designed to specifically screen ligand conformations with reasonable hydrogen bond interactions with template DNA. The format is as follows (in simple terms, the distance between the H atom of the DNA base and the O atom of the candidate small molecule is less than 4 angstroms):

[0059]

[0060] Next, a binary interaction fingerprint (a list consisting of 0 and 1) is generated using a protein-ligand interaction fingerprint toolkit (e.g., ProLIF, Protein-Ligand Interaction Fingerprint) to identify interactions, such as hydrogen bonds, hydrophobic interactions, and electrostatic interactions, that occur between the ligand and the amino acid residues of the protein. Then, the molecular interaction fingerprint generated for each small molecule is compared with the interaction fingerprint of the original nucleotide, and the degree of similarity between the two is calculated, and the molecules with a degree of similarity in the top 20% and screened by LigGrep are further screened to enter the third round of screening.

[0061] In this embodiment, TC similarity (Tanimoto coefficient similarity) is used to compare the similarity of two molecular fingers, and the calculation formula is as follows:

[0062]

[0063] In the formula, TC(A, B) represents the similarity of two molecules, A and B represent the two molecular fingerprints as two vectors, A·B represents the dot product of the vectors, and |A| and |B| represent the modulus of the vectors.

[0064] S5, using a double-stranded prediction model to predict the double-stranded structure of the screened candidate molecules, and based on the prediction results, analyzing to obtain the RNA small molecule evolutionary precursor.

[0065] Specifically, double-stranded structure model prediction. The RNA nucleotide precursor to be screened is one of the largest functions, which is to retain the ability to transmit genetic information, and therefore needs to be able to form a relatively stable double-stranded structure. Therefore, in this embodiment, the small molecules screened in the fourth step are further calculated for double-stranded structure model prediction. A nucleotide analog double-stranded system calculation package (e.g., pNAB, The proto-Nucleic-Acid Builder) is used to construct a double-stranded helix model of the screened nucleic acid analog. By using the parameters of the RNA A-type helix model (helix rise height 2.55 angstroms, helix twist angle 32.69, tilt angle 22.62) as a guide, it is tested whether the nucleic acid analog can adopt a helical conformation similar to RNA. Finally, the molecules screened in this embodiment are selected as RNA small molecule evolutionary precursor candidates that can form stable double-stranded structures and have the potential to retain the ability to transmit genetic information.

[0066] To sum up, by means of the technical scheme of the present application, the present application innovatively combines the front technologies of evolutionary biology and computational biology, and constructs a comprehensive and targeted RNA precursor screening method in combination with some characteristics of RNA nucleotides. Firstly, the evolutionary data of bacterial transcription enzymes are systematically collected and an evolutionary tree is constructed, and advanced protein structure prediction technology is used to model the three-dimensional structure of the widely screened amino acid sequences. This step not only lays a solid evolutionary biology foundation for subsequent molecular docking experiments, but also ensures the wide range and depth of the screened molecules in the evolutionary level. Further, in the field of computational biology, the traditional molecular docking technology is used to calculate the energy binding of the screened protein models and the nucleotide analogue small molecule database, and the molecular fingerprint screening and double-stranded model prediction model screening methods are further designed according to the specific binding characteristics of RNA nucleotides and transcription enzymes, so as to realize the fine screening of small molecules. Through the three-layer screening mechanism, the small molecules successfully screened not only exhibit strong binding ability in the wide spectrum (have good average docking scores with more than 80 transcription enzymes), but also have a highly similar binding mode of molecules and proteins (molecular action fingerprint is similar) to the original nucleotide (NTP) and protein, and can form a stable double-stranded structure (double-stranded structure energy is low). This innovative screening method greatly improves the screening efficiency and accuracy of RNA precursors, and provides new scientific basis for the discovery and application of RNA precursors.

[0067] The above merely describes preferred embodiments of the present application but should not be used to limit the present application, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. An RNA evolution precursor screening method based on evolutionary biology and computational biology, characterized in that, The method comprises the following steps: S1, constructing a phylogenetic tree according to the amino acid sequences of bacterial transcription enzymes, screening the amino acid sequences of bacterial transcription enzymes with broad spectrum characteristics, and predicting the three-dimensional structure of the dimer formed by the bacterial transcription enzymes; S2, selecting ribose isomers and modern nucleotide analogs, constructing a complete nucleotide analog small molecule database by base phosphate replacement and spatial isomerization change method; S3, using molecular docking software to dock the predicted transcription enzyme dimer with nucleoside analogs, predicting the binding mode, and screening small molecules with excellent binding mode; S4, screening molecules similar to the original nucleotide-protein interaction through molecular interaction relationship fingerprint analysis technology, and taking them as potential candidate molecules; S5, using a double-chain prediction model to predict the double-chain structure of the screened candidate molecules, and analyzing the RNA small molecule evolutionary precursor based on the prediction result.

2. The RNA evolution precursor screening method based on evolutionary biology and computational biology according to claim 1, characterized in that, The method comprises the following steps: S11, collecting a plurality of related amino acid sequences of bacterial transcription enzymes, and screening a preset number of amino acid sequences of bacterial transcription enzymes according to the sequence length and labeled sequence; S12, aligning the screened protein amino acid sequences using multi-sequence alignment software, and constructing a phylogenetic tree based on maximum likelihood method and evolutionary model; S13, screening the amino acid sequences of bacterial transcription enzymes with broad spectrum characteristics, and predicting the protein three-dimensional structure model of the screened amino acid sequences by Alphafold2.

3. The method of claim 2, wherein the method is based on the evolutionary biology and computational biology of RNA evolution precursors. The related amino acid sequences of bacterial transcription enzymes are 20,000, the preset number of amino acid sequences of bacterial transcription enzymes is 2,300, and the amino acid sequences of bacterial transcription enzymes with broad spectrum characteristics are 80.

4. The method of claim 2, wherein the method is based on the evolutionary biology and computational biology of RNA evolution precursors. The method comprises the following steps: S121, aligning the screened protein amino acid sequences using multi-sequence alignment software, and inputting the sequence comparison result into bioinformatics software; S122, selecting JTT evolutionary model, setting parameters, and running maximum likelihood algorithm according to the set parameters to obtain a phylogenetic tree.

5. The method of claim 2, wherein the method is based on the evolutionary biology and computational biology of RNA evolution precursors. The method comprises the following steps: S131, based on the topological structure, branch length and bootstrap support rate information of the phylogenetic tree, evaluating the reliability of the phylogenetic tree and the evolutionary relationship between different sequences; S132, according to the evaluation result, screening the amino acid sequences of bacterial transcription enzymes with broad spectrum characteristics, and saving the screened amino acid sequences in a preset format; S133, inputting the amino acid sequence file in the preset format into the locally running Alphafold2 environment, and using the Alphafold2 model to predict the protein three-dimensional structure model of the screened amino acid sequences.

6. The RNA evolution precursor screening method based on evolutionary biology and computational biology according to claim 1, wherein, The selected ribose isomer and modern nucleotide analogues are used to construct a complete nucleotide analogue small molecule database by base phosphate replacement and spatial isomerization, including the following steps: S21, selecting the isomer of ribose, and replacing the base and phosphate group to generate part of the small molecule structure space; S22, based on the chemical information toolkit, combined with the selected modern nucleotide analogues, a complete nucleotide analogue small molecule database is constructed.

7. The RNA evolution precursor screening method based on evolutionary biology and computational biology according to claim 1, wherein, The predicted transcription enzyme dimer is docked with nucleoside analogues using molecular docking software to predict the binding mode, and small molecules with excellent binding mode are screened, including the following steps: S31, using molecular docking software to dock the predicted transcription enzyme dimer with nucleoside analogues to predict the binding mode; S32, score the docking changes by scoring function, and select small molecules with excellent binding mode according to the scoring results.

8. The RNA evolution precursor screening method based on evolutionary biology and computational biology according to claim 7, characterized in that, The number of search spaces in each docking process is 24.

9. The RNA evolution precursor screening method based on evolutionary biology and computational biology according to claim 1, wherein, The molecular interaction relationship fingerprint analysis technology is used to screen molecules similar to the original nucleotide-protein interaction and as potential candidate molecules, including the following steps: S41, using the screening tool to identify the docking conformation, and according to the filter, screening ligand conformations with reasonable hydrogen bond interactions with the template DNA; S42, based on the protein-ligand interaction fingerprint toolkit, generate binary interaction fingerprints to identify the interactions between ligands and protein amino acid residues; S43, compare the molecular interaction fingerprints generated for each small molecule with the interaction fingerprints of the original nucleotide, calculate the similarity between the two, and select a preset number of molecules as potential candidate molecules according to the similarity results.

10. The RNA evolution precursor screening method based on evolutionary biology and computational biology according to claim 1, characterized in that, The double-stranded prediction model is used to predict the double-stranded structure of the screened candidate molecules, and the RNA small molecule evolution precursor is analyzed based on the prediction results Including the following steps: S51, by using the parameters of RNA A-type helix model as guidance, wherein the parameters include helix rise height 2.55 angstrom, helix twist angle 32.69, and tilt angle 22.62; S52, according to the test results of whether the nucleic acid analogue can adopt the RNA-like helical conformation, the RNA small molecule evolution precursor is analyzed.

Citation Information

Patent Citations

  • Application of bacterium nucleic acid fingerprinting feature spectral library to identification and classification

    CN102851747A

  • Simulation prediction method for identifying different promoters through directed evolution of monomer polymerase

    CN113764039A