Sugar molecule-based aptamer de novo design virtual screening method and system

By constructing a virtual screening system for de novo design of aptamers for sugar molecules and using computer automation to build a three-dimensional structural database of nucleic acid aptamers, the problems of cumbersome and blind screening processes in traditional screening are solved, achieving efficient and accurate screening of nucleic acid aptamers, which is suitable for a variety of research needs.

CN120913700APending Publication Date: 2025-11-07INSTITUTE OF PROCESS ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410545809.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-06
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

The screening process for nucleic acid aptamers of sugar molecules in existing technologies is cumbersome, time-consuming, costly, and prone to blindness. Traditional methods cannot effectively guide the binding mechanism between sugar molecules and nucleic acid aptamers.

Method used

A de novo virtual screening method based on glycomolecule-based aptamer design was adopted. A three-dimensional structure database of nucleic acid aptamers was automatically constructed by computer to achieve automated screening, including three-dimensional structure library construction, aptamer library screening and experimental verification. Molecular docking and molecular dynamics simulations were performed using Python, Mxfold, Amber, Autodock Vina and Amber algorithms to reduce experimental dependence.

Benefits of technology

It achieves efficient and accurate nucleic acid aptamer screening, shortens screening time, reduces costs, improves screening targeting and success rate, is suitable for various research needs, and conforms to the concept of sustainable development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913700A_ABST
    Figure CN120913700A_ABST
Patent Text Reader

Abstract

The invention discloses an aptamer de novo design virtual screening method and system based on sugar molecules, and belongs to the field of biology. Comprising the following steps: constructing a three-dimensional structure library: constructing an initial aptamer library and a target sugar molecule library; aptamer library screening, wherein the aptamer library screening comprises initial nucleic acid aptamer molecule docking and filtering, optimized nucleic acid aptamer library construction, optimized nucleic acid aptamer molecule docking and molecular dynamics simulation multi-round filtering; and experimental verification: detecting the binding capacity of the potential nucleic acid aptamer and the sugar molecules. Compared with the prior art, the method has the advantages that the nucleic acid aptamer three-dimensional structure database can be independently constructed, and consumption of a large number of traditional wet experiments in several months is replaced; the screening process is more efficient and controllable, virtual screening of aptamers designed from the beginning of targeted sugar molecules is achieved, the accuracy is higher, and the screening time and cost are effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biology, more particularly, to a method and system for de novo design of virtual screening of glycomer-based aptamer. BACKGROUND

[0002] Currently, researches show that the development process of diseases is related to the glycomers on the surface of proteins, such as the tumor marker CA19-9 of pancreatic cancer, the tumor marker CA125 of ovarian cancer, and the tumor marker CA153 of breast cancer, so that early warning, real-time monitoring and prognosis of diseases can be achieved by detecting the types and quantities of glycomers on the surface of proteins. Nucleic acid aptamer was proposed in 1990, which is ribonucleic acid or single-stranded deoxyribonucleic acid. Nucleic acid aptamer is called "chemical antibody", which can bind to various targets with high specificity and high affinity. Compared with traditional detection of glycomers by lectin, nucleic acid aptamer is more complex and diverse, and can more accurately recognize glycomers; at the same time, the three-dimensional structure size of nucleic acid aptamer is relatively small compared with the structure of glycomer, and is more matched in three-dimensional binding conformation. Therefore, nucleic acid aptamer provides an important tool for accurate detection of glycomers.

[0003] However, at the same time, there are still many problems in the screening of nucleic acid aptamer for glycomer. In the existing traditional nucleic acid aptamer screening method, the exponential enrichment of ligand system evolution technology, the glycomer and nucleic acid aptamer library need to be fixed on the magnetic beads, and through multiple rounds of positive screening, reverse screening, separation, amplification and sequencing steps. This screening method has the problems of complicated process, long time and high cost, and the principle of this traditional screening method is through the black box screening method of incubating a large number of nucleic acid aptamers and target molecules, which increases the blindness and failure rate of screening.

[0004] In the related art, as disclosed in Chinese patent document CN110111849A, a nucleic acid aptamer computer-aided screening method based on a high-performance computing platform and a nucleic acid aptamer are disclosed, the nucleic acid aptamer and the target protein are constructed by using the high-performance computing platform, and then the binding capacity and the interaction mechanism of the two are predicted by molecular docking, which is more targeted and instructive than traditional aptamer screening experiments. However, the platform only uses molecular docking to predict the binding capacity of the nucleic acid aptamer and the target protein, which has low accuracy and does not consider the influence of water molecules and salt ions on the complex under dynamic conditions. The platform uses molecular docking software, which cannot realize automatic batch processing, increases the operation steps and reduces the efficiency of the computer. At the same time, in the complex of the nucleic acid aptamer and the target protein, the nucleic acid aptamer is the ligand and the target protein is the receptor, while in the complex of the nucleic acid aptamer and the sugar molecule, the nucleic acid aptamer is the receptor and the sugar molecule is the ligand, the binding mode and the binding mechanism of the two are completely different, which cannot guide the computer-aided screening of the nucleic acid aptamer for the sugar molecule.

[0005] In summary, how to overcome the complexity and blindness of the aptamer screening of the sugar molecule is a problem to be solved in the prior art. SUMMARY

[0006] 1. Technical problem to be solved

[0007] In view of the problem of how to reduce the complexity and blindness of the aptamer screening of the sugar molecule in the prior art, the present application provides a method and system for de novo design virtual screening of aptamer based on sugar molecules, which can realize automatic screening of potential aptamer of the sugar molecule by computer without traditional screening, has higher accuracy, and effectively reduces screening time and cost. The present application can realize self-construction of a nucleic acid aptamer three-dimensional structure database, replacing the traditional wet experiment consumption for several months. The screening process is more efficient and controllable, realizing de novo design of the aptamer for the target sugar molecule, which has higher accuracy and effectively reduces the screening time and cost.

[0008] 2. Technical solution

[0009] The object of the present application is achieved by the following technical solution.

[0010] A method for de novo design virtual screening of aptamer based on sugar molecules, comprising the following steps,

[0011] Three-dimensional structure library construction: constructing an initial aptamer library and a target sugar molecule library;

[0012] Aptamer library screening: aptamer library screening includes initial nucleic acid aptamer molecular docking filtering, construction of an optimized nucleic acid aptamer library, optimized nucleic acid aptamer molecular docking, and multiple rounds of molecular dynamics simulation filtering;

[0013] Experimental verification: detect the binding ability of potential aptamer and sugar molecules.

[0014] Further, the step of constructing the initial aptamer library specifically comprises: using the random number generation function of Python3, setting the base length and number of aptamers, and outputting the initial sequence library of aptamers;

[0015] Input the initial sequence of aptamers, set the judgment conditions: the proportion of A and T and the proportion of C and G, and output the sequence library filtered by sequence;

[0016] Input the sequence filtered sequence library, use Mxfold to predict the secondary structure, and output the secondary structure representation and energy of each sequence;

[0017] Input the secondary structure representation and energy of the aptamer, set the judgment condition: the number of paired bases, according to the judgment ability, output the sequence library filtered by secondary structure;

[0018] Input the secondary structure filtered sequence library, use clustering analysis algorithm, output high difference sequence and corresponding secondary structure;

[0019] Input the sequence and corresponding secondary structure, output the tertiary structure of the nucleic acid aptamer by 3dRNA algorithm.

[0020] Further, set the judgment condition: the proportion of A and T is not more than 25%, and the proportion of C and G is not more than 50%, output the sequence library filtered by sequence.

[0021] Further, set the judgment condition: the number of paired bases is more than 50%, and the energy is less than-5.7kcal / mol, output the secondary structure filtered sequence library.

[0022] Further, the step of constructing the target sugar molecule library specifically comprises: inputting the molecular formula of the sugar molecule target, and outputting the three-dimensional structure of the sugar molecule target by Amber algorithm;

[0023] The initial nucleic acid aptamer molecule docking filtering step specifically comprises: inputting the initial aptamer library and the three-dimensional structure of the sugar molecule, using the molecular docking algorithm, calculating the binding ability of each aptamer and the sugar molecule, and outputting the scoring value of the binding conformation of each aptamer;

[0024] Input the aptamer number, secondary structure and scoring value, sort them from low to high according to the scoring value, set that the secondary structure of the aptamer only contains one and two stem loop structures, and output the aptamer sequence number filtered by scoring value.

[0025] Further, the optimization of the nucleic acid aptamer library construction step is specifically: input the three-dimensional structure of the nucleic acid aptamer obtained by filtering, calculate the three-dimensional structure change of the nucleic acid aptamer in the production process by using molecular dynamics Amber simulation;

[0026] The production trajectory is imported, the number and binding energy of the aptamer are input, the three-dimensional structure at the moment of the lowest energy is extracted, and the three-dimensional steady-state structure of the aptamer is output.

[0027] Further, the optimization of the nucleic acid aptamer molecule docking step is specifically: input the three-dimensional structure of the optimized aptamer library and sugar molecules, calculate the binding ability of each aptamer and sugar molecules, and output the binding conformation of sugar molecules for each aptamer.

[0028] Further, the molecular dynamics simulation multi-round filtering step is specifically: input the binding conformation obtained by molecular docking, calculate the binding energy of each aptamer and sugar molecules, and output the binding energy of each nanosecond production process;

[0029] The aptamer number and binding energy are input and sorted, and the aptamer sequence with the strongest binding energy is output.

[0030] Further, the experimental verification step is specifically: synthesizing the output aptamer sequence, verifying the binding energy in vitro by using colloidal gold method, and outputting the binding ability.

[0031] The experimental and simulation predicted binding energy is input and compared and sorted, and the nucleic acid aptamer sequence with the consistent trend is output.

[0032] According to the above-mentioned sugar molecule-based aptamer de novo design virtual screening method system, a three-dimensional structure library construction module is used to construct an initial aptamer library and a target sugar molecule library;

[0033] The aptamer library screening module is used for aptamer library screening, including initial nucleic acid aptamer molecule docking filtering, optimization of nucleic acid aptamer library construction, optimization of nucleic acid aptamer molecule docking, and molecular dynamics simulation multi-round filtering.

[0034] The experimental verification module is used for detecting the binding ability of the potential nucleic acid aptamer and the sugar molecules.

[0035] 3. Beneficial effects

[0036] Compared with the prior art, the advantages of the present application are:

[0037] (1) The nucleic acid aptamer de novo design virtual screening for sugar molecules is realized, the construction process is operated step by step, the computer screening efficiency is improved, and the cost consumption of traditional wet experiments for several months is replaced;

[0038] (2) Open the "black box" of traditional SELEX screening, avoiding the blindness of traditional aptamer screening, and can observe the sequence and structure analysis of aptamer and the interaction with sugar molecules in real time, revealing the three-dimensional structure of aptamer and the binding mechanism of sugar molecule-aptamer complex, including binding mode, dynamic behavior, etc., realizing the pertinence of screening, and greatly reducing the failure rate of traditional SELEX screening;

[0039] (3) Wide application range, which can be adjusted and optimized according to specific research problems to adapt to different research needs, and can be further popularized to the screening of organic molecules and ions and other aptamers.

[0040] (4) The virtual screening of aptamer reduces the dependence on experiments and avoids the need for a large number of high-purity target molecule samples for screening, which is a new type of efficient and environmentally friendly aptamer screening method, in line with the concept of sustainable development. BRIEF DESCRIPTION OF DRAWINGS

[0041] Figure 1 The flow chart of the method for de novo design of aptamer based on sugar molecules according to an embodiment of the present application;

[0042] Figure 2 The scoring distribution diagram of molecular docking of the aptamer and paromomycin according to an embodiment of the present application;

[0043] Figure 3 The schematic diagram of Stem2 in the secondary structure of the aptamer according to an embodiment of the present application;

[0044] Figure 4 The schematic diagram of the change of total potential energy of the aptamer according to an embodiment of the present application;

[0045] Figure 5 The three-dimensional structure diagram of the aptamer according to an embodiment of the present application;

[0046] Figure 6 The three-dimensional structure diagram of the aptamer and paromomycin complex according to an embodiment of the present application;

[0047] Figure 7 The binding energy distribution diagram of the "Stem1" group of aptamer and paromomycin complex according to an embodiment of the present application;

[0048] Figure 8 The binding energy distribution diagram of the "Stem2" group of aptamer and paromomycin complex according to an embodiment of the present application. DETAILED DESCRIPTION

[0049] The present application will be described in detail below in combination with the drawings and specific embodiments.

[0050] Example 1

[0051] In combination Figure 1 The method for designing and virtually screening the sugar molecule-based aptamer from scratch of the present application comprises three-dimensional structure library construction, aptamer library screening and experimental verification. The specific steps are as follows:

[0052] The construction of the initial nucleic acid aptamer library, the construction of the sugar molecule library, the docking and filtration of the initial nucleic acid aptamer molecules, the construction of the optimized nucleic acid aptamer library, the docking of the optimized nucleic acid aptamer molecules, the multi-round filtration of molecular dynamics simulation and experimental verification. The specific steps are as follows:

[0053] Three-dimensional structure library construction:

[0054] The three-dimensional structure library construction is to automatically construct the initial aptamer library and the target sugar molecule library, which is as follows:

[0055] Construction of the initial nucleic acid aptamer library:

[0056] The random number generation (rand) function of Python3 is used to set the base length and number of aptamers, and the initial sequence library of aptamers is outputted;

[0057] The initial sequence of aptamers is taken as the input, and the judgment condition is set according to the high stability principle of C and G rich nucleic acids: the proportion of A and T is not more than 25%, and the proportion of C and G is not more than 50%, and the sequence library filtered by the sequence is outputted; wherein, A, T, C and G are bases;

[0058] The sequence filtered sequence library is taken as the input, and the secondary structure is predicted by using Mxfold, and the secondary structure representation and energy of each sequence are outputted;

[0059] The secondary structure representation and energy of aptamers are taken as the input, and the judgment condition is set to ensure the stability of aptamers in the actual environment: more than 50% of the paired bases and less than -5.7kcal / mol of energy, and the sequence library filtered by the secondary structure is outputted;

[0060] The sequence library filtered by the secondary structure is taken as the input, and the clustering analysis algorithm is used to output 10,000 high-difference sequences and the corresponding secondary structures;

[0061] The 10,000 sequences and the corresponding secondary structures are taken as the input, and the 3dRNA algorithm is used to output the tertiary structure of the nucleic acid aptamer.

[0062] Construction of the target sugar molecule library:

[0063] The molecular formula of the sugar molecule target is taken as input, and the GLYCAM_06j-1 force field of the Amber algorithm is used to describe the structure and dynamic properties of the sugar molecule. The three-dimensional structure of the sugar molecule target can be output by connecting the monosaccharide molecules.

[0064] Library screening of aptamers:

[0065] The library screening of aptamers includes initial nucleic acid aptamer molecular docking filtering, construction of an optimized nucleic acid aptamer library, optimized nucleic acid aptamer molecular docking, and multiple rounds of filtering of molecular dynamics simulation. The specific process is as follows:

[0066] Initial nucleic acid aptamer molecular docking filtering:

[0067] The automatic molecular docking of the target sugar molecule and the initial aptamer library is realized to preliminarily screen aptamers.

[0068] Specifically, the three-dimensional structure of the initial aptamer library and the sugar molecule is taken as input, and the Autodock Vina algorithm is used to calculate the binding ability of each aptamer and the sugar molecule. Autodock Vina generates 20 binding conformations each time, so as to ensure the diversity of the structure. The scoring values of 20 binding conformations of each aptamer are output.

[0069] The aptamer number, secondary structure, and scoring value are taken as input, and the aptamers are sorted from low to high according to the scoring value. The secondary structure of the aptamer is set to contain only one and two stem loop structures, and the aptamer sequence number after scoring value filtering is output.

[0070] Construction of an optimized nucleic acid aptamer library:

[0071] The automatic molecular dynamics simulation is realized to optimize the three-dimensional structure of the preliminary screened aptamer library.

[0072] Specifically, the three-dimensional structure of the filtered nucleic acid aptamer is taken as input, and the molecular dynamics simulation of Amber is used to calculate the three-dimensional structure changes of the nucleic acid aptamer in the production process of 5 nanoseconds.

[0073] The production trajectory is imported, and the three-dimensional structure at the moment of the lowest energy is extracted using the Cpptraj tool by taking the aptamer number and binding energy as input, and the three-dimensional steady-state structure of the aptamer is output.

[0074] Optimized nucleic acid aptamer molecular docking:

[0075] The automatic molecular docking of the target sugar molecule and the optimized aptamer library is realized.

[0076] Specifically, the three-dimensional structure of the optimized aptamer library and sugar molecules is taken as input, the molecular docking Autodock Vina algorithm is used to calculate the binding capacity of each aptamer and sugar molecule, and 20 binding conformations of each aptamer and sugar molecule are output;

[0077] Molecular dynamics simulation multi-round filtering:

[0078] Realize automated multi-round molecular dynamics simulation to screen the potential aptamer library of target sugar molecules.

[0079] Specifically, the binding conformation obtained by molecular docking is taken as input, and the molecular dynamics Amber simulation is used to calculate the binding energy of each aptamer and sugar molecule in batches, and the binding energy of each nanosecond production process is output. The initial structure conformation is too much, and the preliminary rapid screening is carried out. The first round of production time is set to 2 nanoseconds;

[0080] Take the aptamer number and the binding energy at the 2nd nanosecond as input, sort, repeat the molecular dynamics simulation of the binding conformation with strong binding energy, avoid the contingency of the structure, and further observe the stability of the binding. The second round of production time is set to 50 nanoseconds;

[0081] Take the aptamer number and the binding energy at the 50th nanosecond as input, sort, repeat the molecular dynamics simulation of the binding conformation with strong binding energy, and fully observe the three-dimensional structure changes of the sugar molecules and aptamers in the complex, and calculate the binding capacity of the complex. The third round of production time is set to 200 nanoseconds, and the multi-round molecular dynamics simulation screening is completed;

[0082] Take the aptamer number and the binding energy at the 200th nanosecond as input, sort, and output the aptamer sequence with strong and stable binding energy.

[0083] Experimental verification:

[0084] Detect the binding capacity of different potential nucleic acid aptamers and sugar molecules.

[0085] Specifically, the output aptamer sequence is synthesized, the colloidal gold method is used for in vitro affinity energy verification, and the binding capacity is output.

[0086] Take the experimental and simulated prediction binding energy as input, compare and sort, and output the nucleic acid aptamer sequence with consistent trend.

[0087] Example 2

[0088] In a specific embodiment, the specific steps of the method for de novo design and virtual screening of the sugar molecule-based aptamer of the application are as follows:

[0089] Construction of initial nucleic acid aptamer library:

[0090] For obtaining three-dimensional structure of representative sequence.

[0091] Specifically, 10 million 40-base aptamer sequences are randomly generated using Python3, and the rand function is used to represent bases A, T, C and G with numbers 0, 1, 2 and 3.

[0092] According to the principle of primer design, appropriately increasing the proportion of bases G and C helps the stability of single-stranded deoxynucleotides, and the base complementary pairing of G-C is more stable than A-T, and is more likely to form a stem-loop structure, which is beneficial to the binding of the target. Set the proportion of A and T to be no more than 25%, and the proportion of C and G to be no more than 50%, and preliminarily filter 213808 sequences.

[0093] The sugar molecule binds the hairpin structure of the nucleic acid aptamer, and the secondary structure is filtered in the second round. In order to realize continuous operation, the source code of Mxfold is obtained, the secondary structure and free energy of each sequence are obtained, the secondary structure is represented by points and parentheses, among which the base close to the 5' end of the base complementary pairing is represented by "( ", and the base close to the 3' end is represented by ") ", and the base without base complementary pairing is represented by ".". Set the paired bases to be more than 50%, select the sequences less than -5.7 kcal / mol, and perform the second filtering to obtain 151783 sequences.

[0094] Through two rounds of filtering, a large number of sequences are still retained. In order to further reduce the operation time, clustering analysis is adopted, and the sequences are divided into different clusters according to the similarity between the sequences. Under the condition of ensuring diversity, 10,000 representative sequences are extracted.

[0095] Based on Ubuntu system, 10,000 nucleic acid aptamer sequences and secondary structures are used as input by using bash (Bourne Again Shell) command, and 10,000 nucleic acid aptamer three-dimensional structures are generated by 3dRNA source code, and the file format is.pdb.

[0096] Construction of sugar molecule library:

[0097] For obtaining three-dimensional structure of sugar molecule.

[0098] Initial aptamer molecule docking filtering:

[0099] For extracting the score value of each complex and sorting.

[0100] Specifically, the source code of Autodock Vina is used for molecular docking of 10,000 nucleic acid aptamers and paromomycin, and the scoring function is obtained as the condition for preliminary screening. Before Autodock Vina docking, the receptor and ligand need to be pretreated to obtain the coordinates and size of the binding pocket, and generate the docking information file.

[0101] In structure preparation, the script prepare_receptor.py in AutoDockTools based on AD4 force field is used to pretreat the nucleic acid aptamer and paromomycin and convert them into pdbqt files using Python3.

[0102] The global docking method is adopted, and the center of the nucleic acid aptamer is taken as the center of the binding pocket. The script prepare_gpf.py is used to obtain the center coordinates of the nucleic acid aptamer. The docking file is constructed, and the config.txt is batch-made to contain specific docking information, file names of ligand and receptor, binding pocket center coordinates, and binding pocket size fixed at 50. The docking parameters include the control docking detail level exhaustiveness and the output number of binding models num_modes, where exhaustiveness = 10 and num_modes = 20.

[0103] Autodock Vina is used to batch molecular docking of 10,000 nucleic acid aptamers and paromomycin, and the scoring records and conformations of nucleic acid aptamer and paromomycin complexes are saved.

[0104] As shown in Figure 2 , the optimal principle is adopted, and the lowest scoring value of 20 conformations of each complex is taken as the evaluation, and the scoring values of each complex are extracted in turn for sorting.

[0105] Based on the molecular docking structure of the nucleic acid aptamer library, two groups are selected according to the number of stem-loop structures in the secondary structure of the aptamer, where one and two stem-loop structures are named as "Stem1" group and "Stem2" group, respectively, as shown in Figure 3 , 36 and 75 aptamer sequences of "Stem1" group and "Stem2" group are extracted in turn.

[0106] Molecular dynamics simulation multi-round filtering:

[0107] Molecular dynamics simulation is used for the complex of nucleic acid aptamer-sugar molecule to obtain potential nucleic acid aptamers.

[0108] Specifically, to ensure the stability of the three-dimensional structure of the aptamer, the three-dimensional structure of the aptamer is optimized by molecular dynamics simulation, the Tleap tool of Amber2021 is used, a script prepared by an input file is run by a bash command, the aptamer is subjected to a leaprc.DNA.bsc1 force field, TIP3P water molecules and salt ions are added, and.inpcrd and.prmtop input files for molecular dynamics simulation are obtained, wherein the calculation interaction cutoff radius cut=10.0, and the water box size TIP3PBOX=10.0.

[0109] The.inpcrd and.prmtop input files of the aptamer are subjected to molecular dynamics simulation, and energy minimization, heating, balancing and production are sequentially performed. The energy minimization is sequentially fixed for the aptamer, the main chain of the aptamer and unfixed, and the system is optimized by 20000 times of steepest descent method and 10000 times of conjugate gradient method. The heating is performed by a Langevin heat bath, and the temperature is raised from 0K to 300K for 500ps, and the collision frequency is 2.0ps -1 . The balancing is performed by an isothermal-isobaric ensemble (NPT) system, the temperature is 300K, and the balancing is performed for 3ns. The production extends the balancing system for 10ns, and the step length of all processes is 0.002fs.

[0110] As shown in Figure 4 and Figure 5 , the total potential energy of the system of each aptamer in the production process is extracted using a Python3 script, the total potential energy changing over time is extracted using the trajectory processing tool Cpptraj of Amber2021, the sequence number at the lowest time is obtained, and the three-dimensional structure.pdb file of the aptamer is generated;

[0111] The molecular docking process is repeated, the three-dimensional structure of the aptamer optimized by molecular dynamics simulation is batched with paromomycin, the scoring records and binding conformations of the aptamer and paromomycin complex are saved, the binding conformations of each aptamer and paromomycin are extracted and saved separately as.pdbqt files using the vina_split command, and each.pdbqt file of paromomycin is converted to a.pdb file using the babel command.

[0112] The sugar molecules on the phosphate backbone side of the aptamer where base complementary pairing occurs are extracted respectively, as shown in Figure 6 , the number of hydrogen bonds formed between each aptamer and sugar molecules is counted, sorted, and each aptamer is preferentially selected a few sugar molecules with the most hydrogen bonds.

[0113] The.inpcrd and.prmtop input files for molecular dynamics simulation of the system containing aptamer, sugar molecule, water molecule and salt ion were generated in batches using Amber2021 Tleap, the aptamer used leaprc.DNA.bsc1 force field, the paromomycin used GLYCAM_06j-1 force field, and each complex included four system input files of aptamer, paromomycin, aptamer-paromomycin, and aptamer-paromomycin-water molecule-salt ion.

[0114] The input files of each group of aptamer-paromomycin-water molecule-salt ion system were used for molecular dynamics simulation, and energy minimization, heating, equilibration and production were carried out in sequence, the energy minimization was fixed in sequence for aptamer-paromomycin, the main chain of aptamer-paromomycin and no fixation, and the system was optimized by 8000 times of steepest descent method and 4000 times of conjugate gradient method. Heating used Langevin heat bath, and the temperature was raised from 0K to 300K for 500ps, and the collision frequency was 2.0ps -1 .

[0115] Equilibration used isothermal-isobaric ensemble (NPT) system, the temperature was 300K, and the equilibration was carried out for 3ns, and the production was extended for 2ns. The step length of all processes was 0.002fs, and the motion trajectory of each aptamer-sugar molecule complex was obtained.

[0116] As shown in Figure 7 and Figure 8 , according to the binding free energy calculation method MM / PBSA, the input files of aptamer, paromomycin, aptamer-paromomycin, aptamer-paromomycin-water molecule-salt ion and production motion trajectory of each complex were taken as input, the binding energy of the 2ns production process of each complex was calculated in batches, and the extraction and sorting were carried out.

[0117] The sequences with the smallest binding energy were selected in sequence, and the strongest structure conformation corresponding to 26 and 41 aptamers of "Stem1" group and "Stem2" group was selected for the second round of molecular dynamics simulation, and energy minimization, heating, equilibration and production were carried out in sequence, the production time was 50ns, the binding energy of the 1ns, 2ns, 10ns, 20ns, 30ns, 50ns of the production process of each complex was calculated in batches, and the extraction and sorting were carried out, and 9 aptamers with the most stable binding energy were obtained, and the calculation of binding energy was shown in Table 1, wherein Number was the aptamer number; Binding Energy was the binding energy corresponding to the aptamer and paromomycin.

[0118] Table 1 Binding energy of 9 aptamers with the most stable binding energy

[0119]

[0120] To ensure the stability of the binding, the complex of aptamer and sugar molecule with strong binding ability was selected for the third round of molecular dynamics simulation. Energy minimization, heating, equilibrium and production were performed in turn, and the production time was 200 ns. The binding energy of each complex in the production process was obtained every ns. Eight potential aptamers were obtained, and the nucleic acid sequence information and the average binding energy are shown in Table 2. In Table 2, Name is the name of the aptamer; Sequence is the sequence of the aptamer; Binding Energy is the corresponding binding energy of the aptamer and the paromomycin.

[0121] Table 2 Sequence and average binding energy of eight potential aptamers

[0122]

[0123] Experimental verification:

[0124] The binding ability of different potential aptamers and sugar molecules was detected.

[0125] Specifically, the computer screened potential aptamers for experimental verification. Nano gold colorimetric method was used. 50 μL of AuNPs solution was added to a 96-well transparent enzyme plate, followed by the addition of different aptamer solutions of 25 μL. The aptamer needed to be pretreated at 95℃ for 5 min, 4℃ for 10 min and 37℃ for 10 min. Incubate on the plate shaker for 20 min, then add 25 μL of paromomycin solution, incubate on the plate shaker for 60 min, add 10 μL of the best concentration of NaCl solution, mix well, and record the absorbance values of the solution at 520 nm and 650 nm after 10 min of room temperature and light protection. The results are shown in Table 3. In Table 3, Number is the aptamer number; ΔA650 / A520 is the absorbance ratio; Binding Energy is the corresponding binding energy of the aptamer and the paromomycin.

[0126] The experimental results show that the method effectively obtains the aptamer of paromomycin. By comparing the predicted binding energy of molecular dynamics simulation, both have a good trend relationship.

[0127] Table 3 Comparison of binding energy and predicted binding energy detected by nano gold colorimetric method

[0128] Number ΔA650 / A520 Binding Energy (kcal / mol) S2_2300 0.43 -32.39±2.13 S1_383 0.39 -32.32±4.06 S2_892 0.36 -24.27±1.83 S2_922 0.32 -25.77±2.90 S1_779 0.30 -25.92±3.10 S1_858 0.28 -27.03±3.06 S1_795 0.25 -20.58±4.06 S2_2366 0.22 -23.17±2.77

[0129] Binding Figures 1 to 8 The system for de novo design of aptamer based on sugar molecules includes a three-dimensional structure library construction module, an aptamer library screening module, and an experimental verification module. Specifically as follows:

[0130] a three-dimensional structure library construction module for constructing an initial aptamer library and a target sugar molecule library;

[0131] an aptamer library screening module for screening the aptamer library, including initial nucleic acid aptamer molecule docking filtering, optimized nucleic acid aptamer library construction, optimized nucleic acid aptamer molecule docking, and multiple rounds of filtering of molecular dynamics simulation;

[0132] an experimental verification module for detecting the binding ability of a potential nucleic acid aptamer and a sugar molecule.

[0133] In one specific embodiment, the system for de novo design of aptamer based on sugar molecules is as follows:

[0134] The nucleic acid aptamer library construction module: 10 million 40-base aptamer sequences are randomly generated by using Python3, a random number generation (rand) function is used, and the numbers 0, 1, 2, and 3 represent the bases A, T, C, and G, respectively; the first filtering is performed by setting the proportions of A and T to be no more than 25% and the proportions of C and G to be no more than 50%;

[0135] The source code of the secondary structure prediction software Mxfold is used to obtain the secondary structure and free energy of each sequence, the secondary structure is represented by points and parentheses, wherein the base close to the 5' end of the base pair formed by base complementation is represented by "(", the base close to the 3' end is represented by ") ", and the base not subjected to base complementation is represented by "."; the second filtering is performed by setting the paired bases to be more than 50% and selecting low-energy sequences;

[0136] Finally, further reduction is performed by cluster analysis, the sequences are divided into different clusters according to the similarity between the sequences, 10,000 representative sequences are extracted under the condition of ensuring diversity, and the three-dimensional structures of the 10,000 representative sequences are obtained by using the publicly disclosed 3dRNA source code.

[0137] The sugar molecule library construction module: the Tleap tool of the molecular dynamics simulation software Amber2021 is used to assemble sugar units by using the GLYCAM_06j-1 force field parameters to obtain the three-dimensional structure of the sugar molecule.

[0138] The molecular docking filtering module: the pdb files of the nucleic acid aptamer and the sugar molecule are preprocessed and converted into pdbqt files by using Python3 based on the AD4 force field in AutoDockTools;

[0139] Adopt the global docking mode, take the center of the aptamer as the binding pocket center, and construct the docking file; use the source code of Autodock Vina to perform batch molecular docking on 10000 nucleic acid aptamers and sugar molecules, save the scoring records and conformations of 10000 nucleic acid aptamer-sugar molecule complexes;

[0140] Adopt the optimal principle, take the lowest scoring value of each complex 20 conformations as the evaluation, extract the scoring value of each complex in turn, and sort.

[0141] Molecular dynamics simulation multi-round filtering module: extract the nucleic acid aptamer with high scoring in the molecular docking, and use Tleap of Amber2021 to generate.inpcrd and.prmtop input files of molecular dynamics simulation of the system containing nucleic acid aptamer, water molecules and salt ions in batch;

[0142] Perform energy minimization, heating, balancing and production in turn, obtain the motion trajectory of each nucleic acid aptamer; use Python3 script to extract the total potential energy of the system of each nucleic acid aptamer in the production process, use the trajectory processing tool Cpptraj of Amber2021, and generate the three-dimensional structure.pdb file of nucleic acid aptamer according to the order number of the lowest time;

[0143] Repeat the specific steps of molecular docking filtering, save and extract the scoring records and binding conformations of nucleic acid aptamer-sugar molecule complex respectively;

[0144] The sugar molecule is combined on the phosphate backbone side of the nucleic acid aptamer which occurs base complementary pairing, and the sugar molecule binding conformation of each nucleic acid aptamer meeting the above conditions is extracted in turn;

[0145] Count the number of hydrogen bonds formed between each nucleic acid aptamer and sugar molecule, sort, and select several sugar molecules with the most hydrogen bonds for each nucleic acid aptamer;

[0146] Use Tleap of Amber2021 to generate.inpcrd and.prmtop input files of molecular dynamics simulation of the system containing nucleic acid aptamer, sugar molecule, water molecule and salt ion in batch;

[0147] Perform energy minimization, heating, balancing and production in turn, and the production time is 2ns, obtain the motion trajectory of each nucleic acid aptamer-sugar molecule complex; according to the binding free energy calculation method MM / PBSA, calculate the binding energy of the 1ns and 2ns of the production process of each complex in batch, extract and sort;

[0148] The complex of aptamer-glycomolecule with strong binding capacity was selected for the second round of molecular dynamics simulation, and energy minimization, heating, equilibrium and production were carried out in turn. The binding energy of the first 1ns, 2ns, 10ns, 20ns, 30ns and 50ns of the production process of each complex was calculated in batches, and was extracted and sorted.

[0149] In order to ensure the stability of the binding, the complex of aptamer-glycomolecule with strong binding capacity was selected for the third round of molecular dynamics simulation, and energy minimization, heating, equilibrium and production were carried out in turn. The binding energy of each ns of the production process of each complex was obtained, and the potential aptamer was obtained.

[0150] Verification module: 50 μL of AuNPs solution was added to a 96-well transparent enzyme plate, followed by the addition of different aptamer solutions 25 μL respectively. The aptamer needs to be pretreated at 95℃ for 5 min, 4℃ for 10 min and 37℃ for 10 min, and then incubated on a plate shaker for 20 min. Then 25 μL of glycomolecule was added and incubated on a plate shaker for 60 min. 10 μL of the best concentration of NaCl solution was added, mixed well, and the absorbance values of the solution at 520 nm and 650 nm were recorded after 10 min of incubation at room temperature in the dark.

[0151] The application is based on the significant changes of sugar chains in physiological, non-disease or pathological states, and breaks through the key technical difficulties of nucleic acid aptamer in efficient detection of sugar molecules. In the traditional SELEX screening, a few aptamers binding to the target can be obtained only through massive screening in the "black box" of 1016 aptamer library. However, the high proportion of rotatable bonds of sugar molecules undoubtedly increases the blindness and failure rate of a large number of screening. In order to open a new path for efficient screening of "black box", the application proposes and constructs a method and system for de novo design virtual screening of aptamer based on sugar molecules, and independently constructs a nucleic acid aptamer three-dimensional structure database. The calculation simulation can be completed in a few days, replacing the traditional wet experiment consumption of several months. The screening method is computerized, and the research process is automated and highly parallelized, which can perform large-scale simulation and data processing at a faster speed, thereby meeting the needs of simultaneous virtual screening of aptamer library of different sugar molecules, saving time and improving efficiency. The method successfully opens a new path for screening of "black box", reveals the three-dimensional structure of nucleic acid aptamer and the binding mechanism of sugar molecule-nucleic acid aptamer complex, so that the screening process is more efficient and controllable. Without wet experiment, the method solves the multiple steps required in traditional SELEX technology, such as target fixation, aptamer library, separation, amplification and sequencing, thereby reducing the complexity and cost of screening. It is a new type of efficient and environmentally friendly aptamer screening method, which meets the concept of sustainable development. The method is widely applicable to a wide range of targets, and the de novo design virtual screening of aptamer based on sugar molecules can be further applied to the screening of nucleic acid aptamer of organic molecules and ions.

[0152] The application and its embodiments have been described above in a schematic manner, and the description is not limiting, and the application can be implemented in other specific forms without departing from the spirit or essential characteristics of the application. The embodiments shown in the drawings are only one of the embodiments of the application, and the actual structure is not limited thereto, and any reference signs in the claims should not limit the claims. Therefore, if a person skilled in the art is inspired by the application, without departing from the spirit of the application, the similar structure and embodiments of the technical solution can be designed without creativity, which should belong to the protection scope of the patent. In addition, the word "comprising" does not exclude other elements or steps, and the word "one" before the element does not exclude "multiple" elements. The multiple elements stated in the product claim can also be realized by software or hardware. The words "first", "second" and the like are used to indicate names, and do not indicate any specific order.

Claims

1. A method for de novo design of aptamer based on sugar molecules, comprising the steps of, three-dimensional structure library construction: constructing initial aptamer library and target sugar molecule library; aptamer library screening: aptamer library screening includes initial nucleic acid aptamer molecule docking filtering, construction of optimized nucleic acid aptamer library, optimized nucleic acid aptamer molecule docking and multiple rounds of filtering of molecular dynamics simulation. 2.The method for de novo design of aptamer based on sugar molecules according to claim 1, wherein, the step of constructing initial aptamer library specifically comprises: using the random number generation function of Python3, setting the base length and number of aptamer, and outputting the initial sequence library of aptamer; inputting the initial sequence of aptamer, setting the judgment conditions: the proportion of A and T and the proportion of C and G, and outputting the sequence library filtered by sequence; inputting the sequence library filtered by sequence, using Mxfold to predict secondary structure, and outputting the secondary structure representation and energy of each sequence; inputting the secondary structure representation and energy of aptamer, setting the judgment conditions and outputting the sequence library filtered by secondary structure according to the judgment ability; inputting the sequence library filtered by secondary structure, using clustering analysis algorithm, and outputting the sequence with high difference and the corresponding secondary structure; inputting the sequence and the corresponding secondary structure, and outputting the three-dimensional structure of nucleic acid aptamer by 3dRNA algorithm. 3.The method for de novo design of aptamer based on sugar molecules according to claim 2, wherein, when the proportion of A and T is not more than 25% respectively, and the proportion of C and G is not more than 50% respectively, output the sequence library filtered by sequence. 4.The method for de novo design of aptamer based on sugar molecules according to claim 2, wherein, when the paired bases are more than 50% and the energy is less than-5.7kcal / mol, output the sequence library filtered by secondary structure. 5.The method for de novo design of aptamer based on sugar molecules according to claim 2, wherein, the step of constructing target sugar molecule library specifically comprises: inputting the molecular formula of sugar molecule target, and outputting the three-dimensional structure of sugar molecule target by Amber algorithm; the step of initial nucleic acid aptamer molecule docking filtering specifically comprises: inputting the three-dimensional structure of initial aptamer library and sugar molecule, using molecular docking algorithm, calculating the binding ability of each aptamer and sugar molecule, and outputting the scoring value of the binding conformation of each aptamer; inputting the aptamer number, secondary structure and scoring value, sorting from low to high according to the scoring value, setting that the secondary structure of aptamer only contains one and two stem loop structures, and outputting the aptamer sequence number filtered by scoring value. 6.The method for de novo design of aptamer based on sugar molecules according to claim 5, wherein, the step of constructing optimized nucleic acid aptamer library specifically comprises: inputting the three-dimensional structure of nucleic acid aptamer filtered, using molecular dynamics Amber simulation, and calculating the three-dimensional structure change of nucleic acid aptamer in production process; inputting the trajectory, inputting the number and binding energy of aptamer, extracting the three-dimensional structure at the moment of the lowest energy, and outputting the three-dimensional steady-state structure of aptamer.

7. The method for de novo design of a sugar molecule-based aptamer according to claim 5 or 6, wherein, the step of docking the optimized aptamer library with the sugar molecule comprises inputting the optimized aptamer library and the three-dimensional structure of the sugar molecule, calculating the binding ability of each aptamer with the sugar molecule, and outputting the binding conformation of the sugar molecule by each aptamer.

8. The method for de novo design of a sugar molecule-based aptamer according to claim 7, wherein, the step of multiple rounds of filtering by molecular dynamics simulation comprises inputting the binding conformation obtained by molecular docking, calculating the binding energy of each aptamer with the sugar molecule, and outputting the binding energy of each nanosecond production process; inputting the aptamer number and the binding energy and sorting, and outputting the aptamer sequence with the strongest binding energy.

9. The method of de novo design of a glycomolecule-based aptamer by virtual screening according to claim 1 or 8, characterized in that, Further comprising, experimental verification for detecting the binding ability of the potential aptamer with the sugar molecule; the step of experimental verification comprises synthesizing the output aptamer sequence, verifying the in vitro affinity by the colloidal gold method, and outputting the binding ability; inputting the experimental and simulated prediction binding energy, comparing and sorting, and outputting the aptamer sequence with the consistent trend.

10. The system for de novo design of a sugar molecule-based aptamer according to any one of claims 1 to 9, wherein, the system comprises a three-dimensional structure library construction module for constructing the initial aptamer library and the target sugar molecule library; the aptamer library screening module is used for screening the aptamer library, including the initial aptamer molecule docking filtering, the construction of the optimized aptamer library, the optimized aptamer molecule docking, and the multiple rounds of filtering by molecular dynamics simulation; the experimental verification module is used for detecting the binding ability of the potential aptamer with the sugar molecule.

Citation Information

Patent Citations

  • Nucleic acid adapter computer-aided screening method based on high-performance calculating platform, and nucleic acid adapter

    CN110111849A