Nucleic Acid Molecule Structure Screening Method and Apparatus, and Computer Device
Based on the calculation simulation of quantum mechanics and molecular dynamics, using multi-dimensional nucleic acid molecular parameters to search for similar structures in the nucleic acid library and perform probability distribution calculations, the problem of insufficient screening of nucleic acid molecular configuration in the prior art is solved, and the accuracy of screening is improved.
Patent Information
- Application Number
- CN202111655356.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-12-30
AI Technical Summary
When screening the configuration of nucleic acid molecules, the prior art relies on the basic energy distribution of nucleic acid molecules, resulting in inaccurate screening results.
By obtaining the configuration and energy data obtained by nucleic acid molecular parameters, including sequence, length, chemical modification, and computational simulation based on quantum mechanics and molecular dynamics, a similar nucleic acid molecular structure was searched for the nucleic acid library, and the probability distribution calculation was performed based on the stability, energy and experimental data of the similar structures, and the configuration of the nucleic acid molecular that meets the conditions was selected.
The accuracy of nucleic acid molecular structure screening is improved, and by considering the stability of multi-dimensional nucleic acid molecular parameters and similar structures, the uncertainty of purely relying on simulation experimental simulation is reduced.
Smart Images

Figure CN114496067B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of biotechnology, and particularly to a method and apparatus for screening nucleic acid molecular structures and a computer device. Background Art
[0002] Nucleic acids are one of the basic substances in cells. Predicting the molecular structures of nucleic acids helps in the study of their functions.
[0003] Related technologies provide a method for predicting nucleic acid molecular structures based on quantum mechanics and molecular dynamics. This method performs simulation calculations based on quantum mechanics and molecular dynamics to determine multiple nucleic acid molecular configurations to be screened, and then based on the energy distributions of these multiple nucleic acid molecular configurations to be screened, one or more nucleic acid molecular configurations that meet the stability conditions are selected from them for research.
[0004] However, the method for screening nucleic acid molecular configurations in related technologies is mainly based on the basic energy distribution of nucleic acid molecules. The calculation of energy distribution is a method of simulating the natural world through physical simulation experiments. The physical simulation experiments themselves rely on mathematical methods for approximate solutions, resulting in inaccurate screening results. Summary of the Invention
[0005] Embodiments of the present disclosure provide a method and apparatus for screening nucleic acid molecular structures and a computer device, which can improve the accuracy of screening nucleic acid molecular structures. The technical solution is as follows:
[0006] At least one embodiment of the present disclosure provides a method for screening nucleic acid molecular structures, the method including:
[0007] Obtain nucleic acid molecule parameters, where the nucleic acid molecule parameters include: nucleic acid molecule sequence, length of the nucleic acid molecule sequence, chemical modification of the nucleic acid molecule, and multiple nucleic acid molecular configurations to be screened and their corresponding energies and experimental data obtained through computational simulations based on quantum mechanics and molecular dynamics;
[0008] Search for similar nucleic acid molecular structures of each nucleic acid molecular configuration among the multiple nucleic acid molecular configurations to be screened in a nucleic acid library;
[0009] Based on the stability of the similar nucleic acid molecular structures, the energies and experimental data corresponding to the multiple nucleic acid molecular configurations to be screened, calculate the probability distribution of the multiple nucleic acid molecular configurations to be screened;
[0010] According to the probability distribution of the multiple nucleic acid molecular configurations to be screened, screen at least one nucleic acid molecular configuration among the multiple nucleic acid molecular configurations to be screened that meets the stability conditions.
[0011] Optionally, searching for similar nucleic acid molecule structures of each nucleic acid molecule configuration among the multiple nucleic acid molecules to be screened in the nucleic acid library includes:
[0012] Inputting the nucleic acid molecule sequence, the length of the nucleic acid molecule sequence, the chemical modification of the nucleic acid molecule, and the multiple nucleic acid molecule configurations to be screened into a similar nucleic acid molecule search model, where the similar nucleic acid molecule search model is used to respectively determine the first similarity between the nucleic acid molecule sequence and each nucleic acid molecule sequence in the nucleic acid library, determine the difference between the length of the nucleic acid molecule sequence and the length of each nucleic acid molecule sequence in the nucleic acid library, determine the second similarity between the chemical modification of the nucleic acid molecule and the chemical modification of each nucleic acid molecule sequence in the nucleic acid library, select nucleic acid molecules from the nucleic acid library for which the first similarity, the difference, and the second similarity all meet the threshold requirements, successively determine the configurations of the selected nucleic acid molecules and the distances from the multiple nucleic acid molecule configurations to be screened, and obtain the similar nucleic acid molecule structures of each nucleic acid molecule configuration based on the distances;
[0013] Obtain the similar nucleic acid molecule structures of each nucleic acid molecule configuration among the multiple nucleic acid molecules to be screened that are searched by the similar nucleic acid molecule search model from the nucleic acid library.
[0014] Optionally, performing the calculation of the probability distributions of the multiple nucleic acid molecule configurations to be screened based on the stability of the similar nucleic acid molecule structures, the energies corresponding to the multiple nucleic acid molecule configurations to be screened, and experimental data includes:
[0015] Inputting the stability of the similar nucleic acid molecule structures, the energies corresponding to the multiple nucleic acid molecule configurations to be screened, and experimental data into a nucleic acid molecule configuration probability distribution calculation model, where the nucleic acid molecule configuration probability distribution calculation model is used to calculate by normalizing the numerical values of the energy of the same nucleic acid molecule configuration, the stability of the similar nucleic acid molecule structures, and the experimental data of the nucleic acid molecule configuration, and obtain the stability probability of each nucleic acid molecule configuration through the calculation; specifically, the calculation process is as follows: form an array with the normalized numerical values of the energy of the nucleic acid molecule configuration, the stability of the similar nucleic acid molecule structures, and the experimental data of the nucleic acid molecule configuration, and use the stability probability algorithm obtained by training to solve the array to obtain the stability probability of each nucleic acid molecule configuration;
[0016] Obtain the probability distributions of the multiple nucleic acid molecule configurations to be screened calculated by the nucleic acid molecule configuration probability distribution calculation model, where the probability distributions of the multiple nucleic acid molecule configurations to be screened include the stability probability of each nucleic acid molecule configuration.
[0017] Optionally, screening at least one nucleic acid molecule configuration with qualified stability from the multiple nucleic acid molecule configurations to be screened according to the probability distribution of the configurations of the multiple nucleic acid molecules to be screened includes:
[0018] Sorting the multiple nucleic acid molecule configurations to be screened according to the stability probability of the multiple nucleic acid molecule configurations, and converting the ranking into a first score;
[0019] Sorting the multiple nucleic acid molecule configurations to be screened according to the drug properties of the multiple nucleic acid molecule configurations, and converting the ranking into a second score;
[0020] Combining the first score and the second score, and screening at least one nucleic acid molecule configuration with a high score from the multiple nucleic acid molecule configurations to be screened.
[0021] Optionally, the method further includes:
[0022] Cleaning and complementing the experimental data of the nucleic acid molecule;
[0023] Normalizing the experimental data of the nucleic acid molecule after cleaning and complementing, and numericalizing and normalizing the non-numerical experimental data.
[0024] Optionally, the experimental data of the nucleic acid molecule configuration to be screened includes data of biomolecular experiments and animal experiments, and the data of the biomolecular experiments and animal experiments includes at least one of experimental environment, half-life, annealing temperature, and drug enrichment.
[0025] At least one embodiment of the present disclosure provides a nucleic acid molecule structure screening device, and the device includes:
[0026] An acquisition module, configured to acquire nucleic acid molecule parameters, where the nucleic acid molecule parameters include: nucleic acid molecule sequence, length of the nucleic acid molecule sequence, chemical modification of the nucleic acid molecule, and multiple nucleic acid molecule configurations to be screened and corresponding energies and experimental data obtained by computational simulation based on quantum mechanics and molecular dynamics;
[0027] A search module, configured to search for similar nucleic acid molecule structures of each nucleic acid molecule configuration in the multiple nucleic acid molecule configurations to be screened in a nucleic acid library;
[0028] A calculation module, configured to perform probability distribution calculation of the multiple nucleic acid molecule configurations to be screened based on the stability of the similar nucleic acid molecule structures, the energies and experimental data corresponding to the multiple nucleic acid molecule configurations to be screened;
[0029] A screening module, configured to screen at least one nucleic acid molecule configuration with qualified stability from the multiple nucleic acid molecule configurations to be screened according to the probability distribution of the multiple nucleic acid molecule configurations.
[0030] Optionally, the search module is configured to input the nucleic acid molecule sequence, the length of the nucleic acid molecule sequence, the chemical modification of the nucleic acid molecule, and multiple nucleic acid molecule configurations to be screened into a similar nucleic acid molecule search model. The similar nucleic acid molecule search model is configured to respectively determine a first similarity between the nucleic acid molecule sequence and each nucleic acid molecule sequence in the nucleic acid library, determine a difference between the length of the nucleic acid molecule sequence and the length of each nucleic acid molecule sequence in the nucleic acid library, and determine a second similarity between the chemical modification of the nucleic acid molecule and the chemical modification of each nucleic acid molecule sequence in the nucleic acid library. Nucleic acid molecules that meet the threshold requirements for the first similarity, the difference, and the second similarity are selected from the nucleic acid library. The configuration of the selected nucleic acid molecule and the distance between the multiple nucleic acid molecule configurations to be screened are determined in sequence, and a similar nucleic acid molecule structure of each nucleic acid molecule configuration is obtained based on the distance;
[0031] Obtain the similar nucleic acid molecule structure of each nucleic acid molecule configuration among the multiple nucleic acid molecule configurations to be screened, which is searched by the similar nucleic acid molecule search model from the nucleic acid library.
[0032] Optionally, the calculation module is configured to input the stability of the similar nucleic acid molecule structure, the energies corresponding to the multiple nucleic acid molecule configurations to be screened, and experimental data into a nucleic acid molecule configuration probability distribution calculation model. The nucleic acid molecule configuration probability distribution calculation model is configured to calculate the stability probability of each nucleic acid molecule configuration by calculating the normalized values of the energy of the same nucleic acid molecule configuration, the stability of the similar nucleic acid molecule structure, and the experimental data of the nucleic acid molecule configuration. The calculation process is as follows: form an array with the normalized values of the energy of the nucleic acid molecule configuration, the stability of the similar nucleic acid molecule structure, and the experimental data of the nucleic acid molecule configuration, and use the stability probability algorithm obtained by training to solve the array to obtain the stability probability of each nucleic acid molecule configuration;
[0033] Obtain the nucleic acid molecule configuration probability distributions calculated by the nucleic acid molecule configuration probability distribution calculation model for the multiple nucleic acid molecule configurations to be screened. The nucleic acid molecule configuration probability distributions for the multiple nucleic acid molecule configurations to be screened include the stability probability of each nucleic acid molecule configuration.
[0034] At least one embodiment of the present disclosure provides a computer device, which includes a processor and a memory. The memory stores at least one program code, and the program code is loaded and executed by the processor to implement the nucleic acid molecule structure screening method as described above.
[0035] At least one embodiment of the present disclosure provides a computer-readable storage medium, in which at least one piece of program code is stored, and the program code is loaded and executed by a processor to implement the nucleic acid molecule structure screening method as described in any one of the preceding items.
[0036] The beneficial effects brought by the technical solutions provided by the embodiments of the present disclosure are as follows:
[0037] In the embodiments of the present disclosure, based on the computational simulations of quantum mechanics and molecular dynamics, nucleic acid molecule parameters in multiple dimensions are used to search for similar nucleic acid molecule structures in a nucleic acid library, and multiple nucleic acid molecule configuration probability distributions are calculated according to the stability of the similar nucleic acid molecule structures and the energy and experimental data corresponding to the nucleic acid molecule configurations to be screened, and then at least one nucleic acid molecule configuration meeting the conditions of stability is screened out. Since the above solutions refer to the stability of similar nucleic acid molecule structures and simultaneously consider the energy and experimental data corresponding to the nucleic acid molecule configurations to be screened, rather than simply relying on simulation experiment simulations to solve, the accuracy of nucleic acid molecule structure screening is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0039] Figure 1 is a flowchart of a nucleic acid molecule structure screening method provided by an embodiment of the present disclosure;
[0040] Figure 2 is a flowchart of a nucleic acid molecule structure screening method provided by an embodiment of the present disclosure;
[0041] Figure 3 is a schematic structural diagram of a nucleic acid molecule structure screening device provided by an embodiment of the present disclosure;
[0042] Figure 4 is a structural block diagram of a computer device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the embodiments of the present disclosure will be further described in detail below with reference to the drawings.
[0044] Unless otherwise defined, technical or scientific terms used herein shall have the ordinary meaning as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second", "third" and similar terms used in the specification and claims of this patent application of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, terms such as "a" or "an" do not denote a quantity limitation, but mean that there is at least one. Terms such as "comprising" or "including" mean that the elements or items appearing before "comprising" or "including" cover the elements or items listed after "comprising" or "including" and their equivalents, and do not exclude other elements or items.
[0045] Figure 1 is a flowchart of a method for screening nucleic acid molecular structures provided by an embodiment of the present disclosure. Refer to Figure 1 and the method includes:
[0046] 101: Obtain nucleic acid molecule parameters.
[0047] Among them, the nucleic acid molecule parameters include: nucleic acid molecule sequence, length of the nucleic acid molecule sequence, chemical modification of the nucleic acid molecule, and multiple nucleic acid molecule configurations to be screened and their corresponding energies and experimental data obtained from computational simulations based on quantum mechanics and molecular dynamics.
[0048] Among them, the sequence of the nucleic acid molecule refers to its base sequence. The length of the sequence refers to the number or length of the bases of the nucleic acid molecule. Chemical modification refers to the modification at positions such as the bases, middle, pentose, both ends, and phosphate backbone of the nucleic acid molecule. The data obtained from computational simulations based on quantum mechanics and molecular dynamics can also be called computational physics result data. The nucleic acid molecule configurations (mainly including configuration coordinates and molecular size), energy magnitude, etc. obtained by computational physics simulation of nucleic acid molecules. The experimental data includes at least one of experimental environment, half-life, annealing temperature, and drug enrichment.
[0049] Among the above parameters, multiple nucleic acid molecule configurations to be screened and their corresponding energies obtained from computational simulations based on quantum mechanics and molecular dynamics can be obtained through quantum mechanics and molecular dynamics simulation experiments. The experimental data is obtained through biomolecular experiments and animal experiments on the nucleic acid molecule configurations to be screened. The nucleic acid molecule sequence, the length of the nucleic acid molecule sequence, and the chemical modification of the nucleic acid molecule are the basic parameters of the nucleic acid molecules processed in this application.
[0050] 102: Search for similar nucleic acid molecule structures of each nucleic acid molecule configuration among the multiple nucleic acid molecule configurations to be screened in the nucleic acid library.
[0051] A nucleic acid library refers to a database that stores data on the structures of known nucleic acid molecules, including the configurations, stabilities, and various experimental data of nucleic acid molecules.
[0052] 103: Calculate the probability distributions of the configurations of the multiple nucleic acid molecules to be screened based on the stability of the similar nucleic acid molecule structures, the energies corresponding to the configurations of the multiple nucleic acid molecules to be screened, and the experimental data.
[0053] After searching for similar nucleic acid molecule structures in the nucleic acid library, it is also possible to obtain the stability of the similar nucleic acid molecule structures from the nucleic acid library, and calculate the stability probability of the configurations of the nucleic acid molecules to be screened based on the stability of the similar nucleic acid molecule structures.
[0054] Among them, calculating the probability distributions of the configurations of multiple nucleic acid molecules to be screened refers to calculating the stability probabilities of the configurations of multiple nucleic acid molecules to be screened. Stability means that in a determined environment, the structural configuration of a nucleic acid molecule will not or has a very low probability of changing. Here, the very low probability can refer to a probability less than a certain threshold, such as less than 2% or 1%, etc. The stability probability refers to the probability that the configuration of the nucleic acid molecule to be screened will not change. The higher the stability probability, the better the stability.
[0055] 104: Screen at least one nucleic acid molecule configuration that meets the stability condition among the configurations of the multiple nucleic acid molecules to be screened according to the probability distributions of the configurations of the multiple nucleic acid molecules to be screened.
[0056] Here, meeting the stability condition can mean that the stability probability is higher than a set value, or that the stability probability ranks within a set range among the configurations of the multiple nucleic acid molecules to be screened, such as the top 3, etc.
[0057] In the embodiments of the present disclosure, based on the computational simulations of quantum mechanics and molecular dynamics, similar nucleic acid molecule structures are searched in the nucleic acid library using nucleic acid molecule parameters in multiple dimensions, and probability distributions of multiple nucleic acid molecule configurations are calculated according to the stability of the similar nucleic acid molecule structures and the energies and experimental data corresponding to the configurations of the nucleic acid molecules to be screened, and then at least one nucleic acid molecule configuration that meets the stability condition is screened out. Since the above solution refers to the stability of similar nucleic acid molecule structures and simultaneously considers the energies and experimental data corresponding to the configurations of the nucleic acid molecules to be screened, rather than simply relying on simulation experiments for simulation, the accuracy of nucleic acid molecule structure screening is improved.
[0058] Figure 2 is a flowchart of a method for screening nucleic acid molecule structures provided by the embodiments of the present disclosure. See Figure 2 and the method includes:
[0059] 200: Train a similar nucleic acid molecule search model and a nucleic acid molecule configuration probability distribution calculation model.
[0060] Among them, the similar nucleic acid molecule search model is used to respectively determine the first similarity between the nucleic acid molecule sequence and each nucleic acid molecule sequence in the nucleic acid library, determine the difference between the length of the nucleic acid molecule sequence and the length of each nucleic acid molecule sequence in the nucleic acid library, determine the second similarity between the chemical modification of the nucleic acid molecule and the chemical modification of each nucleic acid molecule sequence in the nucleic acid library, select from the nucleic acid library the nucleic acid molecules for which the first similarity, the difference, and the second similarity all meet the threshold requirements, sequentially determine the distance between the configurations of the selected nucleic acid molecules and the configurations of the multiple nucleic acid molecules to be screened, and obtain the similar nucleic acid molecule structures of each nucleic acid molecule configuration based on the distance.
[0061] The nucleic acid molecule configuration probability distribution calculation model is used to calculate the stability probability of each nucleic acid molecule configuration by calculating the normalized values of the energy of the same nucleic acid molecule configuration, the stability of the similar nucleic acid molecule structure, and the experimental data of the nucleic acid molecule configuration; among them, the calculation process is as follows: form an array with the energy of the nucleic acid molecule configuration, the stability of the similar nucleic acid molecule structure, and the normalized values of the experimental data of the nucleic acid molecule configuration, and use the stability probability algorithm obtained by training to solve the array to obtain the stability probability of each nucleic acid molecule configuration.
[0062] Exemplarily, the scheme for training the similar nucleic acid molecule search model is as follows:
[0063] Obtain a first training set and a first test set based on the nucleic acid library;
[0064] Use the first training set and the first test set to train the similar nucleic acid molecule search model.
[0065] Among them, the data types in the training set and the test set are the same as the data types input into the similar nucleic acid molecule search model.
[0066] The similar nucleic acid molecule search model can calculate the first similarity, difference, second similarity, and distance. The first similarity, difference, second similarity, and distance can calculate the similarity of two nucleic acid molecule configurations according to certain weights. Among them, the calculation formulas for the first similarity and the second similarity can be implemented using mature similarity algorithms. Initially, the parameters in the formula for calculating the first similarity, the parameters in the formula for calculating the second similarity, the parameters in the formula for calculating the distance, and the weights when calculating the similarity of two nucleic acid molecule configurations are all initial values. Similar nucleic acid molecule configurations are pre-annotated in the sample set. During the training process, it is determined whether the output and input of the similar nucleic acid molecule search model are similar configurations according to the annotation, so as to realize the feedback optimization of the above parameters and weights. Through continuous feedback optimization, the training of the similar nucleic acid molecule search model is completed. Of course, in the above process, the annotation of the sample set can also be not done, and it is judged by manpower whether the output and input of the similar nucleic acid molecule search model are similar configurations.
[0067] Similarly, the nucleic acid molecule configuration probability distribution calculation model is also trained in a similar way. For example:
[0068] Obtain a second training set and a second test set based on the nucleic acid library;
[0069] Use the second training set and the second test set to train the nucleic acid molecule configuration probability distribution calculation model.
[0070] Among them, the data types in the training set and the test set are the same as the data types input to the nucleic acid molecule configuration probability distribution calculation model.
[0071] The nucleic acid molecule configuration probability distribution calculation model uses a stability probability algorithm to calculate the energy of the nucleic acid molecule configuration, the stability of similar nucleic acid molecule structures, and the experimental data of the nucleic acid molecule configuration to obtain a stability probability. This stability probability algorithm can be to perform a weighted sum of the normalized values of parameters such as energy, the stability of similar nucleic acid molecule structures, and experimental data. Initially, the parameters in the stability probability algorithm (that is, the weighted weights) are initial values, and the parameters in the stability probability algorithm are adjusted through training, so that the stability probability calculated by this model is more accurate.
[0072] During training, the energy of the nucleic acid molecule configuration, the stability of similar nucleic acid molecule structures, and the normalized values of the experimental data of the nucleic acid molecule configuration are combined into an array, and this array is used as the input value and input into the nucleic acid molecule configuration probability distribution calculation model. According to the stability probability output by the nucleic acid molecule configuration probability distribution calculation model and the actual stability of this nucleic acid molecule configuration, the parameters of the stability probability algorithm are adjusted by feedback, so as to optimize the above model parameters. Through continuous feedback iteration, the training of the nucleic acid molecule configuration probability distribution calculation model is completed.
[0073] Exemplarily, the similar nucleic acid molecule search model and the nucleic acid molecule configuration probability distribution calculation model adopt neural networks, decision trees, support vector machines, where the neural network can be a deep neural network.
[0074] 201: Obtain nucleic acid molecule parameters.
[0075] Among them, the nucleic acid molecule parameters include: nucleic acid molecule sequence, the length of the nucleic acid molecule sequence, chemical modifications of the nucleic acid molecule, and multiple nucleic acid molecule configurations to be screened and their corresponding energies and experimental data obtained from computational simulations based on quantum mechanics and molecular dynamics.
[0076] In the embodiments of the present disclosure, the computational simulation based on quantum mechanics and molecular dynamics refers to predicting the nucleic acid molecule configuration through computational physics software. During the prediction process, the software simulates the physical experiments of the nucleic acid molecule configuration, and determines the energy distribution of each nucleic acid molecule configuration through quantum mechanics and molecular dynamics.
[0077] Among them, quantum mechanics is implemented based on the ab initio method (the most primitive input of the initial configuration coordinates of the molecule), and uses density functional theory or Hartree - Fock theory to solve the Schrödinger equation and obtain the wave function of this system. As the energy converges continuously and the gradient descends, the nucleic acid molecule configuration after energy convergence is found.
[0078] Molecular dynamics is based on Newton's mechanical equations and known force fields, and dynamically solves the velocity and position of each atom at the next moment. The coordinates when the energy is stable or converges can be used as the prediction result of the nucleic acid molecule configuration.
[0079] Exemplarily, the above - mentioned computational physics software includes but is not limited to VASP, Gaussian, GAMESS, NWchem, Q - chem, Spartan, openMM, Amber, NAMD, Lammps, Gromacs, CHARMM, etc.
[0080] 202: Input the nucleic acid molecule sequence, the length of the nucleic acid molecule sequence, the chemical modification of the nucleic acid molecule, and the configurations of multiple nucleic acid molecules to be screened into the similar nucleic acid molecule search model.
[0081] Here, when inputting the above parameters, these parameters will be characterized and converted into a format that is easy for the similar nucleic acid molecule search model to recognize and process. For example, these parameters are numerically valued and normalized, and then input into the similar nucleic acid molecule search model.
[0082] 203: Obtain the similar nucleic acid molecule structures of each nucleic acid molecule configuration among the multiple nucleic acid molecule configurations searched by the similar nucleic acid molecule search model from the nucleic acid library.
[0083] Determine the nucleic acid molecule structures similar to each nucleic acid molecule configuration through the similar nucleic acid molecule search model, so as to prepare for determining the stability of each nucleic acid molecule configuration.
[0084] 204: Input the stability of the similar nucleic acid molecule structures, the energies corresponding to the multiple nucleic acid molecule configurations to be screened, and the experimental data into the nucleic acid molecule configuration probability distribution calculation model.
[0085] Similarly, when inputting the above parameters, these parameters can be characterized and converted into a format that is easy for the nucleic acid molecule configuration probability distribution calculation model to recognize and process. For example, these parameters are numerically valued and normalized, and then input into the nucleic acid molecule configuration probability distribution calculation model.
[0086] That is, the method further includes:
[0087] Clean and complete the experimental data of the nucleic acid molecule;
[0088] Normalize the experimental data of the nucleic acid molecule after cleaning and completion, and numerically value and normalize the non-numerical experimental data.
[0089] Among them, the experimental data of the nucleic acid molecule configuration to be screened refers to the data obtained from experiments on nucleic acid molecules with the nucleic acid molecule configuration to be screened. The experimental data includes the data of biomolecular experiments and animal experiments. The data of the biomolecular experiments and animal experiments includes at least one of the experimental environment, half-life, annealing temperature, and drug enrichment degree.
[0090] Among them, the biomolecular experiment can be a polymerase chain reaction (PCR) test, a mass spectrometry test, an electron microscopy structure analysis test, and a cell test. The animal experiment can be a mouse or monkey test, etc. The experimental environment can be water, physiological saline, plasma, etc.
[0091] 205: Obtain the probability distributions of the configurations of the multiple nucleic acid molecules to be screened calculated by the nucleic acid molecule configuration probability distribution calculation model, where the probability distributions of the configurations of the multiple nucleic acid molecules to be screened include the stability probabilities of the configurations of each nucleic acid molecule.
[0092] 206: Screen at least one nucleic acid molecule configuration with qualified stability from the configurations of the multiple nucleic acid molecules to be screened according to the probability distributions of the configurations of the multiple nucleic acid molecules to be screened.
[0093] Optionally, when screening nucleic acid molecule configurations, in addition to sorting the nucleic acid molecule configurations only according to the stability, the drug properties of the nucleic acid molecule configurations can also be referred to.
[0094] Exemplarily, step 206 includes:
[0095] Sort the configurations of the multiple nucleic acid molecules to be screened according to the stability probabilities of the configurations of the multiple nucleic acid molecules to be screened, and convert the ranking into a first score;
[0096] Sort the configurations of the multiple nucleic acid molecules to be screened according to the drug properties of the configurations of the multiple nucleic acid molecules to be screened, and convert the ranking into a second score;
[0097] Based on the comprehensive consideration of the first score and the second score, screen at least one nucleic acid molecule configuration with a high score from the configurations of the multiple nucleic acid molecules to be screened.
[0098] Among them, when converting the first score and the second score, the higher the ranking, the higher the score. For example, the greater (or better) the stability probability, the higher the ranking, and the corresponding first score is also higher. When comprehensively considering the first score and the second score, a weighted form can be used for calculation.
[0099] Among them, the drug properties are obtained through the aforementioned biomolecular experiments and animal experiments, and can include the aforementioned drug enrichment degree, half-life, and can also include nucleic acid molecule specificity, drug effectiveness, and delivery efficiency, etc. When the drug properties include multiple parameters, they can be sorted separately to obtain multiple second scores, and then combined with the first score.
[0100] It should be noted that the above step 206 can also be implemented by a neural network model, which will not be elaborated here.
[0101] In an embodiment of the present disclosure, based on computational simulations of quantum mechanics and molecular dynamics, nucleic acid molecule parameters in multiple dimensions are used to search for similar nucleic acid molecule structures in a nucleic acid library, and probability distributions of multiple nucleic acid molecule configurations are calculated according to the stability of the similar nucleic acid molecule structures and the energies and experimental data corresponding to the nucleic acid molecule configurations to be screened, so as to screen out at least one nucleic acid molecule configuration with qualified stability. Since the above solution refers to the stability of similar nucleic acid molecule structures and simultaneously considers the energies and experimental data corresponding to the nucleic acid molecule configurations to be screened, the accuracy of nucleic acid molecule structure screening is improved.
[0102] Figure 3 FIG. is a schematic structural diagram of a nucleic acid molecule structure screening device provided by an embodiment of the present disclosure. Refer to Figure 3 The device includes: an acquisition module 301, a search module 302, a calculation module 303, and a screening module 304.
[0103] Among them, the acquisition module 301 is configured to acquire nucleic acid molecule parameters, where the nucleic acid molecule parameters include: nucleic acid molecule sequences, lengths of the nucleic acid molecule sequences, chemical modifications of the nucleic acid molecules, and multiple nucleic acid molecule configurations to be screened and corresponding energies and experimental data obtained based on computational simulations of quantum mechanics and molecular dynamics;
[0104] The search module 302 is configured to search for similar nucleic acid molecule structures of each nucleic acid molecule configuration among the multiple nucleic acid molecule configurations to be screened in the nucleic acid library;
[0105] The calculation module 303 is configured to calculate probability distributions of the multiple nucleic acid molecule configurations to be screened based on the stability of the similar nucleic acid molecule structures, the energies and experimental data corresponding to the multiple nucleic acid molecule configurations to be screened;
[0106] The screening module 304 is configured to screen out at least one nucleic acid molecule configuration with qualified stability among the multiple nucleic acid molecule configurations to be screened according to the probability distributions of the multiple nucleic acid molecule configurations to be screened.
[0107] Exemplarily, the search module 302 is configured to input the nucleic acid molecule sequence, the length of the nucleic acid molecule sequence, the chemical modification of the nucleic acid molecule, and multiple nucleic acid molecule configurations to be screened into a similar nucleic acid molecule search model. The similar nucleic acid molecule search model is configured to respectively determine a first similarity between the nucleic acid molecule sequence and each nucleic acid molecule sequence in the nucleic acid library, determine a difference between the length of the nucleic acid molecule sequence and the length of each nucleic acid molecule sequence in the nucleic acid library, determine a second similarity between the chemical modification of the nucleic acid molecule and the chemical modification of each nucleic acid molecule sequence in the nucleic acid library, select from the nucleic acid library the nucleic acid molecules for which the first similarity, the difference, and the second similarity all meet the threshold requirements, sequentially determine the distances between the configurations of the selected nucleic acid molecules and the multiple nucleic acid molecule configurations to be screened, and obtain a similar nucleic acid molecule structure for each nucleic acid molecule configuration based on the distances.
[0108] Obtain the similar nucleic acid molecule structure for each nucleic acid molecule configuration among the multiple nucleic acid molecule configurations to be screened, which is searched by the similar nucleic acid molecule search model from the nucleic acid library.
[0109] Exemplarily, the calculation module 303 is configured to input the stability of the similar nucleic acid molecule structure, the energies corresponding to the multiple nucleic acid molecule configurations to be screened, and experimental data into a nucleic acid molecule configuration probability distribution calculation model. The nucleic acid molecule configuration probability distribution calculation model is configured to calculate, by performing normalization on the energies of the same nucleic acid molecule configuration, the stability of the similar nucleic acid molecule structure, and the experimental data of the nucleic acid molecule configuration, and obtain the stability probability for each nucleic acid molecule configuration through the calculation; wherein, the calculation process is as follows: form an array with the normalized values of the energy of the nucleic acid molecule configuration, the stability of the similar nucleic acid molecule structure, and the experimental data of the nucleic acid molecule configuration, and use the stability probability algorithm obtained through training to solve the array to obtain the stability probability for each nucleic acid molecule configuration.
[0110] Obtain the probability distributions of the multiple nucleic acid molecule configurations to be screened calculated by the nucleic acid molecule configuration probability distribution calculation model, where the probability distributions of the multiple nucleic acid molecule configurations to be screened include the stability probability for each nucleic acid molecule configuration.
[0111] Optionally, the screening module 304 is configured to sort the multiple nucleic acid molecule configurations to be screened according to the stability probabilities of the multiple nucleic acid molecule configurations and convert the ranking into a first score; sort the multiple nucleic acid molecule configurations to be screened according to the drug properties of the multiple nucleic acid molecule configurations and convert the ranking into a second score; and comprehensively consider the first score and the second score to screen at least one nucleic acid molecule configuration with a high score from the multiple nucleic acid molecule configurations.
[0112] Optionally, the device further includes: a data processing module 305, configured to perform data cleaning and completion on the experimental data of the nucleic acid molecule;
[0113] Normalize the experimental data of the nucleic acid molecule after cleaning and completion, and perform numerical conversion and normalization on non-numerical experimental data.
[0114] Optionally, the experimental data of the nucleic acid molecule configuration to be screened includes data from biomolecular experiments and animal experiments, and the data from the biomolecular experiments and animal experiments includes at least one of experimental environment, half-life, annealing temperature, and drug enrichment.
[0115] It should be noted that: when screening the nucleic acid molecule structure by the nucleic acid molecule structure screening device provided in the above embodiment, only the division of the above functional modules is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the nucleic acid molecule structure screening device provided in the above embodiment and the nucleic acid molecule structure screening method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0116] Figure 4 It is a structural block diagram of a computer device provided by an embodiment of the present disclosure. Generally, a computer device includes: a processor 401 and a memory 402.
[0117] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state.
[0118] The memory 402 may include one or more computer-readable storage media, which may be non-transitory. The memory 402 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 is used to store at least one instruction for being executed by the processor 401 to implement the nucleic acid molecule structure screening method provided in the method embodiments of the present application.
[0119] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a kind of computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a disk or an optical disc, etc.
[0120] The above are only optional embodiments of the present disclosure, and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for screening nucleic acid molecular structures, characterized in that The method includes: Obtaining nucleic acid molecule parameters, where the nucleic acid molecule parameters include: nucleic acid molecule sequence, length of the nucleic acid molecule sequence, chemical modification of the nucleic acid molecule, and multiple nucleic acid molecule configurations to be screened and corresponding energies and experimental data obtained through computational simulations based on quantum mechanics and molecular dynamics, and the experimental data includes at least one of experimental environment, half-life, annealing temperature, and drug enrichment degree; Searching for similar nucleic acid molecule structures of each nucleic acid molecule configuration among the multiple nucleic acid molecule configurations to be screened in a nucleic acid library; Performing probability distribution calculation of the multiple nucleic acid molecule configurations to be screened based on the stability of the similar nucleic acid molecule structures, the energies corresponding to the multiple nucleic acid molecule configurations to be screened, and the experimental data, where the probability distribution calculation of the multiple nucleic acid molecule configurations to be screened refers to calculating the stability probabilities of the multiple nucleic acid molecule configurations to be screened; Screening at least one nucleic acid molecule configuration that meets the stability condition among the multiple nucleic acid molecule configurations to be screened according to the probability distribution of the multiple nucleic acid molecule configurations to be screened.
2. The method according to claim 1, characterized in that The searching for similar nucleic acid molecule structures of each nucleic acid molecule configuration among the multiple nucleic acid molecule configurations to be screened in a nucleic acid library includes: Inputting the nucleic acid molecule sequence, length of the nucleic acid molecule sequence, chemical modification of the nucleic acid molecule, and multiple nucleic acid molecule configurations to be screened into a similar nucleic acid molecule search model, where the similar nucleic acid molecule search model is used to respectively determine the first similarity between the nucleic acid molecule sequence and each nucleic acid molecule sequence in the nucleic acid library, determine the difference between the length of the nucleic acid molecule sequence and the length of each nucleic acid molecule sequence in the nucleic acid library, determine the second similarity between the chemical modification of the nucleic acid molecule and the chemical modification of each nucleic acid molecule sequence in the nucleic acid library, select nucleic acid molecules from the nucleic acid library that meet the threshold requirements for the first similarity, the difference, and the second similarity, sequentially determine the configurations of the selected nucleic acid molecules and the distances between the multiple nucleic acid molecule configurations to be screened, and obtain the similar nucleic acid molecule structures of each nucleic acid molecule configuration based on the distances; Obtaining the similar nucleic acid molecule structures of each nucleic acid molecule configuration among the multiple nucleic acid molecule configurations to be screened searched by the similar nucleic acid molecule search model from the nucleic acid library.
3. The method according to claim 1, characterized in that The performing probability distribution calculation of the multiple nucleic acid molecule configurations to be screened based on the stability of the similar nucleic acid molecule structures, the energies corresponding to the multiple nucleic acid molecule configurations to be screened, and the experimental data includes: Input the stability of the similar nucleic acid molecular structure, the energies corresponding to the configurations of the multiple nucleic acid molecules to be screened, and the experimental data into a nucleic acid molecular configuration probability distribution calculation model. The nucleic acid molecular configuration probability distribution calculation model is used to calculate the normalized values of the energy of the same nucleic acid molecular configuration, the stability of the similar nucleic acid molecular structure, and the experimental data of the nucleic acid molecular configuration, and obtain the stability probability of each nucleic acid molecular configuration through the calculation. Wherein, the calculation process is as follows: form an array with the normalized values of the energy of the nucleic acid molecular configuration, the stability of the similar nucleic acid molecular structure, and the experimental data of the nucleic acid molecular configuration, and use the stability probability algorithm obtained by training to solve the array to obtain the stability probability of each nucleic acid molecular configuration; Obtain the probability distributions of the multiple nucleic acid molecules to be screened calculated by the nucleic acid molecular configuration probability distribution calculation model. The probability distributions of the multiple nucleic acid molecules to be screened include the stability probabilities of each nucleic acid molecular configuration.
4. The method according to any one of claims 1 to 3, characterized in that Screening at least one nucleic acid molecular configuration with qualified stability from the multiple nucleic acid molecules to be screened according to the probability distributions of the multiple nucleic acid molecules to be screened includes: Sort the multiple nucleic acid molecules to be screened according to the stability probabilities of the multiple nucleic acid molecules to be screened, and convert the ranking into a first score; Sort the multiple nucleic acid molecules to be screened according to the drug properties of the multiple nucleic acid molecules to be screened, and convert the ranking into a second score. The drug properties include at least one of drug enrichment, half-life, nucleic acid molecule specificity, drug effectiveness, and delivery efficiency; Integrate the first score and the second score, and screen at least one nucleic acid molecular configuration with a high score from the multiple nucleic acid molecules to be screened.
5. The method according to any one of claims 1 to 3, characterized in that The method further includes: Clean and complete the experimental data of the nucleic acid molecule; Normalize the experimental data of the nucleic acid molecule after cleaning and completion, and numericalize and normalize the non-numerical experimental data.
6. The method according to any one of claims 1 to 3, characterized in that The experimental data of the nucleic acid molecule to be screened includes data from biomolecular experiments and animal experiments. The data from biomolecular experiments and animal experiments includes at least one of experimental environment, half-life, annealing temperature, and drug enrichment.
7. A nucleic acid molecule structure screening device, characterized in that, The device includes: An acquisition module, configured to acquire nucleic acid molecule parameters, where the nucleic acid molecule parameters include: nucleic acid molecule sequence, length of the nucleic acid molecule sequence, chemical modification of the nucleic acid molecule, and multiple nucleic acid molecules to be screened and their corresponding energies and experimental data obtained by computational simulation based on quantum mechanics and molecular dynamics. The experimental data includes at least one of experimental environment, half-life, annealing temperature, and drug enrichment; A search module, configured to search for similar nucleic acid molecular structures of each nucleic acid molecular configuration among the multiple nucleic acid molecules to be screened in a nucleic acid library; A calculation module, configured to calculate the probability distribution of the configurations of the multiple nucleic acid molecules to be screened based on the stability of the similar nucleic acid molecule structures, the energies corresponding to the configurations of the multiple nucleic acid molecules to be screened, and experimental data, where the calculation of the probability distribution of the configurations of the multiple nucleic acid molecules to be screened refers to calculating the stability probability of the configurations of the multiple nucleic acid molecules to be screened; A screening module, configured to screen at least one nucleic acid molecule configuration with qualified stability from the configurations of the multiple nucleic acid molecules to be screened according to the probability distribution of the configurations of the multiple nucleic acid molecules to be screened.
8. The device according to claim 7, characterized in that, The search module is configured to input the nucleic acid molecule sequence, the length of the nucleic acid molecule sequence, the chemical modification of the nucleic acid molecule, and the configurations of the multiple nucleic acid molecules to be screened into a similar nucleic acid molecule search model. The similar nucleic acid molecule search model is used to respectively determine the first similarity between the nucleic acid molecule sequence and each nucleic acid molecule sequence in the nucleic acid library, determine the difference between the length of the nucleic acid molecule sequence and the length of each nucleic acid molecule sequence in the nucleic acid library, determine the second similarity between the chemical modification of the nucleic acid molecule and the chemical modification of each nucleic acid molecule sequence in the nucleic acid library, select nucleic acid molecules from the nucleic acid library that meet the threshold requirements for the first similarity, the difference, and the second similarity, sequentially determine the configurations of the selected nucleic acid molecules and the distances between the configurations of the multiple nucleic acid molecules to be screened, and obtain the similar nucleic acid molecule structures of each nucleic acid molecule configuration based on the distances; Obtain the similar nucleic acid molecule structures of each nucleic acid molecule configuration among the configurations of the multiple nucleic acid molecules to be screened searched by the similar nucleic acid molecule search model from the nucleic acid library; The calculation module is configured to input the stability of the similar nucleic acid molecule structures, the energies corresponding to the configurations of the multiple nucleic acid molecules to be screened, and experimental data into a nucleic acid molecule configuration probability distribution calculation model. The nucleic acid molecule configuration probability distribution calculation model is used to calculate by normalizing the numerical values of the energy of the same nucleic acid molecule configuration, the stability of the similar nucleic acid molecule structures, and the experimental data of the nucleic acid molecule configuration, and obtain the stability probability of each nucleic acid molecule configuration through the calculation; where the calculation process is as follows: form an array with the normalized numerical values of the energy of the nucleic acid molecule configuration, the stability of the similar nucleic acid molecule structures, and the experimental data of the nucleic acid molecule configuration, and use the stability probability algorithm obtained by training to solve the array to obtain the stability probability of each nucleic acid molecule configuration; Obtain the probability distribution of the configurations of the multiple nucleic acid molecules to be screened calculated by the nucleic acid molecule configuration probability distribution calculation model, where the probability distribution of the configurations of the multiple nucleic acid molecules to be screened includes the stability probability of each nucleic acid molecule configuration.
9. A computer device, characterized in that, The computer device includes a processor and a memory. The memory stores at least one piece of program code, and the program code is loaded and executed by the processor to implement the nucleic acid molecule structure screening method according to any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that, At least one program code is stored in the computer-readable storage medium, and the program code is loaded and executed by a processor to implement the nucleic acid molecule structure screening method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Computational method for designing chemical structures having common functional characteristics
CA2248426A1
Micromolecule target nucleic acid adapter computer-aided screening method based on high-performance computing platform, and micromolecule target nucleic acid adapter
CN110246538A