Method, apparatus and device for generating a prediction model of ionizable lipid molecules
By obtaining the initial characteristics of ionizable lipid molecules and training sets to generate prediction models, the problems of long cycles and high cost in the design and optimization of ionizable lipid molecules are solved, and the rapid prediction of mRNA-LNPs transfection effect is achieved, and the design screening efficiency and prediction accuracy are improved.
Patent Information
- Application Number
- CN202510268720.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-07
AI Technical Summary
In the prior art, the design and optimization of ionizable lipid molecules mainly rely on experimental screening, with long cycles and high costs, making it difficult to efficiently develop carriers that are suitable for specific mRNA nanodrugs. How to quickly predict the transfection effect of mRNA-LNPs after ionizable lipid molecules has become an urgent problem to be solved.
By obtaining the initial characteristics of ionizable lipid molecules, determining the target characteristics of high-dimensional combinations, building a prediction model, and using the training set to train the model to generate a target prediction model, which is used to predict the transfection effect of ionizable lipid molecules on the construction of lipid nanoparticles for mRNA-mounted.
The design screening efficiency of ionizable lipid molecules has been improved, the R&D speed has been accelerated, the accuracy and stability of the prediction model have been improved, and an efficient tool for the design optimization of ionizable lipid molecules has been provided.
Smart Images

Figure CN119785927B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of machine learning, and specifically relates to a method, device, and equipment for generating a prediction model of ionizable lipid molecules. Background Art
[0002] In recent years, mRNA nanomedicines have become a research hotspot of emerging gene drugs due to their rapid R & D, flexible design, and high therapeutic potential. Compared with traditional small molecule drugs or protein drugs, mRNA nanomedicines have programmability and broad adaptability, and can precisely treat a variety of diseases. However, the effective delivery of mRNA nanomedicines remains one of the key bottlenecks in their clinical applications. Lipid nanoparticles (LNPs), as the main carriers for mRNA delivery, have been widely used in the preparation of mRNA nanomedicines due to their excellent biocompatibility, stability, and targeting properties. Ionizable lipid molecules play a core role in the construction of LNPs, and their physicochemical properties directly affect the encapsulation efficiency, stability, and targeted delivery effect of mRNA. Therefore, optimizing the structure of ionizable lipid molecules to improve the delivery efficiency and therapeutic effect of mRNA nanomedicines has become an important research direction.
[0003] Currently, the design and optimization of ionizable lipid molecules mainly rely on experimental screening, which has disadvantages such as long cycle and high cost, and it is difficult to efficiently develop carriers that meet the requirements of specific mRNA nanomedicines. Therefore, how to quickly predict the transfection effect of mRNA-LNPs after the construction of ionizable lipid molecules has become an important problem to be solved urgently. Summary of the Invention
[0004] This application provides a method, device, and equipment for generating a prediction model of ionizable lipid molecules, so as to realize the prediction of the transfection effect of ionizable lipid molecules on the construction of lipid nanoparticles for carrying mRNA, improve the design and screening efficiency of ionizable lipid molecules, and accelerate the R & D speed.
[0005] This application provides a method for generating a prediction model of ionizable lipid molecules, including:
[0006] Obtain the initial features of ionizable lipid molecules;
[0007] Determine the target features according to the initial features, where the target features are high-dimensional combinations of the initial features;
[0008] Construct a prediction model according to the target features, where the prediction model is used to predict the transfection effect of the ionizable lipid molecules on the construction of lipid nanoparticles for carrying mRNA;
[0009] Obtain a training set;
[0010] Train the constructed prediction model according to the training set to obtain a target prediction model.
[0011] According to the method for generating an ionizable lipid molecule prediction model provided by the present application, the initial features include a plurality of chemical structure features and a plurality of spatial features. Determining the target features according to the initial features includes: combining the initial features through mathematical operations to obtain a plurality of candidate features; recursively selecting features from the candidate features whose correlation with the transfection effect is higher than a preset value; combining the features whose correlation is higher than the preset value to obtain a plurality of feature combinations; determining the candidate features included in the feature combination with the best prediction effect among the plurality of feature combinations as the target features.
[0012] According to the method for generating an ionizable lipid molecule prediction model provided by the present application, training the constructed prediction model according to the training set to obtain a target prediction model includes: randomly selecting a plurality of experimental data from the training set as sub-training sets to obtain a plurality of sub-training sets; respectively training the constructed prediction model according to the plurality of sub-training sets to obtain a plurality of trained prediction models; obtaining a test set, where the test set includes a plurality of experimental data; obtaining the actual transfection effect of each experimental data included in the test set; obtaining the predicted transfection effect of each experimental data predicted by the trained prediction model; and selecting a target prediction model from the plurality of trained prediction models according to the actual transfection effect and the predicted transfection effect.
[0013] According to the method for generating an ionizable lipid molecule prediction model provided by the present application, selecting a target prediction model from the plurality of trained prediction models according to the actual transfection effect and the predicted transfection effect includes: determining the prediction type of each trained prediction model for each experimental data according to the actual transfection effect and the predicted transfection effect; determining the accuracy, precision, and recall rate of each trained prediction model according to the prediction type; scoring each trained prediction model according to the accuracy and the recall rate to obtain a target score; and selecting a target prediction model from the plurality of trained prediction models according to the accuracy, the precision, the recall rate, and the target score.
[0014] According to the method for generating an ionizable lipid molecule prediction model provided by the present application, there are four types of prediction types. Among them, the first prediction type is used to indicate that the actual transfection effect is the first effect and the predicted transfection effect is the first effect. The second prediction type is used to indicate that the actual transfection effect is the first effect and the predicted transfection effect is the second effect. The third prediction type is used to indicate that the actual transfection effect is the second effect and the predicted transfection effect is the first effect. The fourth prediction type is used to indicate that the actual transfection effect is the second effect and the predicted transfection effect is the second effect. The first effect is better than the second effect. Determining the accuracy rate, precision rate, and recall rate of each trained prediction model according to the prediction type of each experimental data includes: performing the following operations for each trained prediction model: determining the accuracy rate according to the number of the third prediction type and the fourth prediction type corresponding to the current trained prediction model, and the total number of the prediction types corresponding to the current trained prediction model; determining the precision rate according to the number of the first prediction type corresponding to the current trained prediction model, and the total number of the first prediction type and the third prediction type corresponding to the current trained prediction model; determining the recall rate according to the number of the first prediction type corresponding to the current trained prediction model, and the total number of the first prediction type and the second prediction type corresponding to the current trained prediction model.
[0015] According to the method for generating an ionizable lipid molecule prediction model provided by the present application, the initial features include spatial features. Obtaining the initial features of the ionizable lipid molecule includes: obtaining multiple molecular conformations of the ionizable lipid molecule; respectively performing alignment processing on the molecular structures of the multiple molecular conformations; performing density projection on the aligned multiple molecular structures to obtain a density map; and extracting the spatial features of the ionizable lipid molecule according to the density map.
[0016] According to the method for generating an ionizable lipid molecule prediction model provided by the present application, the respectively performing alignment processing on the molecular structures of the multiple molecular conformations includes: performing the following operations on the molecular structure of each molecular conformation among the multiple molecular conformations: obtaining a target central atom and two reference atoms from the head group of the current molecular structure; fixing the target central atom at a target coordinate; obtaining a first vector according to the target central atom and the molecular centroid of the current molecular structure; obtaining a second vector according to the two reference atoms; and rotating the current molecular structure so that the first vector is parallel to the first coordinate axis, and the second vector is parallel to the plane corresponding to the first coordinate axis and the second coordinate axis.
[0017] According to the method for generating an ionizable lipid molecule prediction model provided by the present application, the extracting of the spatial features of the ionizable lipid molecule from the density map includes: obtaining a first boundary and a second boundary from the density map, where the density value corresponding to the first boundary belongs to a first threshold, and the density value corresponding to the second boundary belongs to a second threshold; extracting the spatial features of the ionizable lipid molecule according to the first boundary and the second boundary, and the spatial features include at least one of the following: molecular length, molecular width, aspect ratio, cotangent value of the cone angle of the molecular head group, and symmetry of the molecular head group.
[0018] According to the method for generating an ionizable lipid molecule prediction model provided by the present application, after training the constructed prediction model with the training set to obtain a target prediction model, the method further includes: obtaining target spatial features and target chemical structure features of a target ionizable lipid molecule to be predicted; inputting the target spatial features and the target chemical structure features into the target prediction model to obtain a prediction result, and the prediction result is used to indicate the transfection effect of a lipid nanoparticle constructed based on the target ionizable lipid molecule for carrying mRNA.
[0019] The present application also provides an ionizable lipid molecule prediction model generation device, including: a first acquisition unit, configured to acquire initial features of an ionizable lipid molecule; a determination unit, configured to determine target features according to the initial features, and the target features are high-dimensional combinations of the initial features; a construction unit, configured to construct a prediction model according to the target features, and the prediction model is used to predict the transfection effect of the ionizable lipid molecule on a lipid nanoparticle constructed for carrying mRNA; a second acquisition unit, configured to acquire a training set; and a training unit, configured to train the constructed prediction model with the training set to obtain a target prediction model.
[0020] The present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the method for generating an ionizable lipid molecule prediction model as described in any one of the above.
[0021] The present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for generating an ionizable lipid molecule prediction model as described in any one of the above.
[0022] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method for generating an ionizable lipid molecule prediction model as described in any one of the above.
[0023] The method for generating a prediction model of ionizable lipid molecules provided by this application first obtains the initial features of ionizable lipid molecules, then determines the target features based on the initial features, where the target features are high-dimensional combinations of the initial features. Then, a prediction model is constructed based on the target features, and the prediction model is used to predict the transfection effect of the ionizable lipid molecules on constructing lipid nanoparticles for carrying mRNA. Then, a training set is obtained, and finally, the constructed prediction model is trained according to the training set to obtain the target prediction model. This can realize the prediction of the transfection effect of ionizable lipid molecules on constructing lipid nanoparticles for carrying mRNA, improve the design and screening efficiency of ionizable lipid molecules, and accelerate the R & D speed. Brief Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0025] Figure 1 It is a schematic flowchart of a method for generating a prediction model of ionizable lipid molecules provided by this application.
[0026] Figure 2 It is a density projection process diagram provided by this application.
[0027] Figure 3 It is a schematic diagram of the molecular structure provided by this application.
[0028] Figure 4 It is a schematic diagram of spatial feature extraction provided by this application.
[0029] Figure 5 It is a schematic diagram of model verification provided by this application.
[0030] Figure 6 It is a schematic diagram of the structure of a device for generating a prediction model of ionizable lipid molecules provided by this application.
[0031] Figure 7 It is a schematic diagram of the structure of an electronic device provided by this application. Detailed Description of the Embodiments
[0032] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0033] The terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0034] Referring to "embodiments" herein means that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0035] Currently, the design and optimization of ionizable lipid molecules mainly rely on experimental screening, which has disadvantages such as long cycle and high cost, and it is difficult to efficiently develop carriers that meet the requirements of specific mRNA nanodrugs. Therefore, how to quickly predict the transfection effect of mRNA-LNPs after constructing ionizable lipid molecules has become an important problem to be solved urgently.
[0036] In view of the above problems, the embodiments of this application provide a method, device, and equipment for generating a prediction model of ionizable lipid molecules. The embodiments of this application will be introduced in detail below with reference to the accompanying drawings.
[0037] Please refer to Figure 1 , Figure 1 The flowchart of a method for generating a prediction model of ionizable lipid molecules provided by this application. The method for generating the prediction model includes the following steps.
[0038] S101, obtain the initial features of the ionizable lipid molecules.
[0039] Among them, the initial features may include chemical structure features and spatial features. For example, the chemical structure features include the total number of carbon atoms in the ionizable lipid molecular head group (the head group indicates the hydrophilic part surrounded by the linking group), the total number of nitrogen and oxygen atoms in the ionizable lipid molecule, the total number of functional groups modifying the head group (such as -CH2-, -OH, -CH3, each functional group is counted as one unit), the type of head group modification (if it contains a hydrophilic group, the value is 1, otherwise it is 0), and the total number of tails in the ionizable lipid molecule. In a specific implementation, the spatial features may include molecular length, molecular width, aspect ratio, cotangent value of the cone angle of the molecular head group, symmetry of the molecular head group, etc.
[0040] It should be noted that the molecules in this solution all refer to ionizable lipid molecules used to construct lipid nanoparticles for carrying mRNA.
[0041] S102. Determine target features according to the initial features, where the target features are high-dimensional combinations of the initial features.
[0042] Among them, there may be multiple target features, and each target feature may be a mathematical operation combination of multiple sample initial features.
[0043] S103. Construct a prediction model according to the target features, where the prediction model is used to predict the transfection effect of the ionizable lipid molecule on constructing lipid nanoparticles for carrying mRNA.
[0044] Among them, after obtaining the target features, the linear fitting coefficient corresponding to each target feature can also be obtained. In particular, a sparse linear model can be constructed by the Sparse Optimization (SO) method, and the form of the finally constructed prediction model can be as follows:
[0045] P = c1D1 + c2D2 + c3D3 +... + c n D n + c0
[0046] Among them, c1, c2, c3, c n represent linear fitting coefficients, c0 is a constant, D1, D2, D3, D n are target features, and P is the sample band property. In particular, n can be 5, that is, the constructed prediction model is constructed from 5 target features.
[0047] S104. Obtain a training set.
[0048] Among them, the training set includes multiple experimental data, and each experimental data includes the sample initial features and sample band properties of an ionizable lipid molecule. The sample band properties are associated with the actual transfection effect of the ionizable lipid molecule for constructing lipid nanoparticles for carrying mRNA. In a specific implementation, the sample band property P = log10(y true ), where y true is the actual transfection effect. The sample initial features are the above-mentioned initial features, including spatial features and chemical structure features.
[0049] S105, training the constructed prediction model according to the training set to obtain a target prediction model.
[0050] It can be seen that in the embodiments of the present application, first, the initial features of the ionizable lipid molecule are obtained, then the target features are determined according to the initial features, the target features are high-dimensional combinations of the initial features, and then a prediction model is constructed according to the target features. The prediction model is used to predict the transfection effect of the ionizable lipid molecule on constructing lipid nanoparticles for carrying mRNA. Then, a training set is obtained, and finally, the constructed prediction model is trained according to the training set to obtain a target prediction model. In this way, the prediction of the transfection effect of the ionizable lipid molecule on constructing lipid nanoparticles for carrying mRNA can be realized, the design and screening efficiency of the ionizable lipid molecule can be improved, and the R & D speed can be accelerated.
[0051] In a possible embodiment, the initial features include multiple chemical structure features and multiple spatial features. Determining the target features according to the initial features includes: combining the initial features through mathematical operations to obtain multiple candidate features; combining the features with a correlation higher than a preset value to obtain multiple feature combinations; and determining the candidate features included in the feature combination with the best corresponding prediction effect among the multiple feature combinations as the target features.
[0052] Among them, the features with a transfection effect correlation higher than the preset value may refer to the features with the highest transfection effect correlation, and there may be multiple such features. The target features are those with the best combination effect among the candidate features with the highest transfection effect correlation.
[0053] In a specific implementation, the initial features can be combined through mathematical operations (such as +, –, ×, ÷, log, exp, etc.) to generate high-dimensional candidate features. Then, the Sure Independence Screening (SIS) method is used to recursively select the features with a correlation higher than the preset value with the transfection effect or with the above-mentioned P. In particular, the features with a correlation higher than the preset value may refer to the candidate features with the highest correlation.
[0054] For example, the obtained target features include D1 = exp(-|A y1i - A y2i | / R y ), D2 = (L i - W xi ) / A x1 , D3 = (|A y1i - A y2i | × R xi ), 2 , D4 = |N N - N O | / R x , D5 = N head / N c-head × A y1i , where A y1i is a cone angle of the molecular head group under the YZ - plane projection at the second boundary, A y2i is another cone angle of the molecular head group under the YZ - plane projection at the second boundary, R y is the ratio of the molecular length to the width corresponding to the first boundary under the YZ - plane projection, L i is the molecular length, W xi is the length separated by the second boundary under the XZ - plane projection, A x1 is a cone angle of the molecular head group under the XZ - plane projection at the first boundary, R xi is the ratio of the molecular length to the width corresponding to the second boundary under the XZ - plane projection, N N is the total number of nitrogen atoms in the molecule, N O is the total number of oxygen atoms in the molecule, R x is the ratio of the molecular length to the width corresponding to the first boundary under the XZ - plane projection, N head is the total number of functional groups of the modified head group, N c-head is the total number of carbon atoms in the molecular head group. That is, the subscript x represents the features obtained from the XZ - plane projection, and the subscript y represents the features obtained from the YZ - plane projection. The subscript i represents the features extracted according to the second boundary (the inner contour S i ) of the molecule, and the features without the subscript i are extracted according to the first boundary (the outer contour S) of the molecule.
[0055] It can be seen that in this embodiment, first, high - dimensional candidate features are generated based on the initial features, and then the target features are selected based on the correlation with the variables, which can improve the prediction accuracy of the constructed prediction model.
[0056] In a possible embodiment, training the constructed prediction model according to the training set to obtain a target prediction model includes: selecting a plurality of experimental data from the training set as sub-training sets to obtain a plurality of sub-training sets; respectively training the constructed prediction model according to the plurality of sub-training sets to obtain a plurality of trained prediction models; obtaining a test set, where the test set includes a plurality of experimental data; obtaining the actual transfection effect of each experimental data included in the test set; obtaining the predicted transfection effect of each experimental data predicted by the trained prediction model; and selecting a target prediction model from the plurality of trained prediction models according to the actual transfection effect and the predicted transfection effect.
[0057] Among them, the training data in the sub-training set can be randomly selected from the training set, and the training data included in each sub-training can be 90% of the training data included in the training set. That is, 90% of the training data can be randomly selected from the training set as the sub-training set. Each sub-training set can train a prediction model, that is, randomly select a plurality of molecules from the training set as the sub-training set for training, and each different selection can obtain a prediction model.
[0058] It can be seen that in this embodiment, training a plurality of prediction models based on different sub-training sets, and then selecting a target training model from the plurality of trained prediction models based on the transfection effect can enhance the prediction accuracy of the finally obtained prediction model.
[0059] In a possible embodiment, selecting a target prediction model from the plurality of trained prediction models according to the actual transfection effect and the predicted transfection effect includes: determining the prediction type of each trained prediction model for each experimental data according to the actual transfection effect and the predicted transfection effect; determining the accuracy rate, precision rate and recall rate of each trained prediction model according to the prediction type; scoring each trained prediction model according to the accuracy rate and the recall rate to obtain a target score; and selecting a target prediction model from the plurality of trained prediction models according to the accuracy rate, the precision rate, the recall rate and the target score.
[0060] Among them, when obtaining the target score, the target product of the precision rate and the recall rate corresponding to the currently trained prediction model can be obtained first, and then the target sum of the precision rate and the recall rate can be obtained, and the target score is obtained based on the target product and the target sum. The target score (F1) can be calculated by the following formula:
[0061] F1=(2×Precision×Recall) / (Precision+Recall),
[0062] Among them, Precision is the precision rate of the current trained prediction model, and Recall is the recall rate of the current trained prediction model.
[0063] It can be seen that in this embodiment, screening the target prediction model based on the accuracy rate, precision rate, recall rate, and target score can ensure the prediction accuracy of the obtained target prediction model after screening.
[0064] In a possible embodiment, there are four types of prediction types. The first prediction type is used to indicate that the actual transfection effect is the first effect and the predicted transfection effect is the first effect. The second prediction type is used to indicate that the actual transfection effect is the first effect and the predicted transfection effect is the second effect. The third prediction type is used to indicate that the actual transfection effect is the second effect and the predicted transfection effect is the first effect. The fourth prediction type is used to indicate that the actual transfection effect is the second effect and the predicted transfection effect is the second effect, and the first effect is better than the second effect. Determining the accuracy rate, precision rate, and recall rate of each trained prediction model according to the prediction type of each experimental data includes: performing the following operations for each trained prediction model: determining the accuracy rate according to the number of the first prediction type and the fourth prediction type corresponding to the current trained prediction model, and the total number of prediction types corresponding to the current trained prediction model; determining the precision rate according to the number of the first prediction type corresponding to the current trained prediction model, and the total number of the first prediction type and the third prediction type corresponding to the current trained prediction model; determining the recall rate according to the number of the first prediction type corresponding to the current trained prediction model, and the total number of the first prediction type and the second prediction type corresponding to the current trained prediction model.
[0065] Among them, the transfection effect can be divided into two effects, the first effect and the second effect, according to the value corresponding to the transfection effect. That is, if the transfection efficiency is higher than the threshold, it is considered that the transfection efficiency is the first effect, otherwise it is the second effect. The threshold can be 10 5 , that is, if y pre > 10 5 is the first effect. For example, the first effect is used to indicate high efficiency, and the second effect is used to indicate low efficiency.
[0066] The accuracy rate (Accuracy), precision rate (Precision), and recall rate (Recall) can be calculated by the following formulas respectively:
[0067] Accuracy rate: Accuracy = (TP + TN) / (TP + FN + FP + TN),
[0068] Precision rate: Precision = TP / (TP + FP),
[0069] Recall rate: Recall = TP / (TP+FN),
[0070] Among them, TP (True Positive) indicates the number of the first prediction types included in the current training model, FN (False Negative) indicates the number of the second prediction types included in the current training model, FP (False Positive) indicates the number of the third prediction types included in the current training model, and TN (True Negative) indicates the number of the fourth prediction types included in the current training model. That is, if the experimental data of 90 ionizable lipid molecules are included in the sub-training set corresponding to the current training model, where the total number of the first type of prediction is 50 and the total number of the fourth type of prediction is 35, then the corresponding accuracy rate is 94%.
[0071] It can be seen that in this embodiment, determining the accuracy rate, precision rate, and recall rate based on the number of prediction types can ensure the prediction accuracy of the obtained target prediction model after screening.
[0072] In a possible embodiment, the initial features include spatial features, and obtaining the initial features of the ionizable lipid molecules includes: obtaining multiple molecular conformations of the ionizable lipid molecules; respectively performing alignment processing on the molecular structures of the multiple molecular conformations; performing density projection on the aligned multiple molecular structures to obtain a density map; and extracting the spatial features of the ionizable lipid molecules according to the density map.
[0073] Among them, the molecular conformation can be obtained by molecular dynamics simulation. Specifically, it includes first constructing a system containing a single ionizable lipid molecule and a solvent in a simulation box. The solvent can be various common solvents, such as configured with the volume ratio of water and ethanol (EtOH). In particular, pure water solvent is used in this embodiment. Then, considering the ionization situation of the ionizable lipid molecule, the influence of the charged situation on the three-dimensional structure of the molecule is investigated. At this time, to ensure the electrical neutrality of the system, Na⁺ and Cl⁻ ions are introduced. Non-ionizable molecules are used in this embodiment. Finally, the classical force field charmm36 can be used to describe the molecule, and the tip3p water model. Other force fields can also be used for simulation. In this embodiment, 100 ns is simulated, and the trajectory data from 20 ns to 100 ns is selected for analysis. The molecular structure is extracted every 0.04 ns, and a total of 2000 molecular conformations are obtained.
[0074] In a specific implementation, when performing density projection, density projections can be respectively carried out on the XZ and YZ planes to obtain two density maps. The density map on the XZ plane represents the side information of the molecule, and the density map on the YZ plane represents the front information of the molecule. When performing density projection, it specifically includes superimposing and density averaging multiple target molecular structures to obtain an initial density map; performing normalization processing on the initial density map to obtain the density map, and the density value of the central nitrogen atom included in the density map is 1. For example Figure 2 as shown Figure 2 is a density projection process diagram provided by the present application, that is, multiple single-molecule structures are subjected to structure superposition and density averaging, and then two-dimensional density projection maps on the XZ and YZ planes are obtained.
[0075] It can be seen that in this embodiment, after aligning the molecular structures, extracting spatial features based on the obtained density map can improve the accuracy of the extracted spatial features.
[0076] In a possible embodiment, the aligning the molecular structures of the multiple molecular conformations respectively includes: performing the following operations on the molecular structure of each molecular conformation among the multiple molecular conformations: obtaining a target central atom and two reference atoms from the head group of the current molecular structure; fixing the target central atom at a target coordinate; obtaining a first vector according to the target central atom and the molecular centroid of the current molecular structure; obtaining a second vector according to the two reference atoms; rotating the current molecular structure so that the first vector is parallel to the first coordinate axis, and the second vector is parallel to the plane corresponding to the first coordinate axis and the second coordinate axis.
[0077] Wherein, the target central atom is a central nitrogen atom; the reference atoms are nitrogen atoms or carbon atoms connected to nitrogen atoms.
[0078] In a specific implementation, as Figure 3 shown, first, perform molecular dynamics simulation for 100 ns, select the trajectory data from 20 ns to 100 ns for analysis, extract the molecular structure every 0.04 ns, and a total of 2000 ionizable lipid molecular conformations are obtained. Then when aligning the molecular structures, the central nitrogen atom N center of the ionizable lipid molecular head group can be fixed at a specific position in the coordinate system first, and then the vector n1 is defined from N center to the molecular centroid Q, and the molecule is rotated so that n1 is parallel to the Z axis and in the same direction as the positive direction of the Z axis. The vector n2 defined by two other nitrogen atoms N1 and N2 is parallel to the YZ plane. In particular, if the number of nitrogen atoms in the head group is less than 3, two carbon atoms connected to the nitrogen atom are used for substitution.
[0079] It can be seen that in this embodiment, by aligning the atoms in the head group with the molecular centroid, the alignment processing efficiency can be improved.
[0080] In a possible embodiment, the extracting the spatial features of the ionizable lipid molecule according to the density map includes: obtaining a first boundary and a second boundary according to the density map, the density value corresponding to the first boundary belonging to a first threshold, the density value corresponding to the second boundary belonging to a second threshold, and the first threshold being less than the second threshold; extracting the spatial features of the ionizable lipid molecule according to the first boundary and the second boundary, where the spatial features include at least one of the following: molecular length, molecular width, aspect ratio, cotangent value of the cone angle of the molecular head group, and symmetry of the molecular head group.
[0081] Among them, a density matrix can be obtained based on the density map first, and then the first boundary and the second boundary can be determined based on the density matrix. The first boundary S can be the boundary formed by the grid points with density values of 0.1 to 0.15 in the density matrix, which is the outer contour of the molecule. The second boundary S i can be the boundary formed by the density values of 0.4 to 0.45 in the density matrix, which is the inner contour of the molecule.
[0082] In specific implementation, when performing feature extraction, the spatial features can be extracted from the density maps of the XZ and YZ planes respectively. Now, please refer to Figure 4 for an example of the specific method of extracting spatial features from the density map of the YZ plane. The molecular length includes the line segment length L separated by the boundary S of the first preset line in the density map and the length L separated by the boundary S i separated by the boundary S i , and the first preset line can be Y = 3nm. Among them, L is completely equivalent on the XZ and YZ plane projection maps.
[0083] Taking the YZ plane projection as an example, the molecular width includes the line segment length W separated by the boundary S of the second preset line in the density map y and the length W separated by the boundary S i separated by the boundary S yi , and the second preset line can be Z = 2.25nm. Then the corresponding aspect ratios R and R i can be obtained based on the molecular length and width, and R y = L / W y , and R yi = L i / W yi .
[0084] Taking the YZ plane projection as an example, when obtaining the cotangent value of the cone angle, Figure 4 in (a) is used to indicate obtaining the cotangent value of the cone angle with the boundary S. That is, a cotangent value A of the cone angle in the YZ plane y1 refers to the cotangent value of ∠O - H - Q1, that is, Ay1 = cot(∠O-H-Q1), where H can be the intersection point of the first preset line and the boundary S, and Q1 and Q2 are the intersection points of the third preset line (e.g., the line Y = 2.76 nm) and the fourth preset line (e.g., the line Y = 3.24 nm) with the boundary S respectively. Another cotangent value A of the cone angle of this head group y2 refers to the cotangent value of ∠O-H-Q2, that is, A y2 = cot(∠O-H-Q2). Similarly, Figure 4 in (b) is used to indicate obtaining the cotangent value of the cone angle related to the boundary S i That is, a cotangent value A of the cone angle in the Y plane y1i refers to the cotangent value of ∠O-H-Q1 i that is, A y1i = cot(∠O-H-Q1 i ), where H can be the intersection point of the first preset line and the boundary S i and Q1 and Q2 are the intersection points of the third preset line (e.g., the line Y = 2.76 nm) and the fourth preset line (e.g., the line Y = 3.24 nm) with the boundary S i respectively. Another cotangent value A of the cone angle of this head group y2i refers to the cotangent value of ∠O-H-Q2 i that is, A y2i = cot(∠O-H-Q2 i ). In particular, if A1 > A2, then similarly A 1i > A 2i . The symmetry of this molecular head group is ∣A1 - A2∣ and ∣A 1i - A 2i ∣.
[0085] It can be seen that in this embodiment, extracting spatial features based on the obtained density map can improve the accuracy of the extracted spatial features.
[0086] In a possible embodiment, after training the constructed prediction model according to the training set to obtain the target prediction model, the method further includes: obtaining the target spatial features and target chemical structure features of the target ionizable lipid molecule to be predicted; inputting the target spatial features and the target chemical structure features into the target prediction model to obtain a prediction result, and the prediction result is used to indicate the transfection effect of the lipid nanoparticle constructed based on the target ionizable lipid molecule for carrying mRNA.
[0087] Wherein, the target spatial features and target chemical structure features can be the spatial features and chemical structure features described above.
[0088] Next, please refer to Figure 5 to verify the target prediction model constructed in this application.
[0089] Six ionizable lipids with different structures were randomly selected, and the information on their chemical structures (two-dimensional structures) was input into the prediction model. As Figure 5 shown, the output results indicate that the ionizable lipids P1-P5 perform as "highly efficient" molecules, predicting that they are expected to be candidate lipid molecules for mRNA delivery. In contrast, the performance of P6 is classified as an "inefficient" molecule. To verify these predictions, these six ionizable lipid molecules were synthesized and purified, and LNPs loaded with Luciferase mRNA were prepared using them. As Figure 5 The in vitro and in vivo test results show that their mRNA delivery efficiency is consistent with the prediction of the machine learning model. The mRNA transfection efficiency of lipids P1 to P4 is high, while the transfection efficiency of lipid P6 is low.
[0090] It can be seen that the above model significantly improves the stability and accuracy of prediction, providing an efficient tool for the design and optimization of ionizable lipid molecules. When using this model subsequently, only the spatial characteristics and chemical structure characteristics of the ionizable lipid molecules whose transfection effects need to be predicted are required. By inputting the relevant characteristics into the final model, the predicted transfection effect can be calculated and it can be determined whether it is a "highly efficient" molecule.
[0091] Next, a prediction model generation device provided by the present application will be described. The prediction model generation device described below corresponds to the prediction model generation method described above with reference to each other.
[0092] Please refer to Figure 6 , the ionizable lipid molecule prediction model generation device 600 includes: a first acquisition unit 601 for acquiring the initial characteristics of the ionizable lipid molecule; a determination unit 602 for determining target characteristics according to the initial characteristics, where the target characteristics are high-dimensional combinations of the initial characteristics; a construction unit 603 for constructing a prediction model according to the target characteristics, where the prediction model is used to predict the transfection effect of the ionizable lipid molecule on constructing lipid nanoparticles for carrying mRNA; a second acquisition unit 604 for acquiring a training set; and a training unit 605 for training the constructed prediction model according to the training set to obtain a target prediction model.
[0093] In a possible embodiment, the initial features include a plurality of chemical structure features and a plurality of spatial features. In terms of determining the target features according to the initial features, the determining unit 602 is specifically configured to: combine the initial features through mathematical operations to obtain a plurality of candidate features; recursively select features from the candidate features whose correlation with the transfection effect is higher than a preset value; combine the features whose correlation is higher than the preset value to obtain a plurality of feature combinations; and determine the candidate features included in the feature combination with the best prediction effect among the plurality of feature combinations as the target features.
[0094] In a possible embodiment, in terms of training the constructed prediction model according to the training set to obtain a target prediction model, the training unit 605 is specifically configured to: select a plurality of experimental data from the training set as sub-training sets to obtain a plurality of sub-training sets; train the constructed prediction model according to the plurality of sub-training sets respectively to obtain a plurality of trained prediction models; obtain a test set, where the test set includes a plurality of experimental data; obtain the actual transfection effect of each experimental data included in the test set; obtain the predicted transfection effect of each experimental data predicted by the trained prediction model; and select a target prediction model from the plurality of trained prediction models according to the actual transfection effect and the predicted transfection effect.
[0095] In a possible embodiment, in terms of selecting a target prediction model from the plurality of trained prediction models according to the actual transfection effect and the predicted transfection effect, the training unit 605 is specifically configured to: determine the prediction type of each trained prediction model for each experimental data according to the actual transfection effect and the predicted transfection effect; determine the accuracy rate, precision rate and recall rate of each trained prediction model according to the prediction type; score each trained prediction model according to the accuracy rate and the recall rate to obtain a target score; and select a target prediction model from the plurality of trained prediction models according to the accuracy rate, the precision rate, the recall rate and the target score.
[0096] In a possible embodiment, there are four types of the prediction types, wherein the first prediction type is used to indicate that the actual transfection effect is the first effect and the predicted transfection effect is the first effect, the second prediction type is used to indicate that the actual transfection effect is the first effect and the predicted transfection effect is the second effect, the third prediction type is used to indicate that the actual transfection effect is the second effect and the predicted transfection effect is the first effect, and the fourth prediction type is used to indicate that the actual transfection effect is the second effect and the predicted transfection effect is the second effect, and the first effect is better than the second effect. In terms of determining the accuracy, precision, and recall rate of each trained prediction model according to the prediction type of each experimental data, the training unit 605 is specifically configured to: perform the following operations for each trained prediction model: determine the accuracy according to the number of the first prediction type and the fourth prediction type corresponding to the current trained prediction model, and the total number of the prediction types corresponding to the current trained prediction model; determine the precision according to the number of the first prediction type corresponding to the current trained prediction model, and the total number of the first prediction type and the third prediction type corresponding to the current trained prediction model; determine the recall rate according to the number of the first prediction type corresponding to the current trained prediction model, and the total number of the first prediction type and the second prediction type corresponding to the current trained prediction model.
[0097] In a possible embodiment, the sample initial features include spatial features. In terms of obtaining the initial features of the ionizable lipid molecules, the first obtaining unit 601 is specifically configured to: obtain multiple molecular conformations of the ionizable lipid molecules; respectively perform alignment processing on the molecular structures of the multiple molecular conformations; perform density projection on the aligned multiple molecular structures to obtain a density map; extract the spatial features of the ionizable lipid molecules according to the density map.
[0098] In a possible embodiment, when respectively performing alignment processing on the molecular structures of the multiple molecular conformations, the first obtaining unit 601 is specifically configured to: perform the following operations on the molecular structure of each molecular conformation in the multiple molecular conformations: obtain a target central atom and two reference atoms from the head group of the current molecular structure; fix the target central atom at a target coordinate; obtain a first vector according to the target central atom and the molecular centroid of the current molecular structure; obtain a second vector according to the two reference atoms; rotate the current molecular structure so that the first vector is parallel to the first coordinate axis, and the second vector is parallel to the plane corresponding to the first coordinate axis and the second coordinate axis.
[0099] In a possible embodiment, in terms of extracting the spatial features of the ionizable lipid molecules according to the density map, the first acquisition unit 601 is specifically configured to: obtain a first boundary and a second boundary according to the density map, where the density value corresponding to the first boundary belongs to a first threshold, the density value corresponding to the second boundary belongs to a second threshold, and the first threshold is less than the second threshold; extract the spatial features of the ionizable lipid molecules according to the first boundary and the second boundary, and the spatial features include at least one of the following: molecular length, molecular width, aspect ratio, cotangent value of the cone angle of the molecular head group, and symmetry of the molecular head group.
[0100] In a possible embodiment, the prediction model generation device 600 further includes a prediction unit. After training the constructed prediction model according to the training set to obtain a target prediction model, the prediction unit is configured to: obtain the target spatial features and target chemical structure features of the target ionizable lipid molecule to be predicted; input the target spatial features and the target chemical structure features into the target prediction model to obtain a prediction result, and the prediction result is used to indicate the transfection effect of the lipid nanoparticle constructed based on the target ionizable lipid molecule for carrying mRNA.
[0101] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of the electronic device provided in the present application. As Figure 7 shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 complete mutual communication through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute the method for generating an ionizable lipid molecule prediction model, and the method includes: obtaining the initial features of the ionizable lipid molecules; determining target features according to the initial features, where the target features are high-dimensional combinations of the initial features; constructing a prediction model according to the target features, and the prediction model is used to predict the transfection effect of the ionizable lipid molecules on the lipid nanoparticles constructed for carrying mRNA; obtaining a training set; training the constructed prediction model according to the training set to obtain a target prediction model.
[0102] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0103] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to implement the method for generating an ionizable lipid molecule prediction model provided by the above-mentioned various methods. The method includes: obtaining the initial features of the ionizable lipid molecule; determining the target features according to the initial features, where the target features are high-dimensional combinations of the initial features; constructing a prediction model according to the target features, and the prediction model is used to predict the transfection effect of the ionizable lipid molecule on constructing a lipid nanoparticle for carrying mRNA; obtaining a training set; and training the constructed prediction model according to the training set to obtain a target prediction model.
[0104] In yet another aspect, the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, it is used to implement the method for generating an ionizable lipid molecule prediction model as described in any one of the above. The method includes: obtaining the initial features of the ionizable lipid molecule; determining the target features according to the initial features, where the target features are high-dimensional combinations of the initial features; constructing a prediction model according to the target features, and the prediction model is used to predict the transfection effect of the ionizable lipid molecule on constructing a lipid nanoparticle for carrying mRNA; obtaining a training set; and training the constructed prediction model according to the training set to obtain a target prediction model.
[0105] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0106] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A method for generating a prediction model of ionizable lipid molecules, characterized in that, Including: Obtaining the initial features of the ionizable lipid molecule, where the initial features include multiple chemical structure features and multiple spatial features. The spatial features are obtained as follows: obtaining the molecular conformation through molecular dynamics simulation, respectively aligning the molecular structures of the multiple molecular conformations, performing density projection on the aligned multiple molecular structures to obtain a first density map representing the side information of the ionizable lipid molecule and a second density map representing the front information of the ionizable lipid molecule, and respectively extracting the spatial features of the ionizable lipid molecule from the first density map and the second density map. The spatial features include at least one of the following: molecular length, molecular width, aspect ratio, cotangent value of the cone angle of the molecular head group, and symmetry of the molecular head group; Determining the target features according to the initial features, where the target features are high-dimensional combinations of the initial features; Constructing a prediction model according to the target features, where the prediction model is used to predict the transfection effect of the ionizable lipid molecule on constructing lipid nanoparticles for carrying mRNA; Obtaining a training set; Training the constructed prediction model according to the training set to obtain a target prediction model.
2. The method according to claim 1, wherein The initial features include multiple chemical structure features and multiple spatial features. The determining the target features according to the initial features includes: Combining the initial features through mathematical operations to obtain multiple candidate features; Recursively selecting features from the candidate features whose correlation with the transfection effect is higher than a preset value; Combining the features whose correlation is higher than the preset value to obtain multiple feature combinations; Determining the candidate features included in the feature combination with the best prediction effect among the multiple feature combinations as the target features.
3. The method according to claim 1, wherein The training the constructed prediction model according to the training set to obtain a target prediction model includes: Selecting multiple experimental data from the training set as sub-training sets to obtain multiple sub-training sets; Respectively training the constructed prediction model according to the multiple sub-training sets to obtain multiple trained prediction models; Obtaining a test set, where the test set includes multiple experimental data; Obtaining the actual transfection effect of each experimental data included in the test set; Obtaining the predicted transfection effect of each experimental data predicted by the trained prediction model; Selecting a target prediction model from the multiple trained prediction models according to the actual transfection effect and the predicted transfection effect.
4. The method according to claim 3, wherein The selecting a target prediction model from the multiple trained prediction models according to the actual transfection effect and the predicted transfection effect includes: Determining the prediction type of each trained prediction model for each experimental data according to the actual transfection effect and the predicted transfection effect; Determining the accuracy, precision, and recall rate of each trained prediction model according to the prediction type; Scoring each trained prediction model according to the accuracy and the recall rate to obtain a target score; Selecting a target prediction model from the multiple trained prediction models according to the accuracy, the precision, the recall rate, and the target score.
5. The method according to claim 4, characterized in that, The prediction types include four types: the first prediction type is used to indicate that the actual transfection effect is the first effect and the predicted transfection effect is the first effect; the second prediction type is used to indicate that the actual transfection effect is the first effect and the predicted transfection effect is the second effect; the third prediction type is used to indicate that the actual transfection effect is the second effect and the predicted transfection effect is the first effect; the fourth prediction type is used to indicate that the actual transfection effect is the second effect and the predicted transfection effect is the second effect; the first effect is better than the second effect. Determining the accuracy rate, precision rate, and recall rate of each trained prediction model according to the prediction type includes: Performing the following operations on each trained prediction model: Determining the accuracy rate according to the number of the first prediction type and the fourth prediction type corresponding to the current trained prediction model, and the total number of the prediction types corresponding to the current trained prediction model; Determining the precision rate according to the number of the first prediction type corresponding to the current trained prediction model, and the total number of the first prediction type and the third prediction type corresponding to the current trained prediction model; Determining the recall rate according to the number of the first prediction type corresponding to the current trained prediction model, and the total number of the first prediction type and the second prediction type corresponding to the current trained prediction model.
6. The method according to claim 1, characterized in that Performing alignment processing on the molecular structures of the multiple molecular conformations respectively includes: Performing the following operations on the molecular structure of each molecular conformation in the multiple molecular conformations: Obtaining a target central atom and two reference atoms from the head group of the current molecular structure; Fixing the target central atom at the target coordinates; Obtaining a first vector according to the target central atom and the molecular centroid of the current molecular structure; Obtaining a second vector according to the two reference atoms; Rotating the current molecular structure so that the first vector is parallel to the first coordinate axis, and the second vector is parallel to the plane corresponding to the first coordinate axis and the second coordinate axis.
7. The method according to claim 1, characterized in that, Extracting the spatial features of the ionizable lipid molecule according to the density map includes: Obtaining a first boundary and a second boundary according to the density map, the density value corresponding to the first boundary belongs to a first threshold, the density value corresponding to the second boundary belongs to a second threshold, and the first threshold is less than the second threshold; Extracting the spatial features of the ionizable lipid molecule according to the first boundary and the second boundary.
8. The method according to claim 1, characterized in that After training the constructed prediction model according to the training set to obtain the target prediction model, the method further includes: Obtaining the target spatial features and target chemical structure features of the target ionizable lipid molecule to be predicted; Inputting the target spatial features and the target chemical structure features into the target prediction model to obtain a prediction result, and the prediction result is used to indicate the transfection effect of the lipid nanoparticle constructed based on the target ionizable lipid molecule for carrying mRNA.
9. An apparatus for generating a prediction model of ionizable lipid molecules, characterized in that, Including: A first acquisition unit, configured to acquire initial features of ionizable lipid molecules, where the initial features include a plurality of chemical structure features and a plurality of spatial features. In terms of acquiring spatial features, the first acquisition unit is further configured to: obtain molecular conformations through molecular dynamics simulation, respectively perform alignment processing on the molecular structures of the plurality of molecular conformations, perform density projection on the aligned plurality of molecular structures to obtain a first density map representing the side information of the ionizable lipid molecule and a second density map representing the front information of the ionizable lipid molecule, and respectively extract the spatial features of the ionizable lipid molecule according to the first density map and the second density map, where the spatial features include at least one of the following: molecular length, molecular width, aspect ratio, cotangent value of the cone angle of the molecular head group, and symmetry of the molecular head group; A determination unit, configured to determine target features according to the initial features, where the target features are high-dimensional combinations of the initial features; A construction unit, configured to construct a prediction model according to the target features, where the prediction model is used to predict the transfection effect of the ionizable lipid molecule on constructing lipid nanoparticles for carrying mRNA; A second acquisition unit, configured to acquire a training set; A training unit, configured to train the constructed prediction model according to the training set to obtain a target prediction model.
Citation Information
Patent Citations
Prediction method and device for mRNA vaccine transfection efficiency and medium
CN115527606A
Therapeutic effect prediction model training method, therapeutic effect prediction method and electronic equipment
CN115631850A
Ionizable lipid compound with three-dimensional space structure as well as preparation method and application of ionizable lipid compound
CN119320334A