Ionizable lipid and design method and related application thereof
By constructing a property prediction model for lipid nanoparticles and a deep learning generation model, the structure of ionizable lipid molecules is optimized, solving the problem of low screening efficiency in existing technologies. This enables efficient screening and development of ionizable lipid molecules with excellent performance, supporting the optimization of nucleic acid drug delivery systems.
Patent Information
- Application Number
- CN202510938096.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-24
AI Technical Summary
Existing methods for screening ionizable lipids rely on trial-and-error experiments, resulting in low success rates, long cycles, high material consumption, and difficulty in fully exploring the chemical space.
Construct a lipid nanoparticle property prediction model, optimize the ionizable lipid molecular structure through a multi-task learning framework and deep learning generative model, and achieve efficient screening.
This significantly improves the development efficiency and success rate of ionizable lipids, provides an optimized approach for nucleic acid drug delivery systems, and generates novel ionizable lipid molecules with excellent performance.
Smart Images

Figure CN120833848A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of nucleic acid delivery, in particular, to ionizable lipids, design methods thereof and related applications. BACKGROUND
[0002] Messenger ribonucleic acid (mRNA) as an important active pharmaceutical ingredient in modern medicine can be delivered into the body to prevent and treat various diseases, showing great application potential.
[0003] Lipid nanoparticles (LNP) is one of the key systems for mRNA delivery at present, and its core feature is to construct an mRNA delivery system with ionizable lipids as the main component. The molecular structure of ionizable lipids usually contains nitrogen atoms, which can ionize and carry positive charges in an acidic environment, but remain uncharged in a neutral environment. It is this ionizability that enables LNP to efficiently encapsulate mRNA during preparation and promote mRNA escape from endosomes after entering cells, thereby achieving the effect of the drug. In addition, compared with lipids with constant positive charges, LNP prepared from ionizable lipids shows higher safety.
[0004] Currently, the screening of ionizable lipids mainly relies on the trial-and-error method, that is, by synthesizing a large number of candidate lipids, using these lipids to prepare RNA (such as mRNA or small interfering RNA)-loaded LNP (RNA-LNP), and testing their delivery efficiency, from which the ionizable lipids with the best delivery efficiency are screened. However, this research method has significant limitations. For example, the success rate of lipid screening is low, the research period is long, and a large amount of experimental materials and animal resources are consumed. In addition, due to the complexity of the chemical structure space, the efficiency of screening solely relying on trial-and-error experiments is low, and it is difficult to fully explore the potential chemical space.
[0005] Therefore, it is urgent to develop a rapid, efficient and low-cost ionizable lipid screening system to accelerate the research process of ionizable lipids. This not only promotes the development of lipid nanoparticle systems, but also provides important support for mRNA delivery technology, which has important significance for the deep mining of ionizable lipids and the development of related fields.
[0006] In view of this, the present application is proposed. SUMMARY
[0007] The present application provides ionizable lipids, design methods thereof and related applications.
[0008] The present application is implemented as follows:
[0009] In a first aspect, the embodiments of the present application provide a method for constructing a lipid nanoparticle property prediction model, which comprises the following steps: obtaining a pre-trained model pre-trained by using an ionizable lipid structure dataset; constructing a multi-task learning framework according to the pre-trained model, fine-tuning the multi-task learning framework by using an LNP property related dataset, and obtaining a multi-task learning model capable of predicting LNP properties.
[0010] In a second aspect, the embodiments of the present application provide a method for predicting properties of lipid nanoparticles, which comprises: inputting a chemical structure of an ionizable lipid to be predicted into a model constructed by the method for constructing described in the foregoing embodiments, to obtain a prediction result.
[0011] In a third aspect, the embodiments of the present application provide a method for designing ionizable lipids, which comprises the following steps: pre-training a deep learning generation model by using an ionizable lipid structure dataset, generating an initial generation model capable of generating new ionizable lipid molecular structures; optimizing and feeding back the initial generation model to guide the generation of ionizable lipid molecular structures meeting preset LNP properties; and obtaining candidate molecules based on the generation model after optimization and feedback.
[0012] In a fourth aspect, the embodiments of the present application provide ionizable lipid molecules, which are designed by the method for designing described in the foregoing embodiments.
[0013] In a fifth aspect, the embodiments of the present application provide a lipid nanoparticle or a nucleic acid drug delivery system containing the lipid nanoparticle, wherein the lipid nanoparticle contains the ionizable lipid molecules described in the foregoing embodiments.
[0014] In a sixth aspect, the embodiments of the present application provide use of the ionizable lipid molecules described in the foregoing embodiments in the preparation of a lipid nanoparticle or a nucleic acid drug delivery system.
[0015] The present application has the following beneficial effects:
[0016] The embodiments of the present application innovatively construct a method for designing ionizable lipids, which can efficiently explore chemical space, generate new ionizable lipid molecules with excellent performance, greatly improve the development efficiency and success rate of ionizable lipids, and provide a new approach for optimizing nucleic acid drug delivery systems. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be considered as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0018] Figure 1 is a schematic diagram of the design method of the ionizable lipid provided by the present application, including the schematic diagram of LNP data construction, LNP property prediction method and ionizable lipid generation method;
[0019] Figure 2 is a schematic diagram of FormulationLNP architecture and training provided by the present application;
[0020] Figure 3A is a data enhancement attention mechanism analysis provided by the present application;
[0021] Figure 3B is an important ionizable lipid molecule substructure for predicting the influence in the in vivo delivery efficiency prediction task provided by the present application;
[0022] Figure 3C is an important ionizable lipid molecule substructure for predicting the influence in the apparent pKa prediction task provided by the present application;
[0023] Figure 4 is the result of the uncertainty measurement method provided by the present application;
[0024] Figure 5 is a schematic diagram of the ionizable lipid generation framework LipidGPT provided by the present application;
[0025] Figure 6 is the prediction performance of the LNP property prediction model in the ionizable lipid generation provided by the present application;
[0026] Figure 7 is the verification result of the ionizable lipid synthesis accessibility scoring tool SAscoreLNP provided by the present application;
[0027] Figure 8 is the 1 H-NMR spectrum of AL-1;
[0028] Figure 9 is the 1 H-NMR spectrum of AL-2;
[0029] Figure 10 is the 1 H-NMR spectrum of AL-3;
[0030] Figure 11 is the 1 H-NMR spectrum of AL-4;
[0031] Figure 12 is the 1 H-NMR spectrum of AL-5;
[0032] Figure 13H-NMR spectrum of AL-6 1 H-NMR spectrum of AL-6
[0033] Figure 14 H-NMR spectrum of AL-7 1 H-NMR spectrum of AL-7
[0034] Figure 15 H-NMR spectrum of AL-8 1 H-NMR spectrum of AL-8
[0035] Figure 16 H-NMR spectrum of AL-9 1 H-NMR spectrum of AL-9
[0036] Figure 17 H-NMR spectrum of AL-10 1 H-NMR spectrum of AL-10
[0037] Figure 18 H-NMR spectrum of AL-11 1 H-NMR spectrum of AL-11
[0038] Figure 19 H-NMR spectrum of AL-12 1 H-NMR spectrum of AL-12
[0039] Figure 20 H-NMR spectrum of AL-13 1 H-NMR spectrum of AL-13
[0040] Figure 21 H-NMR spectrum of AL-14 1 H-NMR spectrum of AL-14
[0041] Figure 22 H-NMR spectrum of AL-15 1 H-NMR spectrum of AL-15
[0042] Figure 23A Particle size, PDI, EE of the newly synthesized ionizable lipids LNP provided by the application;
[0043] Figure 23B Apparent pKa of the newly synthesized ionizable lipids LNP provided by the application;
[0044] Figure 24A Whole body biofluorescence signal intensity of mice at 4 hours after administration of the newly synthesized ionizable lipids LNP provided by the application;
[0045] Figure 24B Whole body biofluorescence signal intensity of mice within 4-48 hours after administration of the newly synthesized ionizable lipids LNP provided by the application;
[0046] Figure 24C Representative bioluminescence images of mice in each group at 4 hours after administration of the newly synthesized ionizable lipids LNP provided by the application. DETAILED DESCRIPTION
[0047] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below. If specific conditions are not specified in the embodiments, the conventional conditions or the conditions recommended by the manufacturers are adopted. If the manufacturers of the reagents or instruments are not specified, the conventional products that can be purchased in the market are adopted.
[0048] DEFINITIONS
[0049] If the terms "first", "second", "third" and the like are used herein, they are only used for distinguishing description, and cannot be understood as indicating or implying relative importance.
[0050] The term "LNP property-related dataset" herein is the same as "dataset of physicochemical properties and biological indicators of LNPs".
[0051] In one aspect, the embodiments of the present application provide a method for constructing a lipid nanoparticle property prediction model (LNP property model), which comprises the following steps:
[0052] Obtaining a pre-trained model pre-trained by using the ionizable lipid structure dataset, and obtaining molecular representation of ionizable lipids based on the pre-trained model;
[0053] Constructing a multi-task learning framework according to the pre-trained model, fine-tuning the multi-task learning framework by using the LNP property-related dataset, and obtaining a multi-task learning model capable of predicting LNP properties.
[0054] In some embodiments, the ionizable lipid structure dataset and the LNP property-related dataset can be obtained by querying public databases, and LNP-related patents and / or scientific literatures can be obtained. The obtained documents are parsed and data is extracted.
[0055] In some embodiments, the public databases include but are not limited to SciFinder, Web of Science, PubMed, Google Scholar, public patent databases, and professional databases related to drug delivery.
[0056] In some embodiments, the method for obtaining the LNP property-related dataset further comprises pre-processing the obtained data (to improve data uniformity): converting the chemical structure of the ionizable lipid into a standard SMILES representation, including but not limited to desalting, duplicate data deduplication, and the like; and / or, standardizing other physicochemical property and biological index data other than the chemical structure of the ionizable lipid, including but not limited to: selecting data with consistent administration routes, consistent measurement indicators, mRNA expression levels comparable to standard LNP, and formulation compositions identical or equivalent to standard LNP, and the like.
[0057] Through the above method for preparing the LNP property-related dataset, LNP-related research data can be systematically collected, organized and standardized, providing a training data basis for ionizable lipid design based on artificial intelligence, thereby improving the prediction accuracy and practical value of the artificial intelligence model.
[0058] In some embodiments, the LNP property includes any one or more of in vivo delivery efficiency, apparent pKa value, particle size, polydispersity index, and encapsulation efficiency.
[0059] In some embodiments, the LNP property-related dataset includes any one or more of the LNP property, LNP formulation, type of protein encoded by the nucleic acid drug encapsulated by the LNP, experimental animal species, administration route, administration dose, and nucleic acid drug expression level.
[0060] In some embodiments, the nucleic acid drug includes RNA and / or DNA.
[0061] In some embodiments, the RNA includes mRNA and / or siRNA.
[0062] In some embodiments, the LNP formulation includes the chemical structure of the ionizable lipid and / or the type of helper lipid.
[0063] In some embodiments, the method for constructing the pre-trained model further comprises: expanding the data amount of the ionizable lipid structure dataset according to a data augmentation strategy; pre-training the pre-trained model using the expanded ionizable lipid structure dataset, with the loss value on the validation set as the target and the condition of no decrease in the validation set loss for 3-10 consecutive training cycles as the convergence condition, to obtain the trained pre-trained model. Pre-training can enable the model to learn the general representation of ionizable lipid molecules, providing a basic representation capability for subsequent tasks.
[0064] In some embodiments, during the pre-training process, a data augmentation strategy is used to expand the data amount of the ionizable lipid structure dataset.
[0065] In some embodiments, the data augmentation strategy includes randomly masking molecular graph partial structure and / or SMILES string enumeration. Randomly masking molecular graph partial structure refers to randomly masking part of the nodes or edges in a molecular graph (consisting of atom nodes and chemical bond edges) to force the model to reconstruct the masked part through contextual information. Molecular functions are often determined by local functional groups (such as hydroxyl groups determining polarity), and the masking operation can train the model to infer local features from global structure, enhancing tolerance to structural variations. SMILES string enumeration refers to the use of the property that the same molecule can have multiple valid SMILES representations to generate different SMILES strings representing the same molecule by changing the starting point and traversal direction of the molecular graph. This method can increase the diversity of training data while maintaining the invariance of molecular structure, thereby improving the robustness of the model to different SMILES representations and enabling the model to focus on learning the structural properties of the molecule rather than relying on a specific SMILES string representation order.
[0066] In some embodiments, the pre-trained model includes any one or more of a combination of RoBERTa, molecular graph attention network (MolGAT), graph convolutional neural network (GCN), message passing neural network (MPNN), and Transformer model.
[0067] In some embodiments, the multi-task learning framework consists of a shared encoding layer and task-specific output layer.
[0068] In some embodiments, the shared encoding layer is from a pre-trained model, including but not limited to a pre-trained model directly used or adapted through fine-tuning to meet the prediction requirements. In some embodiments, the step of fine-tuning the multi-task learning framework using the LNP property-related dataset includes fixing the weight setting or dynamically adjusting the training weights between multiple tasks based on task loss to optimize the prediction performance of key properties.
[0069] In some embodiments, the step of fine-tuning the multi-task learning framework using the LNP property-related dataset includes using a data augmentation strategy to expand the amount of data in the LNP property-related dataset. The method of data augmentation strategy includes randomly masking molecular graph partial structure and / or SMILES string enumeration.
[0070] In some embodiments, the construction method further includes performing explainability analysis on the fine-tuned multi-task learning model to identify key substructures of ionizable lipids that significantly affect the prediction results.
[0071] In some embodiments, the explainability analysis includes merging atoms or groups in a molecular structure into larger fragments to enhance the chemical meaning of the explanation and reveal the degree of attention paid by the model to molecular features through result visualization.
[0072] In some embodiments, the method of explainability analysis includes a toolkit based on attention mechanism analysis methods.
[0073] In some embodiments, the toolkit includes at least one of SHAP and LIME.
[0074] In some embodiments, the construction method further includes evaluating the uncertainty of the prediction results of the multi-task learning model using an uncertainty measurement method.
[0075] In some embodiments, the uncertainty measurement method includes but is not limited to calculating the standard deviation by integrating data enhancement results to quantify prediction uncertainty, using Monte Carlo dropout to evaluate the prediction stability of the model under different sampling conditions, and combining multiple model prediction results based on deep ensemble learning methods for comprehensive analysis. Through uncertainty measurement, potential risks in prediction results can be effectively identified, providing a basis for subsequent experimental verification and model optimization.
[0076] Through the above LNP property prediction model construction method, various advanced artificial intelligence technologies can be effectively utilized to overcome the data limitations and prediction challenges in traditional ionized lipid design, providing strong computational tool support for the rational design of LNP delivery systems, and significantly improving the development efficiency and success rate of new ionized lipids.
[0077] In another aspect, the embodiments of the present application also provide a lipid nanoparticle property prediction device, which comprises:
[0078] An acquisition module is configured to acquire the ionizable lipid structure dataset and the LNP property related dataset according to any of the preceding embodiments.
[0079] A pre-training module is configured to pre-train a pre-trained model using the ionizable lipid structure dataset according to any of the preceding embodiments, to obtain a pre-trained pre-trained model.
[0080] A multi-task learning module is configured to construct a multi-task learning framework according to the pre-trained model and fine-tune the multi-task learning framework using the LNP property related dataset according to any of the preceding embodiments, to obtain a multi-task learning model capable of predicting LNP properties.
[0081] The module in the embodiments of the present application can be stored in the memory in the form of software or firmware, or solidified in the operating system (OS) of the electronic device provided in the present application, and can be executed by the processor in the electronic device. At the same time, the data, program code and the like required for executing the above-mentioned module can be stored in the memory.
[0082] In another aspect, the embodiments of the present application provide a prediction method for predicting the properties of lipid nanoparticles, comprising: inputting the chemical structure of an ionizable lipid to be predicted into the model constructed by the construction method of any of the preceding embodiments to obtain a prediction result.
[0083] In another aspect, the embodiments of the present application provide a design method for ionizable lipids, comprising the following steps:
[0084] The deep learning generation model is pre-trained using the ionizable lipid structure dataset to generate an initial generation model capable of generating new ionizable lipid molecular structures;
[0085] The initial generation model is optimized and fed back to guide the generation of ionizable lipid molecular structures that meet the preset LNP properties.
[0086] The generation model after optimization and feedback is used for single sampling or multiple sampling to obtain candidate molecules.
[0087] In some embodiments, the deep learning generation model comprises any one or more of the following: an encoder (VAE), a generative adversarial network (GAN), and a generative pre-training transformer network (GPT).
[0088] In some embodiments, before pre-training, the design method further comprises preprocessing the ionizable lipid structure dataset: removing molecules with SMILES string length exceeding 140, and / or deleting ionizable lipid molecules in the ionizable lipid structure dataset that appear in the LNP property-related dataset.
[0089] In some embodiments, the optimization and feedback method comprises any one or more of the following: reinforcement learning, Bayesian optimization, and bootstrap fine-tuning. The optimization and feedback can iteratively optimize the parameters of the generation model and gradually improve the performance of the target properties of the generated molecules.
[0090] In some embodiments, the step of optimizing the initial generative model comprises: obtaining an initial generative result of the initial generative model; using the LNP property prediction model to predict the LNP property of the initial result, and screening ionizable lipids that meet the preset LNP property as a fine-tuning dataset; using the fine-tuning dataset to optimize the initial generative model, and obtaining a generative model capable of generating ionizable lipids meeting the preset LNP property.
[0091] In some embodiments, the LNP property prediction model is constructed using any one or more of deep learning, random forest, LightGBM, support vector machine, and multilayer perception.
[0092] In some embodiments, the step of using the LNP property prediction model to predict the LNP property of the initial result comprises: using multiple LNP property prediction models to perform integrated prediction of the LNP property of the initial result. The multiple LNP property prediction models include LightGBM, support vector machine, and multilayer perception.
[0093] In some embodiments, the LNP property includes any one or more of in vivo delivery efficiency, apparent pKa value, particle size, polydispersity index, and encapsulation efficiency.
[0094] Sampling refers to the process of drawing a sample from the probability distribution of the generative model, corresponding to "In the generation process, starting from the initial token "C" representing a carbon atom (as it is the most common token in ionizable lipids), the model predicts the subsequent tokens one by one to construct the molecule. To control the diversity and novelty of the generated molecules, temperature sampling is implemented during token prediction...", single sampling: the model generates one candidate lipid molecule structure for each run; multiple sampling: the model is run multiple times, each time generating a different candidate molecule, and a set of candidate molecules is ultimately obtained.
[0095] In some embodiments, the LNP property prediction model is constructed using the construction method of any of the preceding embodiments.
[0096] In some embodiments, the design method further comprises the step of further screening the candidate molecules.
[0097] In some embodiments, the screening comprises filtering the candidate molecules to obtain new ionizable lipid molecules as a first screening result.
[0098] The "new ionizable lipid molecule" refers to an ionizable lipid that is not present in the entire ionizable lipid structure dataset constructed in the application. The specific implementation is: first, convert the molecule into a canonical SMILES representation, then compare the candidate molecule with the pre-constructed ionizable lipid structure database, and exclude the structures already existing in the database, which is completed using the code written by python.
[0099] In some embodiments, the screening further comprises: using a lipid nanoparticle property prediction model to predict the candidate molecules or the first screening result, screening ionizable lipid molecules that meet the target lipid nanoparticle properties as the second screening result.
[0100] In some embodiments, the screening further comprises: performing structural clustering or similarity analysis on the candidate molecules, the first screening result or the second screening result to obtain different sub-classes, and selecting ionizable lipid molecules from different sub-classes as the third screening result.
[0101] In some embodiments, the screening further comprises: performing synthesis difficulty scoring on the candidate molecules, the first screening result, the second screening result or the third screening result, and screening molecules with low synthesis difficulty for synthesis.
[0102] In some embodiments, the synthesis scoring tool is based on a machine learning model, a chemical rule engine or other algorithms, which can comprehensively score the synthesis steps, raw material availability, chemical reaction conditions, etc. of the molecule, and output the synthesis feasibility score result.
[0103] Through the above-mentioned design method of ionizable lipids, multi-objective optimization of ionizable lipids is realized, and the first generative AI framework specially designed for ionizable lipids is developed. The method can efficiently explore the chemical space, generate new ionizable lipid structures with excellent performance, greatly improve the development efficiency of ionizable lipids, and provide strong technical support for the optimization of nucleic acid drug delivery systems.
[0104] The schematic diagram of the ionizable lipid design method provided by the embodiments of the application includes the schematic diagram of LNP data construction, LNP property prediction method and ionizable lipid generation method, which can be specifically referred to as Figure 1 .
[0105] On the other hand, the embodiments of the application also provide an ionizable lipid design device, which comprises:
[0106] The initial generation module is configured to implement the pre-training of the deep learning generation model using the ionizable lipid structure dataset in any of the preceding embodiments, to generate an initial generation model capable of generating new ionizable lipid molecule structures.
[0107] an optimization feedback module configured to implement the optimization feedback on the initial generated model to guide the generation of the structure of the ionizable lipid molecule satisfying the preset LNP property according to any one of the preceding embodiments;
[0108] a screening module configured to implement the single sampling or multiple sampling by the generated model after the optimization feedback according to any one of the preceding embodiments to obtain the candidate molecule.
[0109] In another aspect, embodiments of the present application also provide an electronic device including a processor and a memory, the memory being configured to store a program, when the program is executed by the processor, causing the processor to implement the method for constructing the LNP property prediction model according to any one of the preceding embodiments, the method for predicting the LNP property according to any one of the preceding embodiments, or the method for designing the ionizable lipid according to any one of the preceding embodiments.
[0110] The electronic device can include a memory, a processor, a bus and a communication interface, which are directly or indirectly electrically connected to each other to realize the transmission or interaction of data. For example, these elements can be electrically connected to each other through one or more buses or signal lines.
[0111] The memory can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM) and the like.
[0112] The processor can be an integrated circuit chip with signal processing capability. The processor 120 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; or can be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component.
[0113] The electronic device can be a server, a cloud platform, a mobile phone, a tablet computer, a notebook computer, an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a personal digital assistant (PDA), a wearable electronic device, a virtual reality device, and the like, and thus the type of the electronic device is not limited in the embodiments of the present application.
[0114] In another aspect, the embodiments of the present application further provide a computer readable medium, which stores a computer program. The computer program is executed by a processor to implement the method for constructing a lipid nanoparticle property prediction model according to any of the preceding embodiments, the method for predicting a lipid nanoparticle property according to any of the preceding embodiments, or the method for designing an ionizable lipid according to any of the preceding embodiments.
[0115] In some embodiments, the computer readable medium can be a general storage medium, such as a mobile disk, a hard disk, etc.
[0116] In another aspect, the embodiments of the present application further provide an ionizable lipid molecule, which is designed by the method for designing an ionizable lipid according to any of the preceding embodiments.
[0117] In some embodiments, the ionizable lipid molecule is selected from any one of AL-1, AL-2, AL-3, AL-4, AL-5, AL-6, AL-7, AL-8, AL-9, AL-10, AL-11, AL-12, AL-13, AL-14, and AL-15.
[0118] The structural formula of AL-1 is as follows: The structural formula of AL-2 is as follows: The structural formula of AL-3 is as follows: The structural formula of AL-4 is as follows: The structural formula of AL-5 is: The structural formula of AL-6 is:
[0119]
[0120] The structural formula of AL-7 is:
[0121]
[0122] The structural formula of AL-8 is: The structural formula of AL-9 is: The structural formula of AL-10 is:
[0123]
[0124] The structural formula of AL-11 is:
[0125]
[0126] The structural formula of AL-12 is:
[0127]
[0128] The structural formula of AL-13 is:
[0129] The structural formula of AL-14 is:
[0130] The structural formula of AL-15 is:
[0131] In another aspect, the embodiments of the present application also provide a lipid nanoparticle or a nucleic acid drug delivery system containing the lipid nanoparticle, wherein the lipid nanoparticle contains the ionizable lipid molecule according to any of the preceding embodiments.
[0132] In some embodiments, the nucleic acid drug delivery system comprises a nucleic acid drug and the lipid nanoparticle encapsulating the nucleic acid drug.
[0133] In some embodiments, the nucleic acid drug comprises RNA and / or DNA.
[0134] In some embodiments, the RNA comprises mRNA and / or siRNA.
[0135] In addition, the embodiments of the present application provide use of the ionizable lipid molecule according to any of the preceding embodiments in the preparation of a lipid nanoparticle or a nucleic acid drug delivery system.
[0136] The features and characteristics of the present application are further described in detail below in conjunction with the embodiments.
[0137] Example 1
[0138] (1) Construction method of dataset
[0139] In this embodiment, the United States Patent and Trademark Office (USPTO) patent database is selected as the main data source, and the specific operation steps are as follows:
[0140] 1.1, Batch download patent text data from USPTO, get XML format of granted patents and published application files;
[0141] 1.2, Parse the downloaded patent data, and identify the patents related to LNP and nucleic acid delivery technology from the patent documents through keyword screening and patent classification code filtering. Set the screening keyword combination, including but not limited to "lipid nanoparticle", "mrna", "sirna", "cationic lipid", "ionizable lipid", "structure", etc.; The set patent classification codes include but are not limited to A61K, C07C, C07D, C07F, etc.
[0142] 1.3, Extract the molecular structure information in the patent file by using multiple methods, and standardize the structures. In this embodiment, the molecular structure is extracted by the mol file attached to the patent or the optical character recognition (OCR) technology, converted to SMILES representation, and standardized by using the RDKit library, including conversion to canonical SMILES, removal of salt and solvent molecules, standardization of aromatic representation, neutralization of formal charge, and standardization of stereochemical representation. Through structure duplication checking and cleaning, the ionizable lipid structure dataset is finally obtained.
[0143] 1.4, In this embodiment, XML parsing code is used to extract text and table data from patent documents, and unit conversion and standardization processing are performed on the data to extract LNP formula composition information (ionizable lipid structure, ionizable lipid, auxiliary lipid, cholesterol and PEG lipid ratio) and experimental result data (particle size, polydispersity coefficient, encapsulation efficiency, transfection efficiency, mRNA species and proportion, drug administration scheme, test object, measurement result and measurement time, etc.). Standardize the data, organize the data according to the molecular layer, formula layer and performance layer, and construct the physicochemical property and biological index dataset of LNP (LNP property related dataset).
[0144] In this embodiment, the physicochemical property and biological index dataset of LNP includes: in vivo delivery efficiency dataset and apparent pKa dataset.
[0145] The method for constructing the in vivo delivery efficiency dataset specifically includes: screening LNP formulation data containing in vivo experimental results. Due to the methodological differences between different sources of data and the variability between laboratories, potential bias may be introduced. In order to reduce these problems and ensure data consistency, the data is processed according to the following rules: a. Only LNP data administered by intravenous injection is retained; b. Experimental data using MC3-LNP formulation as a control is retained. Specifically, the control group MC3-LNP should meet one of the following conditions: it has a standard formulation containing MC3 (ionizable lipid), cholesterol, DSPC (auxiliary lipid), and PEG2000-DMG (PEG lipid) with a molar ratio of 50:38.5:10:1.5; or it has the same composition as other ionizable lipid formulations; c. Using the delivery efficiency of MC3-LNP as a threshold, samples with delivery efficiency equal to or better than MC3-LNP are classified as positive samples; samples with delivery efficiency lower than MC3-LNP are classified as negative samples. After the above processing, the final in vivo delivery efficiency dataset contains 820 data, of which 358 are positive samples and 462 are negative samples. In this dataset, the ionizable lipid is represented in the form of a SMILES string, and other formulation and experimental information is represented in 19 table features.
[0146] The method for constructing the apparent pKa dataset includes: screening LNP formulation data with apparent pKa experimental results to construct the dataset. Previous studies have shown that an apparent pKa value between 6 and 7 may be the optimal range. This embodiment classifies pKa data into the following three categories and further processes the apparent pKa dataset as follows: pKa values less than 6 are classified as category 1; pKa values between 6 and 7 are classified as category 2; pKa values greater than 7 are classified as category 3. After the above processing, the final apparent pKa dataset contains 656 data, of which category 1 contains 109, category 2 contains 449, and category 3 contains 98. In this dataset, the ionizable lipid is represented in the form of a SMILES string, and other formulation and experimental information is represented in 12 table features.
[0147] Embodiment 2
[0148] Based on the ionizable lipid structure dataset and the physicochemical properties and biological indicators dataset of LNP constructed in Embodiment 1, this embodiment provides a method for constructing a lipid nanoparticle property prediction model, which is suitable for execution in a computing device and includes the following steps, which can be referred to Figure 2 .
[0149] (1) Using self-supervised pre-training technology, pre-training the to-be-pre-trained model using the ionizable lipid structure dataset to obtain a pre-trained model. The pre-trained model used is RoBERTa.
[0150] The ionizable lipid structure in the ionizable lipid structure dataset is converted into its simplified molecular linear input specification (SMILES) by the RDKit package in Python, and is processed by tokenization and featureization. The tokenization process specifically includes: splitting the SMILES string according to chemical bonds and molecular fragments to generate basic units with chemical significance (such as atoms, bond types, ring structures, etc.); based on the SMILES basic units appearing in the entire dataset, a vocabulary of chemical fragments is constructed to ensure that the vocabulary can cover all basic units in the molecular structure; the split SMILES string is mapped to the index value in the vocabulary to generate a serialized representation of the molecular structure, which is convenient for input into the deep learning model for training.
[0151] RoBERTa is used as the pre-training model, which contains 6 layers and 12 attention heads, forming 72 different attention mechanisms. The model learns the context information of the token sequence to capture the relationship between chemical fragments in the molecular structure and the characteristics of the whole molecule.
[0152] The pre-training configuration is as follows: the dataset is divided into 90% (training set) and 10% (validation set); 15% of the token mask is used for each input SMILES string; the batch size is set to 128; SMILES enumeration is used for data augmentation to expand the pre-training dataset. The loss value on the validation set is used as the target, and the validation set loss does not decrease for 3-10 consecutive training periods as the convergence condition, and the trained pre-training model is obtained.
[0153] (2) Based on the pre-training model obtained in step (1), a multi-task learning framework is constructed, and the LNP physicochemical property and biological index dataset constructed in embodiment 1 is used to fine-tune the multi-task learning framework to obtain a multi-task learning model (FormulationLNP) capable of predicting LNP properties. LNP properties include predicting LNP particle size, PDI, mRNA delivery efficiency, apparent pKa, etc.
[0154] FormulationLNP integrates a hard-shared multi-task learning framework and a pre-training model for molecular representation. The RoBERTa layer in FormulationLNP is shared by two downstream tasks through hard sharing, which is used to output the molecular embedding of ionizable lipids. After the RoBERTa layer, several fully connected (FC) layers are implemented to input table information specific to each task and predict the results of each task respectively. The weighted loss between the two tasks is calculated by an uncertainty weighting method.
[0155] The in vivo delivery efficiency dataset and the apparent pKa dataset in Example 1 were used as fine-tuning datasets to fine-tune the multi-task learning framework. The fine-tuning configurations are as follows: Adam optimizer was adopted with a learning rate of 1e-5; the datasets were divided into 90% (training set) and 10% (test set); the batch size was set to 64; the best training epoch was determined by 10-fold cross-validation on the training set; early stopping strategy was implemented, and the training was terminated if the validation loss did not decrease for five consecutive epochs; the model was retrained on the whole training set using the best training epoch to obtain the final model, and then the performance was verified on the test set. The whole fine-tuning process was repeated for 10 independent experiments, and the average prediction performance on the test set was reported as the model performance.
[0156] The model performance was evaluated by various evaluation metrics, including accuracy, precision, recall, and area under the receiver operating characteristic curve (ROC-AUC), and compared with other baseline models. Table 1 and Table 2 show the average prediction performance of the Formulation LNP model trained on the in vivo delivery efficiency dataset and the apparent pKa dataset based on Example 1, respectively, on the respective test sets, which outperforms the baseline models, demonstrating the effectiveness and practicality of the method in LNP property prediction.
[0157] Table 1. Comparison of model performance in predicting LNP in vivo delivery efficiency
[0158]
[0159]
[0160] Table 2. Comparison of model performance in predicting LNP apparent pKa
[0161]
[0162] In this example, the model was subjected to explainability analysis using attention mechanism and SHAP toolkit, identifying key molecular substructures that have important influence on the prediction results. The attention mechanism analysis results are shown in Figure 3A , the key molecular substructures in the in vivo delivery efficiency prediction task are shown in Figure 3B , and the key molecular substructures in the apparent pKa prediction task are shown in Figure 3C . By visualizing the attention weight distribution, the molecular regions that the model focuses on when making predictions were determined, providing important insights into the structure-function relationship. SHAP analysis further quantifies the contribution of each structural feature to the prediction results, providing molecular-level guidance for the rational design of ionizable lipids.
[0163] In this embodiment, the standard deviation is calculated by integrating the data enhancement results to quantify the prediction uncertainty, which is done as follows: by calculating the standard deviation of the prediction probability of each molecule under different SMILES representations, it is used as a quantitative indicator of uncertainty: the lower the standard deviation, the higher the prediction confidence, the better the prediction accuracy. The molecules in the test set are sorted by prediction standard deviation from small to large, and the cumulative accuracy of the top n samples is calculated. Figure 4 It is shown that the cumulative accuracy decreases as more high-standard-deviation data is included, verifying the effectiveness of the method. This method provides important insights into prediction reliability, allowing low-uncertainty and well-behaved compounds to be prioritized for experimental verification in large-scale virtual screening, thereby accelerating the experimental screening process.
[0164] Example 3
[0165] A method of designing an ionizable lipid, the method being adapted to be executed in a computing device, comprising the steps of.
[0166] Step one: dataset construction. This method uses the method in Example 1 to construct the ionizable lipid dataset. For the ionizable lipid structure dataset constructed in Example 1, first, the molecules with SMILES string length exceeding 140 are removed, and then further based on the normalized SMILES string representation, the ionizable lipid molecules in the ionizable lipid structure dataset that appear in the LNP property-related dataset are deleted, obtaining the final pre-training dataset for model training. This operation ensures that the pre-training dataset does not contain molecular structures with known specific delivery efficiency and apparent pKa properties, avoiding information leakage leading to model evaluation bias.
[0167] Step two: model training. This method uses a two-stage training strategy:
[0168] (a) pre-training phase: pre-training the model using the pre-training dataset to obtain an initial generative model (performing unconditioned generation of ionizable lipids);
[0169] (b) guided conditional generation phase: optimizing the initial generative model for feedback to guide the generation of ionizable lipid molecular structures that meet the target lipid nanoparticle properties (conditional generation based on bootstrapping fine-tuning).
[0170] As Figure 5As shown, in the present embodiment, a transformer-based generative model framework named LipidGPT (Lipid GPT) is constructed. The core architecture of LipidGPT contains 8 decoder layers, each layer contains 8 self-attention blocks, and the embedding dimension is 256. Each decoder block integrates a masked self-attention mechanism and a feed-forward network to capture the dependencies in the molecular sequence. The masked self-attention layer outputs a 256-dimensional vector, which is expanded to 1024-dimensional through GELU activation in the feed-forward network, and then projected back to 256-dimensional for the next layer. The calculation method of attention weight is to multiply the Query vector (Q) and the Key vector (K) between them, and introduce a scaling factor (d_k, equal to the dimension of the key vector) for scaling, and then process through the softmax function. These attention weights are then applied to the Value vector (V) to obtain the weighted sum, forming the output of the attention mechanism:
[0171]
[0172] In addition, position encoding is added to the input embedding to provide the model with the position information of each token in the sequence.
[0173] In the framework, each input SMILES string is encoded using the SMILES tokenizer. To pretrain LipidGPT, the dataset is divided into a training set (90%) and a validation set (10%). Then, the SMILES enumeration method is adopted on the training set and the validation set to expand the data by generating valid non-canonical SMILES representations of the same molecular structure. The model is trained using the Adam optimizer with a learning rate of 6x10 -4 The loss on the validation set is monitored to stop training early if no reduction in validation loss is observed for 3 consecutive epochs. During the training phase, the model is optimized to perform autoregressive prediction, i.e., receiving a partial SMILES sequence x <t as input and predicting the probability distribution of the subsequent token x t . The objective function uses cross-entropy loss to quantify the difference between the probability distribution p θ (x t |x <t ) predicted by the model and the true token y t,v . The objective function is:
[0174]
[0175] where T is the total number of tokens in the sequence (sequence length), v is a token in the vocabulary set, V is the complete vocabulary set containing all possible tokens, and yt,v is the real token represented by the one-hot encoding vector (the value at the real token position is 1, and the value at other positions is 0), x t is the token at position t in the sequence, x <t is the subsequence of tokens before position t, p θ (x t = v | x <t ) is the probability predicted by the model that the token at position t is v given the preceding tokens.
[0176] During the generation process, starting from the starting token "C" representing a carbon atom (because it is the most common token in ionizable lipids), the model predicts the subsequent tokens in turn to construct the molecule. In order to control the diversity and novelty of the generated molecules, temperature sampling is implemented during the token prediction process. The probability distribution is adjusted by a temperature parameter T: (3);
[0177] where z i is the predicted value for token i; z j represents the predicted scores for all possible tokens j in the vocabulary; p i is the sampling probability of token i given by LipidGPT; T (T > 0) is the temperature parameter, a lower value (T < 1) will produce more conservative but structurally reliable molecules, while a higher value (T > 1) will produce more diverse structures, with increased novelty but possibly reduced effectiveness.
[0178] In this embodiment, the guided conditional generation is implemented based on bootstrapping fine-tuning, and the specific steps include:
[0179] ① Initial sampling: use the pre-trained model to generate the first batch of ionizable lipid molecule structures;
[0180] ② Performance evaluation: use various tools to evaluate the key parameters of the generated molecules, including evaluating the intrinsic pKa, in vivo delivery efficiency, apparent pKa, and other key parameters of the generated molecules, and screening the molecules that meet the preset conditions to construct the fine-tuning data set;
[0181] ③ Fine-tuning: use the constructed fine-tuning data set to fine-tune the model to update the model parameters and further optimize the conditional generation capability;
[0182] (iv) Iterative generation: using the updated model to generate a new batch of candidate molecules, and repeating the performance evaluation and fine-tuning steps until the preset termination condition is met. In this embodiment, the termination condition can include but is not limited to the following cases: the number of generated molecules reaches the expected target, or the performance indicators of the generated molecules meet certain requirements.
[0183] To guide the in vivo delivery efficiency and apparent pKa evaluation in the guided condition generation, multiple LNP property prediction models are built for integrated prediction to improve hit accuracy. The characterization of molecules uses molecular fingerprints, molecular descriptors, and graph structure characterization methods based on graph neural networks (GNN), and algorithms include LightGBM (LGBM), support vector machines (SVM), and multiple layer perceptrons (MLP), etc. The model performance ROC-AUC and PR-AUC are as shown in Figure 6
[0184] In this embodiment, the integrated prediction refers to considering the results of multiple prediction models to improve the accuracy and reliability of the prediction. Specifically, for the evaluation of in vivo delivery efficiency and apparent pKa, the decision criteria for integrated prediction are as follows:
[0185] (I) In vivo delivery efficiency evaluation: an ionizable lipid molecule is determined to have excellent in vivo delivery efficiency only if all prediction models predict that the in vivo delivery efficiency of the ionizable lipid molecule prepared into LNP according to the standard LNP formulation (molar ratio of 50:10:38.5:1.5 of target lipid, DSPC, cholesterol and PEG2000-DMG) is better than that of the reference molecule MC3. This determination adopts a "one-vote veto" mechanism, that is, as long as one model predicts that the result does not meet the requirements, the molecule is determined to not meet the in vivo delivery efficiency condition.
[0186] (II) Apparent pKa evaluation: an ionizable lipid molecule is determined to have an ideal apparent pKa value only if all prediction models predict that the apparent pKa of the ionizable lipid molecule prepared into LNP according to the standard LNP formulation (molar ratio of 50:10:38.5:1.5 of target lipid, DSPC, cholesterol and PEG2000-DMG) is within the range of 6 to 7. Again, a "one-vote veto" mechanism is adopted to ensure that the selected molecules have highly reliable apparent pKa prediction results.
[0187] (III) Intrinsic pKa evaluation: A generated ionizable lipid molecule is determined to have a suitable intrinsic pKa value if and only if all the prediction models predict that the intrinsic pKa of the molecule is within the range of 6 to 10.
[0188] Through the above strict integrated prediction judgment criteria, the embodiment can effectively screen ionizable lipid molecules that simultaneously satisfy the in vivo delivery efficiency better than MC3, the apparent pKa within the range of 6 to 7, and the intrinsic pKa within the range of 6 to 10, significantly improving the reliability of the screening results and the success rate of subsequent experiments.
[0189] In the embodiment, the integrated prediction results are directly used to construct the fine-tuning dataset, and only molecules that pass the integrated prediction judgment are included in the fine-tuning dataset for subsequent optimization of the generation model. This screening mechanism ensures the high quality of the feedback signal, effectively guides the generation model to optimize towards the target properties, and improves the hit rate of target molecules.
[0190] Step three: screening of candidate molecules. Based on the model generated in step two, a number of candidate ionizable lipid molecules are generated, and high-scoring candidate molecules are obtained through screening for experimental verification. In the embodiment, a systematic molecular screening step is adopted, including:
[0191] (A) Using the fine-tuned LipidGPT model to perform 1,000 samplings to obtain 990 molecules with effective chemical structures;
[0192] (B) Screening out lipids from the 990 molecules that do not appear in the known ionizable lipid structure dataset to obtain 962 novel ionizable lipids;
[0193] (C) Performance prediction (synchronous step two integrated prediction) is performed on the 962 novel ionizable lipids to screen out ionizable lipid molecules with intrinsic pKa between 6 to 10, in vivo delivery efficiency better than MC3, and apparent pKa between 6 to 7, obtaining 889 lipids meeting the conditions;
[0194] (D) Calculating the ECFP fingerprint of the 889 lipids meeting the conditions and the Tanimoto similarity of all positive ionizable lipids in the in vivo delivery efficiency dataset;
[0195] (E) Filtering out lipids with Tanimoto coefficient greater than 0.8 to ensure molecular diversity;
[0196] (F) Using K-Means clustering method to classify the lipids obtained in step (E) into 8 clusters;
[0197] (G) Randomly selecting 5 lipids from each cluster, a total of 40 lipids for synthesis accessibility evaluation;
[0198] (H) Synthetic feasibility score of the 40 lipids using the SAscoreLNP scoring tool and further evaluation by a chemistry expert;
[0199] (I) Structural modification of some lipids by a chemistry expert based on raw material availability and re-evaluation of the modified molecules;
[0200] (J) Final selection of 14 lipids for synthesis.
[0201] The present example provides a method for the development of a scoring tool, SAscoreLNP, for the evaluation of the synthetic feasibility of ionizable lipids, based on the SAscore development method published by Ertl and Schuffenhauer (Ertl P, Schuffenhauer A. Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. J Cheminform. 2009 Jun 10; 1(1): 8. doi: 10.1186 / 1758-2946-1-8) with improvements specifically optimized for the characteristics of ionizable lipids. The method comprises the following steps:
[0202] (1) Representative samples were selected from the library of ionizable lipid structures as external validation; (2) Two synthetic experts scored the representative samples for synthetic difficulty, with a four-point scale (1, 2, 3, 4), where higher scores indicate greater synthetic difficulty; expert scores take into account factors such as number of synthetic steps, expected yield of key reactions, commercial availability of starting materials, harshness of reaction conditions, and difficulty of product purification; (3) ECFP molecular fingerprints were calculated for the dataset of ionizable lipid structures constructed in Example 1 using the RDKit toolkit, along with molecular complexity metrics including number of ring systems, ring size, degree of ring fusion, number and type of functional groups, number of chiral centers, and molecular symmetry. Fragment contribution calculation: based on the ECFP molecular fingerprints, molecules were decomposed into fragments, and by analyzing large-scale ionizable lipid synthesis literature, the frequency of each fragment in the synthesized molecules was calculated, and rare fragments were assigned a higher synthetic difficulty weight, and common fragments were assigned a lower synthetic difficulty weight; (4) Synthetic accessibility scores were calculated based on molecular complexity and ECFP fragment contribution; (5) The scores were normalized to the range of 0 to 4, consistent with the expert scoring scale; (6) Two versions of the scoring tool were developed, SAscoreLNPr2 based on ECFP4 molecular fingerprints with a radius of 2, and SAscoreLNPr3 based on ECFP6 molecular fingerprints with a radius of 3; (7) External validation was used to calculate the correlation between model predictions and actual synthetic difficulty, and the results are shown in Figure 7
[0203] The SAscoreLNP scoring tool developed by the above method can objectively and accurately evaluate the synthetic feasibility of ionizable lipid molecules, providing a scientific basis for candidate molecule screening, and significantly improving the efficiency of new ionizable lipids from design to synthesis. In this example, the SAscoreLNP scoring tool was successfully applied to the candidate molecule screening process, and lipids with a score less than 2 were considered to be low in synthetic difficulty. After preliminary calculation and screening, the remaining candidate lipids were handed over to chemical experts for further evaluation of their synthetic feasibility. During the evaluation process, due to the limitation of the availability of raw materials, structural modifications were made to some ionizable lipids. The intrinsic pKa, in vivo delivery efficiency, and apparent pKa of the molecules after structural modification were re-predicted. Although some modified molecules were predicted to be unqualified, these molecules were still synthesized to verify the accuracy of the prediction model. Finally, 14 lipids were selected for synthesis.
[0204] Example 4: Ionizable lipids obtained based on the ionizable lipid design method
[0205] In this example, a plurality of useful ionizable lipid molecules obtained by the design method of ionizable lipids in Example 3 are listed, and the corresponding synthesis and characterization information thereof are provided.
[0206] (1) AL-1:
[0207] The synthetic route of AL-1 is shown as follows:
[0208]
[0209] The specific preparation method is as follows:
[0210] A mixture of triethanolamine (1a, 1.4919 g, 10 mmol), 2-heptylundecanoic acid (1b, 7.1120 g, 25 mmol), 1,3-dicyclohexyl carbodiimide (DCC) (6.1899 g, 30 mmol) and 4-dimethylaminopyridine (DMAP) (0.2443 g, 2 mmol) was dissolved in anhydrous dichloromethane (DCM) (100 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was removed under reduced pressure to remove the solvent, and the residue was purified by silica gel chromatography (petroleum ether (PE): ethyl acetate (EA) = 5:1) to obtain compound 1 (4.4651 g, yield 65%) as a yellow oil.
[0211] Compound 1 (1.0232 g, 1.5 mmol), 4-dimethylamino butyric acid hydrochloride (1c, 0.3352 g, 2 mmol), DCC (0.6190 g, 3 mmol) and DMAP (0.0611 g, 0.5 mmol) were dissolved in a mixed anhydrous solvent of DCM / N,N-dimethylformamide (DMF) (volume ratio 4:1, 25 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was removed under reduced pressure to remove the solvent, and the residue was purified by silica gel chromatography (DCM: methanol (MeOH) = 40:1) to obtain AL-1 (0.8709 g, yield 73%) as a yellow oil. The 1 The H-NMR spectrum is as shown in Figure 8 .
[0212] (2) AL-2:
[0213] The synthetic route of AL-2 is shown as follows:
[0214]
[0215] The specific preparation method is as follows:
[0216] A mixture of 8-bromooctanoic acid (2a, 4.4622 g, 20 mmol), heptadecan-9-ol (2b, 5.1294 g, 20 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (6.1899 g, 30 mmol) and 4-dimethylaminopyridine (DMAP) (0.2443 g, 2 mmol) was dissolved in anhydrous dichloromethane (DCM) (100 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was removed under reduced pressure and the residue was purified by silica gel chromatography (petroleum ether (PE): dichloromethane (DCM) = 8: 1) to obtain compound 2 (6.6308 g, yield 74%) as a colorless liquid.
[0217] A mixture of compound 2 (2.3078 g, 5 mmol), 3-aminopropanol (2c, 0.7511 g, 10 mmol) and N-ethyldiisopropylamine (DIPEA, 0.6462 g, 5 mmol) was dissolved in isopropyl alcohol (Me2CHOH) (50 mL) and stirred at 80°C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (DCM: methanol (MeOH) = 10: 1) to obtain compound 3 (1.4736 g, yield 65%) as a yellow oil. AL-2 was eluted as a yellow oil by-product (0.3506 g, 8%) with DCM / MeOH (30: 1) as the eluent. The H-NMR spectrum of AL-2 is shown in Figure 1. 1 H-NMR spectrum is shown in Figure 1. Figure 9
[0218] (3) AL-3:
[0219] The synthetic route of AL-3 is shown below:
[0220]
[0221] The specific preparation method is as follows:
[0222] A mixture of ethanolamine (3a, 1.2216 g, 20 mmol), compound 2 (2.3078 g, 5 mmol) and N-ethyldiisopropylamine (DIPEA, 0.6462 g, 5 mmol) was dissolved in isopropyl alcohol (50 mL) and stirred at 70°C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (dichloromethane (DCM): methanol (MeOH) = 15: 1) to obtain compound 4 (1.2535 g, yield 57%) as a yellow oil.
[0223] A mixture of 4-pentylcyclohexanol (3b, 0.8514 g, 5 mmol), 4-bromobutyric acid (3c, 0.8350 g, 5 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (2.0633 g, 10 mmol) and 4-dimethylaminopyridine (DMAP) (0.1222 g, 1 mmol) was dissolved in anhydrous dichloromethane (DCM) (40 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was freed from the solvent under reduced pressure, and the residue was purified by silica gel chromatography (petroleum ether (PE): dichloromethane (DCM) = 5:1) to obtain compound 5 (1.2568 g, yield 79%) as a colorless liquid.
[0224] A mixture of compound 4 (0.8835 g, 2 mmol), compound 5 (0.7982 g, 2.5 mmol) and N-ethyldiisopropylamine (DIPEA, 0.6462 g, 5 mmol) was dissolved in isopropanol (30 mL) and stirred at 80°C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (DCM: MeOH = 50:1) to obtain AL-3 (0.7058 g, yield 52%) as a yellow oil. 1 H-NMR spectrum Figure 10 shown.
[0225] (3) AL-4:
[0226] The synthetic route of AL-4 is as follows:
[0227]
[0228] Wherein, the specific preparation method is:
[0229] A mixture of 4-decenoic acid (4a, 0.8512 g, 5 mmol), 7-bromo-1-heptanol (4b, 0.9755 g, 5 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (2.0633 g, 10 mmol) and 4-dimethylaminopyridine (DMAP) (0.1222 g, 1 mmol) was dissolved in anhydrous dichloromethane (DCM) (40 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was freed from the solvent under reduced pressure, and the residue was purified by silica gel chromatography (petroleum ether (PE): dichloromethane (DCM) = 5:1) to obtain compound 6 (1.5995 g, yield 92%) as a colorless liquid.
[0230] A mixture of compound 3 (0.4558 g, 1 mmol), compound 6 (0.6947 g, 2 mmol) and N-ethyldiisopropylamine (DIPEA, 0.5170 g, 4 mmol) was dissolved in isopropanol (20 mL) and stirred at 80 °C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (DCM:MeOH = 50:1) to give AL-4 (0.4289 g, yield 59%) as a yellow oil. The structure of AL-4 was confirmed by1H-NMR spectrum as shown in Figure 1. 1 H-NMR spectrum as shown in Figure 1. Figure 11 .
[0231] (5) AL-5:
[0232] The synthetic route of AL-5 is shown below:
[0233]
[0234] The specific preparation method is as follows:
[0235] A mixture of 4-pentylcyclohexanol (5a, 1.7029 g, 10 mmol) and triethylamine (TEA) (2.0258 g, 20 mmol) was dissolved in anhydrous dichloromethane (DCM) (30 mL), and acryloyl chloride (5b, 1.3576 g, 15 mmol) dissolved in anhydrous dichloromethane (20 mL) was slowly added dropwise under ice bath cooling. Then it was stirred in an ice water bath for 4 hours. The reaction solution was filtered to remove the suspended solids, and then purified by silica gel chromatography (petroleum ether (PE): dichloromethane (DCM) = 3:1) to give compound 7 (0.9517 g, yield 42%) as a colorless liquid.
[0236] A mixture of compound 7 (0.8972 g, 4 mmol) and ethanolamine (5c, 0.6108 g, 10 mmol) was dissolved in isopropanol (20 mL) and stirred at 70 °C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (DCM:MeOH = 30:1) to give compound 8 (1.0545 g, yield 87%) as a yellow oil.
[0237] A mixture of 2-hexyldecanoic acid (5d, 1.7949 g, 7 mmol), 7-bromo-1-heptanol (5e, 1.3657 g, 7 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (3.0950 g, 15 mmol), and 4-dimethylaminopyridine (DMAP) (0.2443 g, 2 mmol) was dissolved in anhydrous dichloromethane (DCM) (40 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was freed from the solvent under reduced pressure, and the residue was purified by silica gel chromatography (PE:DCM = 4:1) to obtain compound 9 (2.4511 g, 81% yield) as a colorless liquid.
[0238] A mixture of compound 8 (1.0635 g, 3.5 mmol), compound 9 (1.7340 g, 4 mmol) and N-ethyldiisopropylamine (DIPEA) (1.0339 g, 8 mmol) was dissolved in isopropanol (30 mL) and stirred at 80°C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (DCM: MeOH = 60:1) to obtain AL-5 (0.9701 g, yield 43%) as a yellow oil. 1 H-NMR spectrum Figure 12 shown.
[0239] (6) AL-6:
[0240]
[0241] The synthetic route of AL-6 is as follows:
[0242]
[0243] Among them, the specific preparation method is:
[0244] A mixture of 2-hexyldecan-1-ol (6a, 1.6972 g, 7 mmol), 7-bromoheptanoic acid (6b, 1.4636 g, 7 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (3.0950 g, 15 mmol) and 4-dimethylaminopyridine (DMAP) (0.2443 g, 2 mmol) was dissolved in anhydrous dichloromethane (DCM) (40 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was freed from the solvent under reduced pressure, and the residue was purified by silica gel chromatography (petroleum ether (PE): dichloromethane (DCM) = 7:1) to obtain compound 10 (2.5410 g, yield 84%) as a colorless liquid.
[0245] A mixture of compound 10 (1.1344 g, 2.5 mmol), pyrrolidin-3-ylmethanol (6c, 0.4046 g, 4 mmol) and N-ethyldiisopropylamine (DIPEA) (1.0339 g, 8 mmol) was dissolved in isopropanol (30 mL) and stirred at 80 °C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (DCM:MeOH = 40:1) to give compound 11 (1.2759 g, yield 70%) as a yellow oil.
[0246] A mixture of compound 11 (1.2759 g, 2.5 mmol), nonanoic acid (6d, 0.6329 g, 4 mmol), DCC (1.0316 g, 5 mmol) and DMAP (0.0611 g, 0.5 mmol) was dissolved in anhydrous dichloromethane (DCM) (30 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was removed under reduced pressure and the residue was purified by silica gel chromatography (DCM:MeOH = 50:1) to give AL-6 (0.6468 g, yield 44%) as a colorless liquid. The H-NMR spectrum of AL-6 is shown in FIG. 1. 1 H-NMR spectrum is shown in FIG. 1. Figure 13
[0247] (7) AL-7:
[0248] The synthetic route of AL-7 is shown below:
[0249]
[0250] The specific preparation method is as follows:
[0251] A mixture of 2-butyl-1-octanol (7a, 1.3044 g, 7 mmol), 4-bromobutyric acid (7b, 1.1690 g, 7 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (3.0950 g, 15 mmol) and 4-dimethylaminopyridine (DMAP) (0.2443 g, 2 mmol) was dissolved in anhydrous dichloromethane (DCM) (40 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was removed under reduced pressure and the residue was purified by silica gel chromatography (petroleum ether (PE): dichloromethane (DCM) = 6:1) to give compound 12 (2.0663 g, yield 88%) as a colorless liquid.
[0252] A mixture of compound 12 (1.6683 g, 5 mmol), 4-amino-1-butanol (7c, 0.1337 g, 1.5 mmol) and N-ethyldiisopropylamine (DIPEA) (0.6462 g, 5 mmol) was dissolved in isopropanol (30 mL) and stirred at 80 °C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (DCM:MeOH = 40:1) to give AL-7 (0.3328 g, yield 38%) as a yellow oil. The H-NMR spectrum of AL-7 is shown in FIG. 1. 1 H-NMR spectrum of AL-7 is shown in FIG. 1. Figure 14
[0253] (8) AL-8:
[0254] The synthetic route of AL-8 is shown below.
[0255]
[0256] wherein the specific preparation method is:
[0257] A mixture of 3-dimethylaminopropylamine (8a, 1.5327 g, 15 mmol), compound 2 (3.4777 g, 7 mmol) and N-ethyldiisopropylamine (DIPEA) (1.0399 g, 8 mmol) was dissolved in isopropanol (40 mL) and stirred at 70 °C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (DCM:MeOH = 20:1) to give compound 13 (1.5054 g, yield 45%) as a white semi-solid.
[0258] A mixture of compound 13 (1.4485 g, 3 mmol), nonanoic acid (8c, 0.7812 g, 5 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (1.0316 g, 5 mmol) and 4-dimethylaminopyridine (DMAP) (0.1222 g, 1 mmol) was dissolved in anhydrous dichloromethane (DCM) (30 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solid was removed by filtration. The filtrate was removed under reduced pressure and the residue was purified by silica gel chromatography (DCM:MeOH = 40:1) to give AL-8 (1.7198 g, yield 92%) as a yellow oil. The H-NMR spectrum of AL-8 is shown in FIG. 2. 1 H-NMR spectrum of AL-8 is shown in FIG. 2. Figure 15
[0259] (9) AL-9:
[0260] The synthetic route of AL-9 is shown below.
[0261]
[0262] wherein the specific preparation method is as follows:
[0263] A mixture of 6-bromohexanoic acid (9a, 1.3654 g, 7 mmol), 2-hexyldecan-1-ol (9b, 1.6972 g, 7 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (3.0950 g, 15 mmol) and 4-dimethylaminopyridine (DMAP) (2443 g, 2 mmol) was dissolved in anhydrous dichloromethane (DCM) (40 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was removed under reduced pressure to remove the solvent, and the residue was purified by silica gel chromatography (petroleum ether (PE): dichloromethane (DCM) = 7:1) to obtain compound 14 (2.6912 g, yield 92%) as a colorless liquid.
[0264] A mixture of compound 14 (0.8390 g, 2 mmol), cyclohexylamine hydrochloride (9c, 1.3564 g, 10 mmol) and N-ethyldiisopropylamine (DIPEA) (1.2924 g, 10 mmol) was dissolved in isopropanol (40 mL) and stirred at 80°C overnight. After the reaction was completed, the solvent was removed under vacuum, and the residue was purified by silica gel chromatography (DCM:MeOH = 50:1) to obtain compound 15 (0.3287 g, yield 38%) as a white semi-solid.
[0265] A mixture of compound 15 (0.3064 g, 0.7 mmol), 2-decyl-oxirane (9d, 0.7373 g, 4 mmol) and N-ethyldiisopropylamine (DIPEA) (0.3877 g, 3 mmol) was dissolved in isopropanol (30 mL) and stirred at 90°C overnight. After the reaction was completed, the solvent was removed under vacuum, and the residue was purified by silica gel chromatography (DCM:MeOH = 50:1) to obtain AL-9 (0.2449 g, yield 56%) as a yellow oil. The H-NMR spectrum of AL-9 is shown in FIG. 1. 1 H-NMR spectrum is shown in FIG. 1. Figure 16
[0266] (10) AL-10:
[0267]
[0268] The synthetic route of AL-10 is shown below:
[0269]
[0270] wherein the specific preparation method is as follows:
[0271] A mixture of compound 16 (0.4177 g, 1 mmol), 2-hexyldecanoic acid (10c, 1.5385 g, 6 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (1.2380 g, 6 mmol) and 4-dimethylaminopyridine (DMAP) (0.1222 g, 1 mmol) was dissolved in anhydrous dichloromethane (DCM) (30 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was removed of solvent under reduced pressure and the residue was purified by silica gel chromatography (DCM:MeOH = 70:1) to obtain AL-10 (0.8365 g, yield 81%) as a white semi-solid. The H-NMR spectrum of AL-10 is shown in FIG. 1.
[0272] A mixture of compound 16 (0.4177 g, 1 mmol), 2-hexyldecanoic acid (10c, 1.5385 g, 6 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (1.2380 g, 6 mmol) and 4-dimethylaminopyridine (DMAP) (0.1222 g, 1 mmol) was dissolved in anhydrous dichloromethane (DCM) (30 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was removed of solvent under reduced pressure and the residue was purified by silica gel chromatography (DCM:MeOH = 70:1) to obtain AL-10 (0.8365 g, yield 81%) as a white semi-solid. The H-NMR spectrum of AL-10 is shown in FIG. 1. 1 The H-NMR spectrum of AL-10 is shown in FIG. 1. Figure 17
[0273] (11) AL-11:
[0274]
[0275] The synthetic route of AL-11 is shown below:
[0276]
[0277] The specific preparation method is as follows:
[0278] A mixture of compound 16 (0.4177 g, 1 mmol), 2-hexyldecanoic acid (10c, 1.5385 g, 6 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (1.2380 g, 6 mmol) and 4-dimethylaminopyridine (DMAP) (0.1222 g, 1 mmol) was dissolved in anhydrous dichloromethane (DCM) (30 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was removed of solvent under reduced pressure and the residue was purified by silica gel chromatography (DCM:MeOH = 70:1) to obtain AL-10 (0.8365 g, yield 81%) as a white semi-solid. The H-NMR spectrum of AL-10 is shown in FIG. 1. 1 The H-NMR spectrum of AL-10 is shown in FIG. 1. Figure 18
[0279] (12) AL-12:
[0280]
[0281] The synthetic route of AL-12 is shown below:
[0282]
[0283] The specific preparation method is:
[0284] A mixture of 2-butyl-1-octanol (12a, 1.3044 g, 7 mmol), 8-bromooctanoic acid (12b, 1.5618 g, 7 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (3.0950 g, 15 mmol) and 4-dimethylaminopyridine (DMAP) (0.2443 g, 2 mmol) was dissolved in anhydrous dichloromethane (DCM) (40 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was removed under reduced pressure to remove the solvent, and the residue was purified by silica gel chromatography (petroleum ether (PE): dichloromethane (DCM) = 7:1) to obtain compound 17 (2.2030 g, yield 80%) as a colorless liquid.
[0285] A mixture of compound 17 (1.1795 g, 3 mmol), N-methylethylenediamine (12c, 0.0371 g, 0.5 mmol) and N-ethyldiisopropylamine (DIPEA) (0.2585 g, 2 mmol) was dissolved in isopropyl alcohol (20 mL) and stirred at 80°C overnight. After the reaction was completed, the solvent was removed under vacuum, and the residue was purified by silica gel chromatography (DCM:MeOH = 30:1) to obtain AL-12 (0.1414 g, yield 28%) as a yellow oil. The 1 H-NMR spectrum is as Figure 19 .
[0286] (13) AL-13:
[0287] The synthetic route of AL-13 is shown below:
[0288]
[0289] The specific preparation method is:
[0290] A mixture of compound 18 (1.6780 g, 4 mmol), 1-butylamine (13c, 0.0731 g, 1 mmol) and N-ethyldiisopropylamine (DIPEA) (0.6462 g, 5 mmol) was dissolved in isopropanol (30 mL) and stirred at 80 °C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (DCM:MeOH = 60:1) to give AL-13 (0.2176 g, yield 29%) as a yellow oil. The structure of AL-13 was confirmed by1H-NMR spectrum as shown in Figure 6.
[0291] A mixture of compound 18 (1.6780 g, 4 mmol), 1-butylamine (13c, 0.0731 g, 1 mmol) and N-ethyldiisopropylamine (DIPEA) (0.6462 g, 5 mmol) was dissolved in isopropanol (30 mL) and stirred at 80 °C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (DCM:MeOH = 60:1) to give AL-13 (0.2176 g, yield 29%) as a yellow oil. The structure of AL-13 was confirmed by1H-NMR spectrum as shown in Figure 6. 1 H-NMR spectrum as shown in Figure 6. Figure 20 .
[0292] (14) AL-14:
[0293] The synthetic route of AL-14 is shown below:
[0294]
[0295] wherein the specific preparation method is:
[0296] A mixture of 2-heptylundecanoic acid (14a, 2.8448 g, 10 mmol), 8-bromo-1-octanol (14b, 2.0913 g, 10 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (3.0950 g, 15 mmol) and 4-dimethylaminopyridine (DMAP) (0.2443 g, 2 mmol) was dissolved in dry dichloromethane (DCM) (60 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was removed under reduced pressure and the residue was purified by silica gel chromatography (petroleum ether (PE): dichloromethane (DCM) = 9:1) to give compound 19 (3.1900 g, yield 71%) as a colorless liquid.
[0297] A mixture of compound 19 (0.8951 g, 2 mmol) and ethanolamine (14c, 0.7330 g, 12 mmol) was dissolved in isopropanol (30 mL) and stirred at 70 °C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (DCM:MeOH = 10:1) to give compound 20 (0.6834 g, yield 80%) as a yellow oil.
[0298] A mixture of 8-bromooctanoic acid (14d, 2.2311 g, 10 mmol), 2-octanol (14e, 1.3023 g, 10 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (3.0950 g, 15 mmol) and 4-dimethylaminopyridine (DMAP) (0.2443 g, 2 mmol) was dissolved in anhydrous dichloromethane (DCM) (60 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solid was removed by filtration. The filtrate was removed under reduced pressure and the residue was purified by silica gel chromatography (petroleum ether (PE): dichloromethane (DCM) = 5:1) to give compound 21 (2.5344 g, yield 76%) as a colorless liquid.
[0299] A mixture of compound 20 (0.6416 g, 1.5 mmol), compound 21 (1.6766 g, 5 mmol) and N-ethyldiisopropylamine (DIPEA) (0.6462 g, 5 mmol) was dissolved in isopropanol (30 mL) and stirred at 90 °C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (DCM:MeOH = 30:1) to give AL-14 (0.7050 g, yield 69%) as a yellow oil. The H-NMR spectrum of AL-14 is shown in FIG. 1. 1 H-NMR spectrum is shown in FIG. 1. Figure 21
[0300] (15) AL-15:
[0301] The synthetic route of AL-15 is shown below:
[0302]
[0303] wherein the specific preparation method is:
[0304] A mixture of compound 19 (0.8951 g, 2 mmol) and 4-amino-1-butanol (15a, 1.0694 g, 12 mmol) was dissolved in isopropanol (30 mL) and stirred at 70 °C overnight. After the reaction was completed, the solvent was removed under vacuum and the residue was purified by silica gel chromatography (DCM:MeOH = 10:1) to give compound 22 (0.6553 g, yield 72%) as a yellow oil.
[0305] A mixture of 10-bromodecanoic acid (15b, 1.2558 g, 5 mmol), heptadecan-9-ol (14c, 1.2824 g, 5 mmol), 1,3-dicyclohexylcarbodiimide (DCC) (1.6506 g, 8 mmol) and 4-dimethylaminopyridine (DMAP) (0.1222 g, 1 mmol) was dissolved in anhydrous dichloromethane (DCM) (40 mL) and stirred at room temperature overnight. After the reaction was completed, the suspended solids were removed by filtration. The filtrate was removed under reduced pressure to remove the solvent, and the residue was purified by silica gel chromatography (petroleum ether (PE): dichloromethane (DCM) = 10:1) to obtain compound 23 ((2.0681 g, yield 84%) as a colorless liquid.
[0306] A mixture of compound 22 (0.6381 g, 1.4 mmol), compound 23 (1.4689 g, 3 mmol) and N-ethyldiisopropylamine (DIPEA) (0.6462 g, 5 mmol) was dissolved in isopropyl alcohol (30 mL) and stirred at 90°C overnight. After the reaction was completed, the solvent was removed under vacuum, and the residue was purified by silica gel chromatography (DCM:MeOH = 30:1) to obtain AL-15 (0.7068 g, yield 57%) as a yellow oil. The H-NMR spectrum of AL-15 is shown in Figure 1. 1 H-NMR spectrum is shown in Figure 1. Figure 22 .
[0307] Example 5: Experimental verification of ionizable lipids obtained based on the ionizable lipid generation method
[0308] In this example, when testing the in vivo effects of LNP of AL-1 to AL-15 of Example 4, ionizable lipids DLin-MC3-DMA (abbreviated as MC3) and SM102 were selected as controls.
[0309] The specific experimental method is as follows:
[0310] Characterization of LNP:
[0311] Luciferase mRNA was synthesized from a linearized DNA template by T7 RNA polymerase-mediated in vitro transcription. During the synthesis process, modified uridine-5'-triphosphate (1-methyl pseudouridine) was introduced into the mRNA molecule at a specific ratio, and this modification can significantly reduce the immunogenicity of mRNA. At the same time, Cap1 structure was also used in the synthesis process, and the introduction of this 5' end cap structure is mainly to improve the translation efficiency of mRNA.
[0312] The preparation of LNP adopted microfluidic technology, by precisely controlling the mixing process of lipids and mRNA, to ensure the uniformity and high efficiency of encapsulation of nanoparticles. During the preparation process, first, four key ingredients were co-dissolved in the alcohol phase according to the specific molar ratio: ionizable lipid, DSPC (distearoylphosphatidylcholine), cholesterol and DMG-PEG (dimyristoylglycerol-polyethylene glycol), the molar ratio of which is 50:10:38.5:1.5, and the total lipid concentration is about 10 mM. The aqueous phase is composed of mRNA dissolved in 25 mM sodium acetate buffer at pH 5.0. Subsequently, the aqueous phase and the alcohol phase are mixed by a microfluidic device at a volume ratio of 3:1, the total flow rate is 12 mL / min, and the molar ratio of ionizable lipid to nucleotide is 6:1. After mixing, the preparation is dialyzed in a 20 mM Tris-acetate buffer (pH ~ 7.5) through a Slide-A-Lyzer dialysis bag (10 kDa molecular weight cut-off, Thermo Fisher) to remove the alcohol solvents used during preparation. After dialysis is completed, the solution is concentrated using an Amicon ultrafiltration centrifuge tube (10 kDa molecular weight cut-off, Millipore), and finally filtered through a 0.22 μm filter to ensure the sterility and uniformity of the final preparation.
[0313] The particle size and polydispersity index (PDI) of LNP were measured by Zetasizer Pro instrument of Malvern company. The encapsulation efficiency (EE, %) of mRNA in LNP was determined by Quant-iT RiboGreen RNA assay kit (Invitrogen, Cat# R11490). The concentration of mRNA was quantified by Stunner analyzer of Unchained Labs company. The apparent pKa of LNP was determined by fluorescence-based TNS (6-(4-methylaminophenyl) naphthalene-2-sulfonic acid) assay.
[0314] Figure 23A The particle size, PDI, EE of newly synthesized ionizable lipid LNP are shown; Figure 23B The apparent pKa is shown. The results show that, in addition to AL-10 LNP, all other ionizable lipid LNPs exhibit suitable particle size, low PDI value and high mRNA encapsulation rate.
[0315] In vivo experiments:
[0316] In this example, BALB / c mice were used as experimental animal models, and lipid nanoparticles (LNPs) encapsulating firefly luciferase (FLuc) mRNA were administered via intravenous (IV) or intramuscular (IM) injection, with a dose of 5 pg mRNA per mouse. At specific time points (4 hours, 24 hours, and 48 hours) after administration, mice received 100 pL of D-luciferin (30 mg / mL in PBS) as a substrate via intraperitoneal injection, and the luminescent signal in vivo was detected using an IVIS Spectrum imaging system (PerkinElmer). The results of in vivo fluorescence signal measurement are shown in FIG. 8. Figures 24A-24C
[0317] The experimental results show that AL-2, AL-3, AL-4, AL-6, AL-12, AL-14, and AL-15 exhibit mRNA delivery efficiency comparable to that of SM-102 LNPs. For those lipids that were predicted to be unqualified after chemical modification (AL-9, AL-11, AL-13), they indeed showed lower in vivo delivery efficiency in the experimental results, highlighting the accuracy of the prediction model.
[0318] The similarity between the AI-designed lipids and the true positive samples was further calculated. The results showed that 7 out of 15 lipids had a similarity lower than 0.7 with the true positive samples. It is particularly noteworthy that AL-6, a lipid generated entirely by the method in Example 2 without chemical modification, exhibited the lowest similarity (0.52). These results not only verify the effectiveness of the AI-driven lipid design method, but also indicate that significant success has been achieved in developing new ionizable lipids for mRNA delivery.
[0319] In addition, the results revealed a significant correlation between the apparent pKa and in vivo delivery performance. Except for AL-1, all LNP formulations with an apparent pKa value within the optimal range (6-7) exhibited significantly higher delivery efficiency in vivo. This finding highlights the importance of pKa as a key parameter for predicting LNP delivery efficiency, providing valuable guidance for future design.
[0320] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for constructing a lipid nanoparticle property prediction model, characterized in that: The method comprises the following steps: obtaining a pre-training model pre-trained by using an ionizable lipid structure dataset; constructing a multi-task learning framework according to the pre-training model, fine-tuning the multi-task learning framework by using an LNP property related dataset, and obtaining a multi-task learning model capable of predicting LNP properties.
2. The construction method according to claim 1, characterized in that, The pre-training model comprises any one or a combination of multiple of the following: RoBERTa, a molecular graph attention network, a graph convolutional neural network, a message passing neural network, and a Transformer model; Optionally, the method for constructing the pre-training model comprises: expanding the data amount of the ionizable lipid structure dataset according to a data enhancement strategy; pre-training the pre-training model by using the expanded ionizable lipid structure dataset, and obtaining a trained pre-training model; Optionally, the data enhancement strategy comprises randomly masking part of the molecular graph structure and / or enumerating a SMILES string; Optionally, the LNP properties comprise any one or multiple of the following: in vivo delivery efficiency, apparent pKa value, particle size, polydispersity index, and encapsulation efficiency; Optionally, the LNP property related dataset comprises any one or multiple of the following: the LNP properties, an LNP formulation, a type of nucleic acid drug carried by the LNP, a type of experimental animal, a route of administration, a dose of administration, and an expression level of the nucleic acid drug; Optionally, the LNP formulation comprises a chemical structure of the ionizable lipid and / or a type of auxiliary lipid.
3. A method of predicting properties of a lipid nanoparticle, characterized in that, The method comprises the following steps: inputting a chemical structure of an ionizable lipid to be predicted into the model constructed by the method for constructing according to claim 1 or 2, to obtain a prediction result.
4. A method of designing an ionizable lipid, characterized in that, The method comprises the following steps: pre-training a deep learning generation model by using an ionizable lipid structure dataset, to generate an initial generation model capable of generating new ionizable lipid molecular structures; optimizing and feeding back the initial generation model, to guide the generation of ionizable lipid molecular structures meeting preset LNP properties; performing single sampling or multiple sampling by using the generation model after the optimization and feedback, to obtain candidate molecules.
5. The method of designing according to claim 4, wherein, The deep learning generation model comprises any one or multiple of the following: an encoder, a generative adversarial network, and a generative pre-training transformer network; Optionally, before the pre-training, the design method further comprises pre-processing the ionizable lipid structure dataset: removing molecules with a SMILES string length exceeding 140, and / or deleting ionizable lipid molecules in the ionizable lipid structure dataset that appear in the LNP property related dataset; Optionally, the method for optimizing and feeding back comprises any one or multiple of the following: reinforcement learning, Bayesian optimization, and bootstrap fine-tuning; Optionally, the step of optimizing and feeding back the initial generation model comprises: obtaining an initial generation result of the initial generation model; performing LNP property prediction on the initial result by using a lipid nanoparticle property prediction model, and screening ionizable lipids meeting preset LNP properties as a fine-tuning dataset; optimizing and feeding back the initial generation model by using the fine-tuning dataset, to obtain a generation model capable of generating ionizable lipids meeting preset LNP properties; Optionally, the LNP property prediction model is constructed by using any one or more of deep learning, random forest, LightGBM, support vector machine, and multilayer perceptron algorithms. Optionally, the LNP property includes any one or more of in vivo delivery efficiency, apparent pKa value, particle size, polydispersity index, and encapsulation efficiency. Optionally, the LNP property prediction model is constructed by using the construction method of claim 1 or 2.
6. The design method according to claim 4 or 5, characterized by, The design method further comprises a step of further screening the candidate molecules. Optionally, the screening comprises filtering the candidate molecules to obtain new ionizable lipid molecules as a first screening result. Optionally, the screening further comprises predicting the candidate molecules or the first screening result by using a LNP property prediction model, screening ionizable lipid molecules meeting a target LNP property as a second screening result. Optionally, the screening further comprises performing structural clustering or similarity analysis on the candidate molecules, the first screening result, or the second screening result to obtain different sub-classes, and selecting ionizable lipid molecules from the different sub-classes as a third screening result. Optionally, the screening further comprises scoring the candidate molecules, the first screening result, the second screening result, or the third screening result for synthesis difficulty, and screening molecules with low synthesis difficulty for synthesis.
7. An ionizable lipid molecule characterized in that, The ionizable lipid molecules are designed by the design method of any one of claims 4-6.
8. The ionizable lipid molecule of claim 7, wherein, The ionizable lipid molecules are selected from any one of AL-1, AL-2, AL-3, AL-4, AL-5, AL-6, AL-7, AL-8, AL-9, AL-10, AL-11, AL-12, AL-13, AL-14, and AL-15. AL-1 has the structural formula: The structural formula of AL-2 is: The structural formula of AL-3 is: AL-4 has the structural formula: AL-5 has the structural formula: AL-6 has the structural formula: The structural formula of AL-7 is: AL-8 has the structural formula: AL-9 has the structural formula: The structural formula of AL-10 is: The structural formula of AL-11 is: The structural formula of AL-12 is: The structural formula of AL-13 is: The structural formula of AL-14 is: The structural formula of AL-15 is:
9. A lipid nanoparticle or a delivery nucleic acid drug containing the lipid nanoparticle, characterized in that: The LNP contains the ionizable lipid molecules of claim 7 or 8. Optionally, the delivery nucleic acid drug comprises a nucleic acid drug and the LNP encapsulating the nucleic acid drug. Optionally, the nucleic acid drug comprises RNA and / or DNA. Optionally, the RNA comprises mRNA and / or siRNA.
10. Use of the ionizable lipid molecules of claim 7 or 8 in the preparation of a LNP or in the delivery of a nucleic acid drug.