Method, program, and apparatus for predicting delivery carrier composition
The method employs machine learning to predict delivery carrier compositions for target cells, addressing the complexity of conventional design methods by optimizing delivery amounts, thus enhancing efficiency and reducing reliance on trial and error.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- KK TOSHIBA
- Filing Date
- 2024-11-26
- Publication Date
- 2026-06-05
AI Technical Summary
Conventional methods for designing delivery carriers are complex and require significant trial and error to optimize constituent components for specific target cells, lacking a systematic approach.
A method, program, and apparatus using machine learning to predict a delivery carrier composition that satisfies a target value for the delivery amount to target cells, utilizing a delivery carrier prediction tool constructed with training datasets of different component compositions, non-target cell membrane compositions, and non-target cell delivery amounts.
Significantly reduces the need for experimental trials and human intervention, enabling efficient design of delivery carriers tailored for target cells by predicting optimal compositions.
Smart Images

Figure 2026092463000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method, program, and apparatus for predicting a delivery carrier composition.
Background Art
[0002] As a method for introducing an active agent into cells, a technique of using a delivery carrier encapsulating the active agent is known. Also, in this technique, it is known that the constituent components of the delivery carrier that are easily taken up differ depending on the type of cell. That is, it means that by optimizing the constituent components of the delivery carrier, excellent delivery efficiency and directivity with respect to a specific type of cell (hereinafter referred to as "target cell") can be achieved. Therefore, in order to obtain a delivery carrier suitable for target cells, it is necessary to design the constituent components of the delivery carrier so as to be appropriate conditions for the target cells.
[0003] Conventional design of delivery carriers has been carried out by manufacturing a large number of types of delivery carriers with different constituent components of the delivery carrier, measuring and confirming their introduction properties into target cells, inferring the contribution degree of each constituent component of the delivery carrier, and adjusting the constituent components based on human experience and skills. Therefore, the conventional design method is very complicated and requires a great deal of trial and error, and there has been a demand for a method, program, and apparatus that can more simply design a delivery carrier suitable for target cells.
Summary of the Invention
Problems to be Solved by the Invention
[0004] An object of the present invention is to provide a method, program, and apparatus that can more easily design a delivery carrier suitable for target cells.
Means for Solving the Problems
[0005] The method according to the embodiment is a method for designing a delivery carrier to obtain a delivery carrier composition that satisfies a target value for a predetermined delivery amount of a delivery substance to a target cell. The method comprises the steps of preparing a delivery carrier prediction tool constructed by machine learning using a delivery carrier composition dataset of multiple types of delivery carriers exhibiting different component compositions, a non-target cell membrane composition dataset of cell membrane lipid compositions of multiple types of cells other than target cells, and a non-target cell delivery amount dataset of delivery amounts of delivery substances by multiple types of delivery carriers to multiple types of cells other than target cells as training datasets, and obtaining a delivery carrier composition that satisfies a target value based on the cell membrane lipid composition of the target cell using the delivery carrier prediction tool.
[0006] The program according to this embodiment is a program that causes a computer to predict the composition of a delivery carrier that satisfies a target value for the amount of activator delivered to target cells. This program is constructed by machine learning using a delivery carrier composition dataset of multiple types of delivery carriers exhibiting different component compositions, a non-target cell membrane composition dataset of multiple types of cells other than target cells relating to cell membrane lipid composition, and a non-target cell delivery amount dataset of multiple types of delivery carriers relating to the amount of delivered substances delivered to multiple types of cells other than target cells as training datasets. This program causes a computer to predict the composition of a delivery carrier that satisfies a target value for the amount of activator delivered to target cells, based on the target cell membrane lipid composition dataset of target cells.
[0007] The apparatus according to this embodiment is a device for predicting the composition of delivery carriers that satisfies a target value for the amount of activator delivered to target cells. The apparatus comprises: an input unit configured to receive the cell membrane lipid composition of target cells from outside the device and transmit it to the calculation unit; a calculation unit configured to calculate the composition of delivery carriers that satisfies a target value for the amount of activator delivered to target cells based on the cell membrane lipid composition of target cells, using a delivery carrier prediction tool constructed by machine learning that uses a delivery carrier composition dataset of multiple types of delivery carriers exhibiting different component compositions, a non-target cell membrane composition dataset of multiple types of cells other than target cells, and a non-target cell delivery amount dataset of the amount of delivered substances delivered by multiple types of delivery carriers to multiple types of cells other than target cells as training datasets; and an output unit configured to display the calculation results of the calculation unit. [Brief explanation of the drawing]
[0008] [Figure 1] Figure 1 is an explanatory diagram showing an example of the configuration of the delivery carrier composition prediction device according to the embodiment. [Figure 2] Figure 2 is an explanatory diagram showing an example of the structure of a target cell lipid dataset. [Figure 3] Figure 3 is an explanatory diagram showing an example of the composition of a non-target cell lipid dataset. [Figure 4] Figure 4 is an explanatory diagram showing an example of the structure of a dataset related to the amount of activator delivered to non-target cells, where (a) is an example of a delivery carrier composition dataset and (b) is an example of a non-target cell delivery amount dataset. [Figure 5] Figure 5 is a flowchart showing an example of a delivery carrier design method. [Figure 6] Figure 6 is a flowchart showing an example of a delivery carrier design method. [Modes for carrying out the invention]
[0009] The following describes a method, program, and apparatus for predicting the delivery carrier composition of an embodiment, with reference to the drawings. Each figure is a schematic diagram of the embodiment to facilitate understanding, and its shape, dimensions, ratios, etc., may differ from the actual product. These can be appropriately modified in the following description and in reference to known technology.
[0010] • Delivery carrier composition prediction device A method for predicting the composition of a delivery carrier is carried out using an embodiment of the delivery carrier composition prediction device 1 illustrated in Figure 1. This prediction device 1 and prediction method are used to obtain a delivery carrier composition that satisfies a predetermined target value for the introduction rate of an activator to target cells. That is, this embodiment obtains a delivery carrier composition dataset 14 (hereinafter referred to as the "target carrier composition dataset") that satisfies a predetermined target value for the introduction amount, as will be described later.
[0011] In this specification, "delivery carrier" refers to any carrier capable of delivering activators into various types of cells. The delivery carrier may be any carrier used in a typical drug delivery system.
[0012] For example, delivery carriers are lipid nanoparticles known to form nano-sized capsules. Here, lipid nanoparticles (LNPs) are roughly spherical structures consisting of a lipid membrane formed by the arrangement of multiple lipid molecules via non-covalent bonds, such as liposomes, micelles, nanoemulsions, or solid lipid nanoparticles. The lipid membrane may be a lipid monolayer or a lipid bilayer. Furthermore, the lipid membrane may consist of a single layer or a multilayer. Nucleic acids may be encapsulated as an activator in the lumen of the center of the lipid nanoparticle.
[0013] Alternatively, the delivery carrier may be in the form of a cationic polymer, such as polyethyleneimine, or a cell membrane-permeable peptide. However, if the desired form of the delivery carrier has been determined, the method, program, and apparatus for predicting the delivery carrier composition of this embodiment may be implemented using only the dataset relating to that desired form of delivery carrier.
[0014] As shown in Figure 1, the prediction device 1 has an input unit 2, a calculation unit 3 connected to the input unit 2, and an output unit 4 connected to the calculation unit 3. The connections between the input unit 2, the calculation unit 3, and the output unit 4 can be made in any manner as long as information communication is possible.
[0015] In this specification, the prediction device 1 is defined as a single arithmetic device configured with an input unit 2, a calculation unit 3, and an output unit 4 as separate units. However, the description of the prediction device 1 may be interpreted as a single arithmetic system (hereinafter referred to as the "prediction system") composed of a combination of the input unit 2, the calculation unit 3, and the output unit 4, each being a separate device or system. In other words, a prediction system having a similar configuration to the prediction device 1 and exhibiting similar technical effects is also included within the scope of disclosure in this specification.
[0016] The following describes in detail each component of the prediction device 1. The input unit 2 is a unit that functions as an interface for reading data from an external source in the prediction device 1. As will be described in more detail later, the data input to the input unit 2 (hereinafter referred to as "input data") includes at least a dataset 10 relating to the lipid composition of the cell membrane of target cells (hereinafter referred to as the "target cell lipid dataset"). The input data may be entered by the user, or it may be data represented by signals or numerical values output from various units. If the input data is entered by the user, the input unit 2 is a user interface, such as a keyboard and mouse provided by a computer.
[0017] Alternatively, the input unit 2 may be various data acquisition units, in which case the input data may be cell membrane lipid data measured and input using the data acquisition unit. As the data acquisition unit, various known chemical analyzers, electron microscopes, measuring devices for measuring rubber properties, etc., can be used. Or, if the input data is input by an external data acquisition device, the input unit 2 may be a calculation processing unit that executes a program to automatically digitize the signal obtained from the external data acquisition device.
[0018] The arithmetic unit 3 is a group of various known arithmetic units capable of executing a predetermined program. Figure 1 shows the case where the arithmetic unit 3 is a group of arithmetic units, and the arithmetic unit 3 is a group of arithmetic units consisting of at least an auxiliary storage unit 5, a main storage unit 6, and an arithmetic processing unit 7 that are connected to each other.
[0019] The auxiliary storage unit 5 is a known auxiliary storage unit (e.g., an HDD) and stores a predetermined program. The predetermined program includes at least a delivery carrier prediction tool (hereinafter also simply referred to as the "prediction tool") 11 that predicts the lipid composition of delivery carriers suitable for the lipid composition of the cell membranes of various cells. The prediction tool is a combination of multiple programs and may include a program 12 that causes the calculation unit 3 to execute an algorithm (hereinafter referred to as the "delivery amount prediction algorithm") that predicts the amount of activator delivered to target cells using candidate composition data of delivery carriers and feature data of the lipid composition of target cells, and a program 13 that causes the calculation unit 3 to execute an algorithm (hereinafter referred to as the "composition optimization algorithm") that optimizes the composition of delivery carriers based on the predicted amount of activator delivered. Furthermore, as will be described in detail later, the prediction tool is an AI tool constructed by various machine learning methods using a training dataset containing data on cells other than target cells (hereinafter referred to as "non-target cells") as training data. This AI tool is trained using various machine learning techniques to infer the relationship between the composition of the delivery carrier, the cell membrane lipid composition of the cells to which the activator is delivered by the delivery carrier, and the amount of activator delivered to the target cells. Non-target cells include target cells collected under different culture conditions, at different harvest times, and from different donors.
[0020] The main memory unit 6 is a known main memory unit (e.g., memory), and the arithmetic processing unit 7 is a known central processing unit (CPU). Using these, data processing specified by a predetermined program 12 can be executed. In the prediction device 1, the main memory unit 6 is connected to the auxiliary memory unit 5 and the arithmetic processing unit 7 in a manner that allows for mutual communication. Therefore, it is possible to configure the device so that a predetermined program stored in the auxiliary memory unit 5 is read into the main memory unit 6 and started, transmitted to the arithmetic processing unit 7 to execute data processing, and the calculation results of the arithmetic processing unit 7 are received by the main memory unit 6. Here, the calculation results of the arithmetic processing unit 7 include at least a target carrier composition dataset 14.
[0021] The calculation result obtained by the calculation unit 3 is transmitted to the output unit 4. The output unit 4 is a unit or device that functions as an interface for outputting and displaying the target carrier composition data set 14 outside the prediction device 1, and is, for example, a display or the like.
[0022] (Data set) Regarding various input data input to the prediction device 1 via the input unit 2, it will be described in detail below. In FIGS. 2 to 4, for the sake of explanation, each data set is shown in tabular form, but each data set may be in any data format or data structure as long as it has various data elements described below. For example, the data set may be matrix data or vector data. In this case, each label of the row data and column data in FIGS. 2 to 4 may be configured as a data set as separate data from the matrix data or vector data.
[0023] As described above, the input data includes at least the target cell lipid data set 10. As shown in FIG. 2, the target cell lipid data set 10 is a data set that includes at least a data element of the name of the lipid compound constituting the cell membrane of the target cell and a data element regarding its composition amount. This data set 10 may be configured to include a plurality of sets of combinations of data elements of the cell membrane lipid composition ratio of the target cell. The combination of a plurality of sets of data elements here refers to, for example, data of lipid composition ratios obtained by applying different measurement methods to the same target cell, membrane composition data of target cells prepared under different culture conditions, membrane composition data of target cells with different lots, etc., and target cells collected from different donors. By configuring the data set 10 with a combination of a plurality of sets of data elements, the target carrier composition data set 14 can be obtained while taking into account minute errors regarding the cell membrane lipid composition of the target cell.
[0024] The composition amount of each lipid compound in the dataset 10 may be set as a ratio with the total amount of lipid compounds constituting the cell membrane being 100 parts by mass. The composition amount of each lipid compound may be a measured value of lipid composition obtained by various known lipid quantification methods. For example, the target cell lipid dataset 10 may be a measured value obtained by an external measuring device of the prediction device 1, or it may be data determined by referring to publicly available literature. If the input unit 2 is equipped with a data acquisition unit, the target cell lipid dataset 10 is a measured value obtained by the data acquisition unit. Alternatively, the weight ratio or molar ratio of each lipid compound to the total cell membrane components may be set as the composition amount of each lipid compound. In this case, the dataset 10 may also include data on cell membrane components other than lipid compounds (e.g., membrane proteins).
[0025] To more accurately understand the characteristics of the cell membrane lipid composition of target cells, it is preferable that the number of lipid compounds constituting this dataset 10 be relatively large. The number of lipid compounds should be at least 20, and preferably 100 or more.
[0026] ·Design method As illustrated in Figure 5, the procedure for designing delivery carriers involves first preparing a delivery carrier prediction tool constructed by machine learning using delivery carrier composition data for multiple types of delivery carriers exhibiting different component compositions, cell membrane composition data for multiple types of non-target cells' cell membrane lipid compositions, and delivery amount data for the amount of delivered substances delivered by the multiple types of delivery carriers to the multiple types of non-target cells, as training data. Next, the cell membrane lipid composition of target cells is applied to the delivery carrier prediction tool as input data to obtain a delivery carrier composition that satisfies the target value.
[0027] Next, we will explain the design method using the design apparatus described above. As illustrated in Figure 6, a design apparatus equipped with a delivery carrier prediction tool is prepared (S100). Next, the lipid composition of the cell membrane of the target cell is input using the input unit 2 of the prediction apparatus 1 (S110). Furthermore, the calculation unit 3 is made to execute each step by reading and executing the program set that executes the prediction tool 11 stored in the auxiliary storage unit 5 of the calculation unit 3 (S120~S150). Finally, a target carrier composition dataset 14 that yields an excellent amount of activator delivery is obtained and output (S160), and the process ends. The contents of each step (S100~S160) are described in detail below.
[0028] In the step of inputting the lipid composition of the cell membrane of the target cells (S110), the input unit 2 of the prepared prediction device 1 is used to input the target cell lipid dataset 10. If the calculation unit 3 has an auxiliary storage unit 5 and a main storage unit 6, the input dataset 10 is stored in the auxiliary storage unit 5 or the main storage unit 6.
[0029] After step (S110), the calculation unit 3 reads the target cell lipid dataset 10 and the prediction tool 11 (step (S120)). For example, if the calculation unit 3 has an auxiliary storage unit 5, a main storage unit 6, and a calculation processing unit 7, the prediction tool stored in the auxiliary storage unit 5 is read into the main storage unit 6, and data processing by the calculation processing unit 7 is started.
[0030] After step (S110), the calculation unit 3 performs data processing to generate a dataset of candidate compositions of delivery carriers (step (S130)). A prediction tool is used to generate the candidate composition data. Specifically, it is generated by combining arbitrary components of many types of delivery carriers recorded in the dataset used for machine learning in the construction of the prediction tool, and determining the amount of each component. For example, if delivery carrier A and delivery carrier B are stored in the dataset used for machine learning in the construction of the prediction tool, component X of delivery carrier A and components Y and Z of delivery carrier B are selected to generate a composition containing components X, Y and Z in arbitrary amounts, which is then designated as a candidate composition.
[0031] After step (S130), the calculation unit 3 applies program 12 for executing the delivery volume prediction algorithm of the prediction tool to the generated dataset of candidate delivery carrier compositions and performs data processing (step (S140)). Examples of delivery volume prediction algorithms include Lasso, Elastic Net, kernel ridge regression (KRR), Gaussian process regression (GRR), and neural networks.
[0032] After step (S140), the calculation unit 3 optimizes the composition of the delivery carrier based on the predicted amount of activator delivered (step (S150)). Specifically, in step (S150), program 13 is used to cause the calculation unit 3 to execute a composition optimization algorithm, and composition data of a delivery carrier that predicts a larger amount of activator delivered is selected from the candidate composition dataset generated in step (S140).
[0033] For example, program 13 (hereinafter also referred to as the "composition optimization program") which executes the composition optimization algorithm consists of: a candidate generation program which causes the computer to generate multiple candidate compositions of delivery carriers; a delivery amount calculation program which causes the computer to calculate a predicted value for the amount of activator delivered by each candidate composition of the delivery carrier based on the multiple candidate compositions generated and the cell membrane lipid composition of the target cell; and a determination program which causes the computer to compare each predicted value with a target value, determine whether each predicted value satisfies the target value, and select the candidate composition for which a predicted value that satisfies the target value has been calculated.
[0034] The composition optimization algorithm is a method that, for example, ranks candidates based on the magnitude of the predicted amount of activator delivered and selects candidate compositions of delivery carriers that represent the predicted amount of activator delivered from 1st place up to a predetermined rank. In other words, depending on the settings of the composition optimization algorithm, the composition data of the delivery carrier selected in step (S150) may not be limited to one, but may include multiple.
[0035] Alternatively, the composition optimization algorithm is a method that sets a predetermined target value and determines and selects a delivery carrier that shows a predicted value of activator delivery amount that approximates the target value as having a high activator delivery amount. Specifically, program 13 causes the calculation unit 3 to compare the calculated delivery amount prediction value with the target value, determines whether it is within the set acceptable range, and executes to select candidate composition data of a delivery carrier within the acceptable range. The acceptable range can be set arbitrarily, but for example, it can be set to a range of ±10% or ±5% relative to the target value.
[0036] Alternatively, a composition optimization algorithm is a method for selecting a delivery carrier that exhibits an activator delivery amount approximating the target value based on an evaluation using an evaluation function. Various known evaluation functions such as mean absolute error, mean squared error, mean absolute error rate, and mean squared error rate can be used as evaluation functions to assess the degree of approximation between the target value and the activator delivery amount. By using an evaluation function as an index for evaluating the activator delivery amount approximating the target value, it is possible to understand in more detail the degree to which differences in the composition of the delivery carrier affect the activator delivery amount.
[0037] Alternatively, the composition optimization program 13 may be a program that searches for composition data of delivery carriers that show an activator delivery amount that approximates the target value by repeatedly performing the steps of generating candidates (S130) and calculating the activator delivery amount (S140), and obtains a target carrier composition dataset 14. For example, an acceptable range of ±10% or ±5% of the target value may be set, and each step (S130, S140) is repeated until the calculated activator delivery amount falls within that acceptable range. That is, in step (S130) of a certain iteration, candidates may be generated based on an evaluation index of the activator delivery amount calculated in the previous or cumulative iteration.
[0038] The composition optimization algorithm, which repeatedly executes steps (S130) and (S140), is a method that uses a generative model obtained through machine learning. Examples of machine learning methods include Gaussian process regression, optimization algorithms, multiple regression analysis, random forests, recurrent neural networks (RNNs), and deep neural networks (DNNs) such as generative adversarial networks (GANs) and conditional generative adversarial networks (CGANs). Examples of optimization algorithms include genetic algorithms (GAs), Bayesian optimization (BOs), and gradient descent.
[0039] Here, we will explain an example of using a genetic algorithm in machine learning. First, a dataset is generated that contains a large number of candidate composition data that satisfy the constraints imposed by the constraint conditions. Next, this dataset is used as the initial current generation dataset, and genetic operations (crossover, mutation) are performed on the composition data selected from the current generation dataset based on an evaluation index (roulette selection, tournament selection, elite selection, etc.). Next, a next-generation dataset is created by aggregating multiple composition data generated by the genetic operations, and an evaluation index is calculated for the created next-generation dataset. Then, the created next-generation dataset is treated as the current generation dataset, and the process of selection, genetic operations, creation of the next-generation dataset, and calculation of the evaluation index is repeated. The dataset with the maximum number of generations obtained through this repetition becomes the dataset containing the optimal solution composition data. This dataset containing the optimal solution composition data is selected as the target carrier composition dataset 14 obtained in steps (S130) to (S150).
[0040] After step (S150), the calculation unit 3 outputs the selected target carrier composition dataset 14 (step (S160)). In step (S160), the output unit 4 performs data processing to output the target carrier composition dataset 14. The output by the output unit 4 allows the prediction device 1 to obtain the target carrier composition dataset 14.
[0041] As described above, the prediction device 1, which receives the target cell lipid dataset 10 from the input unit 2, is configured in its calculation unit 3 to load the prediction tool stored in the auxiliary storage unit 5 and the target cell lipid dataset 10 into the main storage unit 6, and to process them in the calculation processing unit 7 to obtain a target carrier composition dataset 14, thereby designing the composition of the delivery carrier.
[0042] With such a prediction device 1, when designing a delivery carrier, one only needs to refer to the composition data obtained using the prediction device 1. This significantly reduces the number of times experimental data needs to be obtained by preparing target cells and a wide variety of delivery carriers. Furthermore, it eliminates the influence of human experience and skill, allowing for the design of a desirable delivery carrier composition. Therefore, it greatly contributes to improving the efficiency of development and improvement work to obtain a desired delivery carrier.
[0043] Next, modified examples of the design apparatus and design method of the embodiment will be described.
[0044] • The process of building a prediction tool As a variation, the delivery carrier design method may include a step (S10) of constructing a prediction tool before step (S110). Furthermore, the delivery carrier design apparatus may store a program in the arithmetic unit 3 that executes a machine learning algorithm for constructing the prediction tool.
[0045] More specifically, the prediction tool may consist of multiple programs constructed by machine learning using various datasets stored in the auxiliary storage unit 5 within the calculation unit 3 of the prediction device 1. In other words, the predetermined program stored in the auxiliary storage unit 5 may be a program that reads the various datasets stored in the auxiliary storage unit 5 and executes a machine learning algorithm in the calculation unit 3, and this program may be used to construct multiple programs that constitute the prediction tool.
[0046] In step (S10), the dataset used for machine learning of the prediction tool includes at least a dataset relating to the amount of activator delivered to non-target cells. Specifically, the dataset relating to the amount of activator delivered to non-target cells includes at least a dataset 15 of the cell membrane composition of non-target cells (hereinafter referred to as the "non-target cell membrane composition dataset") and a dataset 16 of the amount of activator delivered observed by contacting non-target cells with delivery carriers of multiple different compositions (hereinafter referred to as the "non-target cell delivery amount dataset"). In other words, when the prediction tool is constructed by machine learning within the prediction device 1, the auxiliary storage unit 5 may store the non-target cell membrane composition dataset 15 and the non-target cell delivery amount dataset 16 in addition to the target cell lipid dataset 10.
[0047] Furthermore, if the data acquisition unit is provided as the input unit 2 of the prediction device 1, the calculation unit 3 (more specifically, the auxiliary storage unit 5) of the prediction device 1 prepared in step (S110) does not need to store the non-target cell membrane composition dataset 15 and the non-target cell delivery amount dataset 16. For example, after preparing a prediction device 1 equipped with the data acquisition unit as the input unit 2, the datasets 15 and 16 can be acquired by the data acquisition unit and stored in the auxiliary storage unit 5. In other words, in the prediction method, the datasets 15 and 16 may be introduced from the input unit 2 to the prediction device 1 as input data.
[0048] The non-target cell membrane composition dataset 15 consists of at least the names or serial numbers of each lipid compound constituting the cell membrane of non-target cells, and data elements relating to the composition of each lipid compound. For example, each lipid compound is assigned a unique serial number (in the example in Figure 3, the labels for each column: No. 1 to No. 300), and the dataset consists of data elements representing the composition of each. Furthermore, it is preferable that the non-target cell membrane composition dataset 15 includes multiple sets of data elements relating to the composition of each lipid compound for each of several different types of non-target cells. In this case, the data elements of the non-target cell membrane composition dataset 15 also include the names or identification numbers of the non-target cells. For example, the non-target cell membrane composition dataset 15 consists of data elements such that each non-target cell is assigned a unique serial number (in the example in Figure 3, the labels for each row: No. 1 to No. 100).
[0049] To obtain feature extraction data that reflects the characteristics and trends of the cell membrane lipid composition of non-target cells, it is preferable that the number of lipid compound types and non-target cell types shown in this dataset 15 are relatively large. The number of lipid compound types shown in the data elements of dataset 15 is at least 20, preferably 100 or more. Also, the number of non-target cell types shown in the data elements of dataset 15 is at least 20, preferably 50 or more. Therefore, it is preferable that the data in dataset 15 be 20-dimensional or more, preferably 50-dimensional or more.
[0050] Furthermore, the composition of each lipid compound in dataset 15 may be expressed as a percentage (by weight or molar ratio) of the total lipid compounds in the cell membrane, or as a percentage of a specific lipid compound. Alternatively, it may be expressed as a percentage of all constituent materials that make up the cell membrane, and the percentage may be set as a weight or molar ratio relative to the total amount of all constituent materials that make up the cell membrane. In this case, dataset 15 may also include data on cell membrane components other than lipid compounds (e.g., membrane proteins).
[0051] The composition of each lipid compound in dataset 15 may be measured values of lipid composition obtained by various known quantitative methods. For example, the non-target cell membrane composition dataset 15 may be measured values obtained by an external measuring device of the prediction device 1, or it may be data determined by referring to publicly available literature. If the input unit 2 is equipped with the data acquisition unit, then dataset 15 is measured values obtained by the data acquisition unit.
[0052] The non-target cell delivery amount dataset 16 consists of at least the following data elements: sequential numbers assigned to the composition of multiple types of delivery carriers, the composition ratio of each component of the delivery carrier indicated by each sequential number, the name and identification number of the non-target cell that came into contact with each delivery carrier, and the amount of activator delivered into each non-target cell by each delivery carrier. Furthermore, the non-target cell delivery amount dataset 16 consists of, for example, two data structures (I) and (II).
[0053] For example, data structure (I) consists of data elements representing sequential numbers of the delivery carrier composition, data elements representing the composition ratio of each component of the delivery carrier indicated by each sequential number, and data elements representing the names and identification numbers of all components of the delivery carrier. If such data structure (I) is referred to as a "delivery carrier composition dataset," then it can be understood that the non-target cell delivery volume dataset 16 includes a delivery carrier composition dataset.
[0054] The data structure (II) consists of data elements representing sequential numbers of the composition of multiple types of delivery carriers, data elements representing the names and identification numbers of non-target cells that come into contact with the delivery carriers, and data elements representing the amount of activator delivered into non-target cells by each delivery carrier.
[0055] For example, the data structure (I) of the non-target cell delivery volume dataset 16 consists of data elements with sequential numbers of the delivery carrier composition (labels for each row in Figure 4(a): No.1 to No.40), data elements with names or identification numbers of each component of the delivery carrier (labels for each column in Figure 4(a): No.1 to No.20), and data elements with composition ratios of each component of the delivery carrier (numerical values in the table in Figure 4(a)).
[0056] In the data structure (I), the data element for the composition of each delivery carrier is the proportion of each component to the total amount, for example, by weight or mole fraction. The data exemplified in Figure 4(a) shows the composition ratios of multiple types of lipid nanoparticles (No. 1 to No. 100). Note that all lipid nanoparticles No. 1 to No. 100 in Figure 4(a) consist of multiple types of lipid compounds, so the composition of the delivery carrier is its lipid composition ratio. Alternatively, for example, if the lipid nanoparticles contain components other than lipid compounds (e.g., transmembrane proteins), the total weight or amount of all constituent substances may be set to 100, and the weight ratio or mole fraction of each lipid compound may be calculated and set as the data structure (I) for dataset 16.
[0057] The data structure (I) of dataset 16 may include data elements that show information other than the constituent components of each delivery carrier. For example, data structure (I) may include data on various physical properties such as the average particle size of the delivery carrier for each delivery carrier with a serial number, or it may include data on the manufacturing method of the delivery carrier and the environmental conditions during manufacturing (e.g., temperature during manufacturing).
[0058] The sequential numbered data elements of the delivery carrier composition in data structure (II) are identical to the sequential numbered data elements of the delivery carrier composition in data structure (I). That is, data structure (II) includes data elements relating to the amount of activator delivered to non-target cells when using the delivery carrier composition of each sequential number shown in data structure (I). Preferably, data structure (II) includes multiple sets of data for the amount of activator delivered for multiple different types of non-target cells. That is, preferably, data structure (II) consists of data elements that include the names or identification numbers of multiple types of non-target cells (in the example of Figure 4(b), the labels for each column: No.1 to No.50).
[0059] The amount of activator delivered into non-target cells by a delivery carrier is a measurement obtained, for example, by bringing the delivery carrier into contact with non-target cells to deliver the activator into the non-target cells, and then quantifying the amount of activator in the non-target cells. Such measurement values may be obtained by a data acquisition unit provided as input unit 2 in the prediction device 1. Alternatively, the measurement values may include, as part of the measurement values, measurement values obtained by manufacturers of various formulations containing delivery carriers during the manufacturing process or post-manufacturing inspections on the manufacturing line, measurement values obtained from laboratory data such as the results of numerous experiments and tests in research and development or computer simulation results, or data reported in publicly available literature.
[0060] The measured amount of activator delivered into non-target cells is obtained using various known quantitative methods suitable for the type of activator used. Specifically, when nucleic acids are used as activators, multiple types of delivery carriers containing nucleic acids and non-target cells are prepared, and the non-target cells and delivery carriers are brought into contact in different combinations to deliver nucleic acids into the non-target cells. Furthermore, a fluorescent label that specifically reacts with the delivered nucleic acid is added to the non-target cells as a detection reagent, and the amount of nucleic acid introduced into each non-target cell is measured. The measured values obtained by such quantitative methods may be constructed as a database and set as dataset 16.
[0061] Various machine learning algorithms can be used as machine learning techniques to build prediction tools. Examples include Gaussian process regression, multiple regression analysis, random forests, and deep neural networks (DNNs) such as recurrent neural networks (RNNs) and generative adversarial networks (GANs). This machine learning is used to train the model to infer the relationship between the composition of the delivery carrier, the cell membrane lipid composition of the cells to which the activator is delivered by the delivery carrier, and the amount of activator delivered to the target cells.
[0062] Feature extraction As a further variation, the delivery carrier design method may reduce the dimensionality of the composition of the delivery carrier and target cell components, and further execute an algorithm for extracting features (hereinafter referred to as the "feature extraction algorithm"). Alternatively, the delivery carrier design device may store a program for executing the feature extraction algorithm (hereinafter referred to as the "feature extraction program") in its calculation unit.
[0063] In step (S130), the molecular structure of the delivery carrier is often relatively complex, making it difficult to grasp the relationship between the overall molecular structure of the delivery carrier and its chemical properties. However, the characteristic components of the delivery carrier are relatively easy to understand. Therefore, by representing the overall characteristics of the delivery carrier with the characteristics of its components, it becomes possible to concisely and accurately grasp the composition of the delivery carrier and its relationship to it. Then, using the non-target cell delivery volume dataset obtained by utilizing the understood relationship, it becomes possible to efficiently obtain a delivery carrier composition that satisfies the predetermined activator delivery volume.
[0064] Furthermore, the construction of the prediction tool in step (S10), the generation of candidate compositions using the prediction tool in steps (S110) to (S130), and the acquisition of the target carrier composition dataset may incur excessive computational load. For example, the number of lipid compounds that make up a cell membrane can generally reach several hundred. Also, since the composition of delivery carriers can be artificially designed, there can be a large number of composition patterns. In such cases, the number of data elements in each of the above datasets, and consequently the number of dimensions of the datasets, becomes enormous, and handling each of the above datasets requires a great deal of computational resources.
[0065] On the other hand, when there are a very large number of lipid compounds in the cell membrane and a large compositional pattern of delivery carriers, the information often includes lipid compounds and components of delivery carriers that contribute little to the amount of activator delivered by the carriers. It is not necessary to apply the calculations in steps (S10) and (S110) to (S130) to such information, and these steps can be omitted.
[0066] Therefore, in steps (S10) and (S110), it is preferable to remove less necessary information from each dataset used in the calculation. Specifically, the computational load may be reduced by calculating features. This condition can be defined by the feature extraction program.
[0067] Specifically, a feature extraction program is executed to reduce the dimensionality of the target cell lipid dataset 10, and further data processing is performed in the calculation unit 3 to obtain a dataset 17 of features of the cell membrane lipid composition of target cells (hereinafter referred to as the "target cell lipid feature dataset"). Computational resources can be saved by using dataset 17 instead of dataset 10 for subsequent calculations (e.g., prediction of activator delivery amount). The obtained dataset 17 may be stored in the auxiliary storage unit 5 within the calculation unit 3. Furthermore, by executing the feature extraction program and processing the data in the calculation unit 3, a dataset 18 of features of the cell membrane lipid composition of non-target cells (hereinafter referred to as the "non-target cell lipid feature dataset") may be obtained, and dataset 18 may be used instead of dataset 16 for subsequent calculations. Alternatively, the obtained dataset 18 may be configured to be stored in the auxiliary storage unit 5.
[0068] The feature extraction program may be stored in the auxiliary memory unit 5, or it may be included as one of the programs that make up the prediction tool. The feature extraction algorithm may be an algorithm that extracts features of a reduced-dimensional composition, that is, an algorithm in which dimensionality reduction and feature extraction are integrated, or dimensionality reduction and feature extraction may be separate algorithms. Examples of feature extraction algorithms include variational autoencoders (VAE), t-distribution type stochastic nearest neighbor embeddings (t-SNE), gradient boosting (e.g., XGBoost), or feature selection methods using random forests (e.g., Boruta).
[0069] Dimensionality reduction is preferably achieved by setting certain constraints and generating a dataset of constituent components that satisfy these constraints. For example, when performing dimensionality reduction on the data elements of the cell membrane lipid composition of target cells in dataset 15, the program may be limited to the main components of the cell membrane lipid composition. Specifically, for the cell membrane lipid composition of target cells, multiple types of lipids that are ranked in order of their compositional amount may be considered as the main components and limited accordingly. Alternatively, when generating composition data for candidate delivery carriers, if there is a desired substance that should be the main component of the delivery carrier, candidate composition data may be generated with that substance as a constraint.
[0070] Furthermore, the constraints are not limited to those mentioned above. For example, if the dataset 15 includes manufacturing conditions, and specific manufacturing conditions are desired, those specific manufacturing conditions can also be used as constraints. By constraining the manufacturing conditions in this way, it becomes possible to consider the influence of the delivery carrier's manufacturing conditions on the delivery amount during the process of obtaining the target carrier composition dataset. As a result, it becomes possible to easily understand the manufacturing conditions required to produce the delivery carrier.
[0071] In the design method, the program that executes the feature extraction algorithm may be constructed by a machine learning algorithm at the same time as the step of constructing the prediction tool (S10).
[0072] • Teacher's practice input As a further variation, the delivery carrier design method may provide a dataset of activator delivery amounts obtained by introducing delivery carriers of known compositions into target cells as training data to the prediction tool via the input unit 2. That is, the prediction device may store a program that causes the computer to correct the composition of the delivery carriers with the training data. The purpose of providing this training data is to correct the target carrier composition data obtained by the prediction tool and to make more accurate predictions. The composition of the delivery carriers in the training data may differ from the composition of the delivery carriers in the data elements of the non-target cell delivery amount dataset 16.
[0073] More specifically, this can be applied to a composition optimization algorithm that repeatedly performs steps (S130) and (S140). Since each of the above steps is corrected with training data each time, it is possible to predict the target carrier composition data more accurately.
[0074] While several embodiments of the present invention have been described, these embodiments are presented as examples only and are not intended to limit the scope of the invention. These novel embodiments can be carried out in a variety of other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents. [Explanation of symbols]
[0075] 1...Delivery carrier composition prediction device, 2...Input unit, 3...Calculation unit, 4...Output unit, 5...Auxiliary memory unit, 6...Main memory unit, 7...Arithmetic processing unit, 10…Target cell lipid dataset, 11…Delivery carrier prediction tool, 12…Delivery amount prediction algorithm program, 13…Composition optimization algorithm program, 14…Target carrier composition dataset, 15…Non-target cell membrane composition dataset, 16…Non-target cell delivery volume dataset, 17…Target cell lipid feature dataset, 18…Non-target cell lipid feature dataset
Claims
1. A design method for predicting the composition of a delivery carrier that satisfies the target value of the amount of activator delivered to target cells, A step of preparing a delivery carrier prediction tool constructed by machine learning using a delivery carrier composition dataset of multiple types of delivery carriers exhibiting different component compositions, a non-target cell membrane composition dataset of multiple types of cells other than the target cells relating to the cell membrane lipid composition, and a non-target cell delivery amount dataset of the amount of delivered substances delivered by multiple types of delivery carriers to multiple types of cells other than the target cells as training datasets; and, A method comprising the step of using the delivery carrier prediction tool to obtain a delivery carrier composition that satisfies the target value based on the cell membrane lipid composition of the target cell.
2. The acquisition process described above is: A process for generating multiple candidate compositions of delivery carriers, A step of calculating a predicted value for the amount of activator delivered by the delivery carrier of each candidate composition from the multiple candidate compositions generated and the cell membrane lipid composition of the target cell, A step of comparing the predicted value with the target value and determining whether the predicted value is a delivery carrier composition that satisfies the target value. The method according to claim 1, comprising:
3. The process of preparing the aforementioned delivery carrier prediction tool is as follows: A step of preparing the delivery carrier composition dataset, the non-target cell membrane composition dataset, and the non-target cell delivery volume dataset; and The process involves performing machine learning using the aforementioned delivery carrier composition dataset, the aforementioned non-target cell membrane composition dataset, and the aforementioned non-target cell delivery volume dataset as training datasets to construct the delivery carrier prediction tool. The method according to claim 1, including the method described in claim 1.
4. The method according to claim 3, wherein the machine learning is Gaussian process regression, multiple regression analysis, random forest, or deep neural network (DNN).
5. The method according to claim 1, further comprising the step of extracting characteristic quantities of the cell membrane lipid composition of the target cells.
6. The method according to claim 5, wherein the feature extraction is performed using a feature selection method that employs a variational autoencoder (VAE), a t-distribution stochastic nearest neighbor embedding (t-SNE), gradient boosting, or a random forest.
7. A step of providing a dataset of activator delivery amounts obtained by introducing a delivery carrier of a known composition into the target cells and measuring it, as training data to the delivery carrier prediction tool; and, The method according to claim 1, further comprising the step of correcting the composition of the delivery carrier with the training data.
8. A program for causing a computer to predict the composition of a delivery carrier that satisfies the target value of the amount of activator delivered to target cells, The program is constructed by machine learning using a delivery carrier composition dataset, which comprises a delivery carrier composition dataset, which comprises multiple types of delivery carriers exhibiting different component compositions; a non-target cell membrane composition dataset, which comprises the cell membrane lipid composition of multiple types of cells other than the target cells; and a non-target cell delivery amount dataset, which comprises the amount of delivered substances delivered by multiple types of delivery carriers to the multiple types of cells other than the target cells, as training datasets. A program that causes the computer to predict the composition of a delivery carrier that will satisfy the target amount of activator delivered to the target cells, based on a dataset of the cell membrane lipid composition of the target cells.
9. The program described in claim 8 includes a composition optimization program, and the composition optimization program is The computer is provided with a candidate generation program that generates multiple candidate compositions for delivery carriers, A delivery amount calculation program causes the computer to calculate a predicted value for the amount of activator delivered by the delivery carrier of each candidate composition, based on the multiple candidate compositions generated and the cell membrane lipid composition of the target cells. A determination program causes the computer to compare each of the predicted values with the target value, determine whether each of the predicted values satisfies the target value, and select the candidate composition for which the predicted value that satisfies the target value has been calculated. The program according to claim 8, comprising the following:
10. The program according to claim 9, wherein the composition optimization program causes the computer to repeatedly execute the candidate generation program, the delivery amount calculation program, and the determination program until the predicted value satisfies the target value.
11. The program according to claim 10, wherein the composition optimization program for repeated execution is a genetic algorithm, Bayesian optimization, or gradient descent.
12. A program that causes the computer to perform machine learning to construct the program described in claim 8, The machine learning program uses the delivery carrier composition dataset, the non-target cell membrane composition dataset, and the non-target cell delivery volume dataset as training datasets to cause the computer to infer the relationships between the delivery carrier composition dataset, the non-target cell membrane composition dataset, and the non-target cell delivery volume dataset.
13. The program according to claim 12, wherein the machine learning is Gaussian process regression, multiple regression analysis, random forest, or deep neural network (DNN).
14. The program according to claim 9, wherein the computer is made to execute a feature extraction program for extracting feature quantities of the cell membrane lipid composition of the target cells, and the computer is made to use the obtained feature quantities to calculate the predicted values of the delivery carriers for each of the candidate compositions.
15. The program according to claim 14, wherein the extraction of the aforementioned features is performed by a feature selection method using a variational autoencoder (VAE), a t-distribution stochastic nearest neighbor embedding (t-SNE), gradient boosting, or a random forest.
16. A feature extraction program is executed to extract features from the delivery carrier composition dataset, the non-target cell membrane composition dataset, and the non-target cell delivery volume dataset. The machine learning program according to claim 12, wherein the program causes the computer to infer the relationship between each of the features using each of the features.
17. The program according to claim 16, wherein the extraction of the aforementioned features is performed by a feature selection method using a variational autoencoder (VAE), a t-distribution stochastic nearest neighbor embedding (t-SNE), gradient boosting, or a random forest.
18. The program according to claim 9, which causes the computer to correct the composition of the delivery carrier with training data.
19. A device for predicting the composition of delivery carriers that will satisfy the target value of the amount of activator delivered to target cells, An input unit configured to receive the cell membrane lipid composition of the target cell from outside the device; A calculation unit configured to calculate a delivery carrier composition that satisfies the target value of the amount of activator delivered to the target cell, based on the cell membrane lipid composition of the target cell transmitted from the input unit, using a delivery carrier prediction tool constructed by machine learning that uses delivery carrier composition data for multiple types of delivery carriers exhibiting different component compositions, a non-target cell membrane composition dataset for multiple types of cells other than the target cell, and a non-target cell delivery amount dataset for the amount of delivered substance delivered by multiple types of delivery carriers to multiple types of cells other than the target cell, as training datasets; and Output unit configured to display the calculation result of the aforementioned calculation unit A device equipped with the following features.