Nanometer anti-tumor drug design optimization method and system based on Transform model
Through the nano-anti-tumor drug design optimization method based on the Transformer model, the problem of poor delivery efficiency of nano-drugs in the tumor area is solved, personalized optimization of drug design parameters is achieved, and design efficiency and delivery effect are improved.
Patent Information
- Application Number
- CN202411890216.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-03
AI Technical Summary
The delivery efficiency of nano-anti-tumor drugs in tumor areas is affected by tumor type, cancer stage, and tumor microenvironment heterogeneity, resulting in differences in distribution and accumulation in different patients or in different types of tumors.
The nano-anti-tumor drug design optimization method based on the Transformer model is adopted, and the nano-drug characteristics and delivery-related data are obtained, and the Decoder-only Transformer model is constructed. The loss function and position coding training model are used to generate nano-drug design parameters that meet the target delivery efficiency.
Effectively capture the complex relationship between drug characteristics, generate nanodrug design parameters that meet specific conditions, optimize and customize drug delivery effects, improve design efficiency, reduce experimental costs, and provide new tools for personalized and efficient nanodrug development.
Smart Images

Figure CN120089233A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence-assisted drug design research, and particularly relates to a method and system for optimizing the design of nano anti-tumor drugs based on the Transformer model. Background Art
[0002] Nano anti-tumor drugs encapsulate chemotherapeutic drugs in nano-carriers, which can more effectively target tumor cells and reduce damage to healthy tissues, thereby reducing side effects and increasing the therapeutic effect. Common nano-carriers include liposomes, solid lipid nanoparticles (SLN), nanostructured lipid carriers (NLC), etc. These carriers can improve the bioavailability of drugs, prolong their circulation time in vivo, and achieve controlled release functions.
[0003] However, the delivery efficiency of nano anti-tumor drugs in the tumor region still faces challenges. First, tumor type, cancer stage, and tumor microenvironment heterogeneity can all lead to differences in the distribution and accumulation of nano-drugs in different patients or different types of tumors. Second, the properties of nano-drugs are complex, and the type, shape, particle size, hydrodynamic properties, surface potential, etc. of the nano-carrier can all affect the delivery of nano-drugs. Therefore, a method is needed to systematically study the properties and delivery efficiency of nano anti-tumor drugs. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a method and system for optimizing the design of nano anti-tumor drugs based on the Transformer model, which solves the problems in the prior art.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] A method for optimizing the design of nano anti-tumor drugs based on the Transformer model, comprising the following steps:
[0007] Obtain the characteristics of nano anti-tumor drugs and the data related to their delivery in the tumor region to form a data set;
[0008] Divide the data features into: categorical features, numerical features, and drug delivery efficiency features, and determine the range of variables;
[0009] Perform feature encoding on the data, including: performing one-hot encoding on categorical variables, and performing standardization processing on numerical variables and drug delivery efficiency variables;
[0010] Divide the data set after feature encoding into a training set, a test set, and a validation set;
[0011] Perform feature fusion on the training set, the test set, and the validation set respectively, construct an input sequence, and perform position encoding;
[0012] Construct a Transformer model based on Decoder-only;
[0013] Input the position-encoded training set into the Transformer model for training, enabling the Transformer model to learn to generate a sequence of nano-drug design parameters that meet the target; in the training of the Transformer model, a loss function is introduced to guide the model's learning, a validation set is used for hyperparameter tuning, and a test set is used to evaluate the generalization ability of the model;
[0014] Utilize the trained Transformer model to output the result of a sequence of nano-drug design parameters that meet the target delivery efficiency, including the particle size, surface potential, shape, carrier type, surface modification, and core material of the nano-drug.
[0015] Furthermore, the data related to the characteristics of the nano-antitumor drug include: nano-drug type Type, core material MAT of the nano-drug, delivery strategy TS, shape Shape of the nano-drug, particle size Size of the nano-drug, hydrodynamic diameter HD, and surface potential ZP of the nano-drug;
[0016] The data related to the delivery of the nano-antitumor drug in the tumor region include: cancer type CT, tumor model TM, and data related to drug delivery efficiency; among them, the data related to drug delivery efficiency include: delivery efficiency DE_24 at 24 h, delivery efficiency DE_168 at 168 h, maximum delivery efficiency DE_Max, and delivery efficiency DE_Tlast at the last sampling point.
[0017] Furthermore, the steps of the feature fusion include:
[0018] 1) Concatenate the one-hot encoded vectors of all categorical variables to form an overall sparse vector; project this sparse vector into a dense embedding vector of dimension d model through a linear projection layer;
[0019] 2) For the numerical feature vector [Size, HD, ZP] formed by the standardized numerical variables, use a separate linear projection layer to map the numerical feature vector to the same d model dimension;
[0020] 3) After standardizing all delivery efficiency variables, perform a linear projection to obtain an embedding representation of dimension d model ;
[0021] Furthermore, when performing position encoding, the formula for calculating the position encoding matrix is:
[0022] Even dimensions are calculated through the sine function:
[0023]
[0024] Odd dimensions are calculated through the cosine function:
[0025]
[0026] Among them, PE (pos,2i) is the position encoding for even dimensions, and PE (pos,2i+1) is the position encoding for odd dimensions. pos is the position index, i is the dimension index in the position encoding vector, and d is the feature dimension.
[0027] Furthermore, the Transformer model includes multiple stacked decoder blocks, and the output of each decoder block serves as the input for the next layer of decoder blocks;
[0028] A single decoder block includes: a Masked multi-head self-attention layer, a feed-forward neural network, layer normalization, and residual connections; the input data enters the Masked multi-head self-attention layer to capture the correlations at different positions in the sequence. The output result is added to the input through the residual connection, and then layer normalization is performed; the normalized result enters the feed-forward neural network for non-linear transformation. The output of the feed-forward neural network is added to the input again through the residual connection and undergoes layer normalization processing, and the final output serves as the result of a single decoder block.
[0029] Furthermore, the loss function of the Transformer model is:
[0030]
[0031] In the formula, y t represents the true target value, which is the correct label at time step t; represents the probability that the model generates the output y t at time step t based on the previously generated results and the input F; F is the input feature of the model, including the encodings of categorical variables, numerical variables, and delivery efficiency variables; Y <t represents all the previously generated output sequences before time step t; θ represents the set of all parameters of the model; m represents the length of the generated sequence.
[0032] A nano anti-tumor drug design optimization system based on the Transformer model includes:
[0033] A data acquisition module: acquires data related to the characteristics of nano anti-tumor drugs and their delivery in the tumor region to form a data set;
[0034] Data Classification Module: Classify data features into: categorical features, numerical features, and drug delivery efficiency features, and determine the range of variables;
[0035] Feature Encoding Module: Perform feature encoding on the data, including: performing one-hot encoding on categorical variables, and performing normalization on numerical variables and drug delivery efficiency variables;
[0036] Dataset Partitioning Module: Partition the dataset after feature encoding into a training set, a test set, and a validation set;
[0037] Position Encoding Module: Perform feature fusion on the training set, the test set, and the validation set respectively, construct an input sequence, and perform position encoding;
[0038] Model Construction Module: Construct a Transformer model based on Decoder-only;
[0039] Model Training Module; Input the training set after position encoding into the Transformer model for training, so that the Transformer model learns to generate a sequence of nano-drug design parameters that meet the target; A loss function is introduced during the training of the Transformer model to guide the model to learn, the validation set is used for hyperparameter tuning, and the test set is used to evaluate the generalization ability of the model;
[0040] And, Result Output Module: Use the trained Transformer model to output the result of the nano-drug design parameter sequence that meets the target delivery efficiency, including the particle size, surface potential, shape, carrier type, surface modification, and core material of the nano-drug.
[0041] A computer storage medium stores a readable program that, when run, can execute the above-mentioned nano-antitumor drug design optimization method based on the Transformer model.
[0042] An electronic device includes: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0043] The memory is used to store at least one executable instruction that can perform the operations corresponding to the above-mentioned nano-antitumor drug design optimization method based on the Transformer model.
[0044] A computer program product includes computer instructions that direct a computing device to perform the operations corresponding to the above-mentioned nano-antitumor drug design optimization method based on the Transformer model.
[0045] Advantages of the present invention:
[0046] 1. The Transformer model based on Decoder-only of the present invention can well capture the complex relationships between drug properties, avoiding the limitations of traditional deep learning algorithms in multi-variable complex problems.
[0047] 2. By taking the delivery efficiency as one of the input features and combining autoregressive generation and Masked multi-head self-attention layers, the model can generate nano-drug design parameters that meet specific conditions, thereby optimizing and customizing the drug delivery effect.
[0048] 3. The present invention not only improves the design efficiency but also reduces the experimental cost, providing a new tool for personalized and efficient nano-drug development. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 Flow chart of the method for optimizing the design of nano anti-tumor drugs of the present invention;
[0051] Figure 2 Architecture diagram of the Transformer model based on Decoder-only of the present invention;
[0052] Figure 3 Schematic diagram of the Masked multi-head self-attention layer and its residual connection and layer normalization of the present invention;
[0053] Figure 4 Schematic diagram of the feed-forward neural network and its residual connection and layer normalization of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the protection scope of the present invention.
[0055] As Figure 1 shown, a method for optimizing the design of nano anti-tumor drugs based on the Transformer model includes the following steps:
[0056] S1. Obtain data related to the characteristics of nano-antitumor drugs and their delivery in the tumor region, and form a dataset;
[0057] The data related to the characteristics of the nano-antitumor drugs includes:
[0058] Type of nanoparticles (Type), Corematerials of nanoparticles (MAT), Targeting Strategy (TS), Shape of nanoparticle (Shape), Size of nanoparticles (Size), Hydrodynamic diameter (HD), Zeta potential (ZP) of the nano-drug.
[0059] The data related to the delivery of nano-antitumor drugs in the tumor region includes:
[0060] Cancer Type (CT), Tumor Model (TM), and data related to Delivery efficiency (DE);
[0061] The data related to Delivery efficiency (DE) includes: Delivery efficiency at 24 h (DE_24), Delivery efficiency at 168 h (DE_168), Maximum delivery efficiency (DE_Max), and Delivery efficiency at the last sampling point (DE_Tlast).
[0062] S2. Classify the data features into: categorical features, numerical features, and drug delivery efficiency features, and determine the range of variables;
[0063] The specific content of the data feature classification is shown in Table 1 below:
[0064] Table 1 Data Feature Classification Table
[0065]
[0066] S3. Perform feature encoding on the data, including: performing one-hot encoding on categorical variables, and performing standardization processing on numerical variables and drug delivery efficiency variables;
[0067] 1) Perform one-hot encoding on categorical variables, specifically including:
[0068] Values of the nano - drug type Type: inorganic; organic. If inorganic is selected, the code is [1, 0]; if organic is selected, the code is [0, 1].
[0069] Values of the core material MAT of the nano - drug: gold, dendrimer, liposome, polymer, hydrogel, other materials, others. If MAT takes "gold", the corresponding code is [1, 0, 0, 0, 0, 0, 0], takes "dendrimer" the corresponding code is [0, 1, 0, 0, 0, 0, 0], and so on.
[0070] Values of the delivery strategy TS: passive, active. Represent "passive" as [1, 0], and the corresponding code for "active" is [0, 1].
[0071] Values of the cancer type CT: brain, breast, cervix, colon, liver, lung, ovary, pancreas, prostate, skin. When CT takes "brain", the corresponding code is [1, 0, 0, 0, 0, 0, 0, 0, 0, 0], takes "breast" the corresponding code is [0, 1, 0, 0, 0, 0, 0, 0, 0, 0], and so on.
[0072] Values of the tumor model TM: allograft ectopic (AH), allograft orthotopic (AO), xenograft ectopic (XH), xenograft orthotopic (XO). When the value is AH, the corresponding code is [1, 0, 0, 0], when the value is AO, the corresponding code is [0, 1, 0, 0], and so on.
[0073] Values of the nano - drug shape Shape: spherical, rod - shaped, plate - shaped or others. When Shape takes "spherical", the corresponding code is [1, 0, 0, 0], when the value is "rod - shaped", the corresponding code is [0, 1, 0, 0], and so on.
[0074] 2) Standardize the numerical variables, and then all the standardized numerical variables form a numerical feature vector of length 3 [Size, HD, ZP]; the standardization formula is as follows:
[0075]
[0076] where μ is the mean of the feature and σ is the standard deviation of the feature.
[0077] 3) After standardizing the drug delivery efficiency variable, it forms a numerical feature vector [DE_Tlast, DE_24, DE_168, DE_Max].
[0078] S4. Divide the dataset after feature encoding into a training set, a test set and a validation set;
[0079] Among them, the ratio of the training set, the test set, and the validation set is 7:1.5:1.5.
[0080] S5. Feature fusion is performed on the training set, the test set, and the validation set respectively to construct an input sequence and perform positional encoding.
[0081] The steps of feature fusion include:
[0082] 1) Connect the one-hot encoded vectors of all categorical variables to form an overall sparse vector; project this sparse vector into a low-dimensional (d model ) dense embedding vector through a linear projection layer.
[0083] 2) For the numerical feature vector [Size, HD, ZP] formed by the standardized numerical variables, use a separate linear projection layer to map the numerical feature vector to the same d model dimension to ensure consistency with the categorical variables.
[0084] 3) After standardizing all delivery efficiency variables, perform linear projection to obtain an embedding representation with a dimension of d model .
[0085] Through the above feature fusion, all feature vectors will be transformed into representations of the same dimension d model to make the feature vectors suitable for the Transformer model.
[0086] The process of constructing the input sequence includes:
[0087] Regard all features as a "token" in the input sequence, construct the input sequence, where various features are independent inputs. After fusion, the categorical features are used as the first input token, the numerical features are used as the second input token, and the delivery efficiency features are used as the third input token; therefore, the length of the entire input sequence is 3, and the dimension of each "token" is d model ; that is, the input sequence is F = [f 1 , f 2 , f 3 , where f 1 represents the categorical feature embedding, f 2 represents the numerical feature embedding, and f 3 represents the embedding of the delivery efficiency feature.
[0088] The content of the positional encoding includes:
[0089] The Transformer model is sensitive to position and needs to add positional encoding to each token in the input sequence to introduce sequence information. Define the positional encoding, where each position pos generates a vector, and the value of each dimension is determined by sine and cosine functions. Define the formula to calculate the positional encoding matrix:
[0090] For even dimensions (indexed by 2i), calculate through the sine function:
[0091]
[0092] For odd dimensions (indexed by 2i + 1), calculate through the cosine function:
[0093]
[0094] where, PE (pos,2i) is the positional encoding for even dimensions, PE (pos,2i+1) is the positional encoding for odd dimensions, pos is the position index (starting from 0), i is the dimension index in the positional encoding vector, and d is the feature dimension;
[0095] The final input representation of each token (feature embedding) is the concatenation of its embedding vector and the corresponding positional information, i.e., Input i = [f i , PE i , to ensure that the model can understand the relative positions of these features in the sequence.
[0096] S6. Build a Transformer model based on Decoder - only;
[0097] As Figure 2 shown, the Transformer model includes multiple stacked decoder blocks (Decoder - only). The output of each decoder block serves as the input to the next - layer decoder block. A single decoder block includes: Masked multi - head self - attention layer, feed - forward neural network, residual connection, and layer normalization. The input data enters the Masked multi - head self - attention layer to capture the correlations at different positions in the sequence. The output result of this is added to the input through the residual connection, and then layer normalization is performed to help improve the training stability. The normalized result enters the feed - forward neural network for non - linear transformation. The output of the feed - forward neural network is again added to the input through the residual connection and undergoes layer normalization processing. The final output serves as the result of a single decoder block.
[0098] The Masked Multi-head Self-Attention layer can process multiple attention heads in parallel, capturing various different relationships in the input features. A mask matrix is used to shield the information of subsequent positions, so that when the model generates the t-th token, it can only see the information from 1 to t, and cannot see the information of t+1 and later.
[0099] As Figure 3 shown, when there are h attention heads and the dimension of the data matrix is d, the data matrix processed by each attention head is d / h. Multiple "heads" calculate different attention scores respectively, and then these results will be concatenated and followed by a linear transformation. In this way, the model can capture different types of relationships and dependency information in multiple different representation subspaces. The output of the Masked multi-head attention layer will be directly added to the input to form a residual connection. Subsequently, layer normalization is performed to ensure that the data has stable mean and variance when passing through each layer.
[0100] Feed-Forward Network: After each Masked multi-head self-attention layer passes through the residual connection and layer normalization, a feed-forward neural network is set up to further process the information;
[0101] As Figure 4 shown, the feed-forward neural network usually contains two fully connected layers (one is the activation layer and the other is the output layer) and an activation function (such as ReLU). For the two fully connected layers, the first layer (W 1 ) is regarded as extracting high-level features from the input features, and the second layer (W 2 ) is regarded as combining the high-level features into the output. The feed-forward neural network is used to further process the output after the attention mechanism, increasing the expressive power of the model.
[0102] As Figure 3 shown, the Masked multi-head self-attention layer ensures that the model can only access the current and previous content at time step t. The calculation process of the Masked multi-head attention layer is as follows:
[0103] Query, Key, Value calculation:
[0104] Perform a linear transformation on the input X entering the Masked multi-head attention layer to obtain the query matrix Q, the key matrix K, and the value matrix V:
[0105] Q = XW Q
[0106] K = XW K
[0107] V = XWV
[0108] Among them, W Q , W K , W V is a learnable parameter matrix.
[0109] Calculate the attention score vector A:
[0110]
[0111] Among them, d k is the dimension of the key vector, used for scaling to prevent the gradient from being too large; M is the mask matrix, making its attention score a very small negative number, and then its probability tends to zero in the softmax, used to mask future time steps.
[0112] Attention output: O = AV
[0113] In summary, after passing through a Masked multi-head attention layer, it is expressed as:
[0114]
[0115] In the formula,
[0116] After passing through the Masked multi-head self-attention layer, it enters the residual connection and layer normalization (Add&Norm):
[0117] Z 1 = X + O
[0118]
[0119] Among them, Z 1 is the vector after adding the input and output of the Masked multi-head self-attention layer, X is the input of the Masked multi-head self-attention layer, Norm MMA is the result after layer normalization of the vector obtained by adding the input and output of the Masked multi-head self-attention layer, μ 1 and σ 1 are the mean and standard deviation of Z 1 respectively; γ 1 is used to scale the vector after layer normalization, adjusting the scale of the normalization result; β 1 is used to translate the vector after layer normalization, translating the normalization result to ensure that the model can learn an appropriate offset; γ 1 and β 1 act on Z 1 to ensure that they have appropriate expressive power and information transmission.
[0120] The output of the masked multi-head self-attention layer undergoes residual connection and layer normalization before entering the feed-forward neural network:
[0121] FFN(Norm MMA ) = ReLU(Norm MMA ·W 1 + b 1 )W 2 + b 2
[0122] Where FFN(Norm MMA ) is the output of the feed-forward neural network, W 1 is the weight matrix of the first feed-forward layer, W 2 is the weight matrix of the second feed-forward layer, b 1 is the bias term of the first layer, b 2 is the bias of the second layer.
[0123] Perform residual connection and layer normalization (Add&Norm) on the output of the feed-forward neural network:
[0124] Z 2 = Norm MMA + FFN(Norm MMA )
[0125]
[0126] Where Z 2 is the vector after adding the input and output of the feed-forward neural network, Norm FFN is the input of the feed-forward neural network, Laynorm is the result after layer normalization of the vector after adding the input and output of the feed-forward neural network, μ 2 and σ 2 are the mean and standard deviation of Z 2 respectively; γ 2 is used to scale the vector after adding the output of the feed-forward network and the input, adjusting the scale of the FFN output in the model to make it consistent with the input; β 2 is used to translate and adjust the normalized result so that its output has sufficient offset to better fit the data; γ 2 and β 2 act on Z 2 to ensure that the normalized data still has sufficient expressive power.
[0127] Through such hierarchical stacking, the output of each decoder block serves as the input of the next-layer decoder block until all layers are computed, and finally the output of the model is obtained.
[0128] S7. Input the training set into the Transformer model for training, enabling the Transformer model to learn and generate a sequence of nano-drug design parameters that meet the target; during the training of the Transformer model, a loss function is introduced to guide the model's learning, the validation set is used for hyperparameter tuning, and the test set is used to evaluate the generalization ability of the model.
[0129] In a decoder-only Transformer model composed of 6 decoder blocks (N d = 6), the input features (such as categorical variables, numerical variables, and delivery efficiency variables) are embedded into a sequence of high-dimensional vectors, which are passed layer by layer through multiple decoder blocks. Each block contains a Masked multi-head self-attention layer, a feed-forward neural network, as well as residual connections and layer normalization, for capturing the complex relationships between input features and gradually generating design parameters. In the forward propagation, the model outputs predicted values, which are compared with the true values to calculate the loss function, and then the gradients are calculated layer by layer through backpropagation, and gradient descent is used to update the model parameters, enabling the model to continuously learn and improve, and finally optimizing the generation results to meet a specific delivery efficiency.
[0130] During the training process, the model needs to learn and generate a sequence of nano-drug design parameters that meet the target. Let the generated target sequence be:
[0131] Y = [y 1 , y 2 , ……, y m
[0132] where y t represents the characteristic value generated at the t-th step.
[0133] The cross-entropy loss function is used to measure the difference between the sequence generated by the model and the target sequence.
[0134] Calculate the loss value at each step t
[0135]
[0136] In the formula, y t represents the true target value, which is the correct label at time step t;
[0137] In the formula, represents the probability that the model generates the output y t at time step t based on the previously generated results and the input F; F is the input feature of the model, including the encoding of categorical variables, numerical variables, and delivery efficiency metrics; Y <tDenote all the generated output sequences before time step t; all the parameter sets of the θ model.
[0138] Calculate the loss value of the entire sequence
[0139]
[0140] Here, m represents the length of the generated sequence, that is, the number of features that the model needs to generate step by step;
[0141] In summary, the cross-entropy loss function is defined as:
[0142]
[0143] Use the training set for model learning, and use the validation set for hyperparameter tuning and model selection. After training is completed, use the test set to evaluate the final performance of the model and confirm the generalization ability of the model.
[0144] S8. Utilize the trained Transformer model to output the results of the nano-drug design parameter sequence that meets the target delivery efficiency, including the particle size, surface potential, shape, carrier type, surface modification, and core material of the nano-drug.
[0145] Based on a similar inventive concept, an embodiment of the present invention also provides a computer storage medium storing a readable program that can execute the above-mentioned nano-antitumor drug design optimization method based on a Transformer model when the program runs.
[0146] Based on a similar inventive concept, an embodiment of the present invention provides an electronic device, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0147] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above-mentioned nano-antitumor drug design optimization method based on a Transformer model.
[0148] Based on a similar inventive concept, an embodiment of the present invention also provides a computer program product including computer instructions, and the computer instructions instruct a computing device to perform the operations corresponding to the above-mentioned nano-antitumor drug design optimization method based on a Transformer model.
[0149] Embodiment 2
[0150] Based on the nano anti-tumor drug design optimization method proposed in Example 1, in this example, a nano anti-tumor drug design optimization system based on the Transformer model is proposed, specifically including:
[0151] Data acquisition module: Acquire the characteristics of nano anti-tumor drugs and the data related to their delivery in the tumor area to form a data set;
[0152] Data classification module: Classify the data features into: categorical features, numerical features, and drug delivery efficiency features, and determine the range of variables;
[0153] Feature encoding module: Perform feature encoding on the data, including: performing one-hot encoding on categorical variables, and performing standardization processing on numerical variables and drug delivery efficiency variables;
[0154] Data set division module: Divide the data set after feature encoding into a training set, a test set, and a validation set;
[0155] Position encoding module: Perform feature fusion on the training set, test set, and validation set respectively, construct an input sequence, and perform position encoding;
[0156] Model construction module: Construct a Transformer model based on Decoder-only;
[0157] Model training module; Input the training set after position encoding into the Transformer model for training, so that the Transformer model learns to generate a nano-drug design parameter sequence that meets the target; A loss function is introduced during the training of the Transformer model to guide the model to learn, the validation set is used for hyperparameter tuning, and the test set is used to evaluate the generalization ability of the model;
[0158] And, result output module: Use the trained Transformer model to output the nano-drug design parameter sequence results that meet the target delivery efficiency, including the particle size, surface potential, shape, carrier type, surface modification, and core material of the nano-drug.
[0159] The method of the present invention can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CDROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and to be downloaded through a network and stored in a local recording medium, so that the method described herein can be stored on such a software process on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It will be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.
[0160] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A nano anti-tumor drug design optimization method based on the Transformer model, characterized in that: The following steps are involved: Obtain data related to the characteristics of nano anti-tumor drugs and their delivery in the tumor area to form a data set; Divide the data features into: categorical features, numerical features, and drug delivery efficiency features, and determine the range of the variables; Perform feature encoding on the data, including: one-hot encoding of categorical variables, and standardization of numerical variables and drug delivery efficiency variables; Divide the feature-encoded dataset into training set, test set, and validation set; Perform feature fusion on the training set, test set, and validation set respectively, construct the input sequence, and perform position encoding; Build a Decoder-only Transformer model; The position-encoded training set is input into the Transformer model for training, so that the Transformer model can learn to generate a sequence of nanomedicine design parameters that meets the target; a loss function is introduced into the Transformer model training to guide model learning, a validation set is used for hyperparameter adjustment, and a test set is used to evaluate the generalization ability of the model; Using the trained Transformer model, the sequence results of nanodrug design parameters that meet the target delivery efficiency are output, including the particle size, surface potential, shape, carrier type, surface modification and core material of the nanodrug.
2. A method for designing and optimizing nano anti-tumor drugs based on a Transformer model according to claim 1, characterized in that: The data related to the characteristics of the nano anti-tumor drug include: nano drug type Type, nano drug core material MAT, delivery strategy TS, nano drug shape Shape, nano drug particle size Size, hydrodynamic diameter HD and nano drug surface potential ZP; The data related to the delivery of nano anti-tumor drugs in the tumor area include: cancer type CT, tumor model TM and drug delivery efficiency related data; the drug delivery efficiency related data include: delivery efficiency DE_24 at the 24th hour, delivery efficiency DE_168 at the 168th hour, maximum delivery efficiency DE_Max and delivery efficiency DE_Tlast at the last sampling point.
3. A method for designing and optimizing nano anti-tumor drugs based on a Transformer model according to claim 1, characterized in that: The step of feature fusion includes: 1) Connect the one-hot encoded vectors of all categorical variables to form an overall sparse vector; project this sparse vector into d through a linear projection layer model Dense embedding vector of dimension; 2) Use a separate linear projection layer to map the numerical feature vector [Size, HD, ZP] formed by the standardized numerical variables to the same d model Dimension; 3) After standardizing all delivery efficiency variables, linear projection is performed to obtain a dimension d model The embedded representation of .
4. A method for designing and optimizing nano anti-tumor drugs based on a Transformer model according to claim 3, characterized in that: When position encoding is performed, the formula for calculating the position encoding matrix is: Even dimensions are calculated using the sine function: Odd dimensions are calculated using the cosine function: Among them, PE (pos,2i) For even-numbered dimensions, PE (pos,2i+1) is the position encoding of odd dimensions, pos is the position index, i is the dimension index in the position encoding vector, and d is the feature dimension.
5. A method for designing and optimizing nano anti-tumor drugs based on a Transformer model according to claim 1, characterized in that: The Transformer model consists of multiple stacked decoder blocks, and the output of each decoder block serves as the input of the decoder block in the next layer; A single decoder block includes: Masked multi-head self-attention layer, feedforward neural network, layer normalization and residual connection; the input data enters the Masked multi-head self-attention layer to capture the correlation of different positions in the sequence, and the output result is added to the input through the residual connection, followed by layer normalization; the normalized result enters the feedforward neural network for nonlinear transformation, and the output of the feedforward neural network is again added to the input through the residual connection, and after layer normalization, the final output is the result of a single decoder block.
6. A method for designing and optimizing nano anti-tumor drugs based on a Transformer model according to claim 5, characterized in that: The loss function of the Transformer model is: In the formula, y t represents the true target value, which represents the correct label at time step t; Represents the output y generated by the model at time step t based on the previously generated results and input F t The probability of delivery efficiency; F is the input feature of the model, including the encoding of categorical variables, numerical variables and delivery efficiency variables; Y <t represents all generated output sequences before time step t; the set of all parameters of the θ model; m represents the length of the generated sequence.
7. A nano anti-tumor drug design optimization system based on the Transformer model, characterized in that: include: Data acquisition module: obtains the characteristics of nano anti-tumor drugs and data related to their delivery in the tumor area to form a data set; Data classification module: divides data features into: categorical features, numerical features and drug delivery efficiency features, and determines the range of variables; Feature encoding module: feature encoding of data, including one-hot encoding of categorical variables and standardization of numerical variables and drug delivery efficiency variables; Dataset partitioning module: divides the feature-encoded dataset into training set, test set, and validation set; Position encoding module: perform feature fusion on the training set, test set, and validation set, construct input sequences, and perform position encoding; Model building module: build a Decoder-only based Transformer model; Model training module: input the position-encoded training set into the Transformer model for training, so that the Transformer model can learn to generate a nanomedicine design parameter sequence that meets the target; introduce a loss function in the Transformer model training to guide model learning, use the validation set to adjust the hyperparameters, and use the test set to evaluate the generalization ability of the model; And, the result output module: using the trained Transformer model, output the sequence results of nanodrug design parameters that meet the target delivery efficiency, including the particle size, surface potential, shape, carrier type, surface modification and core material of the nanodrug.
8. A computer storage medium storing a readable program, characterized in that: When the program is running, it can execute the nano anti-tumor drug design optimization method based on the Transformer model as described in any one of claims 1 to 6.
9. An electronic device, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the nano anti-tumor drug design optimization method based on the Transformer model as described in any one of claims 1-6.
10. A computer program product comprising computer instructions, characterized in that: The computer instructions instruct the computing device to perform operations corresponding to the nano anti-tumor drug design optimization method based on the Transformer model as described in any one of claims 1-6.
Citation Information
Cited By
Deep learning-based polypeptide self-assembly nanofiber design system and method
CN122224292A
Polypeptide self-assembly nanofiber design system and method based on deep learning
CN122224292B