Paclitaxel drug process parameter optimization method, device and system, and storage medium
Through the AI-enabled method, genetic algorithms and generative adversarial networks are used to optimize the BP neural network, and reverse search for the optimal parameter combination, solving the problem of low efficiency in optimization of process parameters of paclitaxel drug in the existing technology, achieving higher binding rate and production efficiency.
Patent Information
- Application Number
- CN202510119384.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively optimize the paclitaxel drug process parameters, resulting in low binding rate and production efficiency of the drug, and the traditional methods are inefficient, time-consuming and inflexible.
Using AI-enabled methods, data enhancement is performed by obtaining the basic data set of paclitaxel drug production, genetic algorithm is used to optimize the initial weight and bias of the BP neural network, and virtual samples are generated by the generation of adversarial network GAN, training the optimized BP neural network, and inversely searching for the optimal parameter combination in the prediction results of the optimized BP neural network.
Effectively modeling and finding the relationship between drug process parameters improves the drug binding rate and production efficiency, reduces production costs, and solves the problem that traditional methods are prone to fall into local optimal solutions.
Smart Images

Figure CN119993327A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine learning, and in particular relates to a method and device, a system, and a storage medium for optimizing process parameters of a paclitaxel drug. Background Art
[0002] Paclitaxel is a widely used chemotherapy drug for the treatment of various cancers, especially for the treatment of various malignant tumors, including breast cancer, ovarian cancer, non-small cell lung cancer, gastric cancer, prostate cancer, etc. It is often used as a monotherapy or in combination with other chemotherapy drugs, radiotherapy, etc. to improve the therapeutic effect. Paclitaxel was originally extracted from the bark of the yew tree, and later it was mass-produced through chemical synthesis. Although paclitaxel is effective in treating cancer, it may also cause a series of adverse reactions, including but not limited to bone marrow suppression, neurotoxicity, gastrointestinal reactions, allergic reactions, etc. Therefore, when using paclitaxel for treatment, doctors will weigh the pros and cons and design an individualized treatment plan based on the patient's specific situation and cancer type, while closely monitoring the patient's response and side effects.
[0003] Although paclitaxel has significant therapeutic effects, its low solubility and toxicity when used directly limit its application. The combination of paclitaxel and human albumin is an important drug interaction process, in which paclitaxel is the active ingredient of the drug and human albumin is an excipient. This combination has many benefits, such as improving the solubility and stability of the drug, reducing toxicity and side effects, and improving the utilization rate of the drug. Human albumin is an important protein in the human body, accounting for a large proportion of the plasma, and has multiple physiological functions, such as maintaining blood osmotic pressure, transporting nutrients and drugs, regulating acid-base balance, anti-oxidation, and participating in immune response.
[0004] The combination of human albumin and paclitaxel improves the solubility and stability of paclitaxel, allowing the drug to be more effectively distributed and exert its effects in the body. In clinical applications, this combined preparation significantly reduces the inhibition of paclitaxel on bone marrow hematopoietic function, as well as damage to the digestive tract and liver function, and reduces the risk of red blood cell and platelet decline. The combined preparation also performs well in improving the utilization rate of paclitaxel, allowing patients to obtain better treatment effects while reducing side effects.
[0005] Existing technologies usually rely on manual experience and relatively fixed processes. The optimization process is cumbersome and it is difficult to comprehensively improve the binding rate of drugs. Traditional experiments adjust a single variable by controlling other variables. This method of locally adjusting a variable makes it difficult to obtain better experimental parameters. The traditional experimental preparation process is quite cumbersome, and the experimental data that can be obtained each time is extremely limited, while traditional models require a large amount of data to establish the model. In the case of scarce experimental data, it is difficult to establish a high-accuracy model to predict better experimental parameters. There are complex nonlinear relationships between experimental parameters. Using traditional models to predict is prone to fall into local optimal solutions, and it is difficult to obtain a better global optimal solution. The traditional method of obtaining optimal parameters is to obtain the best output through a large amount of input. This method is inefficient, time-consuming and inflexible. Summary of the invention
[0006] The technical problem to be solved by the present invention is to provide a method and device, system and storage medium for optimizing the process parameters of paclitaxel drug process, to optimize the process parameters of paclitaxel (albumin-bound) drug production for injection, and to find a parameter group that satisfies the particle size and particle size distribution and has a high binding rate. Through AI, enterprises can be empowered to produce, improve drug production efficiency and reduce enterprise production costs.
[0007] To achieve the above object, the present invention adopts the following technical solution:
[0008] A method for optimizing process parameters of paclitaxel medicine, comprising:
[0009] Obtain basic data sets for paclitaxel drug production;
[0010] Perform data enhancement on the basic data set to generate a virtual sample augmented data set;
[0011] Optimize the initial weights and biases of the BP neural network through genetic algorithms;
[0012] The optimized BP neural network is trained by using virtual samples to expand the data set;
[0013] The genetic algorithm is used to reversely search and optimize the optimal parameter combination in the prediction results of the BP neural network.
[0014] Preferably, the human albumin incubation temperature, incubation time, emulsifying disperser speed, albumin aqueous solution feeding speed, and paclitaxel feeding speed are set, and the experiment is carried out according to a preset experimental plan, wherein the preset plan is 27 groups of experiments with 5 factors and 3 levels; the basic data set is obtained by experimental evaluation indicators such as particle size, particle size distribution, and binding rate.
[0015] Preferably, the basic data set is enhanced by generating an adversarial network (GAN) to generate a virtual sample augmented data set.
[0016] The present invention also provides a method for optimizing the process parameters of paclitaxel medicine, comprising:
[0017] An acquisition module is used to obtain the basic data set for paclitaxel drug production;
[0018] The preprocessing module is used to perform data enhancement on the basic data set and generate a virtual sample augmented data set;
[0019] The optimization module is used to optimize the initial weights and biases of the BP neural network through genetic algorithms;
[0020] A training module, used to train the optimized BP neural network by using virtual sample to expand the data set;
[0021] The search module is used to use the genetic algorithm to reversely search for the optimal parameter combination in the prediction results of the optimized BP neural network.
[0022] Preferably, the human albumin incubation temperature, incubation time, emulsifying disperser speed, albumin aqueous solution feeding speed, and paclitaxel feeding speed are set, and the experiment is carried out according to a preset experimental plan, wherein the preset plan is 27 groups of experiments with 5 factors and 3 levels; the basic data set is obtained by experimental evaluation indicators such as particle size, particle size distribution, and binding rate.
[0023] Preferably, the preprocessing module performs data enhancement on the basic data set by generating an adversarial network (GAN) to generate a virtual sample augmented data set.
[0024] The present invention also provides a paclitaxel drug process parameter optimization system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a paclitaxel drug process parameter optimization method when executed by the processor.
[0025] The present invention also provides a storage medium, on which a computer program is stored, and the computer program executes the method for optimizing the process parameters of paclitaxel medicine when it is run.
[0026] The present invention has the following technical effects:
[0027] 1. When traditional machine learning is difficult to characterize the complex nonlinear relationship between parameters, the present invention uses deep learning methods combined with genetic algorithm optimization to effectively model and find the relationship between parameters, solving the problems of slow convergence and easy falling into local optimal solutions of BP neural network, while retaining the self-organizing ability and nonlinear mapping ability of BP neural network; at the same time, the combination of the two enables the model to obtain deeper feature information, handle the nonlinear relationship between parameters, and thus obtain more accurate predictions.
[0028] 2. The present invention adopts data enhancement for small sample data. When the production data of paclitaxel drug is limited, the adversarial network GAN is used to generate realistic virtual samples, expand the original data set, and improve the robustness of the model and its predictive ability.
[0029] 3. The present invention uses the excellent parallel search capability of genetic algorithms. After the model is established, after the parameter range and prediction target are given, the algorithm is used to reversely search for the optimal parameter group, ultimately achieving the effect of optimizing the process parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0031] Figure 1 This is a flow chart of a method for optimizing process parameters of paclitaxel drug according to an embodiment of the present invention;
[0032] Figure 2 To generate a flowchart for adversarial networks;
[0033] Figure 3 Flowchart of BP neural network optimized by genetic algorithm. DETAILED DESCRIPTION
[0034] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0035] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0036] Embodiment 1:
[0037] like Figure 1 As shown, the embodiment of the present invention provides a method for optimizing process parameters of paclitaxel drug, comprising:
[0038] Step 1: Determine the experimental plan through CCD central composite design. After determining the number of key parameters and the adjustable range, obtain the number of levels and factors of the experimental design. After setting various experimental parameters, obtain the final experimental results. In theory, the more data and the smaller the experimental error, the better the training effect of the model.
[0039] The specific experimental scheme is as follows: set the human albumin incubation temperature, incubation time, emulsification dispersion machine speed, albumin aqueous solution feeding speed, paclitaxel feeding speed, and conduct the experiment according to the preset experimental scheme. The basic data set required for the model is obtained by experimental evaluation indicators of particle size, particle size distribution and binding rate. Among them, the preset scheme is 5 factors and 3 levels, with a total of 27 groups of experiments; 3 levels means that each parameter of the experiment has 3 different types of settings, such as the human albumin incubation temperature is set to 65℃, 67.5℃, and 70℃, a total of three different temperatures; 5 factors means that there are 5 variables, such as human albumin incubation temperature, incubation time, emulsification dispersion machine speed, albumin aqueous solution feeding speed, paclitaxel feeding speed, a total of 5 experimental parameters; 27 groups of experiments mean that 27 experiments need to be done according to the experimental scheme designed by CCD.
[0040] Step 2: Generate adversarial networks (GANs) to perform data enhancement on the basis of the basic data set and generate virtual samples to expand the data set, such as Figure 2 as shown in .
[0041] The specific process is as follows: After generating high-dimensional random noise, it is input into the generator after the data preprocessing step. The generator consists of a fully connected layer and an activation function ReLU. After the virtual sample is generated by the generator, it is spliced with the real data and the label value is added. The real data is represented by 1 and the virtual data is represented by 0. After the splicing is completed, it is sent to the discriminator. The discriminator has a deeper network structure. Finally, Sigmoid outputs a probability value. The discriminator will distinguish which parts of the data are real and which parts of the data are virtually generated. The two are constantly confronting each other. The generator constantly tries to generate realistic virtual samples to deceive the discriminator, while the discriminator tries to distinguish between real data and virtual data as much as possible, and gradually improves their respective capabilities. Finally, realistic virtual samples can be generated in high-dimensional space through the trained generator, so as to alleviate the problem of insufficient training samples for the next step of BP neural network. The generated data is judged to be realistic through the correlation coefficient heat map.
[0042] Step 3: Optimize the initial parameters of the BP neural network through genetic algorithm, such as Figure 3 As shown, the BP neural network model is:
[0043] y=f w,b (x)
[0044] where fw,b (x) is a function determined by weight w and bias b, defined as:
[0045] f w,b (x) = f(w2f(w1x+b1)+b2)
[0046] Where w1 is the weight matrix between the input layer and the hidden layer, w2 is the weight matrix between the hidden layer and the output layer, b1 is the bias vector between the input layer and the hidden layer, and b2 is the bias vector between the hidden layer and the output layer.
[0047] The optimization process of genetic algorithm can be expressed as follows: after the initial population is randomly generated, the weights and biases of the neural network are gradually optimized through operations such as selection, crossover, and mutation, so that the optimization problem can be expressed as
[0048] MIN w,b L(f w,b (x),y)
[0049] Where L is the loss function. The population is continuously iterated until an individual that can be used as an approximate optimal solution to the problem is generated in the population. After combining with the BP neural network, the loss function, i.e., the mean square error, is used as the optimization target. This model can be expressed as:
[0050] MIN w,b MSE(f(w2f(w1x+b1)+b2),y)
[0051] Where MSE is the mean square error loss function in the BP neural network, x is the input vector, and y is the target output vector.
[0052] A genetic algorithm is used to continuously optimize the weights w and biases b in the BP neural network. The mean square error (MSE) is used as the fitness function. Each individual represents a set of weights and biases. After selection, crossover, mutation, etc., it is continuously iterated to approach the global optimal solution, making the model's predicted value closer to the true value.
[0053] Step 4: After the BP neural network with initial weights and biases optimized by genetic algorithm, the GAN-expanded data set is used to train the model, and 10% of the original data without duplication is used as the test set of the model. After adjusting the parameters for the experimental data many times. In the process of model establishment, many targeted optimizations are added at the same time, including: 1. Adding an early stopping mechanism. When the performance of the model in the validation set does not improve for a long time, stop training to minimize the probability of overfitting. 2. Adding a model mechanism for selecting the optimal loss value can save the best model in training and reduce the randomness during training. Before this, the model of the last training is selected, and the model of the last training is not necessarily the model with the lowest loss value. 3. Adding a visualization process, when generating virtual samples, add a correlation coefficient heat map to analyze its correlation. After the model training is completed, the loss value changes of the training set and the validation set are output. After the model test is completed, a line chart comparing the true value and the predicted value is output. 4. Adding a dynamic adjustment learning rate mechanism. If the loss value does not decrease after exceeding the set patience value, the learning rate is appropriately reduced.
[0054] Finally, you can see a line chart comparing the true value and the predicted value on the test set. The prediction accuracy of the model can be judged by calculating the average accuracy of the prediction. If the effect is good, save the model file and proceed to the next step.
[0055] Step 5: The BP neural network trained through the above steps has a strong nonlinear mapping ability. Use the adjusted genetic algorithm to reverse search for the best input parameter group. The specific steps are as follows: set the number and range of input parameters, adjust the working goal of the genetic algorithm to find input parameters that are as close to the target value as possible, set the population size, number of iterations, crossover probability, and mutation probability. Load the trained BP neural network model and normalized parameters to ensure the consistency of input data. Given the input parameters X = [x1, x2, ..., x n ], after being standardized, it is input into the pre-trained BP neural network model, and the output prediction value is:
[0056]
[0057] Where: scaler(X) represents the normalization of the input data X. f is the trained neural network model, including weights W and bias b, which is the output value predicted by the model.
[0058] The fitness evaluation is performed by calling the evaluate_target function, using the pre-trained model to calculate the difference between the output value of the parameter combination and the target value, and then determining its fitness score. The formula is as follows:
[0059]
[0060] where ytarget Output value for the set target
[0061] At the same time, this method uses custom crossover and mutation operations, and controls the range of mutation parameters through the custom_mutate function. The formula is as follows:
[0062] x′ i =x i +Δx,Δx~N(0,σ 2 )
[0063] Where △x is a random disturbance term that obeys a Gaussian distribution, and σ is the standard deviation of the Gaussian distribution.
[0064] When generating offspring, the check_bounds pruning function is introduced to ensure that all generated parameter combinations are within the set reasonable range. The formula is as follows:
[0065]
[0066] where bounds min is the lower limit of the parameter x, bounds max is the upper limit of parameter x
[0067] Through continuous iteration of genetic algorithm, the optimal parameter combination X is found. * To minimize the fitness function, the overall objective formula is as follows:
[0068] X * =argminF target (X)
[0069] The final optimization results show the top five best parameter combinations, whose outputs are close to the target values. After multiple iterations, individuals that meet the particle size and particle size distribution and whose binding rate is close to the target value are searched. These individual values are the predicted optimal parameter groups for the next step of experimental verification.
[0070] Step 6: After the model gives a suitable prediction result, further experiments are conducted to verify the results. If the experiment meets the expected results, the next step of amplification experiment is carried out. If the experimental effect is poor, the model is further optimized based on the test results, and the sample size of the data set is increased to improve the model training effect.
[0071] The present invention has the following technical effects:
[0072] 1. The parameter setting in the traditional drug production process usually relies on the experiment of controlling a single variable (One-Variable-at-a-Time, OVAT), that is, adjusting a certain parameter one by one while keeping all other parameters unchanged, and observing its influence on the process results. Although this method is simple and intuitive, it also has some significant disadvantages, such as low experimental efficiency, ignoring the interaction between parameters, non-optimal parameter combination, and waste of resources. The present invention proposes a BP neural network based on genetic algorithm optimization to predict a group of more optimal process parameters, thereby solving various problems of obtaining parameters by controlling a single variable before by focusing on the nonlinear relationship between parameters. Genetic algorithm is used to make up for the situation that BP neural network is easy to fall into the local optimal solution, and the model loss value is reduced as much as possible. The two complement each other, thereby obtaining the characteristic relationship between various process parameters and drug evaluation indicators. Finally, the effect of the optimal parameter group is verified by actual production, and more effective process parameters are provided for enterprises to produce paclitaxel drugs.
[0073] 2. When it is difficult to obtain a data set and the amount of data is small, a data enhancement method using a generative adversarial network (GAN) is used specifically for the expansion of paclitaxel drug experimental data. Since paclitaxel drug experiments are usually limited by the difficulty of sample acquisition and insufficient data, the present invention uses a generative adversarial network to automatically generate high-quality, realistic virtual experimental data, effectively expanding the scale of limited experimental data sets, overcoming the dependence of traditional experimental methods on large-scale real data, and improving the efficiency and accuracy of data analysis and drug screening. Through this method, the generalization ability of the model can be improved under limited sample conditions, providing more effective data support for drug development and experimental research.
[0074] 3. Traditional parameter optimization methods, such as the traversal method, require all possible parameter combinations to be tested and evaluated one by one, resulting in high computational complexity and high time consumption. Especially in high-dimensional space, the computational cost of the traversal method grows almost exponentially. Therefore, traditional methods are difficult to meet the needs of fast and accurate optimization in the case of complex models or a large number of parameter combinations. Therefore, the present invention uses an innovative reverse parameter optimization method, which replaces the traditional fitness evaluation mechanism with a deep learning model by introducing a genetic algorithm and a pre-trained BP neural network, thereby realizing an efficient reverse search for parameters. The fitness of the parameter combination is directly evaluated by the neural network model, which greatly reduces the computational complexity, so that the algorithm can approach the optimal solution in a shorter time. The genetic algorithm is combined with the pre-trained BP neural network, and the traditional fitness evaluation mechanism is replaced by a deep learning model to realize an efficient reverse search of parameter combinations.
[0075] Embodiment 2:
[0076] The embodiment of the present invention also provides a paclitaxel drug process parameter optimization device, comprising:
[0077] An acquisition module is used to obtain the basic data set for paclitaxel drug production;
[0078] The preprocessing module is used to perform data enhancement on the basic data set and generate a virtual sample augmented data set;
[0079] The optimization module is used to optimize the initial weights and biases of the BP neural network through genetic algorithms;
[0080] A training module, used to train the optimized BP neural network by using virtual sample to expand the data set;
[0081] The search module is used to use the genetic algorithm to reversely search for the optimal parameter combination in the prediction results of the optimized BP neural network.
[0082] As an implementation method of an embodiment of the present invention, the incubation temperature, incubation time, emulsifying disperser speed, albumin aqueous solution feeding speed, and paclitaxel feeding speed of human albumin are set, and the experiment is carried out according to a preset experimental scheme, wherein the preset scheme is 27 groups of experiments with 5 factors and 3 levels; the basic data set is obtained by experimental evaluation indicators such as particle size, particle size distribution, and binding rate.
[0083] As an implementation method of an embodiment of the present invention, the preprocessing module performs data enhancement on the basic data set by generating an adversarial network (GAN) to generate a virtual sample expanded data set.
[0084] Embodiment 3:
[0085] The embodiment of the present invention further provides a paclitaxel drug process parameter optimization system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a paclitaxel drug process parameter optimization method when executed by the processor.
[0086] Embodiment 4:
[0087] An embodiment of the present invention further provides a storage medium, wherein a computer program is stored on the storage medium, and the computer program executes a method for optimizing process parameters of a paclitaxel drug when the computer program is run.
[0088] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for optimizing process parameters of paclitaxel drug, characterized in that: include: Obtain basic data sets for paclitaxel drug production; Perform data enhancement on the basic data set to generate a virtual sample augmented data set; Optimize the initial weights and biases of the BP neural network through genetic algorithms; The optimized BP neural network is trained by using virtual samples to expand the data set; The genetic algorithm is used to reversely search and optimize the optimal parameter combination in the prediction results of the BP neural network.
2. The method for optimizing the process parameters of paclitaxel medicine according to claim 1, characterized in that: The human albumin incubation temperature, incubation time, emulsifying disperser speed, albumin aqueous solution feeding speed, and paclitaxel feeding speed were set, and the experiment was carried out according to the preset experimental plan, where the preset plan was 27 groups of experiments with 5 factors and 3 levels. The basic data set was obtained by experimental evaluation indicators such as particle size, particle size distribution, and binding rate.
3. The method for optimizing the process parameters of paclitaxel medicine according to claim 2, characterized in that: The basic data set is enhanced by generating adversarial networks (GANs) to generate virtual sample augmented data sets.
4. A paclitaxel drug process parameter optimization device, comprising: An acquisition module is used to obtain the basic data set for paclitaxel drug production; The preprocessing module is used to perform data enhancement on the basic data set and generate a virtual sample augmented data set; The optimization module is used to optimize the initial weights and biases of the BP neural network through genetic algorithms; A training module, used to train the optimized BP neural network by using virtual sample to expand the data set; The search module is used to use the genetic algorithm to reversely search for the optimal parameter combination in the prediction results of the optimized BP neural network.
5. The device for optimizing process parameters of paclitaxel medicine according to claim 4, characterized in that: The human albumin incubation temperature, incubation time, emulsifying disperser speed, albumin aqueous solution feeding speed, and paclitaxel feeding speed were set, and the experiment was carried out according to the preset experimental plan, where the preset plan was 27 groups of experiments with 5 factors and 3 levels. The basic data set was obtained by experimental evaluation indicators such as particle size, particle size distribution, and binding rate.
6. The device for optimizing process parameters of paclitaxel medicine according to claim 5, characterized in that: The preprocessing module performs data enhancement on the basic data set through the generative adversarial network (GAN) to generate a virtual sample augmented data set.
7. A paclitaxel drug process parameter optimization system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the method for optimizing the process parameters of paclitaxel medicine according to any one of claims 1 to 3 is executed.
8. A storage medium, characterized in that: The storage medium stores a computer program, and the computer program executes the method for optimizing the process parameters of paclitaxel medicine according to any one of claims 1 to 3 when running.
Citation Information
Cited By
Genetic algorithm-based polycaprolactone polyol synthesis path optimization method
CN120808929A