A method and system for intelligent generation of a freeze-drying formulation
By constructing an intelligent method for generating lyophilized formulations, and using interface fragility and backbone fragility indices to generate a candidate space, screening and process window solving are performed together. This solves the problem of unstable lyophilized formulations in the preparation of small-batch personalized mRNA-LNP formulations, and achieves rapid customized formulation generation and stable bioactivity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHINVA MEDICAL INSTR CO LTD
- Filing Date
- 2026-04-08
- Publication Date
- 2026-07-31
AI Technical Summary
In the context of freeze-drying preparation of small-batch personalized mRNA-LNP formulations, existing technologies lack the ability to generate automated, phased, customized freeze-drying formulations. They cannot quickly respond to changes in new mRNA sequences and different proportions of lipid nanoparticle carriers, resulting in the inability to effectively match specific stresses during freeze-drying and leading to instability in the freeze-drying formulation.
A method for intelligently generating freeze-dried formulations is adopted. By acquiring parameters of freeze-dried samples, an interface fragility index and a skeleton fragility index are constructed to generate a candidate space. Screening and joint solution of process windows are performed. Combined with micro-batch verification and iterative correction, a target calibration model is generated to achieve rapid customization of formulations.
It enables rapid cold start of small batches of new samples, accurately quantifies specific stress matching, ensures the stability of the appearance of the freeze-dried cake and its bioactivity after reconstitution, and solves the problem of unstable freeze-drying formulation generation in existing technologies.
Smart Images

Figure CN122493973A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biopharmaceutical preparation technology, specifically to a method and system for intelligent generation of lyophilized formulations. Background Technology
[0002] In the scenario of lyophilization preparation of small-batch personalized mRNA-LNP formulations, it is necessary to rapidly generate formulations that can simultaneously ensure the appearance of the lyophilized cake and the stability of bioactivity after reconstitution, for constantly changing mRNA sequences, different lipid nanoparticle carrier ratios, and a small number of samples for "cold start". The novel technical challenge in this scenario is how to automatically and in stages identify and resist the specific stresses generated during the lyophilization process when the sample size is extremely small and the sequence / carrier parameters are frequently switched, so as to quickly generate customized lyophilized formulations that match the LNP window and avoid a lot of trial and error and reliance on human experience.
[0003] Currently, existing technologies typically attempt to balance cake formation and reconstituted performance by reverting to or referencing existing empirical formulations, using common protective agent combinations, and adjusting freeze-drying parameters. Alternatively, they may employ experimental optimization processes such as designing experiments to gradually adjust the formulation and process. In practice, most solutions first optimize the freeze-drying appearance indicators and then correct them through biological testing after reconstitution. This relies heavily on experimental data and manual judgment, and lacks automated formulation generation and phased modeling methods for small-batch, variable-sequence scenarios.
[0004] In existing technologies, when faced with new mRNA sequences, different LNP vector ratios, and small batches of "cold start" samples, the approach typically reverts to the closest empirical formulation or existing protectant combination, lacking the ability to customize formulations for individualized / small sample sizes. Furthermore, it cannot model and adapt the specific stresses generated during lyophilization in stages, resulting in the inability to quickly, stably, and transferably generate lyophilization formulations for LNPs in small-batch, personalized scenarios. This leads to instability in formulation performance when switching between different sequences / vectors / process conditions, or the need for extensive trial and error. To address these issues, this invention proposes an intelligent lyophilization formulation generation method and system. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method and system for intelligently generating freeze-dried formulations, thereby resolving the problems mentioned in the background section.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for intelligently generating freeze-dried formulations, comprising: Obtain parameters from freeze-dried samples and generate a parameter set; Indices are constructed from the parameter set to obtain the interface vulnerability index and the skeleton vulnerability index; Based on the ratio of the interface vulnerability index to the skeleton vulnerability index, a candidate space is generated; The candidate space is screened to obtain the preferred formulation, and the process window is jointly solved for the preferred formulation to obtain a multi-objective comprehensive function; The multi-objective synthesis function and the preferred formulation were subjected to micro-batch verification to obtain verification feedback data. Based on the verification feedback data, the multi-objective synthesis function is iteratively corrected to obtain the target calibration model, and the target calibration model is robustly solved to obtain the target formulation.
[0007] Preferably, the freeze-dried sample parameters include nucleic acid sequence parameters, lipid nanoparticle process parameters, pretreatment characterization data, thermal characteristic parameters, and freeze-drying stress parameters; The nucleic acid sequence parameters include at least the messenger RNA length, GC content, and the proportion of modified nucleosides; The lipid nanoparticle process parameters include at least the lipid composition ratio and the mixing flow rate; The pretreatment characterization data includes at least the initial particle size, polydispersity index, zeta potential, initial encapsulation efficiency, initial pH, and osmotic pressure of the sample before freezing. The thermal characteristic parameters include at least the glass transition temperature, the eutectic point, and the collapse initiation temperature; The freeze-drying stress parameters include the pH shift values before and after freezing.
[0008] Preferably, the step of constructing indices for the parameter set to obtain the interface fragility index and the skeleton fragility index includes: normalizing the characteristics of each parameter in the parameter set; performing a weighted mapping of the nucleic acid sequence parameters, the preprocessing characterization data, and the freeze-drying stress parameters to obtain the interface fragility index used to characterize the interface protection requirements of lipid nanoparticles; and performing a weighted mapping of the thermal characteristic parameters to obtain the skeleton fragility index used to characterize the requirements for collapse resistance and glass transition support.
[0009] Preferably, generating the candidate space based on the ratio of the interface fragility index and the skeleton fragility index includes: using the interface fragility index and the skeleton fragility index as feature vectors, searching in a historical sample library using local sensitive hashing or weighted cosine similarity, and extracting the stress response mapping relationship of historical samples with a similarity greater than or equal to a preset threshold as a cross-domain migration prior; and generating preliminary formulation combinations in a preset auxiliary material library according to functional partitions based on the ratio relationship and the cross-domain migration priors, wherein the functional partitions include at least independently adjustable interface protection groups, skeleton support groups, and buffer adjustment groups, and the candidate space is composed of multiple preliminary formulation combinations.
[0010] Preferably, the step of screening the candidate space to obtain the preferred formulation includes: performing feedforward thermal verification on each preliminary formulation combination in the candidate space; predicting the collapse initiation temperature of each preliminary formulation combination, and calculating the difference between the collapse initiation temperature and the preset maximum freeze-drying operating temperature; when the difference is greater than or equal to a preset safety margin, determining that the corresponding preliminary formulation combination has passed the verification and is selected as the preferred formulation; and removing preliminary formulation combinations that have not passed the verification.
[0011] Preferably, the process window of the preferred formulation is jointly solved to obtain a multi-objective comprehensive function, which includes: constructing a performance vector that includes cake integrity, freeze-drying window margin, remelting particle size drift, encapsulation loss, and residual moisture; using an integrated regressor to predict the performance vector of the preferred formulation under a preset freeze-drying window, and outputting the expected mean and uncertainty variance of each performance index; and constructing the multi-objective comprehensive function carrying an uncertainty penalty term based on the expected mean and the uncertainty variance.
[0012] Preferably, the step of performing micro-batch verification on the multi-objective synthesis function and the preferred formulation to obtain verification feedback data includes: calculating the overall uncertainty of the preferred formulation; when the overall uncertainty is greater than a preset uncertainty gating threshold, blocking the direct output of the formulation and automatically triggering the micro-batch verification process; selecting a preset number of formulation combinations from the preferred formulations to form a minimum verification set based on the maximum expected information gain or maximum model variance coverage strategy; performing a micro-freeze-drying experiment on the minimum verification set and collecting measured performance parameters as the verification feedback data.
[0013] Preferably, the step of iteratively correcting the multi-objective synthesis function based on the verification feedback data to obtain the target calibration model includes: calculating the micro-batch post-uncertainty after loading the verification feedback data; determining whether the micro-batch post-uncertainty is less than or equal to a preset micro-batch termination threshold; if not, performing Bayesian updates or weighted adjustments on the weights of the cross-domain migration prior according to the verification feedback data, and reconstructing the multi-objective synthesis function using the updated weights, repeatedly executing the micro-batch verification step until the micro-batch post-uncertainty meets the micro-batch termination threshold, and outputting the final target calibration model.
[0014] Preferably, the step of performing robust solution on the target calibration model to obtain the target formulation includes: using a constrained multi-objective evolutionary algorithm to perform global optimization on the target calibration model; outputting at least one set of formulations that meet the safety margin within a set uncertainty perturbation range, as well as the corresponding matching freeze-drying window and risk assessment identifier, as the target formulation.
[0015] The freeze-drying formulation intelligent generation system includes: The data acquisition module is used to acquire parameters of freeze-dried samples and generate parameter sets; The indicator construction module is used to construct indicators from the parameter set to obtain the interface vulnerability index and the skeleton vulnerability index. The candidate generation module is used to generate a candidate space based on the ratio of the interface vulnerability index and the skeleton vulnerability index. A screening feedforward module is used to screen the candidate space to obtain a preferred formulation; The solver module is used to jointly solve the process window of the preferred formula to obtain a multi-objective comprehensive function; The micro-batch verification module is used to perform micro-batch verification of the multi-objective synthesis function and the preferred formulation to obtain verification feedback data; The iterative calibration module is used to iteratively correct the multi-objective comprehensive function based on the verification feedback data to obtain the target calibration model. The results output module is used to perform robust solution on the target calibration model to obtain the target formulation.
[0016] This invention provides a method and system for intelligently generating freeze-dried formulations. It has the following beneficial effects: 1. This invention adopts a technical solution that combines vulnerability index construction with cross-domain migration priors to achieve the technical effect of accurate quantification of specific stress intelligent matching, realize the rapid cold start of personalized formulations for small batches of new samples, and solve the shortcomings of existing technologies that can only return to experience formulations when faced with new samples and lack the ability to generate customized formulations.
[0017] 2. This invention adopts a technical solution of functional partitioning generation and multi-objective joint solution to achieve the technical effect of coordinating appearance requirements with bioactivity, realize a high degree of synergy between cake morphology and reconstitution activity, and solve the shortcomings of existing technologies that optimize appearance and activity separately, resulting in one aspect being neglected and it being difficult to achieve stable balance. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the process of the present invention; Figure 2 This is a schematic diagram of the system framework structure of the present invention. Detailed Implementation
[0019] To enable those skilled in the art to understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort should fall within the scope of protection of the present invention.
[0020] The present invention will now be described in detail with reference to the accompanying drawings: Please see the appendix Figure 1 and attached Figure 2 This invention provides a method for intelligently generating freeze-dried formulations, comprising: Obtain parameters from freeze-dried samples and generate a parameter set; The parameters of the freeze-dried samples include nucleic acid sequence parameters, lipid nanoparticle process parameters, pretreatment characterization data, thermal characteristic parameters, and freeze-drying stress parameters; Nucleic acid sequence parameters include at least messenger RNA length, GC content, and the percentage of modified nucleosides. The process parameters for lipid nanoparticles include at least the lipid composition ratio and the mixing flow rate. Pretreatment characterization data should include at least the initial particle size, polydispersity index, zeta potential, initial encapsulation efficiency, initial pH, and osmotic pressure of the sample before freezing. Thermal characteristic parameters include at least the glass transition temperature, eutectic point, and collapse initiation temperature; Freeze-drying stress parameters include pH shift values before and after freezing; The parameter set was used to construct indices to obtain the interface fragility index and the skeleton fragility index. This included: normalizing the characteristics of each parameter in the parameter set; weighting and mapping the nucleic acid sequence parameters, preprocessing characterization data, and freeze-drying stress parameters to obtain the interface fragility index used to characterize the interface protection requirements of lipid nanoparticles; and weighting and mapping the thermal characteristic parameters to obtain the skeleton fragility index used to characterize the requirements for collapse resistance and glass transition support. Specifically, retrieve the original measurement data of the freeze-dried samples and summarize them to form an initial dataset; For the raw measurement data of different dimensions within the initial dataset, a maximum-minimum mapping rule is used to transform all absolute values, limiting the transformed values to the range of zero to one. Let the current value of a single parameter be x, and the maximum value of this type of parameter in the entire historical sample be x. max The minimum value is x min The normalization calculation formula is expressed as follows: The formula is used to calculate the dimensionless relative numerical eigenvector. A unified relative numerical feature vector is used to assign coefficients representing the importance of various parameters. A first weighted recombination is constructed for nucleic acid sequence parameters, pretreatment characterization data, and freeze-drying stress parameters. For example, when faced with a large pH shift before and after freezing, the system assigns an extremely high risk weight within the first weighted recombination, representing a sharp increase in the risk of interface damage. An independent second weighted recombination is constructed for thermal characteristic parameters; when the collapse initiation temperature is close to the preset operating temperature, the system assigns an extremely high collapse risk weight within the second weighted recombination.
[0021] Retrieve the normalized relative values of the nucleic acid sequence, the relative values of preprocessing characterization, and the relative values of freeze-drying stress. Multiply each of the three types of relative values by the corresponding coefficients within the first weighted recombination, and sum all the multiplication results. Assume the set of interface-related feature values is labeled X. A The corresponding weight set is labeled W. A The formula for the interface fragility index is: V i =∑(W A ×X A The summation of these values yields the final interface fragility index.
[0022] Retrieve the normalized relative values of thermal features. Multiply each extracted relative value of thermal features by the corresponding coefficient in the second weighted set, and sum all the multiplication results. Assume the set of skeleton-related feature values is labeled X. B The corresponding weight set is labeled W. B The formula for the skeleton vulnerability index is: V s =∑(W B ×X B The summation result calculated by this formula is the skeleton vulnerability index.
[0023] The interface vulnerability index and the skeleton vulnerability index are combined to form two-dimensional coordinate axis data, thus completing the construction of the index for the initial parameter set. Based on the ratio of the interface fragility index and the skeleton fragility index, a candidate space is generated, including: using the interface fragility index and the skeleton fragility index as feature vectors, searching in the historical sample library using local sensitive hashing or weighted cosine similarity, and extracting the stress response mapping relationship of historical samples with similarity greater than or equal to a preset threshold as a cross-domain migration prior; based on the ratio relationship and the cross-domain migration prior, generating preliminary formulation combinations in the preset auxiliary material library according to functional partitions, wherein the functional partitions include at least independently adjustable interface protection groups, skeleton support groups, and buffer adjustment groups, and the candidate space is composed of multiple preliminary formulation combinations; Specifically, the interface fragility index and skeleton fragility index obtained in the previous steps are retrieved, and their quantified ratio is calculated. Assuming this ratio is denoted as R, the calculation formula is: R = V i / V s The ratio R can intuitively reflect whether the current sample needs to resist interface fusion and leakage or prevent skeleton collapse during the freeze-drying process.
[0024] The interface fragility index and the skeleton fragility index are packaged together to construct a two-dimensional feature vector V=(V i V sThe two-dimensional feature vector is input into a pre-established historical sample database, and a full database comparison is performed using either a weighted cosine similarity algorithm or a locality-sensitive hashing algorithm. Taking weighted cosine similarity as an example, the similarity score T between the target vector and historical vectors is calculated. sim .
[0025] Set an empirical similarity threshold T0. The system automatically filters out all similarity scores T. sim Historical samples ≥T0 are collected, and outlier samples below this threshold are removed. From the historical samples with high similarity, the "stress response mapping relationship" under specific freeze-drying thermal conditions is extracted, that is, the data pattern of how a specific combination of excipients corresponds to a specific thermal safety margin, as a priori for cross-domain migration.
[0026] The system retrieves independently adjustable interface protection, skeleton support, and buffer adjustment groups from a pre-defined excipient library. Combining the calculated ratio R with the cross-domain migration prior extracted in the third step, it dynamically defines the upper and lower limit concentration ranges for each excipient group. If the ratio R > 1, based on migration priors, the upper limit of the concentration of the interface protection group is adaptively increased, while the concentration range of the skeleton support group is appropriately compressed. If the ratio R < 1, adjust in the opposite direction to expand the concentration range of the skeletal support group.
[0027] Within the defined concentration boundaries of each functional zone, a gridded sampling or orthogonal arrangement algorithm is used to extract excipient components of specific concentrations from the interface protection group, skeleton support group, and buffer adjustment group. The extracted components are then arranged and combined according to a set step size to generate a large number of preliminary formulation combinations.
[0028] All generated preliminary formula combinations are deduplicated and invalidated. All remaining valid preliminary formula combinations are then grouped together to form a complete candidate space. The candidate space is screened to obtain the preferred formulation, and the process window is jointly solved for the preferred formulation to obtain a multi-objective comprehensive function; Perform feedforward thermal verification on each preliminary formulation combination in the candidate space; predict the collapse initiation temperature of each preliminary formulation combination and calculate the difference between the collapse initiation temperature and the preset maximum freeze-drying operating temperature; when the difference is greater than or equal to the preset safety margin, the corresponding preliminary formulation combination is determined to pass the verification and is selected as the preferred formulation; remove the preliminary formulation combinations that fail the verification.
[0029] A performance vector is constructed that includes cake integrity, freeze-drying window margin, reconstitution particle size drift, encapsulation loss, and residual moisture. An integrated regressor is used to predict the performance vector of the optimal formulation under a preset freeze-drying window, and the expected mean and uncertainty variance of each performance index are output. Based on the expected mean and uncertainty variance, a multi-objective comprehensive function with an uncertainty penalty term is constructed. Specifically, the system iterates through each preliminary formulation combination in the candidate space generated by the preceding steps, extracting the types and mass or volume concentrations of excipients in each formulation. With specific concentration values, the system can initiate preliminary simulations of thermal behavior without waiting for actual freeze-drying experiments, significantly reducing trial-and-error costs.
[0030] For each extracted preliminary formulation combination, a pre-defined thermal mapping function is invoked. By calculating the weighted contribution of the glass transition temperature and mass fraction of each component, the collapse initiation temperature corresponding to the preliminary formulation combination is predicted, denoted as T. c .
[0031] Obtain the currently preset maximum freeze-drying operating temperature, i.e., the highest product temperature that the item is allowed to experience in a single drying stage, denoted as T. max The predicted collapse initiation temperature T c With the maximum freeze-drying operating temperature T max Perform the difference calculation; the formula is: ΔT = T c -T max The resulting difference ΔT represents the distance of thermal risk faced by the formula under the existing process settings.
[0032] Preset safety margin threshold ΔT safe For example, it can be set to be greater than or equal to a specific temperature to mitigate the risk of localized temperature fluctuations caused by process equipment. The calculated difference ΔT is then compared with this safety margin ΔT. safe Compare and distinguish: When ΔT≥ΔT safe When the system determines that the preliminary formulation combination has sufficient thermal anti-collapse space and successfully passes the feedforward thermal test, the system retains the preliminary formulation combination and marks it as the preferred formulation. When ΔT < ΔT safe If the system determines that the formula has an extremely high risk of collapse during process execution, it will directly remove the formula from the candidate space.
[0033] The system extracts all validated and optimized formulations, and establishes core evaluation dimensions for the final quality requirements of lyophilized formulations: cake integrity (C), lyophilization window margin (W), reconstitution particle size drift (ΔD), encapsulation loss (ΔE), and residual moisture (R). These five dimensions are combined into a unified performance vector P: P=[C,W,ΔD,ΔE,R]. This comprehensive performance vector overcomes the limitations of previous single-dimensional approaches that focused solely on appearance without considering activity or prioritized encapsulation efficiency, leading to performance collapse.
[0034] Each preferred formulation is matched with one or more sets of preset freeze-drying process window parameters to form a "formulation-process" joint input item.
[0035] The above joint input terms are fed into an ensemble regressor trained on historical data. Unlike traditional single-output models, this ensemble regressor outputs the expected mean and uncertainty variance for each evaluation metric in the performance vector P.
[0036] Retrieve the expected mean and uncertainty variance of each of the above output indicators. First, assign corresponding weight coefficients λ according to the importance of each performance indicator. i The expected mean is weighted and summed; simultaneously, the uncertainty variances of each indicator are fused to calculate the overall uncertainty parameter U, and a penalty weight λ is assigned to it. U Finally, by integrating the two, a multi-objective synthesis function J is constructed, the mathematical expression of which can be described as follows: ; where μ i Let be the predicted expected mean of performance index i; The multi-objective synthesis function and the optimized formulation were validated in a micro-batch to obtain validation feedback data; Calculate the overall uncertainty of the preferred formulation; when the overall uncertainty exceeds the preset uncertainty gate threshold, block the direct output of the formulation and automatically trigger the micro-batch verification process; based on the maximum expected information gain or maximum model variance coverage strategy, select a preset number of formulation combinations from the preferred formulation to form a minimum verification set; perform micro-freeze-drying experiments on the minimum verification set and collect measured performance parameters as verification feedback data; Specifically, the system retrieves the uncertainty variance and expected mean of various performance indicators output by the ensemble regressor for each optimized formulation during the construction of the multi-objective comprehensive function. To eliminate dimensional differences between different evaluation indicators, the overall uncertainty parameter of the optimized formulation is calculated by summing the coefficient of variation or normalized variance. Assuming the overall uncertainty is denoted as U, the calculation formula can be expressed as: μ i Let σ be the expected mean of performance index i. i Let i be the discrete quantity representing the uncertainty of performance index i; A pre-defined uncertainty threshold, denoted as T, is used.U The calculated overall uncertainty U is then compared with the gate threshold T. U Real-time comparison: When U≤T U When U > T, the model's prediction of the formula has extremely high confidence and can be directly approved; when U > T U This indicates that, given the current proportion of new sequences or vectors, relying solely on algorithmic prediction carries a high risk of failure. In this case, a gating blocking command is immediately triggered to intercept the direct output of the formulation, and the downstream micro-batch validation process is automatically activated.
[0037] After the gating block is triggered, the system will not conduct physical experiments on all preferred recipes to avoid high trial-and-error costs. Instead, based on the strategy of "maximum expected information gain" or "maximum model variance coverage," it intelligently selects the most representative preset number of recipe combinations from a massive pool of preferred recipes. Specifically, the system calculates how much the total variance of the prediction model would be reduced if each recipe were actually tested; or it deliberately selects recipes with prediction variance distributions at different extremes. With this intelligent selection strategy based on maximizing information gain or variance coverage, the system can use the minimum number of physical samples to explore the most valuable true boundaries in the current unknown space, forming a minimum validation set.
[0038] The formulation documents within the selected minimum validation set are distributed to the laboratory or micro-pilot workshop. Operators or automated workstations prepare extremely small batches of physical samples according to the formulation documents. These samples are then placed in a micro-controlled cold trap or a small experimental freeze dryer, and the actual freeze-drying process is executed strictly according to the preset freeze-drying process window of the system during the joint solution phase.
[0039] After the micro-freeze-drying experiment was completed, the freeze-dried products underwent physical and biological characterization tests. Instruments were used to accurately measure and record the batch sample's cake integrity score, actual reconstitution particle size drift, actual encapsulation loss rate, and actual residual moisture percentage under real-world conditions. The system then structured and summarized these real-world observations from the laboratory instruments, packaging them to generate validation feedback data. Based on the verification feedback data, the multi-objective synthesis function is iteratively corrected to obtain the target calibration model, and the target calibration model is robustly solved to obtain the target formulation; Calculate the micro-batch post-uncertainty after loading the validation feedback data; determine whether the micro-batch post-uncertainty is less than or equal to the preset micro-batch termination threshold; if not, perform Bayesian update or weighted adjustment of the weights of the cross-domain migration prior based on the validation feedback data, and reconstruct the multi-objective synthesis function using the updated weights, and repeatedly execute the micro-batch validation steps until the micro-batch post-uncertainty meets the micro-batch termination threshold, and output the final target calibration model.
[0040] A constrained multi-objective evolutionary algorithm is used to globally optimize the target calibration model; at least one set of formulations that meet the safety margin within the set uncertainty perturbation range, along with the corresponding matching freeze-drying window and risk assessment label, are output as the target formulations; Specifically, the system receives real-world validation feedback data from micro-freeze-drying experiments. These real physical observations are then used as new known anchor data and re-fed into the ensemble regressor. Based on these newly added real data points, the system recalculates the overall prediction variance of the candidate space or preferred formulation at this point. Let the updated overall uncertainty be denoted as U. post .
[0041] Pre-set a micro-batch termination threshold, assuming it is labeled Tu. final The system will calculate the micro-batch post-uncertainty U. post With the termination threshold Tu final Comparison: When U post ≤Tu final At that time, it is determined that the current model has passed the calibration of a small number of physical experiments and is capable of guiding large-scale production; When U post >Tu final If the current model is found to have a large deviation, then weight updates and the next round of micro-batch validation are required.
[0042] If the termination condition is not met, the system extracts the true values from the validation feedback data and compares them with the expected mean given by the pre-micro-batch ensemble regressor to calculate the prediction error for each dimension. Based on this prediction error, the system uses a Bayesian update algorithm or a weighted adjustment strategy to correct the prior weights of the upstream cross-domain migration. Specifically: if the prior patterns of a historical sample migration lead to a large prediction error, the system will significantly reduce the similarity weight of that historical sample; conversely, it will increase the weight.
[0043] Using the updated prior weights, the system resets and fine-tunes the underlying parameters of the ensemble regressor, and re-outputs the expected mean and uncertainty variance of each indicator, thereby reconstructing a multi-objective integrated function with a new uncertainty penalty term. Then, based on the reconstructed function, a minimum validation set is selected again, and micro-batch validation and error calculation steps are performed iteratively. This continues until, after a certain iteration, U... post Finally, the condition is satisfied that it is less than or equal to Tu. final The system immediately terminates the loop and outputs the model with the currently locked weights as the final target calibration model.
[0044] The system incorporates a constrained multi-objective evolutionary algorithm, using a target calibration model that has already been calibrated with real data. The target calibration model is then used as the fitness evaluation function to perform global optimization within a candidate space that fully satisfies the feedforward thermal safety margin.
[0045] During the global optimization process of the evolutionary algorithm, the system deliberately applies predetermined uncertainty perturbations to the formulation component ratios and freeze-drying process parameters in a virtual environment. The system evaluates whether the performance predictions output by the target calibration model are stable and meet the safety margin after these perturbations. Peak solutions that are "at a single optimal point but have extremely poor anti-interference capabilities" are eliminated, while robust solutions that perform stably within the perturbation range are retained.
[0046] After the evolutionary algorithm converges, the system does not directly provide a unique solution, but outputs at least one set of target recipes that meet the robustness requirements through comprehensive consideration.
[0047] Example 1: This embodiment focuses on a newly developed mRNA vaccine for cancer treatment. Faced with such "cold start" samples, existing technologies can only revert to completely unsuitable empirical formulations, failing to provide targeted stress protection. This embodiment achieves customized formulation generation through feature extraction, cross-domain transfer prior algorithms, and micro-batch iterative calibration.
[0048] Detailed execution steps and procedures The basic parameters of the target sample were obtained as follows: messenger RNA length 2500 nt, GC content 55%, novel lipid composition, initial particle size 85 nm, encapsulation efficiency 92%, glass transition temperature, and pH shift before and after freezing (ΔpH 0.8). The above parameter set was normalized using the minimax mapping method. Taking any feature value x as an example, the maximum value of this feature in the entire historical sample set is x. max The minimum value is x min Normalized feature x norm The formula for calculating is: ; Assign two independent weight coefficient matrices W A and W B W A Corresponding sequence, characterization and stress parameters, W B Corresponding thermal parameters.
[0049] The interface fragility index V is obtained by multiplying the normalized data with the corresponding weights. i =∑(W A ×X A ) and skeleton vulnerability index V s =∑(W B ×X BBecause this sample used a new sequence and had a large pH drift, the interface fragility index V was calculated. i Extremely high, the ratio is calculated to R=1.8, indicating that interface fragility dominates, X A X is the importance coefficient of the interface fragility index. B This represents the importance coefficient of the skeleton vulnerability index.
[0050] (V) i V s Using this as a feature vector, the target vector is not blindly matched against samples with "same formula" in the historical sample database, but rather against samples with "similar stress". A weighted cosine similarity score T is used to calculate the similarity score T between the target vector and historical vectors. sim ; The system sets the similarity threshold to T0 = 0.7. Extract T... sim From historical samples with a value ≥0.7, the stress response mapping relationship of "what auxiliary material combination can maximize the safety margin under high interfacial stress" is extracted as a priori for cross-domain migration. Based on the ratio R=1.8 and migration prior, the system automatically increases the upper limit of the concentration setting for the interface protection group, generating a preliminary formulation combination. Because this sample is entirely new, the integrated regressor outputs a large prediction variance while simultaneously outputting the expected mean performance. The system calculates the overall uncertainty U=0.35. Since U exceeds the preset gating threshold, the system blocks direct formulation output and automatically triggers the micro-batch validation process. Based on the maximum expected information gain strategy, five formulation combinations that can eliminate the model variance blind zone to the greatest extent are selected from the candidate space, forming the minimum validation set. These five formulations are then tested on a micro-freeze-drying experimental platform, and actual reconstitution particle size drift and encapsulation loss are collected as validation feedback data. The system calculates the micro-batch post-uncertainty U after loading real feedback data. post If a deviation exists, the system updates the weights of the cross-domain transfer prior using Bayes' theorem: , Wherein, P(θ|D) is the posterior probability, representing the update result of parameter θ after obtaining the validation feedback data D; P(D|θ) is the likelihood function, representing the probability of observing data D given parameter θ; P(θ) is the prior probability, representing the original knowledge of parameter θ before observing new data; P(D) is the total probability of data D itself occurring; and θ is the parameter to be updated.
[0051] This will lead to a reduction in the weights of historical samples that have significant prediction bias. The multi-objective synthesis function is then reconstructed using the updated weights and solved again until U is reached. post ≤0.12, output the final customized target formula.
[0052] Example 2: This embodiment focuses on a high-concentration mRNA-LNP formulation. In existing technologies, to ensure the appearance of the cake-like form, large amounts of skeletal support agents such as mannitol are often added. However, this exacerbates crystallization stress, leading to LNP rupture after reconstitution and a sharp drop in encapsulation efficiency, resulting in a disconnect and contradiction between cake-like appearance and reconstitution activity. This embodiment achieves a perfect balance between these two aspects through a feedforward thermal verification and a joint solution algorithm.
[0053] Detailed execution steps: The system acquires the initial characterization of the high-concentration sample and calculates the interface fragility index V. i With the skeleton fragility index V S All are in the high-risk range. The system strictly extracts components from the pre-set excipient library according to functional zones: sucrose / trehalose is extracted from the interface protection group, mannitol (limiting dose) is extracted from the framework support group, and histidine is extracted from the buffer regulation group. A vast preliminary formulation combination candidate space is generated in a grid-like combination format. Before conducting any physical experiments, the system invoked thermodynamic models such as the Gordon-Taylor equation to predict the collapse initiation temperature T for each preliminary formulation combination. c Extract the preset maximum freeze-drying operating temperature T. max Calculate the difference between the two: ΔT = T c -T max Set safety margin ΔT safe All preliminary formulation combinations with direct ΔT < 3℃ for the 3℃ system were eliminated, and those that passed the verification were retained as the preferred formulations. To comprehensively consider cake integrity (C), freeze-drying window margin (W), remelting particle size drift (ΔD), encapsulation loss (ΔE), and residual moisture (R), an evaluation system comprising a multidimensional performance vector P=[C,W,ΔD,ΔE,R] is established. The ensemble regressor outputs the expected mean and uncertainty variance of each indicator, and the overall uncertainty U is calculated. The system constructs a multi-objective comprehensive function J: , where μ is the overall uncertainty parameter. i Let be the predicted expected mean of performance index i; Where, λ i λ represents the weighting coefficient for each performance indicator. U This is the uncertainty penalty weight. This function forces the system to simultaneously satisfy the condition of "extremely low uncertainty" when searching for a formula with high encapsulation efficiency and good cake formation; Using the target calibration model as the fitness evaluation environment, the system invokes a non-dominated sorting genetic algorithm for global optimization. The algorithm performs crossover and mutation operations on the population with the preferred formula, and through fast non-dominated sorting and crowding distance calculation, continuously brings the population closer to the optimal Pareto front. This means that the solution found is the theoretically optimal set of solutions that achieves the best pie appearance without sacrificing the encapsulation rate.
[0054] To avoid the selected optimal solution being vulnerable in production, the system applies uncertainty perturbations to the formulations on the Pareto front in a virtual environment. This eliminates peak solutions with poor anti-interference capabilities, ultimately outputting three sets of highly robust target formulations. Simultaneously, the system bundles these three target formulations with matching primary and secondary drying lyophilization curves, along with risk assessment indicators, providing a one-stop solution to the contradiction between appearance and activity in high-concentration formulations.
[0055] Summary of Example 1: This method enables precise customization of personalized stress-resistant formulations for new samples, effectively addressing the limitation of traditional methods that can only revert to fixed, empirical formulations when faced with new samples.
[0056] Summary of Example 2: The optimal formulation and matching freeze-drying window that achieves both excellent skeletal support and high activity retention in one step completely breaks through the technical bottleneck of prior art where appearance and activity are optimized separately.
[0057] Embodiments of the present invention have been presented and described. It will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to the embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for intelligent generation of a lyophilization formulation, the method comprising: receiving a plurality of inputs; and generating a lyophilization formulation based on the plurality of inputs. include: Obtain parameters from freeze-dried samples and generate a parameter set; Indices are constructed from the parameter set to obtain the interface vulnerability index and the skeleton vulnerability index; Based on the ratio of the interface vulnerability index to the skeleton vulnerability index, a candidate space is generated; The candidate space is screened to obtain the preferred formulation, and the process window is jointly solved for the preferred formulation to obtain a multi-objective comprehensive function; The multi-objective synthesis function and the preferred formulation were subjected to micro-batch verification to obtain verification feedback data. Based on the verification feedback data, the multi-objective synthesis function is iteratively corrected to obtain the target calibration model, and the target calibration model is robustly solved to obtain the target formulation.
2. The method of claim 1, wherein, The parameters of the freeze-dried sample include nucleic acid sequence parameters, lipid nanoparticle process parameters, pretreatment characterization data, thermal characteristic parameters, and freeze-drying stress parameters. The nucleic acid sequence parameters include at least messenger length, GC content, and the percentage of modified nucleosides; The lipid nanoparticle process parameters include at least the lipid composition ratio and the mixing flow rate; The pretreatment characterization data includes at least the initial particle size, polydispersity index, zeta potential, initial encapsulation efficiency, initial pH, and osmotic pressure of the sample before freezing. The thermal characteristic parameters include at least the glass transition temperature, the eutectic point, and the collapse initiation temperature; The freeze-drying stress parameters include the pH shift values before and after freezing.
3. The intelligent method for generating freeze-dried formulations according to claim 1, characterized in that, The step of constructing indices for the parameter set to obtain the interface fragility index and the skeleton fragility index includes: normalizing the characteristics of each parameter in the parameter set; weighting and mapping the nucleic acid sequence parameters, the preprocessing characterization data, and the freeze-drying stress parameters to obtain the interface fragility index used to characterize the interface protection requirements of lipid nanoparticles; and weighting and mapping the thermal characteristic parameters to obtain the skeleton fragility index used to characterize the requirements for collapse resistance and glass transition support.
4. The intelligent method for generating freeze-dried formulations according to claim 1, characterized in that, The step of generating a candidate space based on the ratio of the interface fragility index and the skeleton fragility index includes: using the interface fragility index and the skeleton fragility index as feature vectors, searching in a historical sample library using local sensitive hashing or weighted cosine similarity, and extracting the stress response mapping relationship of historical samples with a similarity greater than or equal to a preset threshold as a cross-domain migration prior; and generating preliminary formulation combinations in a preset auxiliary material library according to functional partitions based on the ratio relationship and the cross-domain migration priors. The functional partitions include at least independently adjustable interface protection groups, skeleton support groups, and buffer adjustment groups, and the candidate space is composed of multiple preliminary formulation combinations.
5. The intelligent method for generating freeze-dried formulations according to claim 1, characterized in that, The step of screening the candidate space to obtain the preferred formulation includes: performing feedforward thermal verification on each preliminary formulation combination in the candidate space; predicting the collapse initiation temperature of each preliminary formulation combination and calculating the difference between the collapse initiation temperature and the preset maximum freeze-drying operating temperature; when the difference is greater than or equal to the preset safety margin, determining that the corresponding preliminary formulation combination has passed the verification and is selected as the preferred formulation; and removing preliminary formulation combinations that have not passed the verification.
6. The intelligent method for generating freeze-dried formulations according to claim 1, characterized in that, The process window of the preferred formulation is jointly solved to obtain a multi-objective comprehensive function, which includes: constructing a performance vector that includes cake integrity, freeze-drying window margin, remelting particle size drift, encapsulation loss, and residual moisture; using an integrated regressor to predict the performance vector of the preferred formulation under a preset freeze-drying window, and outputting the expected mean and uncertainty variance of each performance index; and constructing the multi-objective comprehensive function with an uncertainty penalty term based on the expected mean and the uncertainty variance.
7. The intelligent method for generating freeze-dried formulations according to claim 1, characterized in that, The step of performing micro-batch validation on the multi-objective synthesis function and the preferred formulation to obtain validation feedback data includes: calculating the overall uncertainty of the preferred formulation; when the overall uncertainty is greater than a preset uncertainty gating threshold, blocking the direct output of the formulation and automatically triggering the micro-batch validation process; selecting a preset number of formulation combinations from the preferred formulation to form a minimum validation set based on the maximum expected information gain or maximum model variance coverage strategy; performing a micro-freeze-drying experiment on the minimum validation set and collecting measured performance parameters as the validation feedback data.
8. The intelligent method for generating freeze-dried formulations according to claim 1, characterized in that, The step of iteratively correcting the multi-objective synthesis function based on the verification feedback data to obtain the target calibration model includes: calculating the micro-batch post-uncertainty after loading the verification feedback data; determining whether the micro-batch post-uncertainty is less than or equal to a preset micro-batch termination threshold; if not, performing Bayesian updates or weighted adjustments on the weights of the cross-domain migration prior according to the verification feedback data, and reconstructing the multi-objective synthesis function using the updated weights, repeatedly executing the micro-batch verification step until the micro-batch post-uncertainty meets the micro-batch termination threshold, and outputting the final target calibration model.
9. The intelligent method for generating freeze-dried formulations according to claim 1, characterized in that, The step of performing robust solution on the target calibration model to obtain the target formulation includes: using a constrained multi-objective evolutionary algorithm to perform global optimization on the target calibration model; outputting at least one set of formulations that meet the safety margin within a set uncertainty perturbation range, as well as the corresponding matching freeze-drying window and risk assessment identifier, as the target formulation.
10. A freeze-drying formulation intelligent generation system, comprising the freeze-drying formulation intelligent generation method according to any one of claims 1-9, characterized in that, include: The data acquisition module is used to acquire parameters of freeze-dried samples and generate parameter sets; The indicator construction module is used to construct indicators from the parameter set to obtain the interface vulnerability index and the skeleton vulnerability index. The candidate generation module is used to generate a candidate space based on the ratio of the interface vulnerability index and the skeleton vulnerability index. A screening feedforward module is used to screen the candidate space to obtain a preferred formulation; The solver module is used to jointly solve the process window of the preferred formula to obtain a multi-objective comprehensive function; The micro-batch verification module is used to perform micro-batch verification of the multi-objective synthesis function and the preferred formulation to obtain verification feedback data; The iterative calibration module is used to iteratively correct the multi-objective comprehensive function based on the verification feedback data to obtain the target calibration model. The results output module is used to perform robust solution on the target calibration model to obtain the target formulation.