Atractylodin dispersion solubilizing method based on deep neural network and strategy optimization
By employing deep neural networks and strategy optimization methods, the problems of insufficient multimodal information fusion and unstable prediction models during the solubilization process of atractylodes dispersion were solved, enabling accurate estimation of key quality attributes and coordinated protection of equipment safety, thereby improving control efficiency and energy efficiency.
Patent Information
- Application Number
- CN202511759151.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-02-03
AI Technical Summary
Existing technologies for the solubilization of atractylodes dispersions suffer from problems such as insufficient multimodal information fusion, lack of physical consistency and batch adaptability in prediction models, and difficulty in coordinating quality and equipment safety, resulting in unstable prediction of key quality attributes and insufficient equipment safety.
A method based on deep neural networks and policy optimization is adopted. By using continuous-time position encoding of energy sensing and sensor reliability gating weighting, combined with multimodal cross-attention Transformer, key quality attributes are estimated. The state vector is constructed using particle size moment, electric potential, supersaturation and recrystallization risk. Uncertainty prediction is carried out by using a control affine mechanism with mass conservation, non-negativity and monotonic constraints, plus residual neural ordinary differential equations. A dual time-domain model with probabilistic safety constraints and control barrier functions is constructed for control optimization.
It enables online accurate estimation and distribution-level prediction of key quality attributes, ensuring equipment and process safety, reducing energy consumption, adapting to batch differences, and improving control efficiency and consistency.
Smart Images

Figure CN121454949A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of pharmaceutical preparation process control, in particular to a deep neural network and policy optimization based atractylodin dispersion solubilization method. BACKGROUND
[0002] As a poorly soluble active ingredient in traditional Chinese medicine, atractylodin has formed various process paths for its solubilization and dispersion preparation, including high-pressure homogenization, strong shear emulsification, ultrasonic energy input, combined with surfactant and cosolvent system, solid dispersion technology and self-emulsification strategy.
[0003] To improve process visualization and stability, process analysis techniques such as near-infrared spectroscopy, ultraviolet-visible spectroscopy, online particle size analysis, temperature and pressure monitoring are gradually introduced, and chemometrics and soft-sensing models are used to estimate the median particle size, polydispersity index, zeta potential and dissolution behavior; the control side commonly uses open-loop operation based on formula and process curve, supplemented by proportional-integral-derivative control and conventional model predictive control.
[0004] However, the prior art still has the following problems:
[0005] First, the multi-modal information fusion and robustness are insufficient. Different sensing channels have significant differences in sampling frequency, signal-to-noise ratio and drift characteristics. Common methods are mostly simple splicing or linear regression, which are difficult to describe the cumulative effect of mechanical energy, ultrasonic energy and heat, and lack dynamic weighting based on reliability, resulting in unstable prediction of key quality attributes.
[0006] Second, the physical consistency and batch self-adaptation ability of the prediction model are insufficient. Empirical models and black box models often lack quality conservation, non-negativity and monotonicity, and do not form a control-friendly structure expression, which is insufficient in uncertainty propagation, and it is difficult to provide reliable distribution-level prediction under batch difference conditions.
[0007] Third, the closed-loop optimization lacks coordination of quality and safety. Traditional strategies generally use deterministic constraints, which are difficult to meet the probability requirements of key quality attributes reaching the target range, lack systematic barrier function mechanisms for equipment and process safety, and do not fully consider the trade-off between process stage switching and compliance time, action smoothing and energy consumption, which is prone to overshoot or high energy consumption.
[0008] Therefore, a atractylodin dispersion solubilization method that can solve the above problems of the prior art is needed to solve the problem of those skilled in the art. SUMMARY
[0009] One purpose of the present application is to propose a method for solubilizing arctiin dispersion based on deep neural network and policy optimization, aiming at the problems of insufficient multi-modal online monitoring fusion, lack of physical consistency and batch self-adaptation of prediction model, and difficulty in coordinating quality and equipment safety in the prior art, a technical scheme is proposed for key quality attribute estimation by multi-modal cross-attention Transformer driven by energy-aware continuous-time position encoding and sensor reliability gating weighting, combining a state vector composed of particle size matrix, zeta potential, supersaturation and recrystallization risk and subjected to low-rank affine modulation in the batch domain, using a mechanism with affine, quality conservation, non-negativity and monotonicity constraints residual neural ordinary differential equation for uncertainty prediction, and generating control instructions through double-time domain model predictive control of key quality attribute probability safety constraint and control barrier function under the driving of stage determination and closed-loop updating process parameters, The present application has the technical effects of online accurate and stable key quality attribute, equipment and process safety can be proved to meet, smooth control action and energy consumption reduction, and self-adaptation to incoming material differences.
[0010] According to an embodiment of the present application, a method for solubilizing arctiin dispersion based on deep neural network and policy optimization, characterized in that it comprises:
[0011] S1, collecting original multi-modal data, generating batch domain label according to incoming inspection, setting key quality attribute target range, equipment safety boundary, process stage and stage switching condition and current process parameter vector;
[0012] S2, pre-processing and aligning the original multi-modal data, generating continuous-time position encoding combined with the current process parameters and sensor readings, and gating and weighting the multi-modal features based on sensor reliability, outputting the reliability-weighted multi-modal feature sequence and continuous-time position encoding;
[0013] S3, input the reliability-weighted multi-modal feature sequence and continuous-time position encoding into the multi-modal cross-attention Transformer, and call the batch domain label and the current process parameter vector for conditional inference, output the key quality attribute estimation result;
[0014] S4, mapping the particle size distribution matrix according to the key quality attribute estimation result and constructing the state vector, extracting the near-infrared spectrum from the reliability-weighted multi-modal feature sequence, combining the temperature to calculate the supersaturation and entering the state vector, calling the batch domain label to implement low-rank affine modulation on the state vector;
[0015] S5, taking the state vector as the initial condition, combining the current process parameter vector and calling the batch domain label, evolving in the physically constrained neural dynamics model, outputting the predicted state trajectory distribution and key quality attribute prediction distribution through uncertainty propagation;
[0016] S6, constructing the probabilistic safety constraints of the key quality attributes and the control barrier function hard constraints of the limited process variables in the model predictive controller according to the predicted state trajectory distribution and the key quality attribute prediction distribution, the key quality attribute target range, the equipment safety boundary, and the process stage and stage switching condition, setting the time domain length and weight of the double time domain rolling optimization according to the stage determination, and solving and outputting the control instruction vector under the cost function including the compliance time, the control action variation rate, and the energy consumption;
[0017] S7, driving the actuator to update the current process parameter vector according to the control instruction vector, collecting new original multi-modal data for the feature alignment and energy position coding in the next round, and repeating until the key quality attribute target range is stably met within the preset observation time.
[0018] Optionally, step S1 is specifically:
[0019] Sensors and analyzers for online collection are arranged on the preparation equipment, and the sensors and analyzers at least include near-infrared spectrum sensors, ultraviolet-visible spectrum sensors, online particle size analyzers, temperature sensors, pressure sensors, pH meters, torque sensors, ultrasonic power meters, and heating power meters, for forming original multi-modal data;
[0020] A batch domain label is generated according to the raw material incoming inspection results, and the incoming inspection results at least include any one or more of the following information: powder particle size distribution, crystal form information, water content, impurity level, and spectrum fingerprint;
[0021] A preset key quality attribute target range is set, and the key quality attributes at least include the median particle size, the polydispersity index, the zeta potential, and the dissolution curve characteristic value;
[0022] A preset equipment safety boundary is set, and the equipment safety boundary at least includes the allowable range of temperature, pressure, pH, speed, power, valve position, and equipment load;
[0023] Process stages and stage switching conditions are defined, the process stages at least include two or more stages, and the stage switching conditions are determined based on the threshold value, time, or stability of the key quality attributes or process variables;
[0024] A current process parameter vector is set, and the current process parameter vector at least includes the homogenization pressure, the shear rotation speed, the ultrasonic power setting, the temperature setting, the pH setting, the solid content, and the feeding flow rate;
[0025] The original multi-modal data, the batch domain label, the key quality attribute target range, the equipment safety boundary, the process stage definition and the stage switching condition, and the current process parameter vector are output.
[0026] Nomenclature:
[0027] The raw multi-modal data is a set of un-processed signals directly outputted and time-stamped by sensors and analyzers on the manufacturing equipment at their respective native sampling frequencies, including but not limited to measurements of channels such as spectrum, particle characteristics, temperature, pressure, pH, torque, power, valve position, and equipment load, etc.
[0028] The incoming inspection is a receiving inspection and quality characterization process performed on raw material batches before storage or production, whose output is structured inspection results for subsequent modeling and control;
[0029] The batch domain label is a coded representation generated based on incoming inspection results of raw materials to characterize and distinguish differences between different batches of raw materials, including but not limited to discrete batch identification, continuous feature vector, or a combination of the two, and is used to condition model inference and control;
[0030] The target range of key quality attributes is a pre-set qualified interval or target band for each key quality attribute, including upper and lower limits and optional tolerance / preferred interval, used for process decision and optimization constraints;
[0031] The equipment safety boundary is a set of allowed value ranges set for variables related to equipment and process safety, used to ensure that the process does not run out of bounds, covering at least one or more of temperature, pressure, pH, speed, power, valve position, and equipment load;
[0032] The process variable is a general term for measurable or controllable physical and / or calculated quantities reflecting the state of the process and the working condition of the equipment, including but not limited to temperature, pressure, pH, speed, power, flow, torque, valve position, and equipment load;
[0033] The current process parameter vector is an ordered set of process set values and / or actual achieved values used for execution and optimization at the current time, covering at least one or more of homogeneous pressure, shear speed, ultrasonic power setting, temperature setting, pH setting, solid content, and feeding flow.
[0034] Optionally, step S2 is specifically:
[0035] According to the raw multi-modal data, sensor baseline correction and time stamp alignment are performed, and alignment of signals of different sampling frequencies is completed on a unified time grid through baseline subtraction, drift correction, and dimension unification;
[0036] On the aligned time grid, the instantaneous mechanical power is calculated according to the torque and shear rotational speed, and the cumulative mechanical energy is obtained by time integration; the cumulative ultrasonic energy is obtained by time integration according to the ultrasonic power meter reading; the cumulative heat is obtained by time integration according to the heating power meter reading, and the continuous time position code is formed by the time stamp, the cumulative mechanical energy, the cumulative ultrasonic energy and the cumulative heat;
[0037] The sensor reliability weight is assigned to each modality based on the signal-to-noise ratio and drift detection result of the alignment signal, and the abnormality and distortion are suppressed by performing gated weighting on each modality feature according to the sensor reliability weight to generate a reliability-weighted multi-modal feature sequence;
[0038] The reliability-weighted multi-modal feature sequence and the continuous time position code are output.
[0039] Glossary:
[0040] The sensor baseline correction is a process of estimating and eliminating the zero point bias and slow baseline component of the original signal of each sensor / analyser, so that the reference of different channels is consistent without changing the effective change trend;
[0041] The time stamp alignment is a process of mapping multi-channel signals with different sampling frequencies and different starting times to a unified time grid according to their respective time stamps, including but not limited to resampling, interpolation or hold method;
[0042] The baseline deduction is an operation of subtracting the estimated baseline component from the original signal to obtain a corrected signal;
[0043] The drift correction is a process of identifying and compensating for the slow systematic deviation of the signal over time, including zero point, gain or waveband drift;
[0044] The dimension unification is a process of converting the measurement values of each channel to a preset consistent unit of measurement and numerical scale, which may include normalization or standardization when necessary;
[0045] The unified time grid is a set of discrete sampling points of a common time axis for multi-channel signal fusion and subsequent modeling, which can be fixed step or adaptive step;
[0046] The instantaneous mechanical power is a power quantity calculated according to the torque and shear rotational speed, which is used to represent the mechanical input strength;
[0047] The cumulative mechanical energy is an energy quantity obtained by integrating the instantaneous mechanical power with respect to time, which is used to describe the cumulative effect of mechanical energy;
[0048] The cumulative ultrasonic energy is an energy quantity obtained by integrating the ultrasonic power meter reading with respect to time, which is used to describe the cumulative effect of ultrasonic energy input;
[0049] The cumulative heat is the energy quantity obtained by integrating the heating power meter reading with respect to time, used to describe the cumulative effect of heat input;
[0050] The continuous time position encoding is represented by continuous time features consisting of time stamps and at least one of the above cumulative energy quantities, and can be scaled or embedded mapping for time series modeling;
[0051] The signal-to-noise ratio is an estimate of the ratio of signal effective components to noise component intensity within a predefined time window, used to measure channel data quality;
[0052] The drift detection is a recognition process for the change of signal statistical properties over time, including but not limited to trend analysis, change point detection or reference standard comparison;
[0053] The sensor reliability weight is a weight between 0 and 1 assigned to each modality and / or each time point according to the signal-to-noise ratio, drift detection results and optional information, used to quantify the credibility of the data;
[0054] The gated weighting is a process of weighting or masking each modality feature using the sensor reliability weight to suppress the influence of abnormalities and distortions on subsequent models;
[0055] The multi-modal feature sequence is a sequence of pre-processed feature vectors of each modality arranged in time sequence on a unified time grid, containing features directly related to or derived from the original signal;
[0056] The reliability-weighted multi-modal feature sequence is the result of applying gated weighting to the multi-modal feature sequence, used as input for subsequent modeling.
[0057] Optionally, step S3 is specifically:
[0058] The reliability-weighted multi-modal feature sequence and the continuous time position encoding are used to perform cross-modal information fusion and time series modeling in a multi-modal cross-attention Transformer to represent the time evolution and energy accumulation effect of different sensing channels;
[0059] The attention mask and bias are set according to the mechanism coupling relationship, so that the effects of pH and ionic strength on the electromotive potential, the effects of shear and homogenization energy on particle size distribution, and the effects of temperature on solubility and recrystallization risk are strengthened in the attention weight;
[0060] By calling the batch domain label and the current process parameter vector to condition the inference, the model output is self-adaptive to the raw material batch and the current working condition;
[0061] generate critical quality attribute estimation results, the critical quality attribute estimation results including median particle size, polydispersity index, zeta potential and dissolution curve characteristic value;
[0062] output the critical quality attribute estimation results.
[0063] Noun interpretation:
[0064] The multi-modal cross-attention Transformer is a deep network structure based on Transformer, which establishes information interaction and alignment between different sensing modalities through cross-modal attention, supports the combination of self-attention and cross-attention layers, and receives continuous time position encoding as input;
[0065] The cross-modal attention is a mechanism that calculates attention by taking one modality representation as query and another modality representation as key / value, and performs weighted aggregation to realize inter-modal dependency modeling;
[0066] The mechanism coupling relationship is a set of variable interaction relationships determined according to known or recognized physical and chemical mechanisms, which is used as priori in the model to guide attention allocation;
[0067] The attention mask is a structural constraint matrix applied to the attention score matrix, which is used to prohibit or weaken the connection of specific query-key pairs to achieve selective attention;
[0068] The attention bias is an additive or multiplicative adjustment term applied to the attention score according to a predetermined rule, which is used to strengthen or inhibit the attention of a specific coupling path;
[0069] The conditional inference is a way of inference in which batch domain labels and current process parameter vectors are input as conditions or modulation factors during model inference, so that the model output changes adaptively with the batch of raw materials and working conditions;
[0070] The critical quality attribute estimation results are online prediction outputs of the model for critical quality attributes, including point estimates and / or distribution statistics, covering median particle size, polydispersity index, zeta potential and dissolution curve characteristic value;
[0071] The median particle size is the particle size corresponding to the cumulative percentage of 50% of the particle size distribution, which can be defined according to number distribution or volume distribution, preferably volume distribution median D50;
[0072] The polydispersity index is a dimensionless index representing the width of the particle size distribution, commonly represented by PDI measured by dynamic light scattering;
[0073] The zeta potential is the electric potential at the shear surface of the particles in the dispersion system, which is usually represented by ζ potential and can be measured by electrophoretic light scattering;
[0074] The dissolution profile characteristic value is a set of quantitative characterization parameters of the drug dissolution behavior curve, including but not limited to one or more of dissolution percentage at a specific time point, t50, t80, dissolution rate constant, area under the curve AUC or similar factor f2;
[0075] The ionic strength is a comprehensive quantity representing the influence of the total concentration and valence state of ions in a solution on the shielding effect of zeta potential;
[0076] The homogenization energy is a measure of energy input in the homogenization process, which can be obtained by integrating the homogenization-related power with respect to time or estimated according to parameters such as homogenization pressure and flow rate.
[0077] Optionally, step S4 specifically comprises:
[0078] According to the estimation results of the key quality attributes, the median particle size and the polydispersity index are mapped to the zeroth, first and second moments of the particle size distribution through a preset mapping relationship, and the zeta potential is incorporated into the state vector;
[0079] The near-infrared spectrum is extracted from the reliability-weighted multi-modal feature sequence, the solute concentration is inverted using a quantitative correction model, the supersaturation is calculated according to the solubility curve in combination with the temperature, and the supersaturation is incorporated into the state vector;
[0080] A recrystallization risk fast variable is constructed according to the supersaturation and the temperature, and incorporated into the state vector;
[0081] Low-rank affine modulation coefficients are generated using batch domain labels, and affine modulation is performed on the state vector to reflect batch differences;
[0082] And output the state vector.
[0083] Glossary:
[0084] The particle size distribution moment mapping is a process of converting particle size statistical indicators including at least the median particle size and the polydispersity index into particle size distribution moment parameters according to a preset function or rule, which can be realized by analytical relationship, table lookup or data-driven model;
[0085] The particle size distribution moment is a moment quantity obtained by power-weighted integration or summation of a distribution function with respect to particle size as the independent variable according to the selected distribution basis (number, volume or mass), and the k-th moment Mk can be represented by the integral or discrete summation of D^k and the distribution function;
[0086] The zeroth moment of the particle size distribution is a quantity obtained by 0th power-weighted integration or summation of the particle size distribution, which is used to represent the total number / volume / mass scale or normalization constant;
[0087] The first moment of the particle size distribution is a quantity obtained by integrating or summing the particle size to the first power, which is used to represent the average size related information;
[0088] The second moment of the particle size distribution is a quantity obtained by integrating or summing the particle size to the second power, which is used to represent the distribution width, dispersion or dimension information related to the surface area;
[0089] The state vector is a system state representation for dynamic prediction and control optimization, including but not limited to particle size distribution moment parameters, zeta potential, supersaturation and recrystallization risk fast variables, etc.
[0090] The near-infrared spectrum is the spectral data of the sample obtained in the near-infrared wave band, reflecting the absorption / reflection / transmission characteristics of the sample to near-infrared light, preferably covering about 780 nm to 2500 nm and being time-stamped;
[0091] The quantitative correction model is a data-driven quantitative relationship that maps the near-infrared spectrum to the target physicochemical quantity (including solute concentration), which is obtained based on calibration samples and can be realized by partial least squares, regularized regression or neural network, etc.
[0092] The solubility curve is a functional relationship of the equilibrium solubility of the solute with temperature in a given solvent system, which can also include the effects of pressure, pH or ionic strength as needed, and the source can be empirical fitting or mechanism model;
[0093] The supersaturation is an excess measure of the current solute concentration relative to the equilibrium solubility at the same temperature, which can be expressed in the form of relative S=C / C* or difference ΔC=C−C*, where C* is the equilibrium solubility given by the solubility curve;
[0094] The recrystallization risk fast variable is a state component constructed according to the supersaturation and temperature, etc. to measure the crystallization / recrystallization tendency, which has a higher rate of change than the slow variable such as the particle size distribution moment, and is monotonically related to the nucleation / growth driving force;
[0095] The affine modulation is a modulation form of linear transformation and bias superposition on the state vector, which is used to adjust the state representation to match different external conditions or batch characteristics;
[0096] The low-rank affine modulation is an affine modulation form with a limited rank of the linear part, whose linear transformation can be represented as a low-rank factorization (such as A=UV^T, with a rank not exceeding a preset threshold r), and a bias term is used to represent the batch difference with fewer parameters;
[0097] The low-rank affine modulation coefficient is a set of parameters required to realize the low-rank affine modulation, including the low-rank factor of the linear part and the bias vector, which is given by a parameter generation network or a lookup table mechanism according to the batch domain label.
[0098] Optionally, step S5 is specifically:
[0099] Solving the time evolution of state in neural ordinary differential equation with control-affine structure and coupled by mechanism term and residual term, with state vector as initial condition;
[0100] The mechanism term is used to describe the differentiable physical process, and the residual term is used to compensate for the unmodeled dynamic error;
[0101] Calling the current process parameter vector as the control quantity into the affine term, so that the state derivative is affine to the control quantity;
[0102] Calling the batch domain label to implement low-rank affine modulation on model parameters to achieve self-adaptation to raw material batch differences;
[0103] In the numerical solving process, mass conservation, non-negativity and monotonicity constraints are imposed on the state and parameters, wherein non-negativity is maintained by positive value mapping, conservation is maintained by balance constraints or penalty terms, and monotonicity is realized by non-negative weight and cumulative structure;
[0104] Using a numerical solver to obtain the probability distribution of state over time within a predetermined prediction time domain, and performing uncertainty propagation through linearization approximation or random sampling;
[0105] According to the moment mapping and index calculation, the probability distribution of the state is converted into the prediction distribution of the key quality attribute;
[0106] And output the predicted state trajectory distribution and the predicted distribution of the key quality attribute for the probabilistic safety and control barrier function mixed two-time domain model predictive control.
[0107] Noun explanation:
[0108] The control-affine structure is a linear expression of the dynamic equation with respect to the control quantity plus a bias, so that the state derivative can be written as an affine combination of the basis function and the control quantity, which is convenient for integration with model predictive control;
[0109] The neural ordinary differential equation is a continuous-time model that defines the state derivative with a learnable parameterized function and obtains the state trajectory through numerical integration, and the parameters are realized by a neural network;
[0110] The mechanism term is a function part used to describe known or differentiable physical and chemical processes, which satisfies the basic physical consistency and provides an a priori description of the dominant behavior of the system;
[0111] The residual term is a learnable function part used to compensate for unmodeled or difficult to accurately model dynamics, coupled with the mechanism term to improve the overall prediction accuracy;
[0112] The affine term is a linear mapping and bias combined to represent the effect of control variables on the state rate of change in the dynamic equation;
[0113] The control variable is a set of controllable variables entering the dynamic equation and adjustable by actuators, which can be provided by the current process parameter vector;
[0114] The state derivative is the derivative of the state variable with respect to time, which is used to describe the instantaneous rate of change of the state;
[0115] The mass conservation constraint is a constraint that forces the total mass or related conserved quantity to remain constant according to the given balance of income and expenditure or evolve according to a known relationship in numerical solution;
[0116] The non-negativity constraint is a constraint that sets a limit of not less than zero for a specific state or parameter to avoid the occurrence of negative values without physical meaning;
[0117] The monotonicity constraint is a constraint that limits the consistency of the direction of a specific mapping or state change with time or control variable, so that it satisfies the preset monotonic increasing or decreasing relationship;
[0118] The positive value mapping is a way to map any real number to a non-negative number through function transformation, including but not limited to Softplus, exponential or lower limit truncation, etc.;
[0119] The balance constraint is an equation or inequality established according to the conservation of mass or energy, which is used to maintain the compatibility of the input, output and cumulative amount during the solution process;
[0120] The penalty term is a term that imposes an additional cost on the degree of constraint violation in the optimization objective or loss function, which is used to soften the constraint and guide the solution to satisfy the physical consistency;
[0121] The non-negative weight is a constraint that limits the weight parameter to be not less than zero in the model structure, which is used to ensure the monotonicity of the mapping;
[0122] The cumulative structure is a module composed of step-by-step addition or integration, so that the output reflects the cumulative effect of the input and is compatible with the monotonicity constraint;
[0123] The numerical solver is an algorithm and implementation component for numerical integration of ordinary differential equations, including fixed or adaptive step Runge-Kutta, Adams class or equivalent methods;
[0124] The preset prediction time domain is the time interval and discrete step size setting for generating future state trajectories;
[0125] The probability distribution of the state with respect to time is a set of uncertainty distributions of state variables at each time in the prediction time domain, which can be represented by mean, variance or full distribution.
[0126] The linearization approximation is a first-order or equivalent approximation of the nonlinear model at the current operating point to propagate the uncertainty;
[0127] The stochastic sampling is a computational method to estimate the output distribution by Monte Carlo or equivalent sampling of the model and input;
[0128] The uncertainty propagation is a process to transfer the uncertainty of input, parameters and model structure to the state and output and form a distribution description;
[0129] The moment mapping is a process to convert the state distribution or its statistical quantities into the key quality attribute related moments or indices through a pre-set function relationship;
[0130] The index calculation is a process to calculate the key quality attribute corresponding scalar or statistical characteristics based on the state or its derived quantities;
[0131] The predicted state trajectory distribution is a set of probability distribution of state vector over time in the prediction horizon, which is used as the input of subsequent control optimization;
[0132] The predicted key quality attribute distribution is a probability distribution description of the key quality attribute in the prediction horizon, which is used for probabilistic constraint and performance evaluation.
[0133] Optionally, step S6 specifically comprises:
[0134] According to the predicted state trajectory distribution and the predicted key quality attribute distribution, a probabilistic safety constraint is constructed in the model predictive controller for the key quality attribute target range, so that the probability of the key quality attribute being in the target range is not less than a pre-set threshold;
[0135] At the same time, according to the equipment safety boundary, a control barrier function is set for the limited process variable to form a hard constraint, and the limited process variable at least includes one or more of temperature, pH, pressure, speed, power, valve position or equipment load;
[0136] According to the process stage and stage switching condition, a stage judgment is performed on the predicted state trajectory distribution to generate a stage marker, and the time domain length and weight of the double time domain rolling optimization are set according to the stage marker, wherein the short time domain is used for fast approaching the key quality attribute target range, and the long time domain is used for control action smoothing and energy consumption suppression;
[0137] Under the condition of meeting the probabilistic safety constraint and the control barrier function hard constraint, a cost function including the compliance time, control action variation rate and energy consumption is constructed, and a rolling optimization problem is solved to obtain a control instruction vector.
[0138] Nomenclature:
[0139] The model predictive controller is a control module that solves a finite horizon optimization problem based on a prediction model and constraints at each rolling time instant and outputs a control sequence, which is composed of a model, constraints, a cost function, and a solver;
[0140] The probabilistic safety constraint is an opportunity constraint imposed on the probability distribution of a key quality attribute, which is a condition that the probability of being within a target range is no less than a set value, which can be embedded in optimization in the form of an opportunity constraint or a risk metric;
[0141] The preset threshold is the minimum probability value used for the probabilistic safety constraint to determine eligibility, which is set in advance according to process requirements or regulations;
[0142] The limited process variable is a set of process variables that need to meet safety boundaries or operating limits in optimization, including at least one or more of temperature, pH, pressure, speed, power, valve position, and equipment load;
[0143] The control barrier function is a state function constructed for limited process variables, which is used to maintain the forward invariance of the safety set by satisfying the inequality condition, and to ensure process and equipment safety in the form of constraints in optimization;
[0144] The hard constraint is a constraint condition that must be strictly met in optimization solving, and if violated, the solution is determined to be infeasible;
[0145] The process stage is a stage division of the preparation process according to the mechanism or operation target, which is used for stage-by-stage management of optimization strategy and target trade-off;
[0146] The stage switching condition is a criterion for triggering the transition of the process stage from the current stage to the next stage, which is determined based on the threshold, time, or stability of the key quality attribute or process variable;
[0147] The stage determination is a process of identifying the process stage to which the current and future time belongs according to the predicted state trajectory and the stage switching condition;
[0148] The stage marker is an encoded representation of the stage determination result, which is used to drive the time domain length and weight setting of the dual-time-domain rolling optimization;
[0149] The dual-time-domain rolling optimization is an optimization strategy that solves two prediction time domains in parallel or hierarchically at the same rolling time, with a short time domain for fast approximation of the target range and a long time domain for action smoothing and energy consumption suppression;
[0150] The short time domain and the long time domain are two different prediction-control time windows with different coverage lengths set in the dual-time-domain rolling optimization, with the former being shorter and the latter being longer;
[0151] The time domain length is the time range or discrete step number covered by each prediction time domain;
[0152] The weight is a non-negative coefficient used in the cost function to balance different sub-objectives and different time domain contributions;
[0153] The compliance time is a measure of the time required for a key quality attribute to enter and remain within a target range, or its penalty term in the cost function;
[0154] The control action variation rate is a measure of the speed of change of control instructions at adjacent time points, used to constrain and penalize too fast setting adjustments to achieve action smoothing;
[0155] The energy consumption is a measure of energy consumption caused by control actions, usually calculated by integrating relevant power over time and can be decomposed into mechanical, ultrasonic and thermal energy consumption;
[0156] The cost function is an optimization objective function that integrates the objectives of compliance time, control action variation rate and energy consumption, and can include regularization terms;
[0157] The rolling optimization problem is a finite time domain constrained optimization problem that is repeatedly solved over time, with the optimal control at the current time being implemented and updated with new data rolling;
[0158] The control instruction vector is a set of target setting values of each actuator obtained by rolling optimization under satisfied constraints, which is used to update one or more of homogeneous pressure, shear speed, ultrasonic power setting, temperature setting, pH value setting, solid content and feed flow.
[0159] Optionally, step S7 is specifically:
[0160] According to the control instruction vector, the target setting values of each actuator are calculated through the preset instruction-to-actuator mapping relationship, and are issued to the actuators under the constraints of variation rate limitation and amplitude limitation, so that the homogeneous pressure, shear speed, ultrasonic power setting, temperature setting, pH value setting, solid content and feed flow are updated accordingly;
[0161] The instruction issuing time is recorded under the updated working condition, and the actuator feedback or process measurement is collected for execution confirmation, while the original multi-modal data is synchronously collected by sensors and time-stamped to form new collected original multi-modal data;
[0162] According to the comparison between the actuator feedback or process measurement and the target setting value, the actual achieved process variable is calculated, and the actual achieved value is written into the current process parameter vector to form the updated process parameter vector;
[0163] And output the new collected original multi-modal data and the updated process parameter vector for feature alignment and energy position coding in S2.
[0164] Glossary:
[0165] The instruction-to-actuator mapping relationship is a function relationship for converting each component in the control instruction vector into a target setting value acceptable to the corresponding actuator, including unit conversion, calibration compensation, linear or nonlinear characteristic compensation, and communication coding rules;
[0166] The actuator is a general term for controllable components for directly changing process variables or energy inputs, including one or more of homogeneous pumps / valve groups, high-shear stirring drives, ultrasonic transducer power supplies, heating / cooling units, acid-base dosing pumps, feeding pumps, and regulating valves;
[0167] The target setting value is a desired operating point or opening value calculated from the control instruction through the mapping relationship and issued to the corresponding actuator, including position-based or tracking-based setting forms;
[0168] The change rate limit is a constraint on the maximum allowable change amplitude or speed of the target setting value at adjacent control times, to avoid equipment impact and process disturbance caused by sudden changes;
[0169] The amplitude limiting constraint is an upper and lower bound constraint applied to the target setting value or actuator output to ensure that the variable operates within the device allowable range;
[0170] The instruction issuance time is the timestamp recorded by the control system when sending the target setting value to the actuator under the unified time reference, used for time alignment with subsequent feedback and sensor data;
[0171] The actuator feedback is the actual working state quantity returned by the actuator body or its attached measurement device, reflecting the execution result of the target setting value;
[0172] The process measurement is a process-side measurement quantity related to the action of the actuator, used to verify the implementation of the control instruction from the process effect level;
[0173] The execution confirmation is a process of determining the achievement of the target setting value based on actuator feedback and / or process measurement, usually with entering the allowable error band and maintaining the preset duration as the pass condition;
[0174] The synchronous acquisition is a process of acquiring multi-channel sensor data under a unified time base according to a predetermined synchronization strategy, so that the measurement and control events can be time-aligned;
[0175] The actual achieved process variable is the current real implementation value estimated from the actuator feedback or process measurement, used to objectively characterize the working condition and allow deviation from the target setting value;
[0176] The updated process parameter vector is a parameter set formed by writing and replacing the actual achieved process variable into the corresponding component of the original current process parameter vector, for the next round of feature alignment and energy position coding.
[0177] The beneficial effects of the present application are:
[0178] First, online accurate estimation and distribution level prediction of key quality attributes are achieved: through energy-aware continuous-time position coding and sensor reliability-gated weighted multi-modal cross-attention model, combined with control of affine and mechanism plus residual neural differential equation with conservation of mass, non-negativity and monotonicity constraints, the prediction accuracy and robustness of the median particle size, polydispersity index, zeta potential and dissolution curve are significantly improved, which adapts to batch differences and suppresses the influence of transmission drift;
[0179] Second, coordinated provable protection of quality and equipment safety: the probabilistic safety constraints of key quality attributes and the hard constraints of control barrier functions of limited process variables are included in the double-time domain rolling optimization, so that the product quality reaches and stably maintains the target range, while ensuring that variables such as temperature, pressure, pH, speed, power, valve position and equipment load are within the safety boundary, reducing overshoot and risk;
[0180] Third, improve control efficiency and energy efficiency: phase determination driven double-time domain model predictive control considers both fast compliance and smooth action, reducing parameter adjustment and trial and error time, reducing energy input and equipment wear, shortening compliance time and improving batch consistency and production line repeatability. BRIEF DESCRIPTION OF DRAWINGS
[0181] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, together with the embodiments of the present application, to explain the present application, and do not constitute a limitation of the present application. In the drawings:
[0182] Figure 1 A flowchart of a deep neural network and policy optimization-based atractylodin dispersion solubilization method is provided. DETAILED DESCRIPTION
[0183] The present application will now be described in further detail with reference to the drawings. These drawings are simplified schematic diagrams that only schematically show the basic structure of the present application, and thus only show the components related to the present application.
[0184] Reference Figure 1 A deep neural network and policy optimization-based atractylodin dispersion solubilization method, characterized in that it comprises:
[0185] S1, collect original multi-modal data, generate batch domain label according to incoming inspection, set key quality attribute target range, equipment safety boundary, process stage and stage switching condition and current process parameter vector;
[0186] S2, pre-process and align the original multi-modal data, generate continuous time position encoding combined with current process parameters and sensor readings, and gate weight multi-modal features based on sensor reliability, output reliability weighted multi-modal feature sequence and continuous time position encoding;
[0187] S3, input the reliability weighted multi-modal feature sequence and continuous time position encoding into the multi-modal cross attention Transformer, and call the batch domain label and the current process parameter vector for conditional inference, output the key quality attribute estimation result;
[0188] S4, map the particle size distribution matrix according to the key quality attribute estimation result and construct the state vector, extract the near-infrared spectrum from the reliability weighted multi-modal feature sequence, calculate the supersaturation combined with the temperature and enter the state vector, call the batch domain label to implement low rank affine modulation on the state vector;
[0189] S5, take the state vector as the initial condition, combine the current process parameter vector and call the batch domain label, and perform evolution prediction in the physically constrained neural dynamics model, output the predicted state trajectory distribution and key quality attribute prediction distribution through uncertainty propagation;
[0190] S6, according to the predicted state trajectory distribution and the key quality attribute prediction distribution, and calling the key quality attribute target range, the equipment safety boundary and the process stage and stage switching condition, construct the probability safety constraint of the key quality attribute and the control barrier function hard constraint of the limited process variable in the model predictive controller, set the time domain length and weight of the double time domain rolling optimization according to the stage determination, solve and output the control instruction vector under the cost function including the compliance time, control action variable rate and energy consumption;
[0191] S7, update the current process parameter vector according to the control instruction vector to drive the actuator, collect new original multi-modal data for the next round of feature alignment and energy position encoding, and cycle until the key quality attribute target range is stably met within the preset observation time.
[0192] In the specific embodiment, the S1 is specifically:
[0193] The six links include online perception deployment, batch domain label generation, key quality attribute target range setting, equipment and process safety boundary setting, process stage and stage switching condition definition, and current process parameter initialization: first, arrange multi-modal sensors and analyzers on the preparation equipment to form an original data stream, define the transmission channel set as:
[0194] ;
[0195] where is the th sensor channel, is the number of channels, covering near-infrared spectrum, ultraviolet-visible spectrum, online particle size, temperature, pressure, pH, torque, ultrasonic power, and heating power, etc. measurements;
[0196] Let the raw observations with timestamps be denoted as:
[0197] ;
[0198] where is the observation vector at time , is the measurement value of channel at time , and is the set of timestamps, and let the raw multimodal data sequence be denoted as ;
[0199] Secondly, generate batch domain labels according to incoming inspection results (including powder particle size distribution, crystal form information, moisture content, impurity level, and spectral fingerprint, etc.), let the incoming feature vector be , get batch domain labels through encoding mapping , where is the inspection feature summarized by batch, is its dimension, is the vector representation of batch domain labels, is the label dimension, is the encoding function that maps inspection features to labels, denotes the real number field;
[0200] Then set the target range of key quality attributes, define the attribute set , where respectively represent the median diameter of volume distribution, polydispersity index, zeta potential, and selected dissolution curve characteristic value, use interval set to specify the target band, where and are the lower limit and upper limit of attribute respectively; at the same time, set the safety boundary of equipment and process, let the number of limited variables be , the variable is ), use set to depict the allowed range, where and lower and upper bounds of variable , which can correspond to temperature, pressure, pH, rotational speed, power, valve position and equipment load, etc.
[0201] Further define the process stages and stage switching conditions, let the stage set be and , where represents the th stage, is the number of stages, and give the threshold vector related to switching , where is the th switching threshold determined by key quality attributes or process variables, is the stage index, which is used to transition from stage to stage when the threshold is reached;
[0202] Finally, initialize the current process parameter vector and write it into the control system, where is the homogenization pressure setting, is the shear rotational speed setting, is the ultrasonic power setting, is the temperature setting, is the pH setting, is the system solid content, is the feeding flow rate, and is used as output for subsequent steps.
[0203] In this specific embodiment, S2 is specifically:
[0204] For the original multi-modal data, unified preprocessing and time alignment are carried out, and on this basis, energy-aware continuous time position coding and sensor reliability gating weighted feature sequences are generated for subsequent modeling: first, baseline subtraction, drift correction and dimension unification are performed on each sensing channel, signals of different sampling frequencies are resampled to a unified time grid by interpolation or hold method and the time stamps are kept consistent, and at each time, the cumulative mechanical energy is calculated according to the torque sensor and shear rotational speed signal, and the cumulative ultrasonic energy and cumulative heat are calculated according to the ultrasonic power meter and heating power meter signals, to characterize the cumulative effect of energy input, where the cumulative mechanical energy, cumulative ultrasonic energy and cumulative heat are defined as:
[0205] ;
[0206] where is the current time, Let time be the independent variable for integration. For a moment Torque signal, For a moment The angular velocity signal (obtained from the shear rotation speed). For a moment ultrasonic power, For a moment Heating power, for The mechanical energy accumulated before the time, for The accumulated ultrasonic energy before the time, for The heat accumulated before the time;
[0207] Based on this, construct a continuous-time position encoding vector:
[0208] ;
[0209] in for Position encoding vector at time step, square brackets This indicates that a vector is formed by concatenating columns, and the components of the vector can be scaled according to empirical dimensions.
[0210] Subsequently, data quality metrics are calculated for each sensing mode, and reliability weights are formed. Channels with high signal-to-noise ratios and low drift are given greater weights, while abnormal and distorted channels are suppressed, thus providing the modal data. exist The reliability weight at time step is:
[0211] ;
[0212] in For modality exist Reliability weight at any given moment For the Logistic function, and As a weighting factor, In order to be in The modalities evaluated within the nearby sliding window Signal-to-noise ratio estimation In order to be in Modalities obtained by detecting trend changes or change points nearby Drift rating;
[0213] Using reliability weights for gating weighting yields a feature sequence that has undergone noise suppression and drift resistance processing. Defined as follows:
[0214] ;
[0215] wherein is the weighted feature vector of the modality at time ; is the pre-processed feature vector of the modality at time ; Further, the modalities are aggregated on the aligned time grid to form the multi-modal weighted features for subsequent modeling, defined as:
[0216]
[0217] ;
[0218] wherein is the multi-modal weighted feature vector at time ; is the number of modalities involved in fusion, and square brackets represent vector concatenation;
[0219] Finally, the output is outputted on the unified time grid for the multi-modal cross-attention model in step S3, wherein is the set of aligned timestamps.
[0220] In the present specific embodiment, the S3 is specifically:
[0221] The multi-modal weighted feature sequence with reliability gating is inputted with continuous time position encoding, and the multi-modal cross-attention Transformer with mechanism prior is used to complete cross-modal information fusion and time series modeling, and output the online estimation result of the key quality attribute under the conditioning modulation of batch domain label and current process parameter;
[0222] For ease of description, the input on the unified time grid is denoted as:
[0223] ;
[0224] wherein is the time step index, is the number of aligned time steps, is the multi-modal weighted feature vector at the th time, is the continuous time position encoding vector at the th time;
[0225] First, linear embedding is performed on the input to obtain the time series representation:
[0226] ;
[0227] in For the first The embedding vector at each time step, and These are the linear mapping matrices of feature and position encoding, respectively. For bias vector, For the embedded dimension;
[0228] Then construct the queries, keys, and values required for cross-reference, defining:
[0229] ;
[0230] in For the first The query vector at each time point and The first The key and value vector at each time point, and For the corresponding projection matrix, and These are the dimensions of the key and the value, respectively.
[0231] To strengthen the mechanism coupling path, a mechanism mask and a mechanism bias are simultaneously introduced into the attention process, denoted as... For the first For the Attention masking elements Indicates permitted mechanism coupling, This indicates an unacceptable coupling. For the first With the The mechanism of bias between functions Calculations were performed based on the effects of pH and ionic strength on electrokinetic potential, the effects of shear and homogenization energy on particle size distribution, and the effects of temperature on solubility and recrystallization risk.
[0232] ;
[0233] in For the first pH level at all times For the first Time-based ion strength, and The first Accumulated mechanical energy and accumulated ultrasonic energy before the time point For the first Temperature at any given time, these quantities can all be derived from the alignment data in step S2;
[0234] Mask normalization strategy is adopted in computing attention, which combines mask and bias in additive manner into similarity and normalizes on mask support set, obtaining attention weight after mask:
[0235] ;
[0236] wherein is normalized attention weight from to , is coefficient to adjust bias intensity of mechanism, is scaling factor to stabilize attention distribution;
[0237] Context is obtained by weighting and converging value vector with attention weight:
[0238] ;
[0239] wherein is context vector at th time point; to make inference adaptive to raw material batch and current working condition, conditional affine modulation manner is adopted to scale and translate , denoted as is condition vector, is batch domain label vector, is current process parameter vector, is vector splicing;
[0240] define and and implement , wherein and are conditional scaling and translation vectors, and are learnable functions to map condition vector to scaling and translation parameters, is element-wise multiplication;
[0241] Global representation is obtained by time aggregation of time series representation, wherein is spline level representation after aggregation;
[0242] Finally, key quality attribute estimation vector is obtained through linear readout layer:
[0243] ;
[0244] wherein are four predicted key quality attributes, For estimating the median diameter of the volume distribution, For the estimation of the multi-dispersion index, For estimating electromotive force, For the estimation of characteristic values of dissolution curves, To output the mapping matrix, For output bias;
[0245] Furthermore, conventional structures such as multi-head attention and layer normalization, residual connections, and feedforward networks can be used in the implementation without changing the aforementioned inference process with mechanism masking and bias, and subject to batch and conditional constraints.
[0246] In this specific embodiment, S4 specifically refers to:
[0247] Starting with the estimation of key quality attributes, particle size statistics are mapped to particle size distribution moments, and a state vector containing fast variables such as electromotive force, supersaturation, and recrystallization risk is constructed. Simultaneously, low-rank affine modulation is implemented based on batch domain labeling to reflect raw material differences and operational condition adaptation. Specifically: First, the log-normal approximation is used to calculate distribution parameters from the median particle size and polydispersity index under volume distribution normalization conditions, and the zeroth, first, and second moments are given. The position and shape parameters are defined as follows:
[0248] , moment mapping is ;
[0249] in For the estimation of the median diameter of the volume distribution obtained in step S3, The multidispersion index estimate obtained in step S3, The log-position parameter of the log-normal distribution, The log-scale parameter of the log-normal distribution For the natural logarithm, For square root operations, The zeroth moment of the volume distribution (equal to 1 under normalization) For a first-order moment, For the second moment, It is an exponential function;
[0250] Next, near-infrared spectra are extracted from the reliability-weighted multimodal feature sequences, and solute concentration and supersaturation are inverted and calculated. Definition:
[0251] and ;
[0252] in To unify the time on the time grid, for Near-infrared spectral vector at time, For the spectral vector dimension, For the quantitative calibration model trained from the calibration samples, for Solute concentration at any time for Temperature at all times To be at temperature equilibrium solubility, Mapping the solubility curve and for The ratio of supersaturation at any given time;
[0253] To further construct a fast variable for recrystallization risk to reflect rapid sensitivity to supersaturation and temperature, a non-negative Softplus form is used:
[0254] ;
[0255] in for The risk of recrystallization at any moment is a rapid variable. and For risk mapping coefficient, For reference temperature;
[0256] Based on this, construct the unmodulated state vector:
[0257] And incorporate electric potential estimation ;
[0258] in for The state vector at time step For state dimension, square brackets This indicates that a vector is formed by concatenating columns, and the superscript T is the transpose operator;
[0259] Finally, low-rank affine modulation is performed using the batch field marker to reflect batch differences and enhance portability. Definition:
[0260] ;
[0261] in for Modulated state vector at time 10:00 For the batch field tag vector generated in step S1, For marking dimensions, and The low-rank factor matrix of the linear part It is a rank parameter and The batch field marker is generated by the parameter generation function. The obtained coefficient vector, Operators for converting vectors into diagonal matrices The batch field marker is generated by the parameter generation function. The obtained bias vector is the final output. The neural ordinary differential equations in step S5 are used for evolution prediction.
[0262] In this specific embodiment, S5 specifically includes:
[0263] Using the modulated state as the initial condition, the current process parameters are used as control variables, and batch domain labels are invoked to perform evolution prediction and uncertainty propagation in the physically constrained control affine constant differential equation. The obtained state distribution is then converted into a key quality attribute prediction distribution through moment mapping and index calculation for subsequent optimization. The initial time is initially set to... The predicted time domain length is , the length of the walk The discrete time sequence is and and In steps The modulated state vector is the initial value. ,in for The state vector at time step for The state vector after batch affine modulation at each time step, For state dimension, To unify the time grid, the state evolution is described by the neural ordinary differential equations governing the affine structure:
[0264] ;
[0265] in The derivative of the state vector with respect to time, Differential mechanism terms are used to characterize known physicochemical processes. The residual term is used to compensate for unmodeled dynamics and is determined by parameters. control, To make the derivative of the control gain matrix linearly dependent on the control quantity, for The current process parameter vector at time t and For its dimensions, Batch field label vector and Its dimensions;
[0266] And batch adaptation of the residual parameters is performed in a low-rank affine form:
[0267] ;
[0268] in For the set of reference parameters, and For low-rank factor matrices, For parameter dimensions, It is a rank parameter and satisfies The coefficient vector is obtained by generating the parameter from the batch field markers using a parameter generation function. For diagonalization operators;
[0269] During numerical solution, physical consistency constraints are introduced to ensure feasibility and stability. Nonnegativity is achieved jointly through positive value mapping and mass conservation equilibrium constraints, while also taking into account the monotonicity of the cumulative structure.
[0270] ;
[0271] in For internal unconstrained state, For Softplus positive value mapping, For the mass conservation coefficient vector, For total constant, The non-negative weight vector is used for the accumulation structure of the energy input correlation channel. It is represented by its unconstrained parameters;
[0272] Numerical solver Proceeding the state trajectory at discrete moments:
[0273] ;
[0274] in A unified representation for Runge-Kutta or Adams-type solvers and in step size To advance the process, we consider the uncertainties in initial values and parameters and introduce process noise covariance to approximate the distribution propagation using a linearized approach:
[0275] and ;
[0276] in for Time-state covariance, for Time-mean state The Jacobian matrix at the mean of the states, The aforementioned total drift function is... The synthetic representation, For process noise covariance, for Mean of the moment control variable;
[0277] Final with state to key quality attribute differentiable mapping Convert state distribution to key quality attribute distribution and update its statistics at first order approximation With ;
[0278] Where Is Moment key quality attribute vector contains volume distribution median diameter, polydispersity index, zeta potential and dissolution profile eigenvalues, Is differentiable function composed of moment mapping and index calculation, Is Jacobian matrix at state mean;
[0279] Aggregate statistics of discrete moments to form predicted state trajectory distribution And key quality attribute predicted distribution And output And Probabilistic safety constraint combined with control barrier function of two-time domain model predictive control for step S6.
[0280] In the specific embodiment, the S6 is specifically:
[0281] With predicted state trajectory distribution and key quality attribute predicted distribution as input, construct probabilistic safety constraint of key quality attribute and control barrier function hard constraint of limited process variable in model predictive controller at the same time, and set time domain length and weight of two-time domain rolling optimization according to process stage determination, and finally solve to obtain current control instruction vector under cost function including compliance time, control action variability and energy consumption: first define discrete rolling moment sequence as , Wherein Is the Rolling moment, Is the prediction step number;
[0282] Then expressed as:
[0283] ;
[0284] Wherein Is the normal approximation set at each moment, Is normal distribution, Is the mean vector of the Step key quality attribute, Is the covariance matrix and provided by step S5;
[0285] Is the The key quality attributes in the first The random variable in the first step is and the target interval is and the threshold of the qualified probability is The chance constraint is given as to ensure the median diameter of the volume distribution the polydispersity index PDI, the zeta potential and the characteristic values of the dissolution curve The probability of the four key quality attributes in their respective target bands is not less than the threshold, where is the predicted random variable of the first key quality attribute, is the lower limit, is the upper limit, is the minimum qualified probability and can be approximated in the implementation by and the edge statistics of
[0286] At the same time, control barrier function hard constraints are set for the limited process variables with respect to the safety boundary of the device and process, and the index of the limited variable is and is the predicted value of the first limited variable in state and control , the upper and lower limits are defined as and and the discrete barrier function is given as
[0287] and
[0288] In order to ensure the safety of the upper and lower limits at the same time, hard constraints and are imposed and the discrete degenerate condition of forward invariance is introduced, where is the contraction coefficient of the first barrier, is the state vector of the first step and is provided by the prediction of step S5, is the control amount of the first step and is consistent with the actuator setting;
[0289] In order to realize double-time-domain rolling optimization, stage judgment is performed according to the predicted state trajectory distribution to obtain stage markers , where is the process stage index of the first step, For stage determination function, The set of state distributions is provided by step S5, and the lengths of the short-time domain and the long-time domain are adaptively set accordingly. and and its weight in the cost function and ,in and Used to balance "rapid achievement" with "smooth motion and energy consumption suppression";
[0290] Let the optimization variable be the control sequence. And define the control variable. with initial value For the currently implemented process parameter vector, construct a dual time-domain cost function:
[0291] ;
[0292] in To optimize objectives, For the first Step distance penalty For positive part operators, for mean To control the rate of change penalty coefficient, For the second norm, Energy consumption penalty coefficient, For the first The estimated energy consumption consists of mechanical power, ultrasonic power, and heating power. For torque, Angular velocity, For ultrasonic power, This refers to the heating power.
[0293] Considering the above constraints and costs, a rolling optimization problem is presented. This makes it possible for all With all satisfy And for all satisfy and and Simultaneously apply engineering constraints to control the amplitude and rate of change. and ,in For optimal control sequence, and These are the lower and upper bounds for control, respectively. To determine the maximum allowable rate of change, optimization is performed by the solver. Solve online at each rolling time instant and output the first component of the optimal sequence as the current control command vector, denoted as , where is the control command vector issued to actuators.
[0294] In this specific embodiment, the S7 is specifically:
[0295] Take the control command vector as input, convert it to target settings of each actuator through the calibration mapping from command to actuator, and issue it under the constraints of rate limit and amplitude limit, then record the command issuance time and collect actuator feedback or process measurements for execution confirmation, while synchronously collecting raw multi-modal data and updating the current process parameter vector for use in the next round of feature alignment and energy position coding: let the current control step be , and denote the control command vector obtained in step S6 as , where is the control command issued in the th step, is the controllable process parameter dimension;
[0296] Calculate the original target setting according to the command-to-actuator mapping function , where is the original target setting value of each actuator in the th step, includes unit conversion, calibration compensation, and non-linear characteristic compensation;
[0297] To avoid equipment impact and out-of-bound, use step-by-step rate limit and element-level amplitude constraint to get the final issued setting:
[0298] ;
[0299] where is the final target setting of the actuator in the th step, is the executed setting in the th step, is the element-level truncation operator on the vector to limit each component within , is the maximum allowed rate of change vector, is the element-level saturation operator to limit the output between the lower bound and the upper bound ;
[0300] In the case of The instruction issuing time is recorded after the field bus is issued to the homogenizing pump, valve group, high shear stirring drive, ultrasonic power supply, heating and cooling unit, dosing pump and feeding pump, etc. And the feedback vector is collected at the end of the set confirmation window With its timestamp , wherein Is the actual working state quantity returned by the actuator body or the process side in step Define the error vector for execution confirmation And the decision threshold Determined by using the infinite norm The target setting is considered to be achieved, wherein Is the actual achieved value in step The execution error, The vector infinite norm, The allowed error bound;
[0301] The confirmed actual achieved value is written into the current process parameter vector, denoted as , wherein Is the updated process parameter vector in step
[0302] At the same time, new original multi-modal data is collected under a unified time base according to the sensor synchronization strategy and is attached with a timestamp, denoted as:
[0303]
[0304] , wherein Is the newly collected original data set, Is the Multi-channel original measurement vector at time t includes near-infrared spectrum, ultraviolet-visible spectrum, online particle size, temperature, pressure, pH, torque, power, valve position and equipment load, Is the measurement dimension, Is the timestamp set of the newly collected data, and finally output Is used for feature alignment and energy position coding in step S2.
[0305] The above is only the preferred specific implementation of the present application, but the protection scope of the present application is not limited thereto, any skilled person in the art can make equivalent replacement or change according to the technical solution and the inventive concept of the present application within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
[0306] The application constructs stable, physically relevant multi-modal feature input with energy-aware continuous-time position encoding and sensor reliability gated weighting; realizes online accurate estimation of key quality attributes under the mechanism coupling of prior constraints through multi-modal cross-attention Transformer; uses the state vector containing the particle size distribution matrix, zeta potential, supersaturation, and recrystallization risk fast variable as a bridge, uses the neural ordinary differential equation with control affine and quality conservation, non-negativity, and monotonicity constraints for distribution-level prediction, and uses uncertainty propagation to depict the risk and compliance confidence in the probability form on the control side; finally, in the dual-time domain model predictive control with the introduction of key quality attribute probability safety constraints and limited process variable control barrier functions, both fast compliance and action smoothing / energy consumption suppression are considered. The above algorithm combination synergistically enables the key quality attributes such as median particle size, polydispersity index, zeta potential, and dissolution curve to reach the target range within the preset observation time, and the equipment and process variables are kept within the safety boundary, significantly reducing overshoot and batch-to-batch difference, and improving the consistency and energy efficiency of the production line.
[0307] The application makes targeted improvements in algorithm structure to solve the technical problems of "insufficient multi-modal fusion robustness, lack of physical consistency and batch adaptability of prediction model, and difficulty in coordinating quality and equipment safety": continuous-time position encoding explicitly introduces the time integral of mechanical / ultrasonic / thermal energy to characterize the cumulative effect, and sensor reliability gated weighting suppresses the influence of low signal-to-noise ratio and drift; mechanism-coupled attention mask and bias embed mature mechanisms into cross-modal attention to improve generalization and interpretability; batch domain low-rank affine modulation describes raw material differences with low parameter quantity and maintains portability; neural ordinary differential equation with control affine structure maintains friendly optimizability for control variables under physical constraints and supports efficient uncertainty propagation; dual-time domain model predictive control realizes the coordination of quality and safety with chance constraints and control barrier functions, sets weights and time domain length adaptively with stage determination, and realizes the unification of "fast compliance-stable operation-low energy consumption". The above structural improvements further ensure that the technical effect is repeatable, provable, and real-time in actual production environment.
Claims
1. A method for solubilizing atractylodes lancea dispersions based on deep neural networks and strategy optimization, characterized in that, include: S1. Collect raw multimodal data, generate batch domain markers based on incoming inspection, and set the target range of key quality attributes, equipment safety boundaries, process stages and stage switching conditions, and current process parameter vectors. S2. Preprocess and align the original multimodal data, generate continuous time position codes by combining the current process parameters and sensor readings, and perform gating weighting of multimodal features based on sensor reliability, outputting a reliability-weighted multimodal feature sequence and continuous time position codes; S3. Input the reliability-weighted multimodal feature sequence and continuous time location encoding into the multimodal cross-attention Transformer, and call the batch domain label and the current process parameter vector to perform conditional inference, and output the key quality attribute estimation results. S4. Based on the estimation results of key quality attributes, particle size distribution moment mapping is performed and a state vector is constructed. Near-infrared spectra are extracted from the reliability-weighted multimodal feature sequences. The oversaturation is calculated by combining temperature and incorporated into the state vector. The batch domain label is called to perform low-rank affine modulation on the state vector. S5. Using the state vector as the initial condition, combined with the current process parameter vector and calling the batch domain label, perform evolution prediction in the physically constrained neurodynamic model, and output the predicted state trajectory distribution and key quality attribute prediction distribution through uncertainty propagation. S6. Based on the predicted state trajectory distribution and the predicted distribution of key quality attributes, and calling the target range of key quality attributes, equipment safety boundaries, and process stages and stage switching conditions, construct the probabilistic safety constraints of key quality attributes and the hard constraints of the control barrier function of restricted process variables in the model predictive controller. Set the time domain length and weight of the dual time domain rolling optimization according to the stage determination, solve and output the control command vector under the cost function that includes the target time, control action variability and energy consumption. S7. Drive the actuator to update the current process parameter vector according to the control command vector, collect new original multimodal data for the next round of feature alignment and energy position encoding, and repeat until the key quality attribute target range is stably met within the preset observation time.
2. The method for solubilizing atractylodes lancea dispersion based on deep neural networks and strategy optimization according to claim 1, characterized in that, S1 specifically refers to: Sensors and analyzers for online data acquisition are arranged on the preparation equipment. The sensors and analyzers include at least a near-infrared spectroscopy sensor, an ultraviolet-visible spectroscopy sensor, an online particle size analyzer, a temperature sensor, a pressure sensor, a pH meter, a torque sensor, an ultrasonic power meter, and a heating power meter to generate raw multimodal data. A batch field mark is generated based on the raw material incoming inspection results. The incoming inspection results include at least one or more of the following information: powder particle size distribution, crystal form information, water content, impurity level and spectral fingerprint. A target range for key quality attributes is preset, and the key quality attributes include at least the median particle size, polydispersity index, electrokinetic potential, and dissolution curve characteristic values. The equipment safety boundary is preset, which includes at least the allowable range of temperature, pressure, pH, speed, power, valve position and equipment load; Define process stages and stage switching conditions, wherein the process stages include at least two or more stages, and the stage switching conditions are determined based on thresholds, time or stability of key quality attributes or process variables; Set the current process parameter vector, which includes at least homogenization pressure, shear speed, ultrasonic power setting, temperature setting, pH setting, solid content, and feed flow rate; Output raw multimodal data, batch domain markers, target ranges of key quality attributes, equipment safety boundaries, process stage definitions and stage switching conditions, and current process parameter vectors.
3. The method for solubilizing atractylodes lancea dispersion based on deep neural networks and strategy optimization according to claim 1, characterized in that, S2 specifically refers to: Based on the original multimodal data, sensor baseline correction and timestamp alignment are performed. By baseline subtraction, drift correction and dimensional unification, the alignment of signals with different sampling frequencies is completed on a unified time grid. On the aligned time grid, instantaneous mechanical power is calculated based on torque and shear speed, and cumulative mechanical energy is obtained by time integration. Cumulative ultrasonic energy is obtained by time integration based on ultrasonic power meter readings. Cumulative heat is obtained by time integration based on heating power meter readings. A continuous time position code is formed by timestamp, cumulative mechanical energy, cumulative ultrasonic energy, and cumulative heat. Based on the signal-to-noise ratio and drift detection results of the aligned signal, sensor reliability weights are assigned to each mode, and the modal features are gated and weighted according to the sensor reliability weights to suppress anomalies and distortions, generating a reliability-weighted multimodal feature sequence. Output a reliability-weighted multimodal feature sequence and a continuous time location code.
4. The method for solubilizing atractylodes lancea dispersion based on deep neural networks and strategy optimization according to claim 1, characterized in that, S3 specifically refers to: Cross-modal information fusion and temporal modeling are performed in a multimodal cross-attention Transformer using reliability-weighted multimodal feature sequences and continuous temporal position encoding to characterize the temporal evolution and energy accumulation effects of different sensing channels. Attention masks and biases are set according to the mechanism coupling relationship, so that the effects of acidity and ionic strength on electromotive potential, shear and homogeneous energy on particle size distribution, and temperature on solubility and recrystallization risk are enhanced in the attention weights. The inference is conditionalized by calling the batch domain marker and the current process parameter vector, so that the model output is adaptive to the raw material batch and the current operating conditions. Generate key quality attribute estimation results, including median particle size, polydispersity index, electrokinetic potential, and dissolution curve characteristic values; Output the estimation results of the key quality attributes.
5. The method for solubilizing atractylodes lancea dispersion based on deep neural networks and strategy optimization according to claim 1, characterized in that, S4 specifically refers to: Based on the estimation results of key quality attributes, the median particle size and polydispersity index are mapped to the zeroth moment, first moment and second moment of particle size distribution through a preset moment mapping relationship, and the electric potential is incorporated into the state vector. Near-infrared spectra are extracted from the reliability-weighted multimodal feature sequences, solute concentration is inverted using a quantitative correction model, and supersaturation is calculated based on the solubility curve in conjunction with temperature. The supersaturation is then incorporated into the state vector. Based on the supersaturation and temperature, a fast variable for recrystallization risk is constructed and incorporated into the state vector; Low-rank affine modulation coefficients are generated using batch domain labeling, and affine modulation is applied to the state vector to reflect batch differences. It also outputs the state vector.
6. The method for solubilizing atractylodes lancea dispersion based on deep neural networks and strategy optimization according to claim 1, characterized in that, S5 specifically refers to: Using the state vector as initial conditions, the time evolution of the state is solved in a neural ordinary differential equation with a control affine structure and coupled mechanistic and residual terms. Mechanism terms are used to characterize differentiable physical processes, while residual terms are used to compensate for unmodeled dynamic errors. The current process parameter vector is called as the control variable and entered into the affine term, so that the state derivative has an affine relationship with the control variable; The batch domain label is invoked to perform low-rank affine modulation on the model parameters to achieve adaptation to batch differences in raw materials; In the numerical solution process, mass conservation, nonnegativity and monotonicity constraints are imposed on the state and parameters. Nonnegativity is maintained by positive value mapping, conservation is maintained by balance constraints or penalty terms, and monotonicity is achieved by non-negative weights and accumulation structure. A numerical solver is used to obtain the probability distribution of the state over time in a preset prediction time domain, and uncertainty propagation is carried out through linear approximation or random sampling. Based on moment mapping and index calculation, the probability distribution of the state is transformed into the predicted distribution of key quality attributes; It outputs the predicted state trajectory distribution and the predicted distribution of key quality attributes for predictive control of a dual-time-domain model that combines probabilistic safety and control barrier functions.
7. The method for solubilizing atractylodes lancea dispersion based on deep neural networks and strategy optimization according to claim 1, characterized in that, S6 specifically refers to: Based on the predicted state trajectory distribution and the predicted distribution of key quality attributes, a probabilistic safety constraint is constructed in the model prediction controller for the target range of key quality attributes, so that the probability of the key quality attribute being within the target range is not lower than a preset threshold. Simultaneously, based on the equipment safety boundary, control barrier functions are set to form hard constraints for the restricted process variables. The restricted process variables include at least one or more of temperature, pH, pressure, speed, power, valve position, or equipment load. Based on the process stage and stage switching conditions, the predicted state trajectory distribution is staged to generate stage markers, and the time domain length and weight of the dual time domain rolling optimization are set accordingly. The short time domain is used to quickly approach the target range of key quality attributes, and the long time domain is used to control motion smoothing and energy consumption suppression. Under the conditions of satisfying the probabilistic safety constraints and the hard constraints of the control barrier function, a cost function containing the target time, control action variability and energy consumption is constructed and the rolling optimization problem is solved to obtain the control command vector.
8. The method for solubilizing atractylodes lancea dispersion based on deep neural networks and strategy optimization according to claim 1, characterized in that, S7 specifically refers to: Based on the control command vector, the target set value of each actuator is calculated through the preset command-to-actuator mapping relationship, and then sent to the actuator under the change rate limit and amplitude limit constraint, so that the homogenization pressure, shear speed, ultrasonic power setting, temperature setting, pH setting, solid content and feed flow rate are updated accordingly. Under the updated operating conditions, the time of instruction issuance is recorded and actuator feedback or process measurement is collected for execution confirmation. At the same time, raw multimodal data is collected synchronously by sensors and timestamps are attached to form newly collected raw multimodal data. The actual achieved process variables are calculated based on the comparison between actuator feedback or process measurement and the target set value, and the actual achieved values are written into the current process parameter vector to form an updated process parameter vector. The newly acquired raw multimodal data and the updated process parameter vector are output for feature alignment and energy location encoding of S2.