Acanthopanax sessiliflorus flavone online extraction self-optimization method based on reinforcement learning
By using a reinforcement learning-based approach, combining multi-source data to generate fused feature vectors and a virtual metric, the batch variation and safety constraints in the extraction of flavonoids from Acanthopanax senticosus were addressed, enabling real-time monitoring and optimization, and improving extraction stability and safety.
Patent Information
- Application Number
- CN202511545257.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-01-16
AI Technical Summary
Existing technologies cannot address batch-to-batch variations in the origin, harvesting period, and moisture content of Acanthopanax senticosus raw materials. Traditional offline optimization methods struggle to guarantee extraction stability, while existing online spectral monitoring methods lack the ability to comprehensively evaluate multi-component flavonoids and do not adequately consider safety constraints such as solvent flash point and equipment temperature resistance.
By employing a reinforcement learning-based approach, near-infrared spectroscopy, ultraviolet-visible spectroscopy, liquid viscosity, and conductivity data are collected to generate a fused feature vector, construct a virtual measurement device, and combine solvent flash point and equipment temperature threshold to perform linkage optimization of process parameters, thereby achieving real-time monitoring and safety control.
It enables real-time prediction and quantitative evaluation of components such as rutin, quercetin, kaempferol, and hyperoside, improving the stability and safety of the extraction process, reducing operating energy consumption and solvent usage, and possessing self-learning and self-correction capabilities.
Smart Images

Figure CN121350795A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of online extraction control of acanthopanax gracilistylus flavonoids, and in particular to an online extraction self-optimization method for acanthopanax gracilistylus flavonoids based on reinforcement learning. BACKGROUND
[0002] Acanthopanax gracilistylus is a plant of the genus Acanthopanax in the family Araliaceae, and its roots, stems and leaves are rich in flavonoids such as rutin, quercetin, kaempferol and hyperoside, which have pharmacological effects such as antioxidant, anti-fatigue and immune regulation. At present, the industrial extraction of acanthopanax gracilistylus flavonoids mainly adopts ethanol solvent extraction method, and the target components are obtained by controlling process parameters such as temperature, ethanol volume fraction, solid-liquid ratio, soaking time and circulation flow. Traditional process parameter optimization methods include single factor test method, orthogonal test method and response surface method, etc. These methods determine the optimal parameter combination through offline experiments and then apply it to production. In recent years, with the development of process analysis technology, online detection means such as near-infrared spectroscopy and ultraviolet-visible spectroscopy have been gradually applied to the quality monitoring of plant active ingredient extraction process. Some studies use chemometrics methods such as partial least squares regression to establish a quantitative model of spectral data and component content.
[0003] However, the existing technology has the following main problems: first, the traditional offline optimization method cannot cope with the batch differences of acanthopanax gracilistylus raw materials in terms of origin, harvesting period and moisture content, and it is difficult to ensure the stability of the extraction when the preset parameters fluctuate; second, the existing online spectral monitoring methods are mostly limited to tracking a single quality indicator, and lack the ability to comprehensively evaluate the multi-component flavonoids such as rutin, quercetin, kaempferol and hyperoside; third, the existing process optimization algorithms do not fully consider safety constraints such as solvent flash point, equipment temperature resistance and component thermal degradation when adjusting parameters, which may lead to boundary crossing risk. Therefore, it is necessary to develop an online extraction self-optimization method for acanthopanax gracilistylus flavonoids based on reinforcement learning. SUMMARY
[0004] The present application provides an online extraction self-optimization method for acanthopanax gracilistylus flavonoids based on reinforcement learning to reduce batch fluctuations and improve resource utilization efficiency.
[0005] The present application provides an online extraction self-optimization method for acanthopanax gracilistylus flavonoids based on reinforcement learning, comprising: Collecting near-infrared spectroscopy data, ultraviolet-visible spectroscopy data, slurry viscosity data and slurry conductivity data during the extraction process of acanthopanax gracilistylus flavonoids, and generating a fusion feature vector after baseline correction, scattering compensation and feature alignment processing; A mapping relationship of the fusion feature vector to rutin content, quercetin content, kaempferol content and hyperoside content is established by using offline test samples, and a virtual meter is constructed, which outputs a target flavone composition vector and an uncertainty of the target flavone composition vector; A safety domain boundary is constructed according to a solvent flash point temperature, an upper limit of equipment tolerance temperature and a flavone component thermal degradation temperature threshold, and a time-varying safety constraint threshold is calculated in combination with the uncertainty of the target flavone composition vector; A reinforcement learning strategy is adopted to perform linkage optimization on temperature, ethanol volume fraction, solid-liquid ratio, soaking time and circulation flow rate, taking the target flavone composition vector and the time-varying safety constraint threshold as inputs, wherein a reward function of the reinforcement learning strategy is composed of a weighted distance between the target flavone composition vector and a standard flavone composition vector and a solvent consumption energy consumption penalty term, and when an optimization action predicted value reaches the safety domain boundary, the optimization action is rolled back to a safe process parameter; The temperature, ethanol volume fraction, solid-liquid ratio, soaking time and circulation flow rate obtained by the linkage optimization are output to an extraction device for execution.
[0006] The beneficial effects of the technical solutions provided in the application include: (1) By fusing multi-source spectra (near-infrared, ultraviolet-visible) with process parameters (viscosity, conductivity), a fused feature vector is established and a virtual metric is constructed, enabling real-time prediction and quantitative evaluation of the content of key components such as rutin, quercetin, kaempferol, and hyperoside. Compared with traditional extraction processes that rely solely on offline detection, this application can achieve online quantitative monitoring without sampling or machine shutdown, transforming the component control of the extraction process from post-processing adjustments to real-time feedback, significantly improving batch-to-batch stability and quality consistency. (2) This application establishes a safety domain boundary based on solvent flash point, equipment tolerance temperature, and flavonoid thermal degradation threshold, and dynamically calculates time-varying safety constraint thresholds using the uncertainty of the virtual metric, achieving risk-aware safety control. When the predicted value of the reinforcement learning optimization action approaches the safety boundary, the system automatically reverts to conservative process parameters, thereby preventing flavonoid degradation or equipment malfunction due to excessive temperature or solvent concentration, significantly improving the safety and reliability of the extraction process. (3) The reinforcement learning algorithm takes the weighted distance between the target flavonoid composition vector and the standard composition vector as the core optimization objective, and introduces solvent consumption and energy consumption penalty terms. It can optimize multiple process variables such as temperature, ethanol volume fraction, solid-liquid ratio, soaking time and circulation flow rate while ensuring the quality of components. This strategy enables the system to have self-learning and self-correction capabilities. It can autonomously seek optimization according to different raw material characteristics and equipment status, realize the dynamic balance between extraction efficiency and energy consumption, and reduce operating energy consumption and solvent usage. (4) The virtual metric + reinforcement learning dual-layer structure proposed in this application can be directly deployed on the existing extraction equipment data acquisition system without modifying the hardware structure to form an online self-optimization closed loop. This method enables the traditional plant extraction process to have the full-process self-decision-making capability of intelligent perception-prediction-optimization-execution, and provides a generalized and transferable intelligent control framework for the industrial extraction of flavonoids from Acanthopanax senticosus and other Chinese herbal medicine components. Attached Figure Description
[0007] Figure 1 This is a flowchart of an online self-optimization method for extracting flavonoids from Acanthopanax senticosus based on reinforcement learning, provided in the first embodiment of this application. Detailed Implementation
[0008] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.
[0009] The first embodiment of this application provides a self-optimizing method for online extraction of flavonoids from Acanthopanax senticosus based on reinforcement learning. Please refer to [link / reference needed]. Figure 1 This figure is a schematic diagram of the first embodiment of this application. The following is in conjunction with... Figure 1The first embodiment of this application provides a detailed description of an online self-optimization method for extracting flavonoids from Acanthopanax senticosus based on reinforcement learning.
[0010] Step S101: Collect near-infrared spectral data, ultraviolet-visible spectral data, liquid viscosity data, and liquid conductivity data during the extraction of flavonoids from Acanthopanax senticosus. After baseline correction, scattering compensation, and feature alignment, generate a fused feature vector.
[0011] Step S101 involves the acquisition of multi-source process data and the generation of fusion features during the online extraction of flavonoids from Acanthopanax senticosus. This step is the basic data acquisition and feature construction stage of the entire self-optimization control method, and its accuracy directly determines the reliability of subsequent virtual metric modeling and reinforcement learning decision-making.
[0012] During the extraction process, the flavonoid raw material of Acanthopanax senticosus typically undergoes soaking, heating, and solvent circulation, resulting in continuous changes in the chemical composition and physical state of the material over time. To reflect this dynamic process in real time, it is necessary to simultaneously acquire data signals from multiple different sources. Near-infrared spectroscopy signals are used to characterize the vibrational absorption characteristics of functional groups in the extract, especially the stretching vibrations of O–H, C–H, and C=O bonds. These absorption peaks are highly sensitive to changes in the concentration of flavonoid compounds. Ultraviolet-visible spectroscopy is used to characterize the electronic transition characteristics of the conjugated system in flavonoid molecules, reflecting the differences in absorption intensity of components such as rutin, quercetin, and kaempferol. The viscosity data of the extract reflects the dynamic changes in the solid-liquid ratio and solute concentration, and can be obtained in real time using an online rotational viscometer or ultrasonic viscosity sensor. The conductivity data reflects the changes in ion concentration and ethanol volume fraction in the solution, and can be measured using a high-temperature resistant online conductivity sensor.
[0013] During the acquisition process, to ensure data comparability and synchronization, near-infrared and ultraviolet-visible spectral signals must be sampled at the same time intervals, with a recommended sampling period of 5 to 10 seconds. An external trigger synchronization mechanism should be used to ensure that different sensors complete one sampling at the same time point. To prevent signal drift caused by probe contamination, temperature fluctuations, or light source intensity attenuation, all spectral acquisition modules should be calibrated using standard reference samples before daily production. For example, a blank scan can be performed using pure ethanol or a standard silicon wafer that is transparent to the optical system to determine the baseline response.
[0014] After obtaining the original spectral signal, baseline correction needs to be performed. Baseline correction refers to removing the overall signal rise or fall caused by instrument noise or optical path instability, so that the spectral intensity reflects the true absorption characteristics. In practice, this can be achieved through multi-point baseline subtraction or moving average baseline correction. For example, regions with weaker absorption at both ends of the spectrum are selected, a smooth curve is fitted as the baseline, and subtracted from the original spectrum. The corrected spectrum is then subjected to scattering compensation. Scattering compensation is used to eliminate the effects of non-absorption scattering caused by changes in particle concentration, solution turbidity, or bubbles. It can be achieved using the standard normal variable transformation method or the multivariate scattering correction method. The former makes all spectra comparable in intensity by subtracting the spectral mean from each spectral point and dividing by the standard deviation; the latter eliminates systematic bias caused by particle scattering by fitting a linear relationship between the reference spectrum and the sample spectrum.
[0015] After scattering compensation, feature alignment of the multi-source spectral signals is required. Feature alignment refers to the synchronous registration of near-infrared and ultraviolet spectra in the wavelength or wavenumber dimension to eliminate feature misalignment caused by response delays from different sensors. For example, the cross-correlation maximization method can be used, taking the time offset corresponding to the maximum value in the similarity curve of two spectra as the correction basis, and interpolating and resampling the spectral sequence to align key absorption peaks on the time axis. For non-spectral signals, such as viscosity and conductivity, normalization is required to make their dimensions comparable to spectral characteristics. A common method is to subtract the sliding window average from the original value and then divide by the standard deviation to eliminate scale differences.
[0016] After data synchronization, features from different sources are fused. Fusion can be achieved using weighted splicing or principal component extraction (PCE). If weighted splicing is used, appropriate weights are assigned between spectral and process data, for example, based on the sensitivity of each channel to changes in flavonoid concentration, so that the total eigenvector comprehensively reflects both spectral absorption and changes in physical parameters. If PCE is used, principal component analysis is performed on all data matrices to extract the top principal components that explain more than 95% of the total variance, forming a low-dimensional fused eigenvector.
[0017] The generated fusion feature vector is typically a one-dimensional numerical sequence containing several fixed-dimensional data points. Each data point represents the combined optical and physical state information of the extraction system at the current moment. The output of this step is the fusion feature vector dataset. This dataset can not only serve as input for subsequent virtual metric modeling but can also be directly used to analyze dynamic trends during the extraction process. For example, when the extractant color deepens and the absorption at 280 nm in the spectrum increases, the corresponding component of the fusion feature vector will increase, indicating an increase in flavonoid concentration; conversely, when the conductivity decreases significantly, the component representing an increase in ethanol concentration will increase, thus reflecting changes in the solvent ratio.
[0018] To improve data robustness, a sliding time window averaging mechanism can be introduced during implementation, taking a weighted average of several consecutive sampling results to eliminate the impact of occasional fluctuations. For sudden abnormal data, such as sensor communication interruptions or measurements exceeding the equipment's range, these should be automatically discarded, and linear interpolation compensation should be performed using valid data from before and after them to ensure data continuity.
[0019] In summary, the execution process of step S101 starts with real-time acquisition of signals from multiple channels, and through systematic data preprocessing, feature alignment and multi-source fusion, finally generates a fused feature vector that can accurately reflect the key chemical and physical states of the flavonoid extraction process of Acanthopanax senticosus.
[0020] Furthermore, the near-infrared spectral data, ultraviolet-visible spectral data, liquid viscosity data, and liquid conductivity data collected during the extraction of flavonoids from Acanthopanax senticosus are processed through baseline correction, scattering compensation, and feature alignment to generate a fused feature vector, including: Using the sampling time of the conductivity of the liquid as a unified time reference, the sampling frequency of the near-infrared spectrum and the ultraviolet-visible spectrum is resampled so that different sensing channels correspond to the same process state in the same time series, thus obtaining the original multi-source dataset with time alignment. A spectral baseline reference is established using the pure ethanol and deionized water signals from the initial stage of production. The spectral baseline reference is then subtracted point by point from the time-aligned original multi-source dataset to obtain corrected spectral data, which is used to eliminate baseline shifts caused by light source attenuation, optical path drift, and probe temperature changes. In the calibrated spectral data, the range of signal amplitude variation and instantaneous rate of change are calculated according to a fixed time window. When the change in any segment exceeds the reference range under normal operating conditions, it is determined to be a disturbed segment and replaced with a smooth interpolation of the adjacent normal segment, thereby forming disturbed spectral data. By combining the de-disturbed spectral data, viscosity measurement curves, and conductivity change curves, and according to the scattering correction coefficients recorded in the equipment parameter table, the amplitude of the spectral signals in different bands is corrected to ensure that the physical relationship between the spectral intensity and the viscosity and ion concentration of the liquid is consistent, thus obtaining the scattering correction spectrum. In the scattering correction spectrum, the moment when the conductivity curve rises sharply and the moment when the viscosity curve changes from falling to stabilizing are identified. These two moments are determined as characteristic time anchor points. The time axes of spectrum, viscosity and conductivity are uniformly shifted with reference to the characteristic time anchor points to achieve a consistent correspondence of various signals in the chemical reaction stage. The average intensity, peak position shift, and half-peak width of the main absorption peaks are extracted from the time-aligned scattering correction spectrum. The viscosity and conductivity readings of the liquid at the same time point are also extracted simultaneously. The extracted feature values are normalized and scaled to generate a comparable standardized feature set. The weighting coefficients are determined based on the stability of each channel signal in the standardized feature set and its correlation with conductivity. The spectral features, viscosity features and conductivity features are weighted, summed and sequentially concatenated to generate a fused feature vector.
[0021] During the extraction of flavonoids from Acanthopanax senticosus, the optical, fluid, and electrical properties of the extract continuously change over time. To accurately reflect these changes, this step employs multi-channel synchronous acquisition and multi-stage data processing to uniformly correct, compensate, and fuse near-infrared spectroscopy, ultraviolet-visible spectroscopy, and the viscosity and conductivity signals of the extract, ultimately generating a fused feature vector. This vector serves as the input for subsequent virtual metric and reinforcement learning optimization. This process not only resolves the mismatch between multiple source signals in terms of time and scale but also achieves multi-dimensional dynamic characterization of the flavonoid extraction system through the extraction and standardization of key feature parameters.
[0022] During the spectral signal acquisition stage, near-infrared and ultraviolet-visible spectra correspond to different molecular vibrational energy levels. Near-infrared spectroscopy mainly reflects the stretching vibration absorption behavior of C–H, O–H, and C–O bonds, while ultraviolet-visible spectroscopy reflects the π–π* transition characteristics of the benzene ring and conjugated double bond system in flavonoids. In the extraction of flavonoids from *Eleutherococcus senticosus*, major components such as rutin, quercetin, kaempferol, and hyperoside all exhibit significant absorption peaks. The strong absorption bands in the ultraviolet-visible region (wavelength approximately 350–370 nm) typically correspond to the B-ring conjugated system of rutin and quercetin molecules, while the absorption bands in the near-infrared region (wavelength approximately 1450–1950 nm) are related to the O–H stretching vibrations in the solvent molecules and the hydrogen bonding of the flavonoid hydroxyl groups. The system focuses on monitoring signals in these bands during extraction to capture changes in molecular absorption during the dissolution of flavonoids.
[0023] Because different sensors have different sampling rates, the system uses the sampling time of the liquid conductivity as a unified time reference and resamples the near-infrared and ultraviolet spectral signals according to the conductivity time axis. Changes in conductivity directly reflect changes in solute ion concentration, and are therefore selected as the time anchor point for the reaction process. The system establishes a unified time series at fixed intervals (e.g., 0.5 seconds) and supplements the near-infrared and ultraviolet signals with corresponding sampling points using linear interpolation, ensuring that the three signals correspond to the same physical state on the same time scale, thus obtaining a time-aligned original multi-source dataset.
[0024] After obtaining the time-aligned data, to eliminate baseline drift caused by equipment and light source, the system establishes spectral baseline references by scanning with pure ethanol and deionized water separately before extraction. Pure ethanol exhibits a stable broadband near-infrared absorption characteristic, while deionized water has relatively high O–H absorption peaks; the difference between these two is used to construct the spectral energy reference. The system subtracts this baseline reference point-by-point onto the time-aligned spectrum during extraction to obtain corrected spectral data. The physical significance of baseline subtraction lies in eliminating the overall shift caused by light source intensity attenuation, probe temperature changes, and optical path reflections, ensuring that spectral changes are only related to the chemical composition of the extract.
[0025] Within the calibrated spectral data, the system continues to detect transient disturbances caused by stirring, bubbles, or surface fluctuations in the liquid. Specifically, the system calculates the amplitude variation range and instantaneous rate of change of each signal band within a 10-second time window. When the change in any segment exceeds twice the normal reference range, it is determined that the segment is affected by mechanical or fluid disturbance. The system then recovers the smoothed signal through linear interpolation of adjacent normal segments, forming refracted spectral data. In this way, short-term mechanical vibrations do not affect the overall spectral trend, thus ensuring the stability of subsequent peak shape analysis.
[0026] Scattering effects in the spectrum are another common source of distortion. Suspended particles, foam, and solute molecular clusters in the feed solution lead to multiple reflections and scattering of light, manifesting as flattening or peak shifting of absorption peaks. To correct this effect, the system uses real-time viscosity and conductivity data to determine scattering correction coefficients. Increased viscosity indicates enhanced intermolecular interactions in the solution, leading to increased scattering; increased conductivity indicates increased solute concentration, leading to enhanced absorption. The system pre-stores scattering correction coefficients corresponding to different combinations of viscosity and conductivity in the equipment parameter table. During extraction, the system automatically looks up the correction value for the current conditions in the table and adjusts the intensity of each spectral band proportionally to generate a scattering correction spectrum, ensuring that the signal intensity matches the true absorption characteristics of the feed solution.
[0027] In the scattering correction spectroscopy, the system further establishes a unified time scale for chemical reactions. The conductivity curve typically shows a significant increase when flavonoids dissolve in large quantities, while the viscosity curve tends to plateau after the solid-liquid ratio stabilizes. The system automatically identifies the inflection point of conductivity increase and the inflection point of viscosity plateau, and uses the average time of these two points as the characteristic time anchor point. By shifting the time axes of the spectral, viscosity, and conductivity signals around this anchor point as the zero point, the system ensures synchronous correspondence of each signal at the reaction stage. For example, when conductivity increases at 60 seconds and viscosity plateaus at 80 seconds, the system recalibrates the time axis with 70 seconds as the zero point, ensuring consistency of signals from different channels at the same reaction stage.
[0028] After unifying the time axis, the system extracts features from the scattering-corrected spectrum. The "major absorption peak" refers to a signal peak on the spectral curve whose intensity is significantly higher than adjacent bands and corresponds to the characteristic wavelength range of known flavonoid absorption. For example, the major absorption peaks of rutin and quercetin in the ultraviolet region are located at 350–370 nm; kaempferol has a weak absorption band in the visible light boundary region (approximately 400 nm); and hyperoside has a characteristic peak at approximately 1450 nm in the near-infrared region. The system scans the absorption intensity curve within a preset wavelength range, identifies local maxima, defines them as absorption peak apexes, and records the average absorption intensity, peak position shift, and full width at half maximum (FWHM) near the peak apex. The average intensity reflects the magnitude of absorption capacity, the peak position shift reflects changes in the molecular environment, such as solvent polarity or hydrogen bond changes, and the FWHM reflects intermolecular interactions and energy distribution. For example, when the temperature of the extract increases, leading to weakened hydrogen bonds, the peak position may shift to shorter wavelengths by about 2 nm, while an increase in FWHM indicates the presence of more molecular energy level perturbations in the system.
[0029] In addition to spectral characteristics, the system also reads the viscosity and conductivity of the liquid at the same time point. Viscosity data reflects the flow resistance of the liquid and is related to solute concentration and molecular size; conductivity reflects changes in ion concentration. The system stores the characteristic data of each channel synchronously over time. Since spectral intensity, viscosity, and conductivity have different dimensions, normalization and scaling are required for comparison and fusion. Normalization converts all characteristic values into proportional values between zero and one, such as dividing the absorption intensity by the maximum value to reflect relative trends. Scaling amplifies or compresses characteristic values based on historical statistical fluctuation ranges, ensuring that different types of signals have similar sensitivity at the same order of magnitude. For example, if the spectral absorption intensity varies by 0.2, while the conductivity varies by 50 μS / cm, the system scales the conductivity change proportionally to give both equal weight in the fusion process.
[0030] The system then calculates the stability and correlation of each channel feature. Stability is calculated by the average change in feature values over adjacent time periods; the smaller the change, the higher the stability. Correlation is assessed by the degree of synchronization between the feature and conductivity over time, reflecting the feature's sensitivity to the extraction response. For example, when the increase in spectral intensity and the increase in conductivity follow the same trend, the correlation is high; if the two change in opposite directions, the correlation is low. The system determines the weighting coefficients based on these two indicators. Stable and highly correlated features receive higher weights, while features with higher noise or lag receive lower weights.
[0031] Finally, the system sums and fuses the spectral, viscosity, and conductivity features according to weighted coefficients, and then concatenates them in a fixed order to form a fused feature vector. Each dimension of this fused feature vector represents a temporally synchronous and physically complementary process feature. It can simultaneously describe the optical absorption state of molecules, the fluid properties of the feed solution, and the electrochemical behavior of the solute in the flavonoid extraction system of *Eleutherococcus senticosus* at the same time scale. The fused feature vector not only possesses physical interpretability but also maintains a fixed structure and continuous values, allowing it to be directly input into the virtual metric and reinforcement learning control module for real-time estimation of flavonoid component content and guidance for subsequent process optimization decisions.
[0032] Step S102: Using offline test samples, establish the mapping relationship between the fused feature vector and the contents of rutin, quercetin, kaempferol and hyperoside, and construct a virtual metric. The virtual metric outputs the target flavonoid composition vector and the uncertainty of the target flavonoid composition vector.
[0033] The main task of step S102 is to establish a mapping relationship between the fused feature vector generated in step S101 and the actual flavonoid component content data obtained from offline analysis, thereby constructing a virtual measuring instrument that can estimate the content of rutin, quercetin, kaempferol, and hyperoside in real time during the extraction process. The core purpose of this step is to establish a correlation between traditional laboratory chemical analysis results and real-time sensor data, enabling the system to "predict component content without sampling".
[0034] Before starting modeling, two types of datasets need to be prepared. One type is the fused feature vector samples from step S101, which record the preprocessed feature combinations of multi-source signals such as near-infrared spectroscopy, ultraviolet-visible spectroscopy, liquid viscosity, and conductivity at different time points. The other type is the offline test data corresponding to these time points. These test samples are usually taken from representative moments of the extract, such as the initial, middle, and final stages of extraction, and the content of each monomeric flavonoid component is accurately determined by high-performance liquid chromatography (HPLC) or liquid chromatography-mass spectrometry (LC-MS). To ensure the effectiveness of the modeling, each fused feature vector must have a corresponding test result. If test data for a certain time point is missing, linear interpolation can be used to complete it using results from adjacent time points, but the completion marker must be recorded to prevent errors during model training.
[0035] Next, the mathematical form of the mapping relationship needs to be defined. The mapping relationship is essentially a multi-input, multi-output function, with the input being a fused feature vector and the output being the relative content values of the four main flavonoid components and their uncertainties. Since there is usually a highly nonlinear relationship between spectral data and chemical concentrations, simply using linear regression is insufficient to meet the required accuracy. A more feasible approach is to first capture the correlation between high-dimensional features and component concentrations using methods such as principal component regression, partial least squares regression, or random forests, and then introduce residual propagation or sampling statistics mechanisms into the model structure to simultaneously estimate the uncertainty of the prediction results. Taking partial least squares regression as an example, the idea is to simultaneously extract latent variables that can explain the covariance between the input features and the output content. These latent variables can reduce noise interference and improve model stability, and the uncertainty range of each output quantity can be obtained by analyzing the residual distribution of the model on the validation set.
[0036] In practical implementation, all fused feature vectors can first be arranged in chronological order, and a training matrix with a one-to-one correspondence to the test data can be established. Then, the training set and validation set are randomly divided, for example, in an 8:2 ratio. The training set is used to fit the mapping function, and the validation set is used to test the model's generalization ability. For each flavonoid component, the mean absolute error between the predicted value and the measured value can be calculated, and its uncertainty range can be estimated by combining the residual variance. The model's hyperparameters are adjusted through cross-validation until both the error and uncertainty are stable within an acceptable range. Taking rutin as an example, if the average error between the model's predicted rutin content and the measured value is less than 5%, and the prediction uncertainty is less than half of the variance of the same batch of samples, the model can be considered to have the accuracy for production applications.
[0037] The mapping relationship defines the mathematical foundation of the virtual metric, which is the engineered implementation of this mapping relationship. Its input is a fused feature vector, and its output is the relative content and uncertainty of the four main flavonoid components. After the virtual metric is modeled, it needs to be deployed online to receive new fused feature vectors in real time and output prediction results. For this purpose, the model can be stored as a parameter file in the computing unit of the control system. When new spectral data is input, the model will automatically perform forward calculations and output the current target flavonoid composition vector. The target flavonoid composition vector is a four-dimensional vector representing the relative content ratios of rutin, quercetin, kaempferol, and hyperoside, and its unit can be uniformly expressed as a percentage relative to the total solids content. For example, if the system predicts rutin 0.8%, quercetin 0.5%, kaempferol 0.2%, and hyperoside 0.3%, then the vector is denoted as (0.8, 0.5, 0.2, 0.3).
[0038] In addition to the predicted content, the uncertainty of the prediction results also needs to be output. Uncertainty represents the reliability of the model's prediction results under the current input conditions; the larger the value, the more uncertain the model is about the results. This step can use the residual propagation method or the Monte Carlo sampling method to estimate the uncertainty. Taking the residual propagation method as an example, the standard deviation can be obtained based on the residual distribution of each sample on the validation set, and then multiplied by the empirical confidence coefficient during real-time prediction as the uncertainty estimate. For example, at a 95% confidence level, if the standard deviation of the quercetin prediction residual is 0.03%, then the uncertainty of this component can be taken as 0.06%. In this way, the system not only provides the content value but also a range representing the reliability, thus providing a basis for subsequent safety constraint calculations.
[0039] To ensure the stability of the model during long-term operation, a periodic calibration mechanism should be incorporated into the system. As the batches of raw materials extracted, the solvent ratios, and the equipment status change, the model may gradually shift. In practical applications, it can be set to re-extract samples for analysis after every certain number of extraction batches, and then fine-tune and update the model using the new data. During updates, an incremental learning approach can be used, that is, while retaining the core structure of the original model, only the output layer weights are adjusted, allowing the model to absorb new data features while maintaining its ability to recognize old data.
[0040] In industrial production, the real-time response time of virtual measurement devices typically needs to be controlled within the range of several seconds. If the acquisition cycle is ten seconds, the model prediction time should not exceed one second to ensure that the reinforcement learning algorithm can obtain the latest flavonoid content information under near real-time conditions. Therefore, it is recommended to load the model onto a local edge computing node during implementation to avoid data transmission to a remote server and causing latency.
[0041] When the virtual measurement unit operates continuously, its output forms a target flavonoid composition curve that changes over time. This curve can be used to determine the dynamic changes in the extraction process. For example, when the rutin content curve stabilizes while other components continue to rise, it indicates that the extraction process is not yet completely finished. When all curves stabilize and the uncertainty decreases significantly, the extraction can be considered to have reached its optimal point.
[0042] In summary, step S102 achieves real-time conversion from spectral signals to chemical components by constructing a high-precision mapping relationship between the fusion feature vector and the flavonoid component content. The introduction of the virtual measuring device frees the extraction process from the lag of traditional offline detection, forming an intelligent measurement mechanism that can be updated over time.
[0043] Furthermore, by utilizing offline test samples, a mapping relationship is established between the fused feature vector and the contents of rutin, quercetin, kaempferol, and hyperoside, and a virtual metric is constructed. The virtual metric outputs the target flavonoid composition vector and its uncertainty, including: Offline analysis and testing were conducted on different batches of Acanthopanax senticosus raw materials under various extraction conditions. In each test, the fusion feature vector and the actual measured content of rutin, quercetin, kaempferol and hyperoside at the corresponding time point were recorded to form a paired sample set containing time series, feature data and component concentration. The paired sample set is divided into several stability layers according to process parameters such as raw material moisture content, crushing particle size and extraction temperature. The change trend of each fusion feature component with the extraction process is calculated in each stability layer, and feature components with consistent change direction and high correlation in different layers are extracted to form a robust feature set. The robust feature set is used to reflect the signal source that is most sensitive to the changes of the main components during the flavonoid extraction process. In the robust feature set, the response relationship between each feature component and the content of each flavonoid component is determined. The response relationship is obtained by weighted smoothing of the multi-point correspondence between the same feature component and the measured content in the offline sample, so that the flavonoid content change trend corresponding to different feature components has continuity and comparability. The direction of increase and decrease and the slope of the response relationship in each value interval are recorded to characterize the sensitivity of the feature. Based on the sensitivity of each feature component, the weighting coefficients are determined and the fused feature vector is projected onto the weighted feature space. The estimated content values of rutin, quercetin, kaempferol and hyperoside at each time point are calculated to obtain the flavonoid content prediction results in time series form. Then, the difference between the prediction results and the offline laboratory measured values is calculated, and the range of variation of the difference in multiple batches of samples is statistically analyzed. The range of variation is used as the uncertainty of the corresponding component estimation. The estimated content values and their corresponding uncertainties are combined in chronological order to generate a flavonoid component estimation structure that can be updated over time. When a new fusion feature vector is received, the flavonoid component estimation structure outputs a target flavonoid composition vector including rutin, quercetin, kaempferol and hyperoside, and simultaneously provides the confidence interval for each component, which is used as a quantitative basis and safety boundary reference for optimization decision in reinforcement learning control.
[0044] In the extraction of flavonoids from Acanthopanax senticosus, different batches of raw materials often exhibit variations in moisture content, particle size, and storage conditions. These factors directly affect the dissolution rate and proportion of each component in the extraction solvent. To accurately estimate the content of major flavonoid components in real time during online extraction, this step constructs a precise mapping relationship between fused feature vectors and offline analytical data. This allows the system to extrapolate the relative content and reliability of each component using spectral and process data without interrupting production. This mapping relationship is not a traditional mathematical model, but rather a data association structure centered on the quantitative correspondence between characteristic components and chemical assay results.
[0045] First, offline analytical data needs to be collected from multiple batches of *Eleutherococcus senticosus* extraction experiments, ensuring that each batch of raw material exhibits representative differences in extraction conditions. Representative differences refer to identifiable variations in at least three dimensions: raw material moisture content, particle size, and extraction temperature. For example, one set of experiments could use raw material with 8% moisture content and pulverized to 60 mesh size, extracted at 85 degrees Celsius; another set could use raw material with 12% moisture content and pulverized to 80 mesh size, extracted at 90 degrees Celsius. For each experiment, the fused feature vector measured at fixed intervals (e.g., every 30 seconds) during extraction should be recorded synchronously, and samples should be taken at the corresponding time points for analysis to determine the actual concentrations of rutin, quercetin, kaempferol, and hyperoside. This method forms a complete paired sample set containing timestamps, feature signals, and component contents.
[0046] After obtaining a sufficient number of paired samples, the sample set needs to be stratified to reduce data fluctuations caused by differences in raw materials. Stratification is based on the similarity of raw material properties and process conditions, typically divided according to parameters such as moisture content, particle size, and extraction temperature. Each stratum is considered a sample set under certain stable conditions. Moisture content and particle size affect the solvent permeation rate, while extraction temperature affects the diffusion coefficient of flavonoid molecules; therefore, stratification ensures a stable relationship between spectral signals and chemical concentrations under similar physical conditions. Within each stable stratum, the changing trends of each component in the fused feature vector over time are calculated, and which feature components exhibit consistent changing directions and high correlations across different strata are identified. Such feature components typically correspond to key absorption peaks or conductivity changes during the flavonoid dissolution process. For example, if the absorption signal at approximately 360 nm in the ultraviolet spectral region increases with increasing flavonoid concentration in all strata, the feature component corresponding to this band is assigned to the robust feature set.
[0047] Once the robust feature set is determined, it is necessary to clarify the response relationship between each feature component and the content of each flavonoid component. The response relationship is defined as the correspondence between the numerical changes of a feature component and the changes in component content across multiple offline samples. This step eliminates the influence of individual outlier samples by performing weighted smoothing on multiple points corresponding to the same feature component and its corresponding measured values, ensuring the numerical continuity and monotonicity of the response relationship. Weighted smoothing refers to assigning higher weight to data closer to the center of most samples when calculating the average correspondence between feature values and content values, while reducing the influence of samples that deviate further. For example, in near-infrared spectroscopy, if the corresponding data points for absorption intensity at 1450 nm and rutin concentration are densely distributed, the weighted values of these points are higher than those of edge samples, resulting in a smooth and continuous response curve.
[0048] Based on the above response relationship, the sensitivity of each characteristic component can be calculated. Sensitivity refers to the relative magnitude of the change in component content caused by a unit change in the value of the characteristic within its range. For example, if an increase of 10 μS / cm in conductivity leads to a 3% increase in quercetin concentration, while an increase in near-infrared absorption intensity in the same proportion only causes a 1% change in concentration, then the conductivity characteristic is more sensitive than the spectral characteristic.
[0049] The system determines weighting coefficients based on these differences in sensitivity to highlight the most responsive characteristic signals. For example, in the extraction experiment of a batch of Acanthopanax senticosus raw materials, three main characteristic components were selected: the intensity change of the 1450 nm absorption peak in the near-infrared spectrum, the amplitude of the absorption signal at 360 nm in the ultraviolet-visible spectrum, and the rate of change of conductivity. Analysis of offline test data revealed that when the rutin concentration increased by 10%, the average increase in the 1450 nm absorption peak was 0.08 normalized units, the increase in the 360 nm signal was 0.05 normalized units, and the increase in conductivity was 0.02 normalized units. Based on the proportion of these response amplitudes, it can be preliminarily determined that the 1450 nm signal has the highest sensitivity to rutin, followed by the ultraviolet signal, while the conductivity characteristic is the weakest.
[0050] However, relying solely on response amplitude overlooks feature stability. Therefore, the system also examines the fluctuation range of each feature across multiple batches of samples. For example, in different batches of experiments, the standard deviation of the 1450 nm signal is 0.03, the standard deviation of the 360 nm signal is 0.01, and the standard deviation of the conductivity signal is 0.02. Features with smaller fluctuations represent higher stability and stronger repeatability. The system comprehensively measures both sensitivity and stability, and a simple principle can be established: features with higher sensitivity and smaller fluctuations have greater weight.
[0051] Taking this as an example, the system sets the sensitivity score of the 1450 nm signal to 1.0, the 360 nm signal to 0.6, and the conductivity signal to 0.3. Then, it divides each of these scores by a stability factor (e.g., standard deviation) to obtain the normalized weight ratios. Through this calculation, the weight of the 1450 nm signal is approximately 0.5, the 360 nm signal approximately 0.35, and the conductivity signal approximately 0.15. Finally, when generating the fused feature vector, the system weights and sums the three types of features according to these ratios, making the reflection of spectral changes on the flavonoid extraction process more prominent, while the influence of conductivity remains as an auxiliary correction.
[0052] After determining the weighting coefficients, the components of the fused feature vector are linearly combined according to their weights to form a weighted feature space. In this space, each time point corresponds to a comprehensive feature value, reflecting the comprehensive indicative strength of the multi-source signals on the component content. Based on this weighted feature space, the estimated contents of rutin, quercetin, kaempferol, and hyperoside at each time point can be calculated. The calculation method can be understood as follows: the system utilizes the known response relationships in offline samples, substitutes the weighted combination value of the current fused feature vector into the response curve, and reads the corresponding component content value.
[0053] To ensure the reliability of the prediction results, a difference analysis is needed between the estimated values and the offline measured values. The system calculates the magnitude of the difference between the two values across multiple batches of data and records the range of variation of this difference across different raw material batches. A larger range of variation indicates higher prediction uncertainty; a smaller range of variation indicates higher prediction accuracy. For example, if the difference between the predicted and measured values of rutin content remains within ±4% across 10 different batches of samples, then 4% can be considered the uncertainty boundary for rutin content.
[0054] After obtaining the estimated content and uncertainty of each component, the system combines these data in chronological order to form a dynamically updated flavonoid component estimation structure. This structure is a data mapping table that records the fusion feature vector, estimated content, and uncertainty at each time point. When a new fusion feature vector is input, the system does not need to recalculate all steps; it only needs to substitute it into the established mapping table to output the target flavonoid composition vector, including rutin, quercetin, kaempferol, and hyperoside, in real time, and simultaneously output the uncertainty range corresponding to each component as the confidence interval of the prediction result.
[0055] In this way, the system achieves continuous quantification and reliability assessment of the flavonoid extraction process. The output target flavonoid composition vector not only reflects the dynamic content changes of each component during extraction, but also provides decision-making basis in the form of confidence intervals. This enables the subsequent reinforcement learning control stage to adjust process parameters based on reliable quantitative information rather than single-point estimation, thereby achieving self-optimizing control while maintaining safety constraints.
[0056] Step S103: Construct a safety domain boundary based on the solvent flash point temperature, the upper limit of the equipment's tolerance temperature, and the thermal degradation temperature threshold of the flavonoid components, and calculate the time-varying safety constraint threshold in combination with the uncertainty of the target flavonoid composition vector.
[0057] The purpose of step S103 is to establish a dynamically adjustable safety constraint mechanism so that reinforcement learning optimization will never generate process instructions that exceed the equipment's tolerance or cause thermal degradation of flavonoid components. To achieve this, it is necessary to first define the physical boundary conditions of the safety domain, and then construct a time-varying safety constraint threshold based on the uncertainty of the target flavonoid composition vector. This step combines engineering safety with data uncertainty, ensuring that optimization control remains within an operable safety range while guaranteeing extraction quality.
[0058] In the extraction of flavonoids from Acanthopanax senticosus, the solvent typically used is a mixture of ethanol and water. The flash point of ethanol is the lowest temperature at which ethanol vapor can ignite instantly upon contact with an ignition source in air; different concentrations of ethanol have different flash points. For example, a 70% (v / v) ethanol solution has a flash point of approximately 23 degrees Celsius, while a 90% (v / v) solution has a flash point of approximately 16 degrees Celsius. The flash point temperature determines the lower limit of fire safety for the equipment during heating; therefore, when determining the safety range, the operating temperature should always be kept below the flash point of the solvent. The upper limit of the equipment's operating temperature refers to the highest safe operating temperature that the heating vessel, circulating pump, and seals of the extraction device can withstand during long-term operation. For example, if the equipment manufacturer provides a temperature limit of 120 degrees Celsius, exceeding this value will lead to seal aging or pipeline cracking. The thermal degradation threshold of flavonoids refers to the temperature at which flavonoid compounds begin to undergo chemical decomposition or isomerization reactions under heating conditions. For example, experiments have shown that rutin decomposes when heated at temperatures above 95 degrees Celsius for 30 minutes, and quercetin begins to deactivate at 100 degrees Celsius. Therefore, the temperature should be controlled to ensure that it does not exceed these thresholds.
[0059] To make these physical parameters calculable in the algorithm, their reference values must first be determined. Solvent flash points can be obtained by consulting the safety data sheet or through experimental measurement; the upper limit of the equipment's operating temperature is provided in the equipment manual; and the flavonoid thermal degradation threshold can be obtained experimentally or from literature. After converting these values to the same temperature unit, they are stored in the computer as constants or dynamic variables.
[0060] When the extraction system is running, the reinforcement learning module needs to confirm whether it is within the safety region before each attempt to adjust the temperature, solvent ratio, or circulation flow rate. To avoid overly rigid safety thresholds that could reduce the efficiency of the optimization algorithm, this invention introduces the concept of time-varying safety constraints, which dynamically adjust the width of the safety boundary based on the uncertainty of the real-time prediction results. Higher uncertainty indicates lower reliability of the model prediction, and the system should maintain a more conservative operating range; lower uncertainty allows for a more relaxed boundary to improve optimization efficiency.
[0061] Specifically, the uncertainty of the target flavonoid composition vector can be regarded as an indicator reflecting the reliability of the current system state. If the average uncertainty of the four components output by the virtual metric is high, for example, exceeding 0.05%, it indicates that the current system state may be affected by external disturbances or the data may deviate from the training range. In this case, a buffer should be added at the safety boundary. For example, when the normal safety upper limit is 95 degrees Celsius, it can be automatically lowered to 90 degrees Celsius as the dynamic upper limit; if the uncertainty is extremely low (such as below 0.01%), the original safety upper limit can be maintained unchanged.
[0062] In practical implementation, time-varying constraint values can be calculated as follows: First, read the average value of the target flavonoid composition vector uncertainty under the current process conditions, denoted as U. Then, compare U with a preset uncertainty threshold to calculate the safety margin adjustment coefficient. If U is large, the safety margin increases, i.e., the safety interval shrinks; if U is small, the safety interval remains unchanged or is moderately widened. For example, set the safety interval to decrease by 5% when U=0.05% and by 10% when U=0.1%. This proportional adjustment method can achieve dynamic safety boundaries in the control logic without complex calculations.
[0063] In addition to temperature, the ethanol volume fraction and circulation flow rate also need to be subject to safety constraints. Excessive ethanol concentration will lower the solution's flash point and increase the risk of fire; therefore, the system should simultaneously check whether the solvent concentration matches the current temperature before each temperature adjustment. When an increase in the ethanol volume fraction is detected, the upper temperature limit should be automatically lowered accordingly. For example, if the ethanol concentration increases from 70% to 90%, the system can proportionally lower the maximum allowable temperature from 95 degrees Celsius to 85 degrees Celsius to maintain an equivalent safety margin. Safety constraints on the circulation flow rate primarily prevent pump overload or boiling shocks. The upper limit is usually determined by the pump's rated flow rate; for example, if the pump's rated flow rate is 3 liters / minute, the system will automatically limit it to within 2.8 liters / minute if the flow rate parameter output by reinforcement learning exceeds this value.
[0064] To avoid unnecessary fluctuations caused by frequent triggering of safety protection by the algorithm, this invention sets up a buffer zone for the safety domain, that is, extending a small warning zone outside the safety boundary. When the optimization result approaches the warning zone, the system first issues a warning signal instead of immediately interrupting the operation. Only when the predicted action exceeds the safety domain boundary does the controller forcibly revert to the safety parameters. This design ensures that the system has both safety protection capabilities and does not lose its self-optimization flexibility due to overly strict safety constraints.
[0065] Through the above design, the definition of the safety domain boundary and the calculation of the time-varying safety constraint threshold can be automated, interpretable, and consistent with the actual operating rules of industrial equipment. Solvent flash point, equipment tolerance temperature, and flavonoid thermal degradation threshold can be directly set based on equipment technical parameters and experimental data. By introducing an uncertainty adjustment coefficient through program logic, the safety domain calculation process in this invention can be accurately implemented. In this way, the system can automatically update the safety range based on real-time data during operation, ensuring that reinforcement learning decisions are always made within the safety range, thereby effectively preventing risks of process overruns, equipment damage, and product quality degradation.
[0066] Furthermore, the step of constructing a safety domain boundary based on the solvent flash point temperature, the upper limit of the equipment's tolerance temperature, and the thermal degradation temperature threshold of the flavonoid components, and calculating a time-varying safety constraint threshold in conjunction with the uncertainty of the target flavonoid composition vector, includes: The flash point temperature of ethanol solvent, the upper limit of the tolerance temperature of the extraction equipment, and the thermal degradation temperature thresholds of rutin, quercetin, kaempferol and hyperoside are respectively input into the safety parameter set, and the minimum value of the three is used as the initial upper limit of the safety temperature. At the same time, the volume fraction of ethanol and the allowable pressure of the equipment are set as the physical boundary parameters of the safety domain. The temperature, pressure and ethanol volume fraction in the extraction system are monitored in real time. When any parameter approaches the upper limit of the initial safe temperature or exceeds the allowable fluctuation range, the deviation ratio between the three types of temperature parameters is calculated and the deviation ratio is output as a safety margin index. The safety margin index is dynamically corrected by utilizing the uncertainty of the target flavonoid composition vector. When the uncertainty increases, the corresponding safety margin is reduced proportionally to reflect the safety contraction requirement caused by the decrease in the reliability of component prediction, thereby obtaining the corrected dynamic safety margin. The initial safety temperature limit, safety margin index, and corrected dynamic safety margin are logically combined to generate a time-varying safety constraint threshold. The time-varying safety constraint threshold defines the maximum allowable process temperature, the upper limit of ethanol volume fraction, and the circulation flow limit during the extraction process, and changes in real time with time and estimated uncertainty. It is used to limit the range of executable parameters in reinforcement learning optimization decision-making and prevent operating conditions from exceeding the safety domain boundary.
[0067] In the extraction of flavonoids from Acanthopanax senticosus, to prevent safety risks such as flash evaporation, overheating, or thermal degradation of components under continuous heating and circulating solvent conditions, it is necessary to establish a safety domain boundary that can be dynamically adjusted according to process changes. This safety domain boundary is jointly defined by three types of physical parameters: solvent flash point temperature, upper limit of equipment tolerance temperature, and thermal degradation temperature threshold of flavonoid components. It is further corrected in real time by incorporating the uncertainty of the target flavonoid composition vector, thereby generating a time-varying safety constraint threshold. The specific implementation method of this step is as follows.
[0068] First, based on experimental data or equipment technical parameters, the flash point temperature of the ethanol solvent, the upper limit of the extraction equipment's tolerance temperature, and the thermal degradation temperature thresholds of rutin, quercetin, kaempferol, and hyperoside should be obtained. The flash point temperature of the ethanol solvent refers to the lowest liquid temperature at which ethanol vapor forms a flammable mixture with air under the current system pressure and solvent concentration conditions. Since the extraction system is often in a semi-closed or pressurized state, the flash point value needs to be corrected according to the partial pressure of ethanol vapor; the correction factor can be determined based on the system pressure and the volume fraction of ethanol. The upper limit of the extraction equipment's tolerance temperature is derived from the manufacturer's catalog or type test report, reflecting the highest operating temperature that the tank, pipelines, and sealing components can safely withstand under continuous operation. The thermal degradation temperature thresholds of flavonoid components are obtained through differential scanning calorimetry or thermogravimetric analysis, used to indicate the initial temperature at which each component undergoes structural damage or property changes upon heating. These three temperature parameters are unified to the same unit and stored in a safety parameter set; the minimum value among the three is taken as the initial safe upper limit temperature. Meanwhile, the upper limit of ethanol volume fraction and the allowable pressure of the equipment are used as other boundary conditions of the safety domain to ensure that heat, safety, solvent concentration and pressure are controlled in a multi-dimensional manner.
[0069] During the extraction process, the system continuously monitors temperature, pressure, and ethanol volume fraction, recording these parameters over time. When any parameter approaches the initial safe temperature limit or exceeds the allowable fluctuation range, the system calculates the deviation ratio. The deviation ratio quantifies the distance between the current parameter and the boundary value, defined as the proportion of the remaining space to the total safety bandwidth. For example, if the safe temperature limit is 96 degrees Celsius, the minimum process temperature is 80 degrees Celsius, and the total safety bandwidth is 16 degrees Celsius, then at the current temperature of 88 degrees Celsius, there are still 8 degrees Celsius away from the limit, resulting in a safety margin of 50%. For the ethanol volume fraction, if the limit is 80%, the minimum design value is 70%, and the current volume fraction is 78%, then the safety margin is 20%. For the system pressure, if the limit is 1 MPa, the inertial operating pressure is 0.4 MPa, and the current pressure is 0.7 MPa, then the safety margin is 50%. The system takes the minimum of these three values as the comprehensive safety margin index, ensuring that the overall safety assessment is constrained by the most unfavorable factor.
[0070] To ensure that the safety margin boundary reflects the reliability of component estimation, the system incorporates the uncertainty of the target flavonoid composition vector into the safety margin calculation. Uncertainty represents the confidence range of the virtual metric's prediction of the concentration of each flavonoid component at the current moment, expressed as a value between zero and one, where a higher value indicates lower confidence. When uncertainty increases, the system proportionally reduces the safety margin to automatically increase the safety distance when estimation accuracy decreases. For example, when the overall safety margin is 20%, and the overall uncertainty is 0.3%, the system reduces the safety margin by 30%, resulting in a corrected dynamic safety margin of 14%. In this way, the operating intensity can be reduced in advance when estimation confidence is insufficient, avoiding over-limit operation caused by inaccurate component change predictions.
[0071] Once the initial safety temperature limit, safety margin index, and corrected dynamic safety margin are determined, the system logically combines these three to form a time-varying safety constraint threshold. The time-varying safety constraint threshold defines the maximum permissible process temperature, the upper limit of ethanol volume fraction, and the circulation flow rate limit during the extraction process, and dynamically adjusts with time and changes in estimated uncertainty. The calculation method can be described as follows: First, multiply the total safety bandwidth of each dimension by the corrected dynamic safety margin to obtain the dynamically available bandwidth. Then, this bandwidth is rolled back from its respective upper limit. Taking temperature as an example, if the total bandwidth is 16 degrees Celsius and the corrected dynamic safety margin is 30%, the dynamically available bandwidth is approximately 5 degrees Celsius. Subtracting 5 degrees Celsius from the 96-degree Celsius safety temperature limit yields approximately 91 degrees Celsius, which is taken as the current maximum permissible temperature. The ethanol volume fraction dimension is handled in the same way; if the total bandwidth is 10% and the corrected dynamic safety margin is 30%, the current permissible upper limit is 77%. For the circulation flow rate, the total bandwidth can be determined based on the difference between the rated flow rate and the minimum stable flow rate, and then rolled back proportionally according to the dynamic safety margin to obtain the current executable upper limit. This method enables all three types of boundary parameters to shrink synchronously based on real-time changes and measurement uncertainties, forming a continuously updatable time-varying safety domain.
[0072] During the reinforcement learning control phase, the system uses time-varying safety constraint thresholds as optimization boundaries to limit the parameter adjustment suggestions generated by reinforcement learning. When the temperature, ethanol volume fraction, or flow rate output by reinforcement learning exceeds the corresponding constraint, the system automatically truncates the value at the threshold boundary and records a safety violation event. If no violation occurs within multiple consecutive cycles, the system can appropriately relax the correction ratio, gradually restoring the safety margin to the normal range. In this way, the system achieves an adaptive balance between process safety and optimized performance, ensuring that the extraction process does not suffer energy efficiency losses due to excessive conservatism, nor does it lead to thermal runaway or solvent overruns due to estimation errors.
[0073] Through the above steps, the solvent flash point temperature, the upper limit of the equipment's tolerance temperature, and the thermal degradation temperature threshold of the flavonoid component are uniformly incorporated into the dynamic safety parameter system. The system combines real-time process variables and component uncertainties to proportionally adjust the safety margin, forming a time-varying safety constraint threshold that changes with time and confidence level.
[0074] Step S104: Using a reinforcement learning strategy, with the target flavonoid composition vector and time-varying safety constraint threshold as input, the temperature, ethanol volume fraction, solid-liquid ratio, soaking time and circulation flow rate are optimized in a linked manner. The reward function of the reinforcement learning strategy consists of the weighted distance between the target flavonoid composition vector and the standard flavonoid composition vector and a solvent consumption energy penalty term. When the predicted value of the optimized action touches the boundary of the safety domain, it falls back to the safe process parameters.
[0075] The core of step S104 lies in constructing a self-learning and self-correcting reinforcement learning decision-making process. This process simultaneously optimizes multiple process variables such as temperature, ethanol volume fraction, solid-liquid ratio, soaking time, and circulation flow rate during extraction. This ensures that, while maintaining safety constraints, the extraction ratio of the main components in Acanthopanax senticosus flavonoids approaches the set ideal target. This step not only achieves multi-parameter coordinated adjustment, which is difficult to achieve with traditional control systems, but also, through the dynamic design of the reward function, balances the optimization objective between quality and energy consumption, ultimately forming an intelligent closed-loop control system that can operate stably for a long period.
[0076] Before proceeding with reinforcement learning optimization, it is necessary to clarify how reinforcement learning algorithms are applied in this invention. Reinforcement learning is a decision-making algorithm based on environment interaction. Through continuous trial, feedback, and correction, it enables the system to learn to take the optimal action under different states. In this step, the environment corresponds to the entire extraction process, including the dynamic changes in equipment status, solution physical properties, and flavonoid content; the state is the system information described by the target flavonoid composition vector output by the virtual metric at each time step and its uncertainty; the action is the combination of process parameters generated by the reinforcement learning algorithm for the current state, namely, the adjustment values of temperature, ethanol volume fraction, solid-liquid ratio, soaking time, and circulation flow rate; the reward is the feedback signal calculated by the system based on the extraction results after executing the action, used to evaluate whether the action promotes the system towards the optimal direction.
[0077] To enable the algorithm to learn a stable policy within a finite time, the scope of the state space and action space needs to be defined first. The state space should cover all observable system features, such as the relative contents of rutin, quercetin, kaempferol, and hyperoside output by the virtual metric, as well as the time-varying safety constraint threshold calculated in step S103. Each state vector contains comprehensive information on flavonoid extraction quality and safety margin. For example, at a certain moment, if the system detects a relative rutin content of 0.75%, quercetin of 0.50%, kaempferol of 0.18%, and hyperoside of 0.32%, and the upper limit of the safe temperature is 92 degrees Celsius, the algorithm will input this information as the current state. The action space is defined as five adjustable control variables, each with an upper limit, a lower limit, and a step size. For example, the temperature range can be set to 70–95 degrees Celsius in 1-degree Celsius increments; the ethanol volume fraction range can be 60%–90% in 2% increments; the solid-liquid ratio range can be 1:8–1:12; the soaking time can be adjusted between 15–45 minutes; and the circulation flow rate can be varied between 1–3 liters per minute. By discretizing these parameters, the reinforcement learning algorithm can explore within a limited combination of actions.
[0078] During algorithm execution, each time step represents a new decision-making cycle. After obtaining the current state from the virtual metric, the algorithm generates an action instruction containing five variables and sends it to the extraction device for execution. After the device executes the instruction, the new state of the extract is fed back to the system via spectroscopy and sensors, and the virtual metric re-predicts the current flavonoid composition vector and its uncertainty. The system compares this new state with the previous target state and calculates a reward value based on the flavonoid component deviation and energy consumption. A higher reward value indicates that the action brings the system closer to the ideal target; a lower or negative reward value indicates that the action deviates from the target or causes an increase in energy consumption.
[0079] The design of the reward function directly determines the direction of algorithm optimization. The reward function of this invention consists of two parts: the first part is the weighted distance between the target flavonoid composition vector and the standard flavonoid composition vector, used to measure extraction quality deviation; the second part is a solvent consumption and energy consumption penalty term, used to limit resource waste during the optimization process. The weighted distance is the difference between the system's current output and the ideal component ratio; the greater the difference, the lower the reward. The introduction of solvent consumption and energy consumption penalty terms is to prevent the algorithm from excessively increasing the temperature or extending the soaking time when pursuing a high extraction rate.
[0080] In the reward function of reinforcement learning, the first step is to calculate the weighted distance between the target flavonoid composition vector and the standard flavonoid composition vector. The basic idea is to compare the predicted content of each major flavonoid component with its standard content, obtaining the difference. Then, different weights are assigned to each component based on its importance in overall efficacy or quality control. Finally, these weighted differences are summed into a total deviation. Specifically, the system sequentially extracts the differences between the predicted and standard values of rutin, quercetin, kaempferol, and hyperoside. For example, if the predicted value of rutin is lower than the standard value by a certain percentage, while quercetin is slightly higher, the system calculates the absolute value of these deviations separately. Then, the system weights these deviations according to pre-set weights. For example, higher weights are assigned to rutin and quercetin, which contribute significantly to efficacy, while lower weights are assigned to minor components like kaempferol and hyperoside. This way, a small deviation in rutin will have a greater impact on the overall deviation than other components. All weighted deviation values are accumulated on the same scale to obtain a single index representing the "degree of difference between the current extraction result and the overall ideal target." The system then determines the reward value based on the magnitude of this indicator. The smaller the deviation, the closer the target component ratio is to the standard ratio, and the higher the reward value; the larger the deviation, the lower the reward value. In this way, the algorithm can obtain continuous and comparable quality feedback in each iteration, thereby guiding the selection of the next action and gradually bringing the flavonoid extraction ratio closer to the standard target.
[0081] The calculation of solvent consumption and energy consumption penalties is another balancing mechanism to prevent reinforcement learning algorithms from solely pursuing extraction efficiency while neglecting economy and safety. In this part, the system monitors process parameters directly related to resource usage, including ethanol volume fraction, circulation flow rate, heating temperature, and soaking time. Each parameter has a design reference range, representing the reasonable range of solvent usage and energy input while ensuring extraction effectiveness. The standard process range is assumed to be: ethanol volume fraction 70% to 80%, temperature 85°C to 90°C, circulation flow rate 1.5 to 2.5 liters per minute, and soaking time 30 minutes. Within this range, the system considers it a "normal economic zone" for energy consumption and solvent use, and no penalty is triggered. When the action parameters generated by the reinforcement learning algorithm exceed this range, the system immediately calculates the deviation and converts it into a corresponding penalty value.
[0082] For example, in a batch extraction, a reinforcement learning algorithm attempts to increase the ethanol volume fraction to 85% to enhance solubility. This value is 5 percentage points higher than the upper limit. The system calculates the deviation coefficient of the excess percentage relative to the upper limit and multiplies this deviation by a preset solvent consumption weight. The solvent consumption weight is an empirical parameter determined by both equipment safety and ethanol cost; in this example, it can be understood as "how much the system's energy consumption and cost will increase for every percentage point increase in ethanol concentration." If this weight is set to a 2% consumption penalty for each percentage point exceeding the limit, exceeding it by 5 percentage points will result in a relative penalty of approximately 10%. The system adds this penalty value to the reward function of the current decision, causing the algorithm to tend to reduce the ethanol concentration in the next cycle to conserve solvent and restore a safety margin.
[0083] Similarly, if a reinforcement learning algorithm decides to raise the temperature to 100°C, which is 10°C higher than the set upper limit of 90°C, the system will calculate the energy consumption deviation based on the actual change in heating energy consumption. Heating energy consumption is usually related to the temperature difference and heating time, so the system will estimate the additional energy expenditure based on the temperature difference exceeding the limit and the length of time that temperature is maintained. In practice, this calculation can be achieved through energy metering devices or energy balance models. For example, when the temperature rises by 10°C, if the system detects that the power consumption of the electric heater increases by about 15% in the same time period, this excess will be equivalently converted into an energy consumption penalty. The weight of the energy consumption penalty is determined based on the rated power of the equipment and energy costs. If it is set that every 1% increase in energy consumption leads to a 0.5% reduction in reward, then a 15% increase in energy consumption will reduce the reward value by about 7.5%. This mechanism ensures that the algorithm does not blindly increase the temperature to pursue the extraction rate, but automatically finds a balance between the increase in high-temperature energy consumption and the improvement in extraction efficiency.
[0084] Regarding circulation flow rate, the system continuously monitors the pump's actual load and fluid resistance. When the flow rate parameter output by reinforcement learning exceeds 2.5 liters per minute, the pump's power demand increases, and fluid turbulence causes some solvent evaporation loss. The system measures changes in pump current or power and compares them to the rated load. If a 10% increase in current is detected, the system considers the additional energy consumption caused by the excessive circulation flow rate to be equivalent to a 10% penalty. When the circulation flow rate remains excessively high for an extended period, the system records this trend and gradually accumulates the penalty value, thereby guiding reinforcement learning to reduce the flow rate in subsequent iterations and restore the pump to operate in its high-efficiency range.
[0085] Meanwhile, variations in soaking time also affect energy consumption. When the algorithm chooses to extend the soaking time to 45 minutes, compared to the standard 30 minutes, the system will include the additional 15 minutes of heating maintenance cost as a penalty. If the average power during the heating phase is 3 kW, then the additional 15 minutes corresponds to approximately 0.75 kWh of energy consumption, and the system will convert this energy difference into a penalty ratio. For example, if 1 kWh of energy is set as a 5% reward / penalty unit, then this operation will introduce an additional penalty value of approximately 3.75%. In this way, the reinforcement learning system will gradually learn to trade off increased extraction efficiency from increased energy consumption, automatically converging to the optimal soaking time.
[0086] Ultimately, the system will aggregate all the aforementioned penalty items into a comprehensive energy consumption and solvent consumption penalty index. This index, along with the extraction quality deviation index, is used to calculate the total reward value. If a decision maintains a flavonoid ratio close to the standard composition while the ethanol volume fraction is only slightly above the upper limit, and the temperature and flow rate remain within a reasonable range, the comprehensive penalty is small, and the reward value remains high. Conversely, if a decision significantly increases the ethanol concentration, temperature, and soaking time in a short period, the comprehensive penalty will increase dramatically, causing the reward value to decrease. This allows the algorithm to automatically avoid such high-energy-consuming schemes in subsequent learning.
[0087] Within the same decision-making cycle, the system accumulates various penalty values to form a comprehensive indicator representing the degree of resource waste. This indicator is then combined with the aforementioned quality deviation indicator to calculate the total reward value. If the current operation ensures that the extraction ratio is close to the standard while the ethanol concentration and temperature are within the energy-saving range, the total penalty is close to zero, and the reward value will be relatively increased. If the extraction ratio is ideal but consumption is too high, the penalty will reduce the reward, causing the algorithm to automatically lower the temperature or shorten the soaking time in the next round to restore energy efficiency balance.
[0088] Reinforcement learning algorithms accumulate experience through numerous interactions during the training phase. Each time a process adjustment is performed, the current state, the action taken, the reward obtained, and the new state are recorded and stored in an experience pool. The algorithm learns by continuously drawing samples from the experience pool, updating its internal policy parameters. The policy update process can employ modern reinforcement learning methods such as deep Q-networks or proximal policy optimization. These methods can handle continuous state and action spaces, enabling the algorithm to converge stably even in complex process environments.
[0089] To ensure the system remains within a safe operating range, the algorithm calls the time-varying safety constraint threshold calculated in step S103 before each decision. When a predicted action is about to exceed the safety domain boundary, the algorithm automatically triggers a protection mechanism, replacing the current action with the most recently verified safe combination of process parameters. For example, if the algorithm generates a temperature of 96 degrees Celsius, while the safety upper limit is 95 degrees Celsius, the system will immediately adjust the temperature to 94 degrees Celsius to prevent equipment overheating or flavonoid degradation. This "fallback to safe process parameters" mechanism remains effective throughout the optimization process, ensuring that the reinforcement learning's exploratory behavior is physically constrained and preventing actual production safety from being compromised by algorithmic misjudgments.
[0090] During long-term operation, reinforcement learning algorithms can gradually approach optimal process conditions by continuously updating reward values and policy functions. Initially, the system may try different parameter combinations with a large exploration range to accumulate experience; once the learning stabilizes, the algorithm gradually reduces the exploration frequency and tends to use existing experience to execute near-optimal decisions. For example, after running ten batches consecutively, the system can learn that a certain combination (such as a temperature of 88 degrees Celsius, an ethanol volume fraction of 75%, a solid-liquid ratio of 1:10, a soaking time of 30 minutes, and a circulation flow rate of 2 liters per minute) can stably output a flavonoid ratio close to the standard composition. At this point, the algorithm will automatically tend to operate within that parameter range.
[0091] For ease of implementation, the reinforcement learning module can be deployed on an industrial control computer or edge server, employing a periodic data update mechanism to acquire state input and generate optimization instructions every few minutes. If an abnormal data communication is detected or the input signal exceeds the physical range, the system automatically pauses the reinforcement learning decision-making process and retains the previous safety parameters, resuming optimization only after the signal recovers. This design ensures that the algorithm maintains stable decision outputs even in real-world production environments with noise, latency, or equipment fluctuations.
[0092] In summary, step S104, by introducing a reinforcement learning strategy, enables the extraction system to possess self-learning, self-optimization, and self-protection functions. It can autonomously decide on the optimal process operation parameters based on the real-time measured target flavonoid composition vector and dynamic safety constraint thresholds, achieving comprehensive optimal control of extraction efficiency, energy consumption, and safety.
[0093] Furthermore, the reinforcement learning strategy employs a target flavonoid composition vector and a time-varying safety constraint threshold as inputs to perform linked optimization of temperature, ethanol volume fraction, solid-liquid ratio, soaking time, and circulation flow rate. The reward function of the reinforcement learning strategy consists of a weighted distance between the target flavonoid composition vector and the standard flavonoid composition vector, plus a solvent consumption energy penalty term. When the predicted value of the optimized action touches the safety domain boundary, it reverts to safe process parameters, including: At the beginning of each extraction cycle, the target flavonoid composition vector and its uncertainty output by the virtual metric are received, and the numerical range of the current time-varying safety constraint threshold is read from the safety constraint module. Both are used as the initial input for optimization calculation to determine the parameter space of the current extraction state. Based on the weighted distance between the target flavonoid composition vector and the standard flavonoid composition vector, the target deviation value for the current period is calculated, and the target deviation value of the previous period is compared with the current value to obtain the changing trend of the flavonoid extraction direction. After obtaining the trend of change, the influence of temperature, ethanol volume fraction, solid-liquid ratio, soaking time and circulation flow rate on the direction of flavonoid extraction is calculated. The influence is then distributed to the adjustable range of each parameter in a relative proportion to generate a set of initial parameter adjustment schemes. For each set of initial parameter adjustment schemes, the execution results in the extraction equipment are monitored in real time, the corresponding solvent consumption and energy consumption are recorded, and the reward value is calculated based on the degree of improvement of the target flavonoid composition vector. The reward value is obtained by subtracting the penalty amount of energy consumption and solvent consumption from the weighted distance of the target flavonoid composition vector. The parameter adjustment scheme with the highest reward value is used as the prediction result of the current optimization action. When the prediction result touches the boundary of the time-varying safety constraint threshold, a safety rollback operation is performed to roll back the corresponding parameter to the safety range according to the proportion of the remaining safety margin. Then, the rolled-back parameter is used as the initial state input for the new cycle to form a continuous self-optimization loop.
[0094] In the reinforcement learning self-optimization process of extracting flavonoids from Acanthopanax senticosus, the implementation logic of the entire step S105 must be based on executable industrial control, so that the system can not only automatically optimize parameters, but also ensure stability, safety and accuracy in continuous production.
[0095] At the beginning of each extraction cycle, the system first receives the target flavonoid composition vector output by the virtual measurement module through the data acquisition module. This vector represents the relative content of the four main flavonoid components—rutin, quercetin, kaempferol, and hyperoside—in the current feed solution. The virtual measurement module is trained by fusing feature vectors and offline laboratory data, enabling it to predict the proportion of each component in real time based on spectral and sensor signals without requiring system shutdown for sampling. Simultaneously, the virtual measurement module also outputs the uncertainty value for each component to quantify the reliability of the prediction results. When environmental disturbances are significant, raw material batch differences are obvious, or spectral signal fluctuations are abnormal, the uncertainty value will increase accordingly, indicating a decrease in the reliability of the prediction. During this stage, the system also reads various parameters of the time-varying safety constraint thresholds from the safety constraint module, including the currently allowed maximum process temperature, the upper limit of ethanol volume fraction, the solid-liquid ratio variation range, and the maximum circulation flow rate. These parameters collectively define the safe operating space of the extraction equipment at the current moment. The system uses the target flavonoid composition vector, uncertainty, and time-varying safety constraint thresholds as inputs for optimization calculations to form a complete extraction state description, used to determine the adjustable parameter range.
[0096] After initial input, the system calculates the target deviation value for the current period based on the weighted distance between the target flavonoid composition vector and the standard flavonoid composition vector. The standard flavonoid composition vector is the ideal component ratio determined by long-term production statistics or pharmacodynamic experiments, reflecting the target direction of product quality control. The weighted distance calculation follows the principle of "key components first," meaning rutin and quercetin have higher weights, while kaempferol and hyperoside have relatively lower weights. The deviation value can be understood as the comprehensive difference between the system and the ideal state. By comparing the current period's deviation with the previous period's deviation, the changing trend of flavonoid extraction direction can be obtained. For example, when the deviation continues to decrease, it indicates that the system is converging towards the target ratio; if the deviation increases instead, it indicates that a certain process variable has deviated from the optimal range and needs to be adjusted in time. Through this trend judgment, the system can identify whether the current process conditions promote flavonoid extraction or lead to component loss.
[0097] After determining the extraction direction trend, the system further analyzes the influence of various process parameters (temperature, ethanol volume fraction, solid-liquid ratio, soaking time, and circulation flow rate) on flavonoid extraction. The determination of the influence magnitude relies on historical periodic disturbance response data, i.e., the change in the flavonoid composition vector after minor adjustments to each parameter in past periods. For example, if a 2°C increase in temperature significantly increases the relative content of rutin and quercetin while energy consumption increases only slightly, it indicates that the temperature parameter has a high positive influence under the current raw material conditions. Conversely, if an increase in ethanol volume fraction increases flavonoid dissolution rate but leads to a significant increase in energy consumption and solvent usage, the system will determine its influence magnitude to be low. The system normalizes the influence magnitudes of the five parameters and proportionally distributes them into adjustable ranges, forming a set of initial parameter adjustment schemes. During this process, the adjustment magnitudes of temperature, ethanol volume fraction, and circulation flow rate are typically limited by safety constraint thresholds, while the adjustments to the solid-liquid ratio and soaking time are relatively slow to maintain system equilibrium.
[0098] Once the initial parameter adjustment plan is generated, the system executes it step by step in the equipment and monitors the results in real time. Temperature changes are controlled by the heating control unit, ethanol volume fraction is adjusted by the solvent replenishment valve, the solid-liquid ratio is controlled by the feed and discharge flow rates, the circulation flow rate is controlled by the variable frequency pump speed, and the soaking time is timed by the control program. After each parameter adjustment, the system immediately collects energy consumption data (such as power meter output), solvent usage (accumulated by the flow meter), and real-time prediction results of flavonoid components, and records them in the data storage unit. At this time, the system calculates the reward value based on the extraction effect and resource consumption. The reward value is defined based on two main parts: first, the degree of convergence of the target flavonoid composition vector to the standard composition vector, i.e., the reduction in weighted distance; and second, the penalty terms for solvent and energy consumption, i.e., the resource cost of improving the extraction effect. For example, when the system increases the flavonoid extraction rate by 5% by increasing the temperature and circulation flow rate, but increases energy consumption by 10% and ethanol consumption by 8%, the system calculates the net reward value after adjusting according to preset weights. If the improvement benefit is greater than the consumption cost, the plan is determined to be a positive and effective plan.
[0099] After all schemes have been executed, the system compares the reward values of each scheme and selects the set of parameters with the highest reward value as the prediction result for the current optimization action. If the value of any parameter in the prediction result touches the boundary of the time-varying safety constraint threshold, such as the temperature reaching the safety upper limit of 90 degrees Celsius or the ethanol volume fraction reaching 80%, the system will immediately trigger a safety rollback operation. The logic for safety rollback is based on the proportion of the remaining safety margin, that is, the difference between the safety upper limit and the current value. For example, when the temperature is only 1 degree Celsius away from the safety upper limit, the system calculates the rollback amount inversely proportional to the safety margin, reducing the temperature to near the center value of the safety range, thereby avoiding the safety risks caused by thermal degradation of flavonoid components or solvent boiling. After the rollback operation, the system rewrites the updated parameter values into the initial input of the next extraction cycle, making the reinforcement learning process a closed loop.
[0100] The core of this self-optimizing cycle lies in continuous learning and correction. In each extraction cycle, the system uses the optimization results and safety feedback information from the previous cycle to revise the parameter adjustment strategy for the next cycle, thereby gradually bringing the extraction process closer to the optimal range. As extraction batches accumulate, the system can automatically identify the impact of different raw material moisture content, particle size, and seasonal differences on the process, ultimately establishing a dynamically stable extraction strategy and achieving the automated control effect of "same quality across different batches".
[0101] Through this series of operations, the present invention achieves intelligent and adaptive processing of flavonoid extraction from Acanthopanax senticosus. While maintaining high extraction rates and high purity, the system achieves an optimal balance between energy consumption and safety, making equipment operation more economical and reliable.
[0102] Step S105: Output the temperature, ethanol volume fraction, solid-liquid ratio, soaking time, and circulation flow rate obtained from the linkage optimization to the extraction equipment for execution.
[0103] Step S105 is the final execution stage of the entire reinforcement learning self-optimization process. Its task is to accurately transmit the combination of process parameters generated by intelligent decision-making to the extraction equipment, enabling the system to automatically adjust its operating state based on real-time optimization results, thereby achieving true closed-loop control. This step involves not only the formatting of parameter output and signal transmission, but also the parameter execution method, response feedback mechanism, and stability assurance measures. Only when these details are fully implemented can the optimization results of the reinforcement learning model be translated into process actions at the physical level, enabling the extraction process of Acanthopanax senticosus flavonoids to achieve continuous, autonomous, and verifiable optimization and adjustment.
[0104] After the reinforcement learning module obtains new process parameters through multiple rounds of calculations, these parameters include specific values for five dimensions: temperature, ethanol volume fraction, solid-liquid ratio, soaking time, and circulation flow rate. The system first converts these values from the internal representation of the calculation module into command signals that the equipment can recognize. Typically, industrial control systems use standardized data communication protocols such as Modbus, OPC UA, or EtherCAT. These protocols ensure the real-time performance and reliability of control signals. For example, when the reinforcement learning output temperature is 88 degrees Celsius, the system sends this target value to the heating control unit in the form of a digital command. Upon receiving the command, the temperature controller in the heating control unit adjusts the heating tube power or the circulation medium flow rate to gradually bring the reactor temperature closer to the set point of 88 degrees Celsius.
[0105] During parameter issuance, the system needs to ensure the synchronous execution of each variable. Temperature and ethanol volume fraction are continuous variables, while the solid-liquid ratio, soaking time, and circulation flow rate simultaneously affect the dynamic distribution of materials. Therefore, control commands must be time-coordinated. In practice, the control system sets a unified time base, such as an execution cycle of 1 second or 5 seconds. At the beginning of each cycle, all control values are refreshed simultaneously to prevent inconsistencies in extraction conditions due to delays. For example, when reinforcement learning determines to adjust the ethanol volume fraction from 75% to 78%, the system will synchronously adjust the opening ratio of the solvent feeding valve and execute it gradually while the temperature is still within the target range to prevent rapid changes in temperature and solvent concentration from causing sudden boiling or foaming.
[0106] Adjusting the solid-liquid ratio is typically achieved by controlling the on / off timing of the feed pump and drain valve. The system calculates the ratio of the current liquid volume to the solid feed material based on the target solid-liquid ratio value generated by reinforcement learning. When the ratio deviates from the target, the system automatically opens the feed pump or drain valve to compensate for the error. For example, if there is excessive residual solvent after the previous extraction, resulting in a solid-liquid ratio of 1:9, while the reinforcement learning output target is 1:10, the system automatically replenishes an appropriate amount of solvent until the ratio reaches the target. The soaking time is controlled by a timing module. When the system enters a new extraction cycle, the soaking timer is reset, and the next stage of the circulating flow operation is triggered after a specified time (e.g., 32 minutes).
[0107] Circulating flow control is typically achieved using a variable frequency pump or flow regulating valve. The system transmits the target flow value output from reinforcement learning to the flow control module. This module reads the real-time feedback signal from the flow meter and forms a closed-loop control by adjusting the pump speed or valve opening to stabilize the actual flow rate near the target value. For example, when the target flow rate is 2.2 liters per minute, the flow control module continuously compares the current flow rate with the target flow rate. When the deviation exceeds a set range (e.g., 0.1 liters per minute), it automatically adjusts the pump speed until the deviation is eliminated. In this way, the output of the reinforcement learning module is precisely implemented in the physical system through an industrial control loop.
[0108] While executing the parameters, the system also needs to monitor feedback signals in real time to determine whether the equipment is responding accurately. Each controlled object is equipped with corresponding sensors or detection devices, such as temperature sensors, spectral monitoring probes, flow meters, and conductivity meters. After receiving the target parameters output by reinforcement learning, the control system continuously reads these sensor signals and calculates the deviation between the current actual value and the target value. If the deviation exceeds the allowable range, for example, if the temperature target is 88 degrees Celsius but the actual value remains below 86 degrees Celsius for more than 2 minutes, the system will determine that the heating unit power is insufficient or the temperature probe is drifting, and will automatically enter the anomaly handling procedure, suspend reinforcement learning updates, and switch to executing preset safety process parameters.
[0109] To prevent device errors due to communication delays or data anomalies, the system provides an acknowledgment response after each command is issued. That is, after the control system sends a command to the execution unit, the execution unit will send a status signal back to the upper control module after completing the command or reaching the set target. The upper control module will not proceed to the next reinforcement learning iteration until it receives the acknowledgment signal. For example, when the new temperature setting is 88 degrees Celsius, the heating unit must return a "target reached" flag signal after reaching the target temperature before the system continues with subsequent decisions. If no acknowledgment signal is received within the set time, the system will automatically roll back to the previous safe parameters to ensure continuous and stable operation.
[0110] This step also involves data recording and traceability mechanisms. Whenever reinforcement learning generates a new combination of parameters and is executed, the system automatically generates a complete runtime log in the data recording module, recording the parameter values, execution time, feedback response, and actual execution results. These records are not only used for subsequent performance evaluation but also provide retraining data for the reinforcement learning model. Through this continuous recording and feedback mechanism, the system can achieve long-term self-learning and performance accumulation, making each parameter execution a new learning sample for the reinforcement learning algorithm.
[0111] To illustrate the actual effect of parameter output and execution, a complete example can be given. Suppose that in the tenth extraction cycle, the reinforcement learning module, after analysis, decides to adjust the temperature to 87 degrees Celsius, the ethanol volume fraction to 78%, the solid-liquid ratio to 1:10, the soaking time to 32 minutes, and the circulation flow rate to 2.2 liters per minute. Upon receiving these parameters, the system first sends them to each execution unit as control signals. The heating system gradually adjusts its power according to the new temperature target, the solvent metering system precisely regulates the ethanol concentration by controlling the proportional valve, the dosing pump regulates solvent replenishment to maintain the solid-liquid ratio, and the circulation pump uses a frequency converter to control the flow rate. During operation, the temperature sensor detects changes in the temperature curve; when the temperature approaches the target value, it automatically enters constant temperature control; the flow meter displays the flow rate in real time, stabilizing within the target range; and the spectral sensor reports changes in the color and absorption intensity of the extract, confirming a gradual increase in flavonoid concentration. When the total soaking time reaches 32 minutes, the system automatically ends the control of this cycle and feeds the results back to the reinforcement learning module as input for the next round of optimization.
[0112] Through the above process, the optimization results output by the reinforcement learning module are fully executed on the device, forming a closed-loop system from data acquisition and intelligent decision-making to physical control. This step can be implemented simply by configuring a programmable control unit and a standardized communication interface on the existing extraction equipment, thus enabling the extraction process of flavonoids from Acanthopanax senticosus to have intelligent characteristics of real-time response, dynamic self-adaptation, and long-term self-optimization.
[0113] A second embodiment of this application provides an electronic device, the electronic device comprising: processor; The memory is used to store a program, which, when read and executed by the processor, executes the self-optimizing method for online extraction of flavonoids from Acanthopanax senticosus based on reinforcement learning provided in the first embodiment of this application.
[0114] The third embodiment of this application provides a computer-readable storage medium storing a computer program thereon. When the program is executed by a processor, it executes a reinforcement learning-based online extraction self-optimization method for flavonoids from Acanthopanax senticosus provided in the first embodiment of this application.
[0115] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.
Claims
1. A method for online extraction self-optimization of flavonoids from S. caudata based on reinforcement learning, characterized in that, The method comprises the following steps: Collecting near-infrared spectrum data, ultraviolet-visible spectrum data, viscosity data and conductivity data of the extractive process of short-stalked ginseng flavonoids, and generating a fusion feature vector after baseline correction, scattering compensation and feature alignment; Using offline test samples to establish a mapping relationship between the fusion feature vector and the contents of rutin, quercetin, kaempferol and hyperoside, and constructing a virtual measurer, which outputs a target flavonoid composition vector and an uncertainty of the target flavonoid composition vector; Constructing a safety domain boundary according to the solvent flash point temperature, the upper limit of the equipment tolerance temperature and the thermal degradation temperature threshold of flavonoid components, and combining the uncertainty of the target flavonoid composition vector to calculate a time-varying safety constraint threshold; Using a reinforcement learning strategy to optimize the temperature, ethanol volume fraction, solid-liquid ratio, soaking time and circulation flow rate in linkage, with the target flavonoid composition vector and the time-varying safety constraint threshold as inputs, wherein the reward function of the reinforcement learning strategy is composed of the weighted distance between the target flavonoid composition vector and a standard flavonoid composition vector and a solvent consumption energy penalty term, and when the predicted value of the optimization action touches the safety domain boundary, the process parameters are returned to the safe state; Outputting the temperature, ethanol volume fraction, solid-liquid ratio, soaking time and circulation flow rate obtained by the linkage optimization to the extraction equipment for execution.
2. The online extraction self-optimization method of the flavonoids in S. macrostachys based on reinforcement learning according to claim 1, characterized in that, The method for collecting near-infrared spectrum data, ultraviolet-visible spectrum data, viscosity data and conductivity data of the extractive process of short-stalked ginseng flavonoids, and generating a fusion feature vector after baseline correction, scattering compensation and feature alignment comprises the following steps: Taking the conductivity sampling time as a unified time reference, resampling the sampling frequencies of the near-infrared spectrum and the ultraviolet-visible spectrum, so that different sensing channels correspond to the same process state at the same time sequence, and obtaining a time-aligned original multi-source data set; Using pure ethanol and deionized water signals at the beginning of production to establish a spectral baseline reference, and subtracting the spectral baseline reference from the time-aligned original multi-source data set point by point to obtain corrected spectral data, which is used to eliminate baseline deviation caused by light source attenuation, light path drift and probe temperature change; In the corrected spectral data, the signal amplitude variation range and the instantaneous variation rate are calculated in a fixed time window, and when the variation of any section exceeds the reference range under normal working conditions, it is determined as a disturbed section, and the adjacent normal section is replaced by smooth interpolation, thereby forming a de-disturbed spectral data; Combining the de-disturbed spectral data, the viscosity measurement curve and the conductivity change curve, and according to the scattering correction coefficients recorded in the equipment parameter table, the amplitude of the spectral signal in different wave bands is corrected, so that the physical relationship between the spectral intensity and the viscosity and ion concentration of the extractive liquid is kept consistent, and a scattering corrected spectrum is obtained; In the scattering corrected spectrum, the time when the conductivity curve rises sharply and the time when the viscosity curve changes from falling to stable are determined as feature time anchor points, and the time axes of the spectrum, viscosity and conductivity are uniformly translated with the feature time anchor points as the reference, so that the signals are uniformly corresponding in the chemical reaction stage. The average intensity, peak position offset and half-peak width of the main absorption peak are extracted from the time-aligned scattering correction spectrum, and the viscosity and conductivity readings of the material liquid at the same time point are extracted synchronously, and the extracted characteristic values are normalized and scaled to generate a comparable standardized feature set; According to the stability of each channel signal in the standardized feature set and its correlation with the conductivity, a weighting coefficient is determined, and the spectral features, viscosity features and conductivity features are weighted and sequentially spliced to generate a fusion feature vector.
3. The online extraction self-optimization method of the flavonoids from S. macrostachys based on reinforcement learning according to claim 1, characterized in that, The mapping relationship between the fusion feature vector and the contents of rutin, quercetin, kaempferol and hyperoside is established by using offline test samples, and a virtual meter is constructed, which outputs the target flavonoid composition vector and the uncertainty of the target flavonoid composition vector, including: Offline test detection is performed on different batches of short-stalked acanthopanax raw materials under various extraction conditions, and the fusion feature vector and the actual measured contents of rutin, quercetin, kaempferol and hyperoside at the corresponding time point are recorded in each detection to form a paired sample set containing time series, feature data and component concentration; The paired sample set is divided into several stability layers according to process parameters such as moisture content of the raw material, particle size after crushing and extraction temperature, the change trend of each fusion feature component with the extraction process is calculated in each stability layer, and the feature components with consistent change direction and high correlation in different levels are extracted to form a robust feature set, which is used to reflect the most sensitive signal source in the flavonoid extraction process; In the robust feature set, the response relationship between each feature component and the content of each flavonoid component is determined, and the response relationship is obtained by weighted smoothing of the multi-point correspondence relationship between the same feature component and the measured content in the offline sample, so that the change trend of the flavonoid content corresponding to different feature components has continuity and comparability, and the increase / decrease direction and slope amplitude of the response relationship in each value interval are recorded to represent the sensitivity of the feature; According to the sensitivity of each feature component, the weighting coefficient is determined and the fusion feature vector is projected into the weighted feature space, the estimated content values of rutin, quercetin, kaempferol and hyperoside at each time point are calculated, and the time series form of the flavonoid content prediction result is obtained, then the difference amplitude between the prediction result and the measured value of the offline test is calculated, and the change range of the difference in a plurality of batches of samples is counted, and the change range is taken as the uncertainty of the estimation of the corresponding component; The estimated content value and its corresponding uncertainty are combined in time sequence to generate a flavonoid component estimation structure that can be updated over time, which outputs the target flavonoid composition vector including rutin, quercetin, kaempferol and hyperoside, and at the same time gives the confidence interval of each component, which is used as a quantitative basis and safety boundary reference for optimization decision in reinforcement learning control.
4. The online extraction self-optimization method of the flavonoids from S. macrostachys based on reinforcement learning according to claim 1, characterized in that, The safety domain boundary is constructed according to the solvent flash point temperature, the upper limit of the equipment tolerance temperature and the thermal degradation temperature threshold of the flavonoid components, and the time-varying safety constraint threshold is calculated combined with the uncertainty of the target flavonoid composition vector, including: The flash point temperature of the ethanol solvent, the upper limit of the tolerance temperature of the extraction equipment, and the thermal degradation temperature threshold of rutin, quercetin, kaempferol, and hyperoside are input into a safety parameter set, and the minimum value of the three is taken as the initial safety temperature upper limit, while the ethanol volume fraction and the equipment allowable pressure are set as the physical boundary parameters of the safety domain; The temperature, pressure, and ethanol volume fraction in the extraction system are monitored in real time, and when any parameter approaches the initial safety temperature upper limit or exceeds the allowable fluctuation range, the deviation proportion among the three types of temperature parameters is calculated, and the deviation proportion is output as a safety margin indicator; The uncertainty of the target flavonoid composition vector is used to dynamically correct the safety margin indicator, and when the uncertainty increases, the corresponding safety margin is proportionally reduced to reflect the safety contraction demand caused by the decrease in the prediction credibility of the components, thereby obtaining the corrected dynamic safety margin; The initial safety temperature upper limit, the safety margin indicator, and the corrected dynamic safety margin are logically combined to generate a time-varying safety constraint threshold, which defines the allowable maximum process temperature, the upper limit of the ethanol volume fraction, and the circulation flow rate limit during the extraction process, and changes in real time with time and estimated uncertainty, and is used to limit the executable parameter range in the reinforcement learning optimization decision to prevent the operating conditions from exceeding the safety domain boundary.
5. The online extraction self-optimization method of the flavonoids from S. macrostachys based on reinforcement learning according to claim 1, characterized in that, The reinforcement learning strategy is adopted to optimize the temperature, ethanol volume fraction, solid-liquid ratio, soaking time, and circulation flow rate in a linked manner, with the target flavonoid composition vector and the time-varying safety constraint threshold as inputs. The reward function of the reinforcement learning strategy is composed of the weighted distance between the target flavonoid composition vector and the standard flavonoid composition vector and the energy consumption penalty term of the solvent consumption. When the optimization action prediction value touches the safety domain boundary, it is returned to the safe process parameters, including: At the beginning of each extraction period, the target flavonoid composition vector and its uncertainty output by the virtual meter are received, and the numerical range of the current time-varying safety constraint threshold is read from the safety constraint module, which are used as the initial input for optimization calculation to determine the parameter space of the current extraction state; The target deviation value of the current period is calculated according to the weighted distance between the target flavonoid composition vector and the standard flavonoid composition vector, and the target deviation value of the previous period is compared with the current value to obtain the change trend of the flavonoid extraction direction; After obtaining the change trend, the influence amplitude of temperature, ethanol volume fraction, solid-liquid ratio, soaking time, and circulation flow rate on the flavonoid extraction direction is calculated, and the influence amplitude is distributed to the adjustable range of each parameter in a relative proportion, thereby generating a set of initial parameter adjustment schemes; For each initial parameter adjustment scheme, the execution result in the extraction equipment is monitored in real time, the corresponding solvent consumption and energy consumption are recorded, and the reward value is calculated according to the improvement degree of the target flavonoid composition vector. The reward value is obtained by subtracting the penalty amount of energy consumption and solvent consumption from the weighted distance of the target flavonoid composition vector. The parameter adjustment scheme with the highest reward value is taken as the prediction result of the current optimization action, and when the prediction result touches the boundary of the time-varying safety constraint threshold, a safety fallback operation is performed, the corresponding parameter is fallen back to the safety range in proportion to the remaining safety margin, and then the fallen back parameter is taken as the initial state input of a new cycle to form a continuous self-optimization cycle.
Citation Information
Cited By
Fine chemical product purity data prediction method based on deep learning
CN122369671A