Intelligent control method, device and equipment for marigold extraction process and medium thereof

By establishing a calibration model between spectral characteristics and concentration, key oxidative byproducts in the marigold extraction process can be monitored and correlated in real time. This solves the problem of lagging quality control, realizes closed-loop control from problem detection to precise inhibition, and improves product quality and stability.

CN121680321APending Publication Date: 2026-03-17QIANHE (ZHANGJIAKOU) BIOTECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511936696.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing technologies cannot monitor the generation of key oxidative byproducts during marigold extraction in real time, resulting in lagging quality control and an inability to achieve a complete control loop from problem detection to root cause identification and precise inhibition, thus affecting product purity and safety.

Method used

By establishing a calibration model between spectral characteristics and by-product concentration, the concentration and formation rate of key oxidation by-products can be monitored in real time. Combined with process parameters, correlation analysis can be performed to intelligently infer the causes of formation and implement precise suppression control.

Benefits of technology

It achieves complete closed-loop control from by-product monitoring to source tracing and inhibition, improving the quality control level and product stability of marigold extraction process, and solving the problem of lag in traditional offline analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121680321A_ABST
    Figure CN121680321A_ABST
Patent Text Reader

Abstract

The invention relates to a marigold extraction process intelligent control method, device and equipment and a medium thereof. The method comprises the following steps: establishing an accurate calibration model between the concentration of a key oxidation byproduct and spectral characteristics offline, and applying the model to an online monitoring system to obtain the concentration data of the byproduct in real time; the instantaneous generation rate is obtained through time sequence differential calculation on the basis of real-time concentration data, intelligent correlation analysis is conducted on the generation rate and extraction process parameters, and therefore the root cause of generation of the by-product is accurately deduced; finally, the targeted suppression operation is automatically triggered according to the diagnosis result, and the control effect is verified by monitoring the generation rate change in real time, so that the complete closed-loop control from byproduct monitoring, tracing to suppression is realized, the hysteresis problem of traditional offline analysis is effectively solved, and the method is suitable for popularization and application. The quality control level and the product stability in the marigold extraction process are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of industrial process control, and particularly relates to a marigold extraction process intelligent control method, device, equipment and medium thereof. BACKGROUND

[0002] In the field of marigold extraction processing, lutein as an important carotenoid is extremely susceptible to oxidation degradation in the production process under the influence of process conditions such as temperature and oxygen, generating key oxidation byproducts such as 5,6-epoxy lutein. These byproducts not only reduce product purity, but also affect the safety and stability of the final product. At present, the industry generally uses offline analysis methods such as high-performance liquid chromatography for quality monitoring. Although this method has high accuracy, it has obvious analysis lag - it often takes several hours from sampling, pretreatment to instrument analysis, which cannot reflect the quality change in the production process in real time, resulting in the inability to achieve true process control.

[0003] Although process analysis techniques such as near-infrared spectroscopy and Raman spectroscopy have been tried for monitoring the extraction process, existing technologies mainly focus on monitoring the content of main components, and there are still obvious deficiencies in detecting trace oxidation byproducts: on the one hand, there is a lack of clear identification of the characteristic spectral fingerprint region of specific key byproducts such as 5,6-epoxy lutein; on the other hand, high-precision chemometric models that can accurately correlate spectral signals with byproduct concentrations have not been established.

[0004] More importantly, the existing technical system can only achieve the function of "monitoring", and cannot intelligently associate the detected byproduct changes with upstream process parameters, so it cannot achieve a complete control closed loop from "finding problems" to "locating sources" to "precise inhibition", which makes the production process still in a passive state of "after-the-fact remediation". When the byproduct is detected to be out of standard, it has often caused batch quality loss.

[0005] Therefore, developing an intelligent control method that can monitor key oxidation byproducts in real time, accurately trace their sources, and automatically implement precise inhibition has become a technical problem to be solved to improve the quality control level of marigold extraction process. SUMMARY

[0006] Therefore, it is necessary to provide a marigold extraction process intelligent control method, device, equipment and medium thereof to solve the above technical problems.

[0007] In a first aspect, the application provides a marigold extraction process intelligent control method, comprising:

[0008] S1, determining the concentration value of the key oxidative byproduct in each sample of a standard sample set containing different concentrations of 5,6-epoxy zeaxanthin, denoted as sample key oxidative byproduct concentration value; collecting sample spectrum data of each sample in the standard sample set, and establishing a calibration model between spectral characteristics and byproduct concentration based on the sample spectrum data and the sample key oxidative byproduct concentration value by chemometrics method; wherein the standard sample set is prepared in an offline accelerated oxidation experiment, and the experiment is carried out on the key oxidative byproducts generated in the marigold extraction process;

[0009] S2, collecting real-time spectrum data of the marigold extraction process, inputting the real-time spectrum data into the calibration model, and outputting the current concentration value of the key oxidative byproduct;

[0010] S3, based on the current concentration value, calculating the instantaneous generation rate of the key oxidative byproduct by time series differentiation; wherein the instantaneous generation rate represents the concentration change amount of the key oxidative byproduct per unit time;

[0011] S4, monitoring the process parameters of the marigold extraction process, correlating the instantaneous generation rate with the process parameters, and obtaining the correlation result; based on the correlation result, intelligently inferring the root cause of the generation of the key oxidative byproduct;

[0012] S5, triggering the inhibition control operation according to the root cause, and monitoring the change of the instantaneous generation rate in real time to verify the control effect.

[0013] In a second aspect, the present application also provides a marigold extraction process intelligent control device for realizing the method described in the first aspect, which comprises:

[0014] The feature correlation modeling module is used for determining the concentration value of the key oxidative byproduct in each sample of a standard sample set containing different concentrations of 5,6-epoxy zeaxanthin, denoted as sample key oxidative byproduct concentration value; collecting sample spectrum data of each sample in the standard sample set, and establishing a calibration model between spectral characteristics and byproduct concentration based on the sample spectrum data and the sample key oxidative byproduct concentration value by chemometrics method; wherein the standard sample set is prepared in an offline accelerated oxidation experiment, and the experiment is carried out on the key oxidative byproducts generated in the marigold extraction process;

[0015] The real-time feature analysis module is used for collecting real-time spectrum data of the marigold extraction process, inputting the real-time spectrum data into the calibration model, and outputting the current concentration value of the key oxidative byproduct;

[0016] The rate dynamic deduction module is used for calculating the instantaneous generation rate of the key oxidative byproduct based on the current concentration value by time series differentiation; wherein the instantaneous generation rate represents the concentration change amount of the key oxidative byproduct per unit time;

[0017] a process correlation inference module, configured to monitor process parameters of the marigold extraction process, and to analyze the instantaneous generation rate in correlation with the process parameters to obtain a correlation result; and based on the correlation result, intelligently infer a root cause of generation of the key oxidation byproduct;

[0018] a control verification execution module, configured to trigger a suppression control operation according to the root cause, and to monitor a change in the instantaneous generation rate in real time to verify a control effect.

[0019] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the intelligent control method of the marigold extraction process according to the first aspect when executing the computer program.

[0020] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the intelligent control method of the marigold extraction process according to the first aspect.

[0021] The intelligent control method, device, equipment and medium of the marigold extraction process according to the above, by establishing an accurate calibration model between the concentration of the key oxidation byproduct and the spectral feature offline, and applying the model to an online monitoring system to obtain byproduct concentration data in real time, based on the real-time concentration data, calculating the instantaneous generation rate through time series differentiation, intelligently analyzing the generation rate in correlation with the extraction process parameters, and thus accurately inferring the root cause of generation of the byproduct, finally triggering a targeted suppression operation automatically according to the diagnosis result, and verifying the control effect by monitoring the change in the generation rate in real time, realizing a complete closed-loop control from byproduct monitoring, tracing to suppression, effectively solving the hysteresis problem of traditional offline analysis, and significantly improving the quality control level and product stability of the marigold extraction process. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments or the related art, the following will briefly introduce the drawings needed to be used in the embodiments or the related art description. Obviously, the drawings in the following description only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor based on these drawings.

[0023] Figure 1 A flowchart of the intelligent control method of the marigold extraction process provided by the present application;

[0024] Figure 2 A flowchart of intelligently inferring the root cause of generation of the key oxidation byproduct in an optional embodiment of the present application;

[0025] Figure 3This is a schematic diagram of the structure of an intelligent control device for the marigold extraction process provided by the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0027] refer to Figure 1 The application presents a flowchart illustrating an intelligent control method for marigold extraction, which includes the following steps:

[0028] S1. Determine the concentration values ​​of key oxidation byproducts in each sample of the standard sample set containing different concentrations of 5,6-epoxylutein, and record them as the key oxidation byproduct concentration values ​​of the samples; collect the sample spectral data of each sample in the standard sample set, and establish a calibration model between spectral characteristics and byproduct concentrations based on the sample spectral data and the key oxidation byproduct concentration values ​​of the samples using chemometric methods; wherein, the standard sample set is prepared in an offline accelerated oxidation experiment, and experiments are carried out on the key oxidation byproducts generated during the extraction of marigolds.

[0029] Specifically, the standard sample set was prepared through offline accelerated oxidation experiments. Using marigold extract as the base liquid, it simulated the core environmental parameters such as temperature and oxygen partial pressure in the actual extraction process. By systematically inducing the generation of key oxidation byproducts through the controlled variable method, a series of samples covering the concentration range that may occur in production were formed.

[0030] The concentrations of key oxidation byproducts were determined using high-performance liquid chromatography (HPLC). A specific column was selected, with a mixed solution of appropriate proportions as the mobile phase. A pH adjuster was added to control the pH value. The column temperature and flow rate were set, and the detection wavelength was determined based on the characteristic absorption peaks of the byproducts. Multiple parallel determinations were performed for each sample, and the average peak area was used to calculate the concentration using a standard curve. The standard curve was plotted by preparing byproduct standard solutions of different concentrations, and the regression equation was obtained through linear regression analysis.

[0031]

[0032] In this formula, y represents the chromatographic peak area, x represents the concentration of the key oxidation byproduct, a is the slope of the regression equation, and b is the intercept term. The formula is obtained by fitting and the linear correlation coefficient is analyzed to verify the reliability of the curve.

[0033] Spectral data acquisition employed a near-infrared spectrometer equipped with an invasive fiber optic probe and a specially structured flow cell. Both ends of the flow cell were sealed with a high-transmittance material to prevent signal attenuation. Baseline correction was performed using a blank extraction solvent before acquisition. Multiple scans of each standard sample were performed and averaged. Environmental temperature and humidity fluctuations were controlled to eliminate external interference.

[0034] Chemometric modeling employs partial least squares regression combined with spectral preprocessing techniques. Preprocessing steps include truncation of characteristic spectral regions of byproducts, first-order derivative processing, smoothing, and transformation of standard normal variables to eliminate baseline drift, matrix effects, and scattering influences. During modeling, the standard sample set is proportionally divided into training and validation sets, and the number of latent variables is optimized through cross-validation. The mathematical principle of the model is based on maximizing the covariance of spectral and concentration data using the partial least squares algorithm. The matrix is ​​solved through singular value decomposition, latent variables are iteratively extracted, and the residual matrix is ​​updated, ultimately establishing a linear mapping relationship. The concentration calculation formula is:

[0035]

[0036] In the formula, c is the concentration of key oxidation byproducts, W is the weight matrix, P is the loading matrix, x is the preprocessed spectral vector, and b is the intercept term. These parameters are determined by fitting the training set data.

[0037] S2. Collect real-time spectral data during the marigold extraction process, input the real-time spectral data into the calibration model, and output the current concentration values ​​of key oxidation byproducts.

[0038] Specifically, the detection probe of the spectral acquisition system is installed at a specific location in the discharge pipe of the extraction tank. This location has a stable fluid state and no dead volume, thus avoiding spectral distortion caused by solid-liquid separation. The probe is connected to the pipe via a sealing gasket, and the insertion depth must ensure full contact with the material. An external insulation layer is wrapped around the probe to reduce the impact of temperature fluctuations on the spectral signal.

[0039] The real-time spectral acquisition parameters are consistent with those of the standard sample acquisition. The data is transmitted to the industrial controller in real time via a specific network protocol, and the transmission latency must be kept within a low range. To adapt to the complex environment of the industrial site, a filtering step is required for real-time spectral preprocessing to eliminate peak noise caused by particulate impurities and bubbles in the pipeline. The total preprocessing time must be strictly controlled.

[0040] The calibration model is invoked through the controller's built-in algorithm module. This module solidifies the partial least squares model parameters. The real-time spectrum is preprocessed and converted into an absorbance vector, which is then input into the model to quickly solve for the concentration value through matrix operations. The concentration calculation formula used is consistent with that in S1, i.e. All parameters have the same meaning, ensuring that the total delay from sampling to concentration output meets the requirements of real-time monitoring.

[0041] A data validity judgment mechanism is also set up. When the absorbance value exceeds the calibration model input range by a certain proportion, or when the spectral correlation coefficient of multiple consecutive acquisitions is lower than the threshold, a "data anomaly" signal is output, and the automatic probe cleaning program is triggered to ensure the reliability of the monitoring data. The concentration output results retain the corresponding decimal places, are displayed in real time on the central control system interface, and are simultaneously stored in the database to provide continuous data support for subsequent generation rate calculations.

[0042] S3. Based on the current concentration value, calculate the instantaneous generation rate of the key oxidation byproducts through time-series differentiation; where the instantaneous generation rate represents the amount of concentration change of the key oxidation byproducts per unit time.

[0043] Specifically, this step quantifies the dynamic trend of key oxidation byproduct formation through time-series differential analysis, providing dynamic response indicators for process control. The instantaneous formation rate is defined as the change in byproduct concentration per unit time, mathematically expressed as:

[0044]

[0045] In the formula, v is the instantaneous generation rate, c is the concentration of the key oxidation byproduct, and t is time. This formula reflects how quickly the byproduct concentration changes over time.

[0046] Data processing is based on the time-series concentration data stored in the S2 stage, with a fixed sampling time interval set to form a time series. ,in For the i-th sampling time, This represents the concentration value at the corresponding time point. To improve the accuracy of the differential calculation, a five-point numerical differential method can be used. This method eliminates the influence of random noise by weighted fitting of five adjacent data points, resulting in a lower truncation error.

[0047] For the midpoint of the time series, the formula for calculating the instantaneous generation rate is:

[0048]

[0049] In the formula, Let i be the instantaneous generation rate at time i. , , , These are the concentration values ​​at times i-2, i-1, i+1, and i+2, respectively. The sampling time interval is defined. For the start and end points of the sequence, the corresponding forward or backward three-point differential formula is used to ensure the integrity of the rate data for the entire time series.

[0050] Before calculation, the units need to be standardized, and the sampling time interval needs to be converted to a time unit that matches the rate unit. To further reduce noise interference, a moving average filter is applied to the calculated rate sequence. The filter formula is as follows:

[0051]

[0052] In the formula, The instantaneous generation rate after filtering. to The calculation method for edge points can be adjusted based on actual conditions, using the rate values ​​at five consecutive time points. A rate threshold criterion is also set, and the process state corresponding to different rate ranges is determined through statistical analysis of multiple batches of industrial production data, such as "rapid generation of by-products," "stable generation rate," and "decreasing by-product concentration."

[0053] S4. Monitor the process parameters of marigold extraction, perform correlation analysis between the instantaneous generation rate and the process parameters, and obtain the correlation results; based on the correlation results, intelligently infer the root cause of the generation of key oxidation byproducts.

[0054] Specifically, the key process parameters affecting lutein oxidation are first selected, including extraction temperature, oxygen partial pressure in the extraction tank, stirring rate, pH value of the extraction solution, and concentration of the extraction solvent.

[0055] Each parameter is monitored using appropriate specialized equipment, which is installed appropriately. A high-precision thermometer is used for temperature monitoring, installed in the middle material layer of the extraction tank; an electrochemical sensor is used for oxygen partial pressure monitoring, installed in the gas phase space at the top of the tank; the stirring rate is calculated using the output frequency of a frequency converter; an online pH electrode is used for pH value monitoring, installed adjacent to the spectral probe on the discharge pipe; and an online refractometer is used for solvent concentration monitoring, installed on the reflux pipe. All parameters are sampled at a consistent frequency and synchronized with the central control system via a specific protocol to ensure timestamp synchronization accuracy and time consistency with concentration and rate data.

[0056] The correlation analysis employs a combination of partial correlation analysis and grey relational analysis. First, all data are mean-normalized and dimensionless to eliminate dimensional differences. The processing formula is as follows:

[0057]

[0058]

[0059] In the formula, Let j be the standardized value of the j-th process parameter at time i. Let be the original value of the j-th process parameter at the i-th time point, and n be the number of data points. Let be the standardized value of the instantaneous generation rate at time i. This is the original value of the instantaneous generation rate at time i.

[0060] Partial correlation analysis is used to eliminate cross-interference between parameters and calculate the net correlation coefficient between the instantaneous generation rate and a single parameter. The formula is:

[0061]

[0062] In the formula, Let be the partial correlation coefficient between the instantaneous generation rate and the j-th process parameter. The simple correlation coefficient between the two is... The correlation coefficient between the instantaneous generation rate and other parameters. Let be the correlation coefficient between the j-th parameter and other parameters.

[0063] Grey relational analysis is used to quantify the nonlinear relationship between parameters and rates. First, the first-order difference between the reference sequence and the comparison sequence is calculated:

[0064]

[0065] In the formula, Let be the absolute value of the difference at time i. Then, solve for the maximum and minimum differences between the two levels, define the resolution coefficient, and calculate the correlation coefficient:

[0066]

[0067] In the formula, Let be the correlation coefficient of the j-th parameter at time i. It is the minimum difference between two levels. The maximum difference between the two levels, The relevance coefficient is used for differentiation. The final correlation coefficient is:

[0068]

[0069] In the formula, Let be the correlation between the j-th process parameter and the instantaneous generation rate.

[0070] The overall correlation score is calculated using a weighted sum, for example:

[0071]

[0072] In the formula, The comprehensive correlation score for the j-th process parameter is determined by weights (0.6 and 0.4) based on the complementarity of the two analysis methods: partial correlation analysis focuses on linear correlation, while grey relational analysis focuses on the overall trend. The root cause inference rule is to select the parameter with the largest comprehensive correlation score. If the current value of this parameter exceeds the optimal range of the process, it is determined to be the root cause; if it is within the optimal range, the influence of the interaction term is further analyzed through multiple linear regression to ensure the accuracy of the inference.

[0073] S5. Trigger the suppression control operation based on the root cause and monitor the change in instantaneous generation rate in real time to verify the control effect.

[0074] Specifically, this step implements precise suppression control based on the root cause and verifies the control effect through real-time monitoring, forming a complete control closed loop. Specific control strategies are designed for different root causes: when the root cause is abnormal extraction temperature, a PID control algorithm is used to adjust the opening of the steam regulating valve, with the control objective being to reduce the temperature to the process optimum median. The standard mathematical expression for a PID controller is:

[0075]

[0076] In the formula, This is the output signal of the controller, i.e., the valve opening. This is the error signal, equal to the difference between the current temperature and the target temperature. For proportional gain, For integral gain, For differential gain, This represents the integral of the error from the initial time to the current time. This indicates the rate of change of error. The opening adjustment needs to be set with a reasonable step size to avoid sudden temperature drops causing the target components to be extracted. Temperature upper and lower limit protection should also be set.

[0077] When the root cause is excessively high oxygen partial pressure, start the inert gas purging system and adjust the purging flow rate through the mass flow controller. The flow rate is determined based on the oxygen partial pressure deviation. At the same time, close the pneumatic ball valves at the inlet and outlet to maintain a slight positive pressure inside the tank to prevent air from seeping in. Stop purging when the oxygen partial pressure drops to the set value.

[0078] When the root cause is a pH deviation from the optimal range, add buffer solution to the extraction tank using a metering pump. The amount added is calculated based on the pH deviation, and the addition rate needs to be controlled to avoid local pH abrupt changes. At the same time, the stirring rate can be appropriately increased to promote mixing.

[0079] When the root cause is abnormal solvent concentration, the corresponding concentration of solvent is added by a variable frequency metering pump. The amount added is calculated based on the concentration deviation and the effective volume of the extraction tank.

[0080] The control effectiveness verification adopts a dual-indicator system. The dynamic indicator is that the instantaneous generation rate drops below a set threshold within a certain period after control is started; the steady-state indicator is that the rate remains stable for a certain period of time, which is considered as the control being effective. If the dynamic indicator fails to meet the standard, the control parameters are automatically adjusted and the control is re-executed; if the standard is still not met after multiple adjustments, an audible and visual alarm is triggered and an alarm signal is sent to the central control system to prompt the operator to intervene.

[0081] During the control process, changes in other process parameters are monitored in real time to ensure that control operations do not affect extraction efficiency and product yield. All actuators are equipped with limit protection and fault self-checking functions to avoid over-adjustment leading to process runaway, ultimately achieving closed-loop control from "problem detection" to "precise suppression".

[0082] The aforementioned intelligent control method for marigold extraction establishes an accurate calibration model offline between the concentration and spectral characteristics of key oxidative byproducts, and applies this model to an online monitoring system to acquire byproduct concentration data in real time. Based on the real-time concentration data, the instantaneous generation rate is calculated using time-series differential calculations. This generation rate is then intelligently correlated with extraction process parameters to accurately infer the root cause of byproduct generation. Finally, targeted inhibition operations are automatically triggered based on the diagnostic results, and the control effect is verified by real-time monitoring of changes in the generation rate. This achieves complete closed-loop control from byproduct monitoring and tracing to inhibition, effectively solving the lag problem of traditional offline analysis and significantly improving the quality control level and product stability of the marigold extraction process.

[0083] In one optional embodiment, based on sample spectral data and the concentration values ​​of key oxidation byproducts in the sample, a calibration model between spectral characteristics and byproduct concentrations is established using chemometric methods, including the following steps:

[0084] S11. Based on the sample spectral data, calculate the Pearson correlation coefficient between the spectral intensity at each wavelength or wavenumber and the concentration value of the key oxidation byproducts in the sample, and screen out the characteristic spectral regions whose absolute value of the Pearson correlation coefficient is greater than the first preset threshold to obtain the candidate feature set.

[0085] Specifically, this step uses Pearson correlation coefficient analysis to screen spectral regions that have a significant linear correlation with the concentration of key oxidation byproducts, providing high-quality feature input for subsequent modeling. Before implementation, the sample spectral data is preprocessed, typically including smoothing and baseline correction. Smoothing can use the Savitzky-Golay algorithm to eliminate random noise, while baseline correction uses polynomial fitting to remove background interference caused by instrument drift, ensuring that the spectral intensity data accurately reflects the characteristic absorption information of the substances.

[0086] The Pearson correlation coefficient is calculated based on the linear correlation analysis between spectral intensity and concentration values, and its core formula is:

[0087]

[0088] In the formula, r represents the Pearson correlation coefficient at the j-th wavelength or wavenumber. Let be the spectral intensity of the i-th standard sample at the j-th wavelength. Let be the average spectral intensity of all samples at the j-th wavelength. Let be the concentration value of the key oxidation byproducts of the i-th standard sample. The concentration is the average of all samples, and n is the total number of samples in the standard sample set. The value of this coefficient ranges from [-1, 1]. The closer the absolute value is to 1, the stronger the linear correlation between spectral intensity and concentration; the closer the absolute value is to 0, the weaker the linear correlation.

[0089] During the calculation process, the above operations are performed one by one for each wavelength or wavenumber, forming a correlation coefficient sequence across the entire spectrum. The screening operation is based on a first preset threshold, which is determined by analyzing the spectral correlation coefficient distribution of blank samples, taking into account the spectral noise level and preliminary experimental results, to ensure that the screened features are not affected by noise. Screening does not involve selecting a single highly correlated wavelength in isolation, but rather identifying continuous characteristic spectral regions: when the absolute value of the correlation coefficient of a certain wavelength is greater than the first preset threshold, it is further determined whether its adjacent wavelengths meet the same conditions. If multiple consecutive wavelengths meet the threshold requirements, the wavelength segment is merged into a single characteristic region; if a single wavelength meets the conditions but adjacent wavelengths do not, it is considered a noise point and eliminated. This region-based screening method effectively avoids the random interference of a single wavelength, enhances the stability and representativeness of the features, and the final set of continuous wavelength segments constitutes the candidate feature set.

[0090] S12. Perform principal component analysis on the candidate feature set and extract the principal component with the highest contribution rate as the comprehensive spectral feature; divide the comprehensive spectral feature and the concentration values ​​of key oxidation byproducts of the sample into training set and validation set; based on the training set, use partial least squares regression to construct the initial calibration model.

[0091] Specifically, this step achieves feature dimensionality reduction through principal component analysis and constructs an initial calibration model using partial least squares regression. The core objective is to simplify the model structure while preserving key information, thereby improving prediction accuracy and generalization ability. First, the candidate feature set is standardized to eliminate dimensional differences in spectral intensity across different wavelengths. The standardization formula is:

[0092]

[0093] In the formula, For the standardized spectral intensity, The standard deviation of the spectral intensity at the j-th wavelength is used to ensure that each wavelength feature has equal weight in the principal component analysis.

[0094] Principal component analysis (PCA) essentially transforms a high-dimensional candidate feature set into a low-dimensional composite variable (i.e., principal components) through orthogonal transformation. The specific steps are: calculating the covariance matrix of the standardized candidate feature set, and the elements of the covariance matrix... The calculation formula is:

[0095]

[0096] In the formula, Let $\mathbf{j}$ be the covariance of the normalized spectral intensities of the $j$-th and $k$-th wavelengths. Let be the mean of the normalized spectral intensity at the j-th wavelength. Eigenvalue decomposition is performed on the covariance matrix to obtain eigenvalues ​​and corresponding eigenvectors. The magnitude of the eigenvalue reflects the original information content carried by the corresponding principal component, while the eigenvector represents the composition direction of the principal component. The principal component with the highest contribution rate is extracted as the comprehensive spectral feature. The contribution rate is calculated as the ratio of a single eigenvalue to the sum of all eigenvalues. The top few principal components whose cumulative contribution rate reaches a preset proportion are selected to ensure that most of the effective information in the original feature set is retained, while simultaneously achieving dimensionality compression and resolving the multicollinearity problem of spectral data.

[0097] The division of samples based on spectral characteristics and the concentration values ​​of key oxidation byproducts follows a random stratification principle to ensure consistent concentration distribution between the training and validation sets, avoiding model evaluation distortion due to differences in data distribution. During the division process, samples are first grouped according to concentration gradients, and then a certain proportion of samples are randomly selected from each group to form the training set, with the remaining samples serving as the validation set. This approach ensures that both sets of data cover the entire concentration range, providing a reliable data foundation for model training and validation.

[0098] The initial calibration model was constructed using partial least squares regression. The advantage of this method lies in its ability to simultaneously consider the variation information of both spectral characteristics (independent variable) and concentration values ​​(dependent variable), achieving the optimal correlation between the two through iterative extraction of latent variables. The modeling process first extracts the first pair of latent variables from the independent variable matrix and the dependent variable vector. (Independent variable side) and (Dependent variable side), making and The covariance reaches its maximum, and the calculation formula is:

[0099]

[0100]

[0101] In the formula, X' is the standardized comprehensive spectral feature matrix, and Y' is the standardized concentration vector. The weight vector is the one on the independent variable side. This represents the weight vector on the dependent variable side. Then, the regression coefficients are calculated, establishing X' pairs. Y' The regression equation is derived, and the residual matrix is ​​solved. If the residuals do not meet the convergence condition, latent variables are repeatedly extracted based on the residual matrix until the number of latent variables meets the preset requirement or the residuals converge. Finally, a linear regression equation between the comprehensive spectral characteristics and concentration values ​​is established using all extracted latent variables, forming the initial calibration model.

[0102] S13. Based on the validation set, calculate the determination coefficient of the initial calibration model; if the determination coefficient is less than the second preset threshold, return to S11 to re-optimize feature selection; otherwise, output the final calibration model.

[0103] Specifically, this step evaluates the performance of the initial calibration model using a validation set, constructing a closed-loop process of "modeling-validation-optimization" to ensure that the final output calibration model has sufficient predictive accuracy and generalization ability. The core evaluation metric is the coefficient of determination, calculated using the following formula:

[0104]

[0105] In the formula, The coefficient of determination is denoted by m, where m is the number of samples in the validation set. To verify the actual concentration value of the i-th sample in the set, Let be the model's predicted concentration value for the i-th sample. This represents the average actual concentration of the validation set samples. This metric reflects the model's ability to explain concentration changes; a value closer to 1 indicates a better fit and higher prediction accuracy, while a value closer to 0 indicates poorer model performance.

[0106] The second preset threshold is set based on industry standards for methodology validation and practical application requirements, referencing the accuracy requirements of similar spectral analysis methods to ensure that the model can meet the quantitative detection needs of trace by-products during marigold extraction. If the calculated coefficient of determination is greater than or equal to the second preset threshold, it indicates that the performance of the initial calibration model has met expectations and can be directly output as the final calibration model; if the coefficient of determination is less than the second preset threshold, return to S11 to re-optimize feature selection.

[0107] The re-optimization mainly includes three aspects: First, adjusting the first preset threshold. If the original threshold is too low, resulting in too many noisy features in the candidate feature set, the threshold can be appropriately increased to remove interfering wavelengths; if the original threshold is too high, resulting in the loss of effective features, the threshold can be decreased to expand the feature range. Second, optimizing the feature region screening rules, such as adjusting the judgment window size for continuous wavelengths, or adding spectral preprocessing steps (such as standard normal variable transformation) to enhance the correlation of effective features. Third, supplementing the coverage of the standard sample set. If the concentration gradient of existing samples is insufficient, leading to inaccurate correlation coefficient calculations, samples with different concentration gradients can be added through offline accelerated oxidation experiments to improve the reliability of feature-concentration correlation. After re-executing the S11 to S13 process, the determination coefficient of the model is re-evaluated until it meets the second preset threshold requirement, ensuring that the final output calibration model can accurately quantify the concentration of key oxidation byproducts.

[0108] In one optional embodiment, the instantaneous formation rate of key oxidation byproducts is calculated by time-series differentiation based on the current concentration value, including the following steps:

[0109] S21. Extract continuous time series data from the real-time output current concentration value, perform exponential weighted moving average processing on the time series data, and obtain a smoothed concentration series.

[0110] Specifically, this step processes real-time concentration data using an exponentially weighted moving average. The core objective is to eliminate random noise interference, preserve the true trend of concentration changes, and provide a stable input for subsequent rate calculations. The current concentration value output in real time is affected by factors such as electromagnetic interference and material flow fluctuations in the industrial environment, and is accompanied by high-frequency noise. Directly using it for differential calculations will lead to distortion of the rate data; therefore, smoothing preprocessing is necessary.

[0111] The Exponentially Weighted Moving Average (EWMA) assigns decreasing weights to concentration data at different times, highlighting the dominant role of recent data while retaining the influence of historical data, making it suitable for dynamic, real-time monitoring scenarios. Its core calculation formula is:

[0112]

[0113] In the formula, The smoothed concentration value at time t. For smoothing coefficients, This represents the original current concentration value at time t. This is the smoothed concentration value from the previous moment. Sampling time interval. Smoothing coefficient. The value of is between 0 and 1, and its magnitude determines the model's response speed to concentration changes and its noise suppression ability: if the concentration changes rapidly and fluctuates significantly during the marigold extraction process, A larger value is needed to quickly capture changes; if the concentration change is gradual, A smaller value is chosen to enhance the noise filtering effect. The actual value is determined through preliminary experiments, with the goal of minimizing the mean square error between the smoothed sequence and the original sequence.

[0114] The extraction of continuous time series data is synchronized with the frequency of spectral acquisition to form... A time series set, where each time point is at an interval of 1. At the initial time (t=0), an initial smoothed concentration value is set, for example, the arithmetic mean of the first 3 to 5 original concentration data, to ensure the continuity of subsequent recursive calculations. During processing, each time a new real-time concentration value is obtained, it is immediately substituted into the formula to update the smoothed concentration sequence, achieving dynamic smoothing. The smoothed data needs to be stored in the buffer in real time to prepare for the next difference calculation.

[0115] S22. Based on the smoothed concentration sequence, the generation rate at each time point is calculated using the central difference method; the formula for calculating the generation rate is:

[0116]

[0117] in, Represents the instantaneous generation rate at time t. This represents the current concentration value at time t. Indicates the sampling time interval.

[0118] Specifically, this step uses the central difference method to perform differential operations on the smoothed concentration sequence, converting discrete concentration data into a continuous generation rate. The core principle is to quantify the rate of change through the concentration differences between adjacent time points. Compared to simple difference, this method effectively reduces endpoint errors and better aligns with the physical definition of instantaneous rate.

[0119] Before calculation, ensure the integrity of the smoothed concentration sequence and the consistency of time intervals. If data loss or abnormal intervals occur, activate the completion mechanism: for a single missing data point, use linear interpolation of adjacent smoothed concentration values ​​to fill the gap; for multiple consecutive missing points, trigger a data anomaly alarm and pause rate calculation until the data returns to normal. In the formula for calculating the generation rate, r(t) represents the generation rate at time t. This represents the smoothed concentration value at time t. This represents the smoothed concentration value over time t-Δt. This represents the sampling time interval. This formula approximates the rate of change of concentration over time by dividing the smoothed concentration difference between the current and previous moments by the time interval. Although it is a simplified central difference form, it can reduce computational latency while ensuring accuracy in industrial scenarios with high real-time requirements.

[0120] For the starting time of the sequence (t≤Δt), since there is a lack of forward data, the forward difference method can be used to supplement the calculation. The formula is as follows:

[0121]

[0122] In the formula, The initial generation rate. This represents the smoothed concentration value of the first sampling point after the initial time. All calculations are implemented in the industrial controller through a built-in algorithm module, ensuring that the time from data input to rate output is controlled within one-tenth of the sampling interval, meeting the requirements for real-time monitoring.

[0123] S23. Normalize the generation rate to obtain the normalized generation rate, which is used as the instantaneous generation rate.

[0124] Specifically, this step normalizes the generation rate to eliminate rate fluctuations caused by differences in concentration magnitude, mapping the rate data to a uniform scale and providing standardized indicators for subsequent correlation analysis of process parameters. Unnormalized rate data is difficult to use directly for trend comparison between different batches or different concentration ranges due to different concentration benchmarks; normalization effectively solves this problem.

[0125] In industrial settings, Z-score normalization is commonly used. This method transforms the rate sequence based on its statistical characteristics, preserving the positive or negative attributes of the rate (reflecting an increasing or decreasing trend in concentration). Its calculation formula is as follows:

[0126]

[0127] In the formula, Let r(t) be the normalized generation rate at time t, and r(t) be the original generation rate at time t. The mean of the generated rate sequence, The standard deviation of the generated rate sequence. Mean. and standard deviation The calculation is based on a sliding time window, the size of which is determined according to the extraction process cycle. It usually covers a complete concentration change stage, which ensures the representativeness of statistical characteristics and can adapt to the dynamic changes in process conditions.

[0128] Before calculation, outlier detection is performed on the generated rate sequence, and the 3σ criterion is used to remove outliers. Extreme data within the range are used to avoid outliers interfering with statistical parameters. If an outlier is detected, it is replaced with a linear interpolation of the adjacent normal rate value. The normalized generation rate typically falls within the range of [-3, 3]. Value changes within this range directly reflect the relative rate of byproduct formation: positive values ​​indicate an increase in concentration, with larger absolute values ​​indicating a faster increase; negative values ​​indicate a decrease in concentration, with larger absolute values ​​indicating a faster decrease. The final output normalized generation rate is the instantaneous generation rate used for subsequent process analysis and is synchronized in real time to the central control system and database.

[0129] refer to Figure 2 In one optional embodiment, the process parameters of the marigold extraction process are monitored, and the instantaneous generation rate is correlated with the process parameters to obtain the correlation results. Based on the correlation results, the root cause of the generation of key oxidative byproducts is intelligently inferred, including the following steps:

[0130] S31. Collect real-time process parameter data to form a multivariate time series dataset; by discretizing the continuous variables in the multivariate time series dataset, each variable is divided into multiple intervals to obtain a discretized multivariate dataset; among which, the real-time process parameter data includes heater surface temperature, tank pressure, oxygen concentration and stirring speed.

[0131] Specifically, this step solves the problem of continuous variables being difficult to directly use for correlation calculations by synchronously collecting process parameters and generation rate data, discretizing them, and converting them into a dataset suitable for probability analysis. The sampling frequency of real-time process parameters and instantaneous generation rate is kept consistent and can be set to the same time interval to ensure that each time point corresponds to a complete parameter combination. The heater surface temperature is collected using a contact platinum resistance thermometer installed in the middle region of the heating tube's outer wall; the tank pressure uses a diffused silicon pressure sensor deployed at the gas phase interface at the top of the extraction tank; the oxygen concentration uses an electrochemical oxygen sensor, with the probe inserted into the upper part of the liquid phase inside the tank; the stirring speed is obtained by converting the output frequency of the frequency converter. All sensor data are transmitted to the data acquisition module via current signals, converted into digital signals, and aligned by timestamps to form a multivariate time series dataset in the following format:

[0132]

[0133] Where T represents the heater surface temperature, P represents the tank pressure, O represents the oxygen concentration, S represents the stirring speed, and r represents the instantaneous generation rate.

[0134] The continuous variable discretization employs an equal-frequency discretization method. This method ensures that each interval contains a similar number of samples, avoiding probability estimation bias caused by uneven data distribution. During implementation, historical data for each parameter is first statistically sorted and divided into multiple intervals with an equal total sample size. The number of intervals is determined based on the parameter sensitivity. For example, heater surface temperature and oxygen concentration significantly affect the oxidation reaction, resulting in 5 intervals; tank pressure and stirring speed have relatively mild effects, resulting in 3 intervals. For instance, heater surface temperature can be divided into five levels: "extremely low, low, medium, high, and extremely high," each corresponding to a continuous temperature range. The boundary of the division is determined by historical normal process intervals and abnormal event trigger thresholds. The ordered nature of variables is preserved during discretization to ensure that interval divisions conform to process logic; for example, "low, medium, and high" stirring speeds correspond to increasing rotational speed ranges. Finally, continuous parameter values ​​are replaced with corresponding interval labels, forming a discretized multivariate dataset, laying the foundation for subsequent mutual information calculation and probabilistic modeling.

[0135] S32. Based on the time series and discretized multivariate dataset of the instantaneous generation rate, calculate the mutual information value between the instantaneous generation rate and each process parameter to characterize the strength of the nonlinear correlation; the formula for calculating the mutual information value is:

[0136]

[0137] in, This represents a discretized set of process parameter values. A discretized set representing the instantaneous generation rate values; Indicates instantaneous generation rate With process parameters Mutual information value; Indicates the values ​​of process parameters With the instantaneous generation rate value The joint probability distribution is obtained by statistically discretizing the multivariate dataset. and The frequencies falling within the corresponding discrete intervals are estimated simultaneously. and These represent the values ​​of the process parameters. and instantaneous generation rate value The marginal probability distribution is obtained by estimating the frequency of each variable falling into its corresponding discrete interval in a statistically discretized multivariate dataset.

[0138] Specifically, this step quantifies the nonlinear correlation strength between the instantaneous generation rate and various process parameters through mutual information values, compensating for the limitation of linear correlation analysis in capturing complex relationships. Mutual information, based on information entropy theory, reflects the degree to which one variable contains information about another. The formula for calculating mutual information includes… Indicates instantaneous generation rate The mutual information value between the process parameter X and the process parameter X is a non-negative real number. The larger the value, the stronger the correlation between the two. X represents the discretized set of process parameter values, such as the set of five interval labels of heater surface temperature. The discrete set representing the instantaneous generation rate values ​​is divided into five intervals based on the normalized rate values: "significantly decreasing, slowly decreasing, stable, slowly increasing, and significantly increasing". This represents the process parameter value x and the instantaneous generation rate value. The joint probability distribution is calculated by statistically analyzing x and x in a discretized multivariate dataset. The proportion of times they appeared simultaneously out of the total sample size; The marginal probability distribution of the process parameter value x is the proportion of the number of times x appears in the dataset to the total number of samples. Indicates the instantaneous generation rate value The marginal probability distribution, calculated in the same way as Consistent.

[0139] The calculation process needs to handle the special case of zero joint probability, which is corrected using the Laplace smoothing method. During joint probability statistics, each possible case is processed accordingly. An extra count is added to the combination to avoid meaningless zero values ​​in logarithmic operations. The smoothed joint probability calculation formula is adjusted as follows:

[0140]

[0141] In the formula, The joint probability after smoothing. For x and The original number of occurrences, where N is the total number of samples, and |X| is the number of elements in the discretized set of process parameters. The number of elements in the discretized set of instantaneous generation rates.

[0142] The marginal probabilities are calculated by summing the smoothed joint probabilities, as shown in the following formula:

[0143]

[0144]

[0145] The mutual information value obtained through the above calculation can objectively reflect the degree of influence of each process parameter on the instantaneous generation rate, providing a quantitative basis for screening key factors.

[0146] S33. Extract a historical multivariate time series dataset from the historical database, which includes historical process parameters, external events, and abnormal generation rate data; wherein, external events include local overheating and oxygen leakage, and abnormal generation rate data is determined based on the historical instantaneous generation rate exceeding a preset threshold.

[0147] Specifically, this step involves selecting representative datasets from historical databases for the construction and training of the Bayesian network. The core requirement is to cover different process states and abnormal scenarios to ensure the model's generalization ability. Historical process parameter data must be consistent with the types of parameters acquired in real time, including complete time-series records of heater surface temperature, tank pressure, oxygen concentration, and stirring speed, covering various operating conditions such as normal operation, minor anomalies, and severe anomalies.

[0148] External event data is extracted from historical production logs and equipment alarm records. Local overheating events are determined based on the heater surface temperature exceeding a preset safety threshold and the duration exceeding a preset time. Oxygen leakage events are determined based on a sudden increase in tank oxygen concentration exceeding a preset percentage of normal fluctuation range, accompanied by synchronous abnormal pressure changes. Each event is tagged with its occurrence time, duration, and termination status, and incorporated into the dataset as binary variables (1 indicates occurrence, 0 indicates non-occurrence). Abnormal generation rate data is determined based on historical instantaneous generation rates. The preset threshold is determined using the 3σ criterion from historical normal batch rate data. When the rate value exceeds... When the range is specified, it is marked as abnormal (1 indicates abnormal, 0 indicates normal). This represents the average rate of normal batches. This represents the standard deviation of the normal batch rate.

[0149] The extracted historical multivariate time series dataset needs to be preprocessed to remove samples with missing timestamps or invalid parameter values; the continuous parameters are discretized in the same way as in S31; the data is divided into fixed-length sample windows in chronological order, each window containing parameter and event data from multiple consecutive time points, corresponding to a rate anomaly state label, forming a sample format suitable for network training.

[0150] S34. Based on the variable relationships of historical process parameters, external events, and abnormal generation rate data, construct the initial structure of a Bayesian network containing visible and hidden nodes.

[0151] Specifically, this step constructs an initial Bayesian network structure based on the causal relationships between variables. By combining a directed acyclic graph with a conditional probability table, probabilistic modeling of complex process systems is achieved. The Bayesian network consists of visible nodes and hidden nodes. Visible nodes directly correspond to observable variables, including key process parameters (heater surface temperature, oxygen concentration, tank pressure, stirring speed), external events (local overheating, oxygen leakage), and abnormal states of the formation rate. Hidden nodes characterize intermediate variables that cannot be directly measured but have a significant impact on the entire marigold extraction system. Based on the marigold extraction process mechanism, "oxidative reaction activity" is introduced as a hidden node. This node comprehensively reflects the degree to which factors such as temperature and oxygen promote the oxidation reaction and is a key intermediate link connecting process parameters and by-product formation.

[0152] The directed edges of the network structure are set according to causal relationships. For example, the edges "heater surface temperature → local overheating, oxygen concentration → oxygen leakage" reflect the direct causal relationship between abnormal parameters and external events; "heater surface temperature → oxidation reaction activity, oxygen concentration → oxidation reaction activity, stirring speed → oxidation reaction activity" reflect the direct influence of process parameters on the reaction process; "local overheating → oxidation reaction activity, oxygen leakage → oxidation reaction activity" reflect how external events affect byproduct formation by aggravating reaction activity; "oxidation reaction activity → abnormal formation rate" is the core causal chain of the entire network, reflecting the direct correlation between changes in reaction activity and abnormal rates; tank pressure, as an auxiliary parameter, indirectly affects reaction activity by influencing the mixing state of materials and is set as the parent node of oxidation reaction activity. During construction, an acyclic structure is ensured. The node connection logic is verified by drawing a causal relationship graph to avoid circular dependencies, forming a preliminary directed acyclic graph structure.

[0153] S35. Using a historical multivariate time series dataset, train the parameters of the initial structure of the Bayesian network using the maximum likelihood estimation method to obtain the trained conditional probability table; based on the conditional probability table and the initial structure of the Bayesian network, obtain the trained Bayesian network model.

[0154] Specifically, this step uses the maximum likelihood estimation method to train the parameters of the initial network structure, solves the conditional probability table of each node, and completes the construction of the Bayesian network model. The core idea of ​​maximum likelihood estimation is to find the parameter values ​​that maximize the probability of occurrence in the historical dataset. For a Bayesian network, the parameters are the conditional probability distribution of each node given the value of its parent node.

[0155] Before training, the historical multivariate time series dataset is converted into a sample matrix that conforms to network nodes. Each row of samples contains discrete values ​​of all visible and hidden nodes. For the value of the hidden node "oxidation reaction activity," it is indirectly inferred through the combination of parent node parameters (heater surface temperature, oxygen concentration, etc.), and divided into three discrete levels: "low activity," "medium activity," and "high activity." Maximum likelihood estimation first calculates the frequency of each node's parent node combination, then statistically analyzes the conditional frequency of different values ​​of the node under each parent node combination, using this as an estimate of the conditional probability. Taking the "oxidation reaction activity" node as an example, if its parent nodes are heater surface temperature (T) and oxygen concentration (O), then the formula for calculating a certain conditional probability P(activity = high | T = high, O = high) is:

[0156]

[0157] In the formula, N(activity=high,T=high,O=high) is the number of samples in the dataset that are high in activity, high in temperature and high in oxygen concentration, and N(T=high,O=high) is the total number of samples that are high in temperature and high in oxygen concentration.

[0158] The likelihood function for the entire network is defined as the product of the joint probabilities of all samples, and the formula is:

[0159]

[0160] In the formula, Let be the likelihood function. Here is the network parameter vector, where M is the number of samples and K is the total number of network nodes. Let be the value of the j-th node in the i-th sample. Let be the set of parent nodes of the j-th node. Let be the conditional probability parameter for the j-th node. To simplify the calculation, take the logarithm of the likelihood function to obtain the log-likelihood function:

[0161]

[0162] By maximizing the log-likelihood function, the conditional probability table for each node is obtained. This table details the probability distribution of each node under all possible combinations of its parent node's values. Combining the trained conditional probability table with the initial network structure forms the trained Bayesian network model. The model's performance is validated through cross-validation to ensure that the accuracy of anomaly inference on the validation set meets the preset standard.

[0163] S36. Select process parameters with mutual information values ​​greater than the third preset threshold as key influencing factors. Input the key influencing factors and abnormal generation rate data into a Bayesian network model. Calculate the posterior probability distribution through probabilistic inference to infer the root cause of the abnormal generation rate. Among them, the key influencing factors include heater surface temperature and oxygen concentration.

[0164] Specifically, this step screens key influencing factors based on mutual information values ​​and uses Bayesian network probabilistic inference to pinpoint the root cause of abnormal generation rates, thus transforming data correlation into causal diagnosis. The screening of key influencing factors uses mutual information values ​​as the core indicator. The third preset threshold is determined by statistically analyzing the distribution of mutual information values ​​for all process parameters, and can be taken as the critical value of the top 20% of mutual information values ​​to ensure that the screened factors have significant correlation strength. Based on the process mechanism and mutual information calculation results, the mutual information values ​​of heater surface temperature and oxygen concentration are consistently higher than other parameters, and are therefore identified as key influencing factors.

[0165] The inference process first inputs the real-time collected discrete values ​​of key influencing factors and anomaly markers (1 indicates anomaly) into the trained Bayesian network model. The model then performs probabilistic inference based on conditional probability tables and Bayes' theorem. The core of Bayesian inference is calculating the posterior probability distribution of each possible cause (external event and hidden node state) given the evidence (key factor state and anomalous state), as shown in the formula:

[0166]

[0167] In the formula, P(C|E) is the posterior probability of cause C under evidence E, P(E|C) is the conditional probability of evidence E given cause C, P(C) is the prior probability of cause C, and P(E) is the marginal probability of evidence E.

[0168] Taking an abnormal generation rate with high heater surface temperature and normal oxygen concentration as an example, the reasoning process is as follows: The model calls the conditional probability table to calculate the posterior probability of possible causes such as "local overheating = yes", "oxygen leakage = no", and "oxidation reactivity = high". Among these, the posterior probability of "local overheating = yes" is significantly higher than that of other causes, therefore the root cause is inferred to be local overheating. If the key factors are high oxygen concentration and normal heater surface temperature, then the posterior probability of "oxygen leakage = yes" is the highest, and it is determined to be the root cause. When the differences in the posterior probabilities of multiple causes are small, the model further outputs the probability ranking of each cause and combines the diagnostic results of similar historical cases to assist in the judgment. The inference results are fed back to the central control system in real time, and at the same time, corresponding process adjustment suggestions are triggered, forming a closed-loop response of diagnosis and control.

[0169] In one alternative embodiment, after the suppression control operation is triggered based on a root cause, the following steps are further included:

[0170] S41. Based on real-time process parameter data and real-time instantaneous generation rate data of key oxidation byproducts, construct a state space including temperature deviation, pressure fluctuation, oxygen concentration change rate and generation rate trend vector; define an action space including PID controller proportional gain adjustment, integral time adjustment, derivative time adjustment and alarm threshold correction value.

[0171] Specifically, this step constructs a core interactive framework for reinforcement learning by quantifying process status and controllable operation, providing a structured input-output carrier for intelligent decision-making.

[0172] The state space represents the real-time operating state of the process in a multi-dimensional vector form, integrating parameter deviations, dynamic change rates, and trend characteristics. Temperature deviation is the normalized difference between the actual heater surface temperature and the process setpoint, calculated using the following formula:

[0173]

[0174] in, Represents temperature deviation. Represents real-time temperature. This represents the optimal setting value. This represents the upper limit of the safe temperature. This represents the lower limit of the safe temperature range.

[0175] Pressure fluctuations are calculated using the coefficient of variation of the tank pressure within the sliding window, and the formula is as follows:

[0176]

[0177] in, Represents pressure fluctuations. The standard deviation of the pressure within the sliding window. The average pressure within the sliding window is divided into different levels by equal-frequency discretization.

[0178] The rate of change in oxygen concentration is calculated based on the first derivative of the exponentially weighted moving average, using the following formula:

[0179]

[0180] in, This represents the rate of change in oxygen concentration. This represents the smoothed oxygen concentration at the current moment. This represents the oxygen concentration after smoothing from the previous sampling time. The sampling interval is represented by multiple levels that are discretized according to the direction and magnitude of change.

[0181] The generation rate trend vector is generated by extracting the normalized generation rate at multiple recent time points, and then using linear fitting to obtain the slope and goodness of fit to construct a two-dimensional feature. Both the slope and goodness of fit are classified into levels according to a set rule.

[0182] The final state vector format is ,in Represents the slope of the generation rate trend. It represents the goodness of fit of the generation rate trend. By discretizing, continuous variables are mapped to a finite set of states, balancing representation accuracy and computational efficiency.

[0183] The action space focuses on directly executable control parameter adjustments, covering core PID control parameters and safety threshold corrections. The PID proportional gain adjustment is a relative adjustment based on the current proportional gain value; the PID integral time adjustment is an absolute adjustment in time units; the PID derivative time adjustment is an absolute adjustment in time units; and the alarm threshold correction value is a relative correction for oxygen concentration and temperature alarm thresholds. The action vector format is as follows: ,in This represents the PID proportional gain adjustment amount. This represents the PID integral time adjustment. This represents the PID derivative time adjustment. This represents the alarm threshold correction value. Each action is accompanied by process constraint verification, and invalid actions are automatically replaced with zero adjustment.

[0184] S42. Based on the current state vector in the state space and the set of candidate actions in the action space, calculate the instantaneous reward value for each action through a reward function; wherein, the reward function simultaneously considers two optimization objectives: the rate of decrease in generation speed and the energy saving rate.

[0185] Specifically, this step designs a reward function based on a "reward-constraint" composite structure, while optimizing the byproduct suppression effect and energy consumption cost, and constructs it with reference to a knowledge-enhanced reinforcement learning framework.

[0186] The generation rate decrease reward is used to quantify the inhibitory effect of an action on byproduct generation; the calculation formula is as follows:

[0187]

[0188] in, Represents a reward for decreasing generation rate. The weighting coefficient representing the reward for decreasing generation rate. Represents the normalized generation rate at the current moment. Represents the normalized generation rate after the action is performed. The maximum value representing the absolute value of the historical generation rate is awarded a positive reward only when the rate decreases. The energy saving reward is related to the change in heating power, calculated using the following formula:

[0189]

[0190] in, Represents energy-saving rewards, The weighting coefficient representing the energy saving reward. Represents the heating power at the current moment. This represents the heating power after the action is performed. This represents the rated heating power; a positive reward is given when energy consumption is reduced.

[0191] To ensure operational safety and stability, three types of penalty terms are introduced. The safety boundary penalty constrains process parameters within a safe range; a fixed penalty is applied when temperature or pressure exceeds the safe range, and a linear penalty is applied as the temperature approaches the boundary. The operational stability penalty penalizes drastic adjustments to PID parameters; the calculation formula is as follows:

[0192]

[0193] in, This represents a penalty for operational stability. This represents the penalty coefficient, used to prevent oscillations in control actions. The mass conservation penalty is calculated based on material balance deviation, using the following formula:

[0194]

[0195] in, Represents the penalty for conservation of mass. Represents the penalty coefficient. This represents the material balance deviation, ensuring the stability of the process system.

[0196] The formula for calculating the overall reward value is:

[0197]

[0198] in, Represents the instant reward value. Represents the penalty at the safety boundary. Weighting coefficient. , It can be dynamically adjusted to increase the generation rate when it is abnormal. Weighting, which increases when energy consumption exceeds the limit. Weights enable dynamic trade-offs among multiple objectives.

[0199] S43. Input the current state vector, the selected action vector, and the corresponding immediate reward value into the Q-learning algorithm, and update the state-action value function using the temporal difference update rule; the update formula for the temporal difference update rule is:

[0200]

[0201] in, Indicates taking an action in state s The expected cumulative reward, For learning rate, For instant reward value, As a discount factor, To perform the action The next state after transitioning to Indicates the next state The maximum expected reward.

[0202] Specifically, a Q-table corresponding to the size of the state space and action space is constructed, with initial values ​​set to non-zero values ​​to avoid initial exploration bias towards zero actions. The state-action mapping maps real-time state vectors and action vectors to Q-table coordinates through discretization indexes, establishing a correspondence between states, actions, and Q-table values.

[0203] In the temporal differential update formula, Q(s,a) represents the expected cumulative reward for taking action a in state s. The learning rate is used to control the update magnitude. Represents the instant reward value. The s' represents the discount factor used to balance near-term and long-term rewards, and s' represents the next state transitioned to after performing action a. This represents the maximum expected reward among all possible actions in the next state s'. The learning rate balances learning speed and stability and can be adjusted by decaying according to changes in the Q-value. The discount factor prioritizes near-term rewards while retaining long-term gains. After executing action a in the next state s', new process data is collected after a sampling interval; if data is missing, forward predictions are used to temporarily fill the gaps.

[0204] To address the issue of sample correlation, an experience pool mechanism is introduced. The combination of state, action, immediate reward value, and the next state generated from each interaction is stored in the experience pool. During each update, samples are randomly drawn from the experience pool to update the Q-table in batches, improving training stability.

[0205] S44. Based on the state-action value function, select the optimal control parameters to adjust the action through an ε-greedy strategy.

[0206] Specifically, the dynamic exploration rate adopts an exponential decay strategy, and the calculation formula is as follows:

[0207]

[0208] in, The exploration rate at time t, Represents the minimum exploration rate. Let represent the initial exploration rate, λ represent the decay coefficient, and t represent the number of iterations. A higher exploration rate is used initially to fully explore the action space. As iterations progress, the exploration rate gradually decays to a minimum, and a small amount of exploration is retained during the stable period to cope with environmental changes.

[0209] The action selection process begins with probability assessment, generating a random number and comparing it with the current exploration rate. If the random number is greater than the exploration rate, the "utilize" operation is executed, querying the Q-table row corresponding to the current state and selecting the action with the largest Q-value; if multiple optimal actions exist, one is randomly selected. If the random number is less than or equal to the exploration rate, the "explore" operation is executed, randomly selecting an action from the action space, but invalid actions that violate process constraints must be filtered out. When the generation rate becomes severely abnormal, a forced switch to "conservative utilization" mode is initiated, reducing the exploration rate and prioritizing actions with higher historical correction success rates to ensure process safety.

[0210] Policy convergence is determined by combining Q-value changes and reward stability. Policy convergence is determined when the maximum change in Q-value over multiple consecutive steps is within a small range, and the average reward value remains stable within a set interval. After policy convergence, the current Q-table is fixed as the control model, and subsequent updates are only performed through online fine-tuning to maintain the model's adaptability.

[0211] The aforementioned intelligent control method for marigold extraction establishes an accurate calibration model offline between the concentration and spectral characteristics of key oxidative byproducts, and applies this model to an online monitoring system to acquire byproduct concentration data in real time. Based on the real-time concentration data, the instantaneous generation rate is calculated using time-series differential calculations. This generation rate is then intelligently correlated with extraction process parameters to accurately infer the root cause of byproduct generation. Finally, targeted inhibition operations are automatically triggered based on the diagnostic results, and the control effect is verified by real-time monitoring of changes in the generation rate. This achieves complete closed-loop control from byproduct monitoring and tracing to inhibition, effectively solving the lag problem of traditional offline analysis and significantly improving the quality control level and product stability of the marigold extraction process.

[0212] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0213] Based on the same inventive concept, this application also provides an apparatus for implementing the intelligent control method for the marigold extraction process described above. The solution provided by this apparatus is similar to the solution described in the above method; therefore, the specific limitations in one or more embodiments of the intelligent control apparatus for the marigold extraction process provided below can be found in the limitations of the intelligent control method for the marigold extraction process described above, and will not be repeated here.

[0214] In one exemplary embodiment, such as Figure 3 As shown, an intelligent control device 30 for marigold extraction process is provided to implement the methods in the above-described method embodiments. The device includes:

[0215] The feature association modeling module 31 is used to determine the concentration values ​​of key oxidation byproducts in each sample of a standard sample set containing different concentrations of 5,6-epoxylutein, and denoted as the key oxidation byproduct concentration values ​​of the sample; it collects the sample spectral data of each sample in the standard sample set, and establishes a calibration model between spectral features and byproduct concentrations based on the sample spectral data and the key oxidation byproduct concentration values ​​of the sample using chemometric methods; wherein, the standard sample set is prepared in an offline accelerated oxidation experiment, and experiments are conducted on the key oxidation byproducts generated during the extraction of marigolds.

[0216] The real-time feature analysis module 32 is used to collect real-time spectral data during the marigold extraction process, input the real-time spectral data into the calibration model, and output the current concentration value of key oxidation byproducts.

[0217] The rate dynamic extrapolation module 33 is used to calculate the instantaneous generation rate of key oxidation byproducts based on the current concentration value through time-series differentiation; wherein, the instantaneous generation rate represents the amount of concentration change of key oxidation byproducts per unit time.

[0218] The process correlation inference module 34 is used to monitor the process parameters of marigold extraction, perform correlation analysis between the instantaneous generation rate and the process parameters, and obtain the correlation results; based on the correlation results, it intelligently infers the root cause of the generation of key oxidation byproducts.

[0219] The control verification execution module 35 is used to trigger the suppression control operation based on the root cause and monitor the change in instantaneous generation rate in real time to verify the control effect.

[0220] Embodiments of this application also provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the aforementioned method embodiments.

[0221] Embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method embodiments.

[0222] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0223] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. A method for intelligent control of the extraction process of marigold, characterized by, The method comprises: S1, determining the concentration value of the key oxidative byproduct in each sample of a standard sample set comprising different concentrations of 5,6-epoxy zeaxanthin, denoted as sample key oxidative byproduct concentration value; collecting sample spectrum data of each sample in the standard sample set, and establishing a calibration model between spectral characteristics and byproduct concentration by chemometrics method based on the sample spectrum data and the sample key oxidative byproduct concentration value; wherein the standard sample set is prepared in an offline accelerated oxidation experiment, and the key oxidative byproduct generated in the marigold extraction process is subjected to experiment; S2, collecting real-time spectrum data of the marigold extraction process, inputting the real-time spectrum data into the calibration model, and outputting the current concentration value of the key oxidative byproduct; S3, based on the current concentration value, calculating the instantaneous generation rate of the key oxidative byproduct by time series differentiation; wherein the instantaneous generation rate represents the concentration change amount of the key oxidative byproduct per unit time; S4, monitoring the process parameters of the marigold extraction process, correlating the instantaneous generation rate with the process parameters, and obtaining a correlation result; based on the correlation result, intelligently inferring the root cause of the generation of the key oxidative byproduct; S5, triggering inhibition control operation according to the root cause, and monitoring the change of the instantaneous generation rate in real time to verify the control effect.

2. The method of claim 1, wherein, The calibration model between spectral characteristics and byproduct concentration is established by chemometrics method based on the sample spectrum data and the sample key oxidative byproduct concentration value, comprising: S11, based on the sample spectrum data, calculating the Pearson correlation coefficient between the spectral intensity at each wavelength or wave number and the sample key oxidative byproduct concentration value, screening out the feature spectral region with an absolute value of the Pearson correlation coefficient greater than a first preset threshold, and obtaining a candidate feature set; S12, performing principal component analysis on the candidate feature set, extracting the principal component with the highest contribution rate as the comprehensive spectral feature; dividing the comprehensive spectral feature and the sample key oxidative byproduct concentration value into a training set and a validation set; based on the training set, using a partial least squares regression method to construct an initial calibration model; S13, based on the validation set, calculating the determination coefficient of the initial calibration model; if the determination coefficient is less than a second preset threshold, returning to S11 to re-optimize feature selection, otherwise outputting a final calibration model.

3. The method of claim 1, wherein, The instantaneous generation rate of the key oxidative byproduct is calculated by time series differentiation based on the current concentration value, comprising: S21, extracting continuous time series data from the current concentration value output in real time, and performing exponential weighted moving average processing on the time series data to obtain a smoothed concentration sequence; S22, based on the smoothed concentration sequence, calculating the generation rate at each time point by central difference method; the calculation formula of the generation rate is: wherein, represents the instantaneous generation rate at time t, represents the current concentration value at time t, represents the sampling time interval; S23, normalizing the generation rate to obtain a normalized generation rate as the instantaneous generation rate.

4. The method of claim 1, wherein, The process parameters of the monitoring process of the marigold extraction process are analyzed, and the correlation result is obtained by correlating the instantaneous generation rate with the process parameters; Based on the correlation result, the root cause of the generation of the key oxidation byproduct is intelligently inferred, including: S31, collect real-time process parameter data to form a multivariate time series data set; by discretizing the continuous variables in the multivariate time series data set, each variable is divided into multiple intervals to obtain a discretized multivariate data set; wherein the real-time process parameter data includes heater surface temperature, tank pressure, oxygen concentration and stirring speed; S32, based on the time series of the instantaneous generation rate and the discretized multivariate data set, the mutual information value between the instantaneous generation rate and each process parameter is calculated to represent the non-linear correlation strength; the calculation formula of the mutual information value is: in, This represents a discretized set of process parameter values. Represents a discretized set of instantaneous generation rate values; Indicates instantaneous generation rate With process parameters Mutual information value; Indicates the values ​​of process parameters With the instantaneous generation rate value The joint probability distribution is obtained by statistically analyzing the discrete multivariate dataset. and The frequencies falling within the corresponding discrete intervals are estimated simultaneously. and These represent the values ​​of the process parameters. and instantaneous generation rate value The marginal probability distribution is estimated by statistically analyzing the frequency of each variable in the discretized multivariate dataset falling into the corresponding discrete interval; S33, extract a historical multivariate time series data set containing historical process parameters, external events and generation rate abnormal data from a historical database; wherein the external events include local overheating and oxygen leakage, and the generation rate abnormal data is determined based on the historical instantaneous generation rate exceeding a preset threshold; S34, based on the variable relationship of the historical process parameters, the external events and the generation rate abnormal data, an initial structure of a Bayesian network containing visible nodes and hidden nodes is constructed; S35, using the historical multivariate time series data set, the initial structure of the Bayesian network is parameter trained by maximum likelihood estimation method to obtain a trained conditional probability table; based on the conditional probability table and the initial structure of the Bayesian network, a trained Bayesian network model is obtained; S36, screen the process parameters with mutual information value greater than a third preset threshold as key influence factors, input the key influence factors and generation rate abnormal data into the Bayesian network model, calculate the posterior probability distribution by probability reasoning, and infer the root cause of the generation rate abnormality; wherein the key influence factors include heater surface temperature and oxygen concentration.

5. The method according to any one of claims 1 to 4, characterized in that, After triggering the inhibition control operation according to the root cause, it further includes: S41, based on real-time process parameter data and real-time instantaneous generation rate data of the key oxidation byproduct, a state space including temperature deviation, pressure fluctuation, oxygen concentration change rate and generation rate trend vector is constructed; define an action space including PID controller proportional gain adjustment, integral time adjustment, differential time adjustment and alarm threshold correction value; S42, based on the current state vector of the state space and the candidate action set of the action space, the immediate reward value of each action is calculated by a reward function; wherein the reward function considers both the generation rate drop amplitude and the energy saving rate as two optimization objectives; S43, input the current state vector of the state space, the selected action vector and the corresponding immediate reward value into the Q-learning algorithm, and update the state-action value function by the time difference update rule; the update formula of the time difference update rule is: wherein, represents the expected cumulative reward of taking action in state s, is the learning rate, is the immediate reward value, is the discount factor, is the next state transitioned to after performing action , and represents the maximum expected reward in the next state . S44. Selecting, based on the state-action value function, an optimal control parameter adjustment action by an ε-greedy policy.

6. A device for intelligent control of the extraction process of marigold, for implementing the method according to any one of claims 1 to 5, characterized in that, The device comprises: The feature correlation modeling module is configured to measure the concentration value of the key oxidative byproduct in each sample of a standard sample set containing different concentrations of 5,6-epoxy zeaxanthin, denoted as a sample key oxidative byproduct concentration value; collect sample spectral data of each sample in the standard sample set, and establish a calibration model between spectral features and byproduct concentrations by chemometrics based on the sample spectral data and the sample key oxidative byproduct concentration value; wherein the standard sample set is prepared in an offline accelerated oxidation experiment by conducting experiments on the key oxidative byproducts generated during the marigold extraction process; The real-time feature analysis module is configured to collect real-time spectral data of the marigold extraction process, input the real-time spectral data into the calibration model, and output the current concentration value of the key oxidative byproduct; The rate dynamic deduction module is configured to calculate the instantaneous generation rate of the key oxidative byproduct by time series differentiation based on the current concentration value; wherein the instantaneous generation rate represents the concentration change amount of the key oxidative byproduct per unit time; The process correlation inference module is configured to monitor the process parameters of the marigold extraction process, perform correlation analysis on the instantaneous generation rate and the process parameters to obtain a correlation result, and intelligently infer the root cause of the generation of the key oxidative byproduct based on the correlation result; The control verification execution module is configured to trigger a suppression control operation according to the root cause and monitor the change of the instantaneous generation rate in real time to verify the control effect. 7.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-6 when the computer program is executed by the processor. The processor executes the computer program to realize the method of any one of claims 1 to 5.

8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Storage device, hydrogen peroxide extraction process control method, device and equipment

    CN114906817A

  • Mornidazole oxidation impurity preparation process supervision system and preparation method thereof

    CN118454600A

  • Data processing method, device and equipment based on artificial intelligence technology and data model

    CN119598410A

  • Bi-BDO fermentation pH dissolved oxygen dynamic optimization method based on online Raman spectrum

    CN121034441A

  • System and method for fault diagnosis based on causal graph model of production process

    CN121092914A