Process control method based on intelligent pulp quality prediction
By integrating multimodal data and extracting features, combined with multi-objective optimization algorithms, intelligent quality prediction and control of the pulp production process were achieved, solving the problem of quality control in traditional pulp production and improving batch stability and operating efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 福建省尤溪永丰茂纸业有限公司
- Filing Date
- 2025-11-13
- Publication Date
- 2026-05-01
AI Technical Summary
In traditional pulp production, quality control struggles to cope with raw material fluctuations and process disturbances, resulting in large fluctuations in key quality indicators and limited batch stability. Furthermore, existing data-driven methods are insufficient to meet the multi-objective and dynamic optimization needs of complex processes.
The process control method based on intelligent pulp quality prediction uses multimodal data integration, feature extraction and multi-objective optimization to dynamically control pulp quality in real time, including data preprocessing, feature engineering, short-term prediction and control strategy generation.
It has enabled accurate prediction and intelligent adjustment of pulp quality, improved the proactive control capability of the process, reduced the probability of anomalies, increased the batch product compliance rate and operational performance, and achieved synergistic optimization of process, energy consumption and operation and maintenance.
Smart Images

Figure CN121119852B_ABST
Abstract
Description
Process Control Method Based on Intelligent Pulp Quality Prediction Technical Field
[0001] This invention relates to the field of process control, and in particular to a process control method based on intelligent pulp quality prediction. Background Technology
[0002] As the paper industry places higher demands on product quality, consistency, and green production, quality control in pulp manufacturing is becoming increasingly challenging. Traditional pulp production relies on manual experience and static process parameter adjustments, making it difficult to respond promptly to raw material fluctuations, process disturbances, and process nonlinearities. This results in large fluctuations in key quality indicators (such as brightness, viscosity, and Kappa value), limiting batch stability. Meanwhile, the Internet of Things (IoT) sensing capabilities on the production floor have significantly improved, with continuously collected multi-source sensor data, process control logs, and business information becoming increasingly abundant, giving rise to a massive multimodal data stream.
[0003] How to fully utilize these heterogeneous field data to break through the bottlenecks of traditional process control based on post-hoc analysis and achieve predictive, proactive, and collaborative quality regulation has become an important direction for modern paper and pulp mills to improve their intelligent manufacturing capabilities and reduce costs while improving quality. In recent years, data-driven machine learning models have been increasingly widely used in industrial forecasting and process control, but relying solely on a single prediction point is still insufficient to meet the actual needs of complex processes with multiple objectives and dynamic optimization. Summary of the Invention
[0004] To address the aforementioned problems, the present invention aims to provide a process control method based on intelligent pulp quality prediction, which enhances the proactive control capability of the process and reduces the probability of anomalies.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] The process control method based on intelligent pulp quality prediction includes the following steps:
[0007] S1: Acquire on-site physical quantities and business data, and preprocess them to obtain a multimodal data stream;
[0008] S2: Based on the multimodal data stream, perform feature engineering to obtain multimodal features, and obtain the feature matrix through feature filtering and dimensionality reduction;
[0009] S3: Based on the feature matrix, predict the short-term trend of key quality and obtain the probability of meeting the standard; and for batches that are about to be completed or are in progress, predict sensitive variables and obtain a list of sensitive variables.
[0010] S4: Based on the probability of achieving the target and the list of sensitive variables, and based on a multi-objective optimization algorithm, obtain a candidate structured control strategy package;
[0011] S5: Perform optimal process control based on the control strategy package and the current process state.
[0012] Furthermore, on-site physical quantities and operational data include process variables, online quality signals, visual images, equipment status, and production plans and raw material ledgers.
[0013] Further preprocessing includes time alignment, outlier identification, and noise suppression, as detailed below:
[0014] All data are resampled to the target time grid while retaining the original resolution index; a step holding or interval assignment strategy is adopted for slow variables; sliding window aggregation is used to preserve the fidelity of fast variables; a causal delay library is established by combining process knowledge, and the library is used to perform lag alignment and effective window truncation on upstream variables, so that the input is more physically consistent with the target quality in time.
[0015] Outlier identification employs multiple parallel strategies: hard thresholding based on physical constraints, sliding window z-score quantiles based on statistics, jump detection based on continuity and derivative, and cross-variable consistency verification; missing bars in spectral and image data are masked at the sample level.
[0016] For process variables with high-frequency noise, low-pass filtering or wavelet denoising is used; for signals with periodic disturbances, periodic components are extracted and subtracted to obtain the net process response; dark current subtraction, whiteboard normalization and drift calibration are performed on spectral data; illumination equalization, background separation and artifact removal are performed on image data; for vibration and current signals, the energy ratio of the original domain and specific frequency bands is preserved to accommodate both diagnostic and control applications.
[0017] Unify the units, dimensions and calibration coefficients of each data source to the factory standard, perform zero-point correction on key measurement points and record the calibration time, maintain measurement reliability labels, normalize the context variables of chemical concentration and raw material moisture content, generate equivalent load and equivalent reactivity physical consistency index, and reduce false differences between different raw material batches.
[0018] A unified data object is established, with time as the primary key. Process variables, online quality estimation, equipment status, spectral and image summary features of the same time slice are bound to the business context. After the above preprocessing is completed, a multimodal data stream is formed that can be used for subsequent modeling and control.
[0019] Furthermore, based on the multimodal data stream, feature engineering is performed to obtain multimodal features, including process time-series features, spectral channel features, image channel features, device health features, contextual features, and cross-modal interaction features, as detailed below:
[0020] Time-series process characteristics: For high-frequency process variables such as temperature, pressure, flow rate, liquid ratio, dosing curve, residence time, pH and ORP, statistical and dynamic features are extracted based on sliding time windows, including mean, standard deviation, extreme values, quantiles, skewness, kurtosis, trend slope, acceleration, volatility and power ratio of short and long windows, lag features are constructed to reflect causal delay, and exponentially weighted moving statistics representing process inertia are also used.
[0021] Spectral channel feature construction: After completing dark current subtraction, whiteboard calibration, drift correction and smoothing, the spectral data are used to extract features that can reflect chemical composition and structural information, including the integrated intensity of the selected band, characteristic peak position and peak area, peak width, peak ratio, second-order guide spectral features, principal component fraction, independent component fraction, and chemical indicators related to residual lignin, hemicellulose and cellulose crystallinity.
[0022] Image channel feature construction: After illumination equalization, background removal and noise suppression through an online vision system, fiber and fine fiber segmentation and instantiation statistics are performed to construct fiber length distribution, width distribution, aspect ratio, curl, interlacing degree, proportion of fine fibers, proportion of aggregates, and porosity morphological features. At the same time, texture and structure descriptors are extracted. For bleaching visual monitoring, the mean and uniformity of brightness, color deviation index and surface defect density are extracted.
[0023] Equipment health feature construction: Equipment health features and control reachability features are constructed from vibration, bearing temperature, motor current, valve opening and feedback, execution lag and dead zone signals, including frequency band energy ratio, kurtosis, spectral peak drift, temperature rise rate, current imbalance factor, valve position command and feedback difference statistics, response time distribution, and saturation duty cycle.
[0024] Contextual feature construction: Introduce raw material and business context: tree species and core-to-edge ratio, raw material moisture content and density, wood chip gradation, chemical purity, ambient temperature and humidity, order quality grade and scheduling constraints, and adopt range assignment and change point strategy;
[0025] Construction of cross-modal interaction features: Based on mechanism and causal prior, a cross-modal feature set is constructed, and a mechanism proxy is formed using material balance and reaction kinetics as a robust structural feature;
[0026] Finally, all candidate features are tagged with their source, processing pipeline, and version, and their physical meaning and expected direction are recorded.
[0027] Furthermore, through feature selection and dimensionality reduction, the feature matrix is obtained, as follows:
[0028] A multi-criteria initial screening method is adopted, including correlation threshold screening, multicollinearity detection, stability selection, and physical rationality.
[0029] Using gradient-based attribution methods and response analysis based on partial dependence local effects, a preliminary importance ranking is formed; combining causal priors and intervention rationality verification, spurious correlation features from common causes or sampling biases are eliminated, and priority is given to retaining those that can be controlled by process variables.
[0030] To address high-dimensional multimodal features and ensure both real-time performance and robustness, a hierarchical dimensionality reduction strategy is adopted.
[0031] Linear dimensionality reduction is applied to the statistical features of highly collinear processes, and the number of principal components retained is limited to achieve a specified cumulative variance explanation rate, while the invertible mapping matrix from the original space to the principal components is preserved.
[0032] Domain-guided subspace compression of spectral and image summarization vectors is preferred over pure black-box compression;
[0033] Grouping, aggregation, and orthogonalization are used to process equipment health and actuator dynamic characteristics to maintain the independence of control feasibility signals; the screened and dimensionality-reduced multimodal features are organized into a unified feature matrix by time slice or batch.
[0034] Furthermore, based on the feature matrix, the short-term trend of key quality is predicted to obtain the probability of meeting the standard, as follows:
[0035] Based on the feature matrix, using the multimodal features of the sliding time window as input, the aligned feature fragments of the past w minutes are spliced into time series samples to predict the trajectory of key quality indicators in the next h minutes.
[0036] A short-term prediction model is constructed using Temporal CNN, and the output includes both point prediction and interval estimation. Multiple confidence quantiles are directly given through quantile regression.
[0037] For each target indicator, a process standard range is maintained, and the conditional distribution obtained from short-term forecasts is used to calculate the instantaneous probability of achieving the target and the worst probability of achieving the target within the next h minutes.
[0038] Furthermore, the short-term prediction model uses 1D dilated convolutions to capture long dependencies and maintain causality, residual blocks to improve stability and trainability, and hierarchical accumulation to cover the receptive field window length, as detailed below:
[0039] Input sequence 1D dilated convolution:
[0040] ;
[0041] in, For the corresponding number The weight matrix for each time offset; b is the bias vector; h tK represents the output of the causal convolution at time step t; K is the kernel length; T is the total number of time steps.
[0042] Expand the receptive field using the void ratio d while keeping the parameters controllable:
[0043] ;
[0044] For L layers, the core length K of each layer, and the expansion rate :
[0045] ;
[0046] Where RF represents the receptive field;
[0047] Residual block given input channel C in Hidden channel C, two layers of causal dilated convolutions with normalization and activation:
[0048] ;
[0049] ;
[0050] ;
[0051] ;
[0052] Where U represents the feature sequence entering the current residual block; This is the one-dimensional causal dilated convolution operator for the first layer; For normalization layer; Z is a non-linear activation function. (1) The activation output after the first convolution in the residual block; S is the residual branch; V is the final output of the residual block; It is a 1×1 causal convolution; Z is the one-dimensional causal dilated convolution operator for the second layer; (2) This is the activation output after the second convolution in the residual block;
[0053] No. Each residual block uses the expansion rate Nucleus length K, forming a sequence:
[0054] ;
[0055] Where Stem(·) is the input projection layer; For the first A residual block, core length K, void ratio L represents the number of residual block stacking layers;
[0056] From the last feature U (L+1)The output is mapped laterally to the future H steps via causal convolution. ;
[0057] Quantile regression head, for each quantile τ∈Q:
[0058] ;
[0059] in, and These represent the weights and biases of the quantile regression head, respectively; Reshape indicates the rearrangement operation. The tensor of the predicted quantile estimates; u T This is the global feature vector output by the timing encoder at the last moment;
[0060] Probabilistic regression head:
[0061] ;
[0062] Among them, W μ ,b μ These are the weights and biases for mean prediction, respectively; W σ ,b σ These are the weights and biases for predicting the standard deviation, respectively. This is the default value; , These are the predicted conditional mean matrix and standard deviation matrix, respectively; softplus is the non-negative activation function.
[0063] For a single quality threshold, with an upper limit U or an interval [L, U], the probability of meeting the standard is:
[0064] ;
[0065] Where y represents the actual value of a certain indicator at a certain future step; Let be the cumulative distribution function of the standard normal distribution; P(y≤U) is the one-sided probability of achieving the standard; P(L≤y≤U) is the two-sided probability of achieving the standard.
[0066] Further obtain the worst-case joint achievement probability over the entire prediction time domain. :
[0067]
[0068] in, This means calculating the joint probability of achieving the target for each time step k in the prediction time domain and taking the minimum value among them; This is the product of the probabilities of achieving the target for all indicators; m is the number of indicators monitored in parallel. This represents the probability that index j will reach its process qualification range at the k-th time step in the future.
[0069] Furthermore, for batches that are about to be completed or are in progress, the following measures are taken to predict sensitive variables: In the batch scenario, the known operating conditions, cumulative dose, stage labels, and intermediate sampling results of the current batch are summarized into a batch feature view to predict the final Kappa, whiteness, viscosity, yield, and drug consumption. The prediction adopts a multi-task framework that shares the underlying feature representation, and after obtaining new online observations, the prediction is updated in a loop using a filtering method to dynamically converge uncertainty and improve the probability of achieving the target. Each prediction also generates a structured list of sensitive variables, including: variable name and level, directionality, relative contribution, adjustment accessibility, estimated effect, and interactive reminders.
[0070] Furthermore, based on the achievement probability and the list of sensitive variables, and using a multi-objective optimization algorithm, a candidate structured control strategy package is obtained, as follows:
[0071] Let the set of controllable variables be U = {u1, ..., u2}. p}, the decision vector is Δu;
[0072] constraint:
[0073] ;
[0074] Among them, u i Let i be the i-th controllable variable; This is the current setting value; These are the upper and lower limits of a controllable variable; This represents the upper limit of the rate of change. To control the step size; A and b are the linear coupling constraint matrix and upper bound vector, respectively; device_health i For device health status trigger conditions; S i Let i be the set of reachable increments for the i-th actuator in the current health state; This represents the increment of the operation on the i-th variable during this period; A vector consisting of the current setpoints of all control variables;
[0075] Define a multi-objective optimization problem using the set of indicators K:
[0076] ;
[0077] ;
[0078] ;
[0079] ;
[0080] ;
[0081] ;
[0082] in, The feasible region is defined for all constraints; J1, J2, J3, J4, and J5 are the joint achievement probability target, operating cost target, expected quality deviation, risk measure, and execution complexity, respectively; c is the linear cost coefficient vector per unit increment; ρ is the quadratic regularization weight. It is the square of the L2 norm; The set of process indicators to be evaluated; w j For indicator weights; y j (Δu) is the random variable that predicts index j under the current action; This is the deviation loss function; This represents the expectation of the predicted distribution; For the degree of violation; For the conditional risk value, the confidence level α∈(0,1) represents the average degree of violation under the worst-case scenario; λ1,λ group , respectively, represent the sparse and group sparse weights; Δu(g) is the action vector of the variable group g; C is the joint probability function; op The objective function is the operating cost;
[0083] Based on the above objective function and constraints, the NSGA-II algorithm is used to find the optimal solution, which is suitable for black-box probabilistic evaluation and non-convex feasible regions.
[0084] Furthermore, based on the control strategy package and the current process state, optimal process control is performed, as follows:
[0085] Using the action of the strategy package as a reference trajectory, the loop layer issues setpoints for each key controllable quantity and uses a servo controller with feedforward and filtering for tracking;
[0086] Online self-tuning is performed based on closed-loop data and tuning rules, and the servo controller parameters are updated periodically. When a change in gain, time delay drift, or sudden increase in noise level is detected, the controller switches to a robust parameter set and gradually interpolates to the new parameters to avoid disturbances caused by step parameter changes. Anti-saturation protection adopts inversion backflow and conditional integration: when the actuator reaches saturation or is limited by the rate, the integral term is frozen and the saturation error is backflowed to the integral state to prevent the accumulation of integrals from causing a large overshoot after saturation. At the same time, anomaly suppression and short-term weight reduction are used for measurement spikes to ensure that the loop is not pulled by noise. For coupled loops, decoupling compensation matrix or sequential loop priority is used.
[0087] The present invention has the following beneficial effects:
[0088] 1. This invention integrates on-site multimodal data and extracts intelligent features to perform global modeling of factors affecting pulp quality in real time and dynamically. Through short-term accurate quality prediction and batch-level variable sensitivity analysis, the probability of key quality indicators meeting standards can be known in advance, which significantly improves the ability to actively control the process and reduces the probability of anomalies.
[0089] 2. This invention introduces a multi-objective optimization framework. In the control strategy generation stage, it takes into account multiple objectives such as product quality compliance, raw material and energy costs, process deviation risks and operational complexity. It also combines the dynamic constraints of sensitive variables to optimize the space and effectively improves the overall operating performance of each batch and operation cycle through the screening of structured strategy packages. This achieves the coordinated optimization of process, energy consumption and operation and maintenance, reduces ineffective or excessive operation, and ensures equipment and environmental safety.
[0090] 3. This invention enables accurate prediction of key quality parameters of pulp and generation of intelligent adjustment strategies, thereby improving the batch product compliance rate, ensuring efficient use of energy and resources, and helping paper manufacturing enterprises upgrade to intelligent, lean and green production. Attached Figure Description
[0091] Figure 1 is a flowchart of the method of the present invention. Detailed Implementation
[0092] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0093] Referring to Figure 1, in this embodiment, a process control method based on intelligent pulp quality prediction is provided, including the following steps:
[0094] S1: Acquire on-site physical quantities and business data, and preprocess them to obtain a multimodal data stream;
[0095] S2: Based on the multimodal data stream, perform feature engineering to obtain multimodal features, and obtain the feature matrix through feature filtering and dimensionality reduction;
[0096] S3: Based on the feature matrix, predict the short-term trend of key quality and obtain the probability of compliance; and for batches that are about to be completed or are in progress, predict sensitive variables, including Kappa, whiteness, viscosity, fiber morphology grade, yield and chemical consumption, and obtain a list of sensitive variables.
[0097] S4: Based on the probability of achieving the target and the list of sensitive variables, and based on a multi-objective optimization algorithm, obtain a candidate structured control strategy package;
[0098] S5: Perform optimal process control based on the control strategy package and the current process state.
[0099] In this embodiment, the on-site physical quantities and business data include process variables (temperature, pressure, flow rate, liquid ratio, chemical addition curve), online quality signals (near-infrared / visible spectrum, online whiteness, viscosity), visual images (microscope fiber images, online bleaching tower views), equipment status (vibration, temperature rise, motor current), and production plans and raw material ledgers (timber source, moisture content, tree species ratio).
[0100] In this embodiment, preprocessing includes time alignment, outlier identification, and noise suppression, as detailed below:
[0101] All data are resampled to the target time grid (e.g., 1 s or 5 s) while retaining the original resolution index; a step-holding or interval assignment strategy is adopted for slow variables (such as laboratory results and batch metadata); sliding window aggregation (mean, extreme values, quantiles) is used to preserve the fidelity of fast variables (such as process variables and online quality signals); a causal delay library is established by combining process knowledge, such as the typical lag interval of the effect of cooking temperature change on Kappa and the response time of bleaching dosing on whiteness. This library is used to perform lag alignment and effective window truncation on upstream variables, so that the input is more physically consistent with the target quality in time.
[0102] Outlier identification employs multiple parallel strategies: hard thresholds based on physical constraints (e.g., directly identifying outliers when the temperature is <0 ℃ or >250 ℃), sliding window z-score quantiles based on statistics, jump detection based on continuity and derivative, and cross-variable consistency verification (e.g., inconsistency between flow rate and valve position); missing bars in spectral and image data are masked at the sample level.
[0103] For process variables with high-frequency noise, low-pass filtering or wavelet denoising is used; for signals with periodic disturbances, periodic components are extracted and subtracted to obtain the net process response; dark current subtraction, whiteboard normalization and drift calibration are performed on spectral data; illumination equalization, background separation and artifact removal are performed on image data; for vibration and current signals, the energy ratio of the original domain and specific frequency bands is preserved to accommodate both diagnostic and control applications.
[0104] Unify the units, dimensions, and calibration coefficients of all data sources to the factory standard. For example, unify temperature to ℃, pressure to MPa, flow rate to t / h or L / min, and use effective alkali equivalent or active chlorine equivalent for chemical dosing. Perform zero-point correction on key metering points and record the calibration time. Maintain metering reliability labels. Normalize the context variables of chemical concentration and raw material moisture content to generate equivalent load and physical consistency indicators of equivalent reactivity, thereby reducing false differences between different batches of raw materials.
[0105] Establish a unified data object with time as the primary key, and bind process variables, online quality estimation, equipment status, spectral / image summary features and business context (raw material batch number, work order, and section operating status) for the same time slice. After completing the above preprocessing, a multimodal data stream is formed that can be used for subsequent modeling and control.
[0106] In this embodiment, feature engineering is performed based on the multimodal data stream to obtain multimodal features, including process time-series features, spectral channel features, image channel features, device health features, contextual features, and cross-modal interaction features, as detailed below:
[0107] Time-series process characteristics (structured variables): For high-frequency process variables such as temperature, pressure, flow rate, liquid ratio, dosing curve, residence time, pH, and ORP, statistical and dynamic features are extracted based on sliding time windows, including mean, standard deviation, extreme values, quantiles, skewness, kurtosis, trend slope of short and long windows, acceleration (derivative / second derivative), volatility (coefficient of variation), and power ratio (energy percentage within the window). Lag features are constructed to reflect causal delay (such as moving average and moving variance), and exponentially weighted moving average (EWMA) is used to represent process inertia.
[0108] Spectral channel feature construction: After completing dark current subtraction, whiteboard calibration, drift correction and smoothing, the spectral data are used to extract features that can reflect chemical composition and structural information, including the integrated intensity of the selected band, characteristic peak position and peak area, peak width, peak ratio, second-order guide spectral features, principal component fraction, independent component fraction, and chemical indicators related to residual lignin, hemicellulose and cellulose crystallinity.
[0109] Image channel feature construction: After illumination equalization, background removal and noise suppression through an online vision system, fiber and fine fiber segmentation and instantiation statistics are performed to construct fiber length distribution, width distribution, aspect ratio, curl, interlacing degree (node density), fine fiber ratio, aggregate ratio, and porosity proxy morphological features. At the same time, texture and structure descriptors (GLCM contrast, energy, homogeneity, LBP histogram, Gabor energy) are extracted. For bleaching visual monitoring, the mean and uniformity of brightness, color deviation index and surface defect density are extracted.
[0110] Equipment health feature construction: Equipment health features and control reachability features are constructed from vibration, bearing temperature, motor current, valve opening / feedback, execution lag and dead zone signals, including frequency band energy ratio (e.g., 1×, 2×, bearing characteristic frequency band), kurtosis, spectral peak drift, temperature rise rate, current imbalance factor, difference statistics between valve position command and feedback (hysteresis, overshoot, hysteresis), response time distribution, and saturation duty cycle;
[0111] Contextual feature construction: Introduce raw material and business context: tree species and core-to-edge ratio, raw material moisture content and density, wood chip gradation (particle size distribution), chemical purity, ambient temperature and humidity, order quality grade and scheduling constraints, and adopt range assignment and change point strategy;
[0112] Cross-modal interaction feature construction: Based on mechanism and causal priors, cross-modal feature sets are constructed. For example, "spectral residual lignin estimation × cooking temperature trend" is a sensitive combination for predicting Kappa; "fiber fineness ratio × screening pressure fluctuation" is associated with washing and screening load and whiteness fluctuation; "equipment hysteresis factor × control setting change rate" is used to predict control accessibility and error upper limit. Material balance and reaction kinetics are used to form mechanism proxy (such as conversion rate, reaction rate constant approximation, heat transfer coefficient proxy) as robust structural features to help the model maintain interpretation and generalization when distribution drift occurs.
[0113] Finally, all candidate features are labeled with their source, processing pipeline, and version, and their physical meaning and expected direction (positive / negative correlation) are recorded.
[0114] In this embodiment, the feature matrix is obtained through feature filtering and dimensionality reduction, as detailed below:
[0115] A multi-criteria initial screening was adopted, including correlation threshold screening (Pearson / Spearman correlation with the target, mutual information), multicollinearity detection (VIF), stability selection (consistency of importance under different time windows / raw material scenarios), and physical rationality (excluding spurious correlations that violate the mechanism).
[0116] Using gradient-based attribution methods and response analysis based on partial dependence local effects, a preliminary importance ranking is formed; combining causal priors and intervention rationality verification, spurious correlation features from common causes or sampling biases are eliminated, and priority is given to retaining those that can be controlled by process variables.
[0117] To address high-dimensional multimodal features and ensure both real-time performance and robustness, a hierarchical dimensionality reduction strategy is adopted.
[0118] Apply linear dimensionality reduction (such as PCA) to the statistical features of highly collinear processes, and limit the number of principal components retained to achieve a specified cumulative variance explanation rate (such as 90%-95%), while retaining the invertible mapping matrix from the original space to the principal components;
[0119] Domain-guided subspace compression (selecting key bands / key morphological statistics) is preferred over pure black-box compression for spectral and image summarization vectors to enhance interpretability and transferability.
[0120] Grouping, aggregation, and orthogonalization are used to process equipment health and actuator dynamic characteristics, maintaining the independence of control feasibility signals. The filtered and dimensionality-reduced multimodal features are organized into a unified feature matrix X by time slice or batch, including rows corresponding to time slice / batch samples and columns corresponding to versioned feature sets (including principal components and mechanism proxies). Simultaneously, feature metadata (name, source, unit / dimension, processing pipeline, importance, directionality, quality weight) and synchronized target labels and context key-value pairs are output.
[0121] In this embodiment, the short-term trend of key quality is predicted based on the feature matrix to obtain the probability of achieving the target, as detailed below:
[0122] Based on the feature matrix, using the multimodal features of the sliding time window as input, the aligned feature fragments (including process statistics, lag terms, spectral / visual summaries, device health factors and context labels) of the past w minutes are concatenated into a time series sample to predict the trajectory of key quality indicators (such as Kappa prior, online whiteness proxy, and online viscosity proxy) in the next h minutes.
[0123] A short-term prediction model is constructed using Temporal CNN, and the output includes both point prediction and interval estimation. Multiple confidence quantiles are directly given through quantile regression.
[0124] For each target indicator, maintain the process standard range, and use the conditional distribution obtained from short-term forecasts to calculate the instantaneous probability of achieving the target and the worst probability of achieving the target within the next h minutes (considering the extreme values on the trajectory).
[0125] In this embodiment, the short-term prediction model uses 1D dilated convolution to capture long dependencies and maintain causality, residual blocks to improve stability and trainability, and hierarchical accumulation to cover the receptive field window length, as detailed below:
[0126] Input sequence (multimodal features aligned by a sliding time window) 1D dilated convolution:
[0127] ;
[0128] in, For the corresponding number The weight matrix for each time offset; b is the bias vector; h t K represents the output of the causal convolution at time step t; K is the kernel length; T is the total number of time steps.
[0129] Expand the receptive field using the void ratio d while keeping the parameters controllable:
[0130] ;
[0131] For L layers, the core length K of each layer, and the expansion rate :
[0132] ;
[0133] Where RF represents the receptive field;
[0134] Residual block given input channel C in Hidden channel C, two layers of causal dilated convolutions (with the same dilation rate d), with normalization and activation:
[0135] ;
[0136] ;
[0137] ;
[0138] ;
[0139] Where U represents the feature sequence entering the current residual block; This is the one-dimensional causal dilated convolution operator for the first layer; For normalization layer; Z is a non-linear activation function. (1) is the activation output after the first convolution in the residual block; S is the residual branch; V is the final output of the residual block; It is a 1×1 causal convolution; Z is the one-dimensional causal dilated convolution operator for the second layer; (2) This is the activation output after the second convolution in the residual block;
[0140] No. Each residual block uses the expansion rate Nucleus length K, forming a sequence:
[0141] ;
[0142] in, For input projection layer; For the first A residual block, core length K, void ratio L represents the number of residual block stacking layers;
[0143] From the last feature U (L+1) The output is mapped laterally to the future H steps via causal convolution. ;
[0144] Quantile regression head, for each quantile τ∈Q (e.g., {0.1, 0.5, 0.9}):
[0145] ;
[0146] in, and These represent the weights and biases of the quantile regression head, respectively; Reshape indicates the rearrangement operation. The tensor of the predicted quantile estimates; u T This is the global feature vector output by the timing encoder at the last moment;
[0147] Probabilistic regression head:
[0148] ;
[0149] Among them, W μ ,b μ These are the weights and biases for mean prediction, respectively; W σ ,b σ These are the weights and biases for predicting the standard deviation, respectively. This is the default value; , These are the predicted conditional mean matrix and standard deviation matrix, respectively; softplus is the non-negative activation function.
[0150] For a single quality threshold, with an upper limit U or an interval [L, U], the probability of meeting the standard is:
[0151] ;
[0152] Where y represents the actual value of a certain indicator at a certain future step; Let be the cumulative distribution function of the standard normal distribution; P(y≤U) is the one-sided probability of achieving the standard; P(L≤y≤U) is the two-sided probability of achieving the standard.
[0153] Further obtain the worst-case joint achievement probability over the entire prediction time domain. :
[0154]
[0155] in, This means calculating the joint probability of achieving the target for each time step k in the prediction time domain and taking the minimum value among them; This is the product of the probabilities of achieving the target for all indicators; m is the number of indicators monitored in parallel. This represents the probability that index j will reach its process qualification range at the k-th time step in the future.
[0156] In this embodiment, for batches that are about to be completed or are in progress, sensitive variables are predicted as follows: In the batch scenario, the known operating conditions, cumulative dosage, stage labels, and intermediate sampling results of the current batch are summarized into a batch feature view, and the final Kappa, whiteness, viscosity, yield, and drug consumption are predicted. The prediction adopts a multi-task framework to share the underlying feature representation, and after obtaining new online observations, the prediction is updated in a loop using a filtering method to dynamically converge uncertainty and improve the probability of achieving the target. Each prediction generates a structured list of sensitive variables, including: variable name and level (e.g., cooking temperature setting, bleaching drug flow rate, liquid ratio, pH target, etc.), directionality (increasing / decreasing will cause the target to rise or fall), relative contribution (based on a comprehensive score of global and local indicators), adjustment accessibility (upper and lower limits and maximum rate of change), estimated effect (the amount of quality target change caused by unit adjustment), and interactive reminders (e.g., "temperature is more sensitive to Kappa only when liquid ratio > threshold").
[0157] In this embodiment, based on the probability of achieving the target and the list of sensitive variables, a candidate structured control strategy package is obtained using a multi-objective optimization algorithm, as follows:
[0158] Let the set of controllable variables be U = {u1, ..., u2}. p (e.g., temperature setting, dosing flow rate, liquid ratio, pH, valve position target, etc.), the decision vector (incremental representation) is Δu;
[0159] constraint:
[0160] ;
[0161] Among them, u i Let i be the i-th controllable variable; This is the current setting value; These are the upper and lower limits of a controllable variable; This represents the upper limit of the rate of change. To control the step size; A and b are the linear coupling constraint matrix and upper bound vector, respectively; device_health i For device health status trigger conditions; S i Let i be the set of reachable increments for the i-th actuator in the current health state; This represents the increment of the operation on the i-th variable during this period; A vector consisting of the current setpoints of all control variables;
[0162] Define a multi-objective optimization problem (Pareto principle) using a set of indicators K (such as Kappa, whiteness, viscosity, yield, and drug consumption):
[0163] ;
[0164] ;
[0165] ;
[0166] ;
[0167] ;
[0168] ;
[0169] in, The feasible region is defined for all constraints; J1, J2, J3, J4, and J5 are the joint achievement probability target, operating cost target, expected quality deviation, risk measure, and execution complexity, respectively; c is the linear cost coefficient vector per unit increment; ρ is the quadratic regularization weight. It is the square of the L2 norm; The set of process indicators to be evaluated; w j For indicator weights; y j (Δu) is the random variable that predicts index j under the current action; This is the deviation loss function; This represents the expectation of the predicted distribution; To determine the degree of violation (e.g., taking a positive value for each indicator's exceedance and then summing the results using weighted averages); For the conditional risk value, the confidence level α∈(0,1) represents the average degree of violation under the worst-case scenario; λ1,λ group , respectively, represent the sparse and group sparse weights; Δu(g) is the action vector of the variable group g; C is the joint probability function; op The objective function is the operating cost;
[0170] Based on the above objective function and constraints, the NSGA-II algorithm is used to find the optimal solution, which is suitable for black-box probabilistic evaluation and non-convex feasible regions.
[0171] In this embodiment, optimal process control is performed based on the control strategy package and the current process state, as follows:
[0172] Using the action of the strategy package as a reference trajectory, the loop layer issues setpoints for each key controllable quantity and uses a servo controller with feedforward and filtering for tracking;
[0173] The loop layer structure includes: measurement filtering (low-pass / Kalman), error shaping (PI main body with anti-integral saturation, optional D feedforward), setpoint acceleration limiting (jerk limiting to suppress equipment shock) and soft constraint limiting (dynamic upper and lower limits are given by equipment health and safety barriers);
[0174] Online self-tuning is performed based on closed-loop data and tuning rules (such as IMC / λ tuning or least squares identification based on step response), and the servo controller parameters are updated periodically. When gain changes, time delay drift, or sudden increases in noise levels are detected, the controller switches to a robust parameter set and gradually interpolates to the new parameters to avoid disturbances caused by step parameter changes. Anti-saturation protection adopts inversion backflow and conditional integration: when the actuator reaches saturation or is limited by the rate, the integral term is frozen and the saturation error is backflowed to the integral state to prevent large overshoot after saturation due to integral accumulation. At the same time, anomaly suppression (amplitude limiting / median filtering) and short-term weight reduction are used for measurement spikes to ensure that the loop is not pulled by noise. For coupled loops (such as temperature × flow rate, dosing × pH), decoupling compensation matrix or sequential loop priority is used.
[0175] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0176] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.
[0177] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0178] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0179] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A process control method based on intelligent pulp quality prediction, characterized in that, The process includes the following steps: S1: Acquire on-site physical quantities and business data, and preprocess them to obtain a multimodal data stream; S2: Based on the multimodal data stream, perform feature engineering to obtain multimodal features, and obtain a feature matrix through feature filtering and dimensionality reduction; S3: Based on the feature matrix, predict the short-term trend of key quality and obtain the worst-case joint compliance probability. For batches that are about to be completed or are in progress, predict sensitive variables and obtain a list of sensitive variables; S4: Based on the worst-case joint compliance probability and the list of sensitive variables, obtain a candidate structured control strategy package based on a multi-objective optimization algorithm; S5: Perform optimal process control according to the control strategy package and the current process state; The prediction of the short-term trend of key quality based on the feature matrix and the acquisition of the worst-case joint compliance probability are as follows: Based on the feature matrix, using the multimodal features of the sliding time window as input, the aligned feature segments of the past w minutes are spliced into a time series sample to predict the trajectory of key quality indicators in the next h minutes; A short-term prediction model is constructed using Temporal CNN, and the output includes both point prediction and interval estimation. Multiple confidence quantiles are directly given through quantile regression; For each target indicator, the process standard interval is maintained, and the conditional distribution obtained from the short-term prediction is used to calculate the instantaneous compliance probability and the probability of compliance in the next h minutes. Worst-case joint achievement probability within minutes; the short-term prediction model uses 1D dilated convolution to capture long dependencies and maintain causality, residual blocks to improve stability and trainability, and hierarchical accumulation of the receptive field coverage window length. After the input sequence passes through the input projection layer, it passes through L residual blocks in sequence to obtain the final features, and then is laterally mapped to the output u of the next H steps through causal convolution. T ; Quantile regression head, for each quantile τ∈Q: ;in, and These represent the weights and biases of the quantile regression head, respectively; Reshape indicates the rearrangement operation. Tensor for predicted quantile estimates; Probability regression head: Among them, W μ ,b μ These are the weights and biases for mean prediction, respectively; W σ ,b σ These are the weights and biases for predicting the standard deviation, respectively. This is the default value; , Here, represents the predicted conditional mean matrix and standard deviation matrix, respectively; softplus is the non-negative activation function; for a single quality threshold, with an upper limit U′ or an interval [L', U′], the probability of meeting the threshold is: Where y represents the true value of a certain indicator at a future step; Φ(·) is the cumulative distribution function of the standard normal distribution; P(y≤U′) is the one-sided probability of achieving the target; P(L'≤y≤U′) is the two-sided probability of achieving the target; further, the worst-to-joint probability of achieving the target is obtained over the entire prediction time domain. : ;in, This means calculating the joint probability of achieving the target for each time step k in the prediction time domain and taking the minimum value among them; This is the product of the probabilities of achieving the target for all indicators; m is the number of indicators monitored in parallel. This represents the probability that index j will reach its process qualification range at the k-th time step in the future.
2. The process control method based on intelligent pulp quality prediction according to claim 1, characterized in that, The on-site physical quantities and business data include process variables, online quality signals, visual images, equipment status, and production plans and raw material ledgers.
3. The process control method based on intelligent pulp quality prediction according to claim 1, characterized in that, The input sequence, after passing through the input projection layer, is sequentially processed through L residual blocks to obtain the final features, which are then laterally mapped to the output u of the next H steps via causal convolution. T Specifically, the input sequence is as follows: 1D dilated convolution: ;in, For the corresponding number The weight matrix for each time offset; b' is the bias vector; h t The output of the causal convolution at time step t; K is the kernel length; T is the total number of time steps; the receptive field is expanded using the dilation rate d while keeping the number of parameters controllable. For the L-layer, the core length K of each layer, and the expansion rate... : Where RF represents the receptive field; the residual block is given by the input channel C. in Hidden channel C, two layers of causal dilated convolutions with normalization and activation: ; ; ; Where U represents the feature sequence entering the current residual block; Z is the one-dimensional causal dilated convolution operator for the first layer; Norm(·) is the normalization layer; φ(·) is the nonlinear activation function; (1) is the activation output after the first convolution in the residual block; S is the residual branch; V is the final output of the residual block; It is a 1×1 causal convolution; Z is the one-dimensional causal dilated convolution operator for the second layer; (2) The activation output after the second convolution in the residual block; 'Each residual block uses an expansion rate d Nucleus length K, forming a sequence: Where Stem(·) is the input projection layer; For the first There are 1 residual block, kernel length K; L is the number of residual block stacking layers; from the last feature U (L+1) The causal convolution is laterally mapped to the output u of the next H steps. T .
4. The process control method based on intelligent pulp quality prediction according to claim 2, characterized in that, The preprocessing includes time alignment, outlier identification, and noise suppression, specifically as follows: all data are resampled to the target time grid while retaining the original resolution index; a step holding or interval assignment strategy is adopted for slow variables; sliding window aggregation is used to preserve the fidelity of fast variables; a causal delay library is established by combining process knowledge, and the library is used to perform lag alignment and effective window truncation on upstream variables, so that the input is more physically consistent with the target quality in time. Outlier identification employs a multi-strategy parallel approach: hard thresholding based on physical constraints, sliding window z-score quantiles based on statistics, jump detection based on continuity and derivatives, and cross-variable consistency verification; sample-level masking is used for missing segments in spectral and image data; low-pass filtering or wavelet denoising is used for process variables with high-frequency noise; for signals with periodic perturbations, periodic components are extracted and subtracted to obtain the net process response; dark current subtraction, whiteboard normalization, and drift calibration are performed on spectral data; illumination equalization, background separation, and artifact removal are performed on image data; for vibration and current signals, the original domain and specific frequency bands are preserved. The process involves: unifying the units, dimensions, and calibration coefficients of various data sources to factory standards; performing zero-point correction on key measurement points and recording calibration times; maintaining measurement reliability labels; normalizing context variables such as chemical concentration and raw material moisture content; generating equivalent loads and physical consistency indicators for equivalent reactivity; reducing spurious differences between different raw material batches; establishing a unified data object with time as the primary key; binding process variables, online quality estimation, equipment status, spectral and image summary features of the same time slice to the business context; and forming a multimodal data stream for subsequent modeling and control after completing the above preprocessing.
5. The process control method based on intelligent pulp quality prediction according to claim 3, characterized in that, The process involves feature engineering based on multimodal data streams to acquire multimodal features, including process time-series features, spectral channel features, image channel features, equipment health features, contextual features, and cross-modal interaction features. Specifically: Time-series process features: For high-frequency process variables such as temperature, pressure, flow rate, liquid ratio, dosing curve, residence time, pH, and ORP, statistical and dynamic features are extracted based on sliding time windows, including mean, standard deviation, extreme values, quantiles, skewness, kurtosis, trend slope, acceleration, volatility, and power ratio for short and long windows. Lag features are constructed to reflect causal delays, and exponentially weighted moving statistics representing process inertia are also included. Spectral channel feature construction: [The text abruptly ends here, likely due to an incomplete sentence or missing information.] After dark current subtraction, whiteboard calibration, drift correction, and smoothing, the spectral data is used to extract features that reflect chemical composition and structural information. These features include the integrated intensity of the selected band, characteristic peak position and area, peak width, peak ratio, second-order guide spectral features, principal component fraction, independent component fraction, and chemical indicators related to residual lignin, hemicellulose, and cellulose crystallinity. Image channel feature construction: After illumination equalization, background removal, and noise suppression using an online vision system, fiber and fine fiber segmentation and instantiation statistics are performed to construct fiber length distribution, width distribution, aspect ratio, crimp, interlacing degree, fine fiber proportion, aggregate proportion, and porosity. Morphological features are extracted, along with texture and structure descriptors. For visual monitoring of bleaching, mean brightness and uniformity, color deviation index, and surface defect density are extracted. Equipment health features are constructed from vibration, bearing temperature, motor current, valve opening and feedback, execution lag and dead zone signals. These features include frequency band energy ratio, kurtosis, spectral peak drift, temperature rise rate, current imbalance factor, difference statistics between valve position command and feedback, response time distribution, and saturation duty cycle. Contextual feature construction: Introduce raw material and business context: tree species and core-to-edge ratio, raw material moisture content and density, wood chip gradation, chemical purity, ambient temperature and humidity, order quality grade and scheduling constraints, and adopt range assignment and change point strategy; Cross-modal interaction feature construction: Based on mechanism and causal prior, a cross-modal feature set is constructed, and a mechanism proxy is formed using material balance and reaction kinetics as a robust structural feature. Finally, all candidate features are tagged with source, processing pipeline and version, and their physical meaning and expected direction are recorded.
6. The process control method based on intelligent pulp quality prediction according to claim 4, characterized in that, The feature matrix is obtained through feature screening and dimensionality reduction, as follows: A multi-criteria initial screening is employed, including correlation threshold screening, multicollinearity detection, stability selection, and physical rationality. A gradient-based attribution method and response analysis based on partial dependence on local effects are used to form a preliminary importance ranking. Combining causal priors and intervention rationality verification, falsely correlated features from common causes or sampling biases are eliminated, prioritizing features that can be controlled by process variables. For high-dimensional multimodal features, to meet real-time requirements and robustness, a hierarchical dimensionality reduction strategy is adopted: linear dimensionality reduction is applied to highly collinear process statistical features, and the number of principal components retained is limited to achieve a specified cumulative variance explanation rate, while retaining the invertible mapping matrix from the original space to the principal components. Domain-guided subspace compression is preferred over pure black-box compression for spectral and image summary vectors; grouping aggregation and orthogonalization are used for device health and actuator dynamic features to maintain the independence of control feasibility signals; and the filtered and dimensionality-reduced multimodal features are organized into a unified feature matrix by time slice or batch.
7. The process control method based on intelligent pulp quality prediction according to claim 1, characterized in that, For batches that are about to be completed or are in progress, the prediction of sensitive variables is as follows: In the batch scenario, the known operating conditions, cumulative dose, stage labels, and intermediate sampling results of the current batch are summarized into a batch feature view, and the final Kappa, whiteness, viscosity, yield, and drug consumption are predicted. The prediction adopts a multi-task framework to share the underlying feature representation, and after obtaining new online observations, the prediction is updated in a loop using a filtering method to dynamically converge uncertainty and improve the probability of achieving the target. Each prediction generates a structured list of sensitive variables.
8. The process control method based on intelligent pulp quality prediction according to claim 1, characterized in that, Based on the worst-case joint achievement probability and a list of sensitive variables, a candidate structured control strategy package is obtained using a multi-objective optimization algorithm, as follows: Let the set of controllable variables be U′′ = {u′1,…,u′ p }, the decision vector is Δu'; constraints: ;in, Let i be the i-th controllable variable; This is the current setting value; These are the upper and lower limits of a controllable variable; This represents the upper limit of the rate of change. To control the step size; A and b are the linear coupling constraint matrix and upper bound vector, respectively; device_health i For device health status trigger conditions; S i Let i be the set of reachable increments for the i-th actuator in the current health state; This represents the increment of the operation on the i-th variable during this period; Let the vector consist of the current setpoints of all controllable variables; define the set of indicators as a multi-objective optimization problem: in, The feasible region is defined for all constraints; J1, J2, J3, J4, and J5 are the joint achievement probability target, operating cost target, expected quality deviation, risk measure, and execution complexity, respectively; c is the linear cost coefficient vector per unit increment; ρ is the quadratic regularization weight. Let K' be the square of the L2 norm; K' be the set of process indicators being evaluated; w j For indicator weights; y j (Δu') is a random variable that predicts index j under the current action; This is the deviation loss function; is the expectation of the predicted distribution; violation(·) is the degree of violation; CVaR α (·) represents the conditional risk value, and the confidence level α∈(0,1) represents the average degree of violation under the worst-case scenario; λ1,λ group These are the sparse and grouped sparse weights, respectively. Let g be the action vector of the variable group; C is the joint probability function; op The objective function is the operating cost. Based on the above objective function and constraints, the NSGA-II algorithm is used to find the optimal solution, which is adapted to black-box probabilistic evaluation and non-convex feasible regions.
9. The process control method based on intelligent pulp quality prediction according to claim 7, characterized in that, The optimal process control based on the control strategy package and the current process state is as follows: using the action of the strategy package as a reference trajectory, the loop layer issues setpoints to each controllable variable and uses a servo controller with feedforward and filtering for tracking; online self-tuning is performed based on closed-loop data and tuning rules, and the servo controller parameters are updated periodically; when a change in gain, time delay drift, or sudden increase in noise level is detected, the controller switches to a robust parameter set and gradually interpolates to the new parameters to avoid disturbances caused by step parameter changes; Anti-saturation protection employs inversion backflow and conditional integration: when the actuator reaches saturation or is limited by the rate, the integral term is frozen and the saturation error is backflowed to the integral state to prevent overshoot caused by integral accumulation after saturation; at the same time, abnormal suppression and short-term weight reduction are used for measurement spikes to ensure that the loop is not pulled by noise, and for coupled loops, decoupling compensation matrix or sequential loop priority is used.
Citation Information
Patent Citations
Full-automatic intelligent continuous rack plating line body detection method and device
CN120275318A
Production data monitoring method and system based on artificial intelligence
CN120765120A