An Adaptive Control Method for Reinforcement Learning-Based Recycling Production Line

By using an adaptive control method based on reinforcement learning, key parameters of the copper salt production line are identified and adjusted in real time, solving the problem that static control laws cannot cope with time-varying parameters, and achieving high-performance robust control and stability of the copper salt production line.

CN121300102BActive Publication Date: 2026-04-03HANGZHOU FUYANG HONGYUAN RENEWABLE RESOURCES CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, regulation schemes based on static control laws cannot effectively cope with the time-varying and non-stationary characteristics of parameters in complex industrial systems caused by differences in raw material composition, equipment aging, and changes in operating points, resulting in decreased controller performance and system instability.

Method used

An adaptive control method based on reinforcement learning is adopted. By identifying key process parameters online, an online identifier combining recurrent neural networks and extended Kalman filters is used to estimate material conversion characteristics and uncertainties in real time. Dynamic parameter adjustment is achieved through a strategy network and value network composed of multilayer sensing mechanisms to realize controller self-tuning.

Benefits of technology

It achieves intelligent and dynamic balancing of control targets under complex working conditions, improves the automation level and response speed of the production process, and can maintain high performance, robustness and stability in uncertain environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121300102B_ABST
    Figure CN121300102B_ABST
Patent Text Reader

Abstract

This invention relates to the field of adaptive control technology, specifically to an adaptive control method for a renewable resource production line based on reinforcement learning. The method includes: identifying key model parameters representing process dynamics in real time and estimating their uncertainties using an online identifier; adaptively tuning controller parameters online using a policy network based on these uncertainties; and continuously learning and self-optimizing the parameter tuning strategy online using a value network and reinforcement learning algorithms to achieve optimal control of the time-varying process. By identifying the uncertainties of key process parameters online and continuously self-tuning the controller parameters using reinforcement learning, high-performance robust control of a time-varying, uncertain production process is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of adaptive control technology, specifically to an adaptive control method for a renewable resource production line based on reinforcement learning. Background Technology

[0002] In complex industrial systems, such as the chemical synthesis process of specific metal salt products, it is crucial to maintain the optimal dynamic performance of key process variables in order to ensure high production efficiency and consistent product quality. These systems guide the physicochemical state of raw materials to a predetermined target by applying precise closed-loop regulation to the controlled object.

[0003] The dominant technology at present is the regulation scheme based on static control law. The internal parameters of this scheme are usually tuned offline once based on an idealized linear time-invariant process mathematical model established under nominal operating conditions. The fundamental limitation of this design paradigm lies in its core assumption: that the dynamic characteristics of the controlled object remain unchanged in long-term operation and can be accurately described by this initial model.

[0004] Its characteristics will vary due to differences in raw material composition from batch to batch, aging and drift of equipment performance, and changes in operating conditions, thus exhibiting obvious time-varying and non-stationary characteristics of parameters. These inevitable parameter variations and unmodeled dynamic phenomena together cause a model mismatch between the mathematical model on which the controller design depends and the actual behavior of the physical process.

[0005] When the mismatch in this model accumulates to a certain extent, the control law with fixed parameters can no longer match the changing controlled object. The direct result is a significant deterioration in the performance of the closed-loop system, which can lead to the deterioration of performance indicators such as increased overshoot, aggravated oscillations, or reduced steady-state accuracy, and may even threaten the stability of the system. The conventional way to solve this problem is to rely on operators to manually and tentatively readjust the parameters. This approach has a significant lag in response and cannot guarantee the optimal level of adjustment.

[0006] Therefore, there is an urgent need in this technical field for an advanced control architecture that can break through the limitations of traditional static control laws. The core capability of this architecture is that it can identify the dynamic characteristics of the controlled object online and in real time, and adjust its own control law autonomously and continuously based on the identification results, that is, to achieve control law self-tuning. Such an adaptive control method can automatically maintain the preset ideal dynamic characteristics and robustness in environments with significant uncertainties and disturbances.

[0007] To address this, the present invention proposes an adaptive control method for a renewable resource production line based on reinforcement learning. Summary of the Invention

[0008] The purpose of this invention is to provide an adaptive control method for a renewable resource production line based on reinforcement learning. By identifying the uncertainty of key process parameters online and using reinforcement learning to continuously self-tune the controller parameters, high-performance robust control of time-varying uncertain production processes can be achieved.

[0009] To achieve the above objectives, the present invention provides the following technical solution:

[0010] An adaptive control method for a renewable resource production line based on reinforcement learning includes:

[0011] The process variables of the copper salt product production line are acquired in real time and constructed into a state vector representing the real-time process state.

[0012] By using an online identifier, the historical sequence of the state vector is analyzed using the gating unit of a recurrent neural network, the material reaction sensitivity parameter that characterizes the material conversion characteristics under the current working condition is identified online, and the uncertainty information of the sensitivity parameter is estimated.

[0013] Based on a strategy network composed of multi-layer sensing mechanisms, nonlinear mapping is performed on uncertain information to output control gain modulation coefficients for compensating for changes in material conversion characteristics. Using the modulation coefficients, the preset performance benchmark controller parameter set and robust benchmark controller parameter set are dynamically weighted and combined to adaptively generate real-time controller parameters.

[0014] The controller output is adjusted based on real-time controller parameters; a value network composed of multi-layer sensing mechanisms is used to evaluate the state value based on the real-time process state, and a timing differential error signal is formed by combining the controller output with the expected state setpoint.

[0015] By utilizing the temporal difference error signal, the control update gradient is determined through a reinforcement learning-based control gradient algorithm, and the network weights of the online self-tuning policy network and value network are synchronously determined.

[0016] Preferably, the process of constructing the state vector representing the real-time process state includes:

[0017] Using an online spectrometer, temperature sensor, and mass flow meter, process variables reflecting the chemical composition of the material, furnace temperature, and feed rate are acquired. The acquired process variables are then filtered and synchronized. The processed process variables are calibrated and used to form the state vector.

[0018] Preferably, the online identifier is a hybrid structure cascaded with a recurrent neural network and an extended Kalman filter, and the process of identifying the material reaction sensitivity parameters includes:

[0019] In each control cycle, the recurrent neural network receives the historical state vector sequence as input, models the temporal dynamics of the process through its internal gating unit structure, and outputs a prior estimate of the material reaction sensitivity parameter at the current moment. The extended Kalman filter receives the prior estimate and simultaneously acquires the current state vector as measurement input. Through internal state transition and measurement update calculations, it corrects the prior estimate and outputs a posterior estimate of the material reaction sensitivity parameter.

[0020] Preferably, the estimation of the uncertainty information of the sensitivity parameter is accomplished by analyzing the posterior state covariance matrix synchronously generated inside the extended Kalman filter, specifically including:

[0021] Identify and locate the position index of the material reaction sensitivity parameter in the internal state vector of the extended Kalman filter, extract the corresponding numerical element from the main diagonal of the posterior state covariance matrix according to the position index, and use the element as the variance estimate of the sensitivity parameter and as the uncertainty information.

[0022] Preferably, the process of generating the real-time controller parameters includes:

[0023] A weighted performance control component and a weighted robust control component are generated. The weighted performance control component is obtained by multiplying each parameter in the performance reference controller parameter set with the control gain modulation coefficient. The weighted robust control component is obtained by multiplying each parameter in the robust reference controller parameter set with the complementary coefficient of the control gain modulation coefficient, and the sum of the modulation coefficient and the complementary coefficient is always one. The parameters at corresponding positions in the weighted performance control component and the weighted robust control component are added one by one to synthesize the final real-time controller parameters.

[0024] Preferably, the performance benchmark controller parameter set is predetermined using this offline tuning method:

[0025] A process model of the production line under nominal operating conditions is established, and the parameter set is calculated by analyzing the characteristic parameters of the open-loop response curve of the model using the Ziegler-Nichols step response method based on the model.

[0026] Preferably, the robust reference controller parameter set is predetermined using an offline tuning method:

[0027] A set of production line uncertainty process models including parameter perturbations and unmodeled dynamics is established. Based on the model set, H-infinity control theory is used to solve for controller parameters that can provide stable closed-loop performance for all models in the model set. This set of parameters is the robust benchmark controller parameter set.

[0028] Preferably, the formation of the timing differential error signal includes:

[0029] The value network calculates the current state value estimate based on the current state vector, and obtains the instantaneous feedback signal generated by the production line operation and the next moment state value estimate after discount factor decay; the current state value estimate, the instantaneous feedback signal, and the decayed next moment state value estimate are combined to obtain the time series difference error signal.

[0030] Preferably, the control gradient algorithm based on reinforcement learning is a deep deterministic policy gradient algorithm, and the online self-tuning of network weights specifically includes two distinct but synchronous update processes:

[0031] The value network weight update is achieved by constructing a value loss function by calculating the mean square error of the temporal difference error signal and adjusting the weights along the negative gradient direction of the loss function. The policy network weight update is achieved by using the output of the value network as an evaluation of the policy performance and adjusting the weights along the gradient direction of the evaluation value relative to the policy network parameters.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0033] 1. This invention estimates the changes and uncertainties in material conversion characteristics in real time through an online identifier, and automatically and in real time adjusts the controller parameters accordingly. This overcomes the shortcomings of existing fixed-parameter controllers, which suffer from performance degradation when operating conditions change and rely on repeated manual retuning. It realizes the self-tuning of the control method and significantly improves the automation level and response speed of the production process.

[0034] 2. This invention does not generate a single control parameter, but rather uses the modulation coefficients output by the strategy network to online weightedly combine two sets of pre-designed offline benchmark parameters, one high-performance and one robust. This enables the control method to favor optimal performance when the process characteristics are relatively stable, and to automatically favor stability and safety when uncertainty increases, thus achieving intelligent and dynamic balance of control objectives under complex operating conditions.

[0035] 3. This invention utilizes a reinforcement learning framework to evaluate the long-term returns of the current control strategy through a value network and generates a time-series differential error signal to guide the weight updates of the strategy network and the value network. This enables the control strategy to continuously learn and evolve from historical operating data, and to autonomously optimize its control behavior to adapt to gradual or unknown changes that may occur during the long-term operation of the production line. It has a long-term optimization capability that traditional adaptive control methods do not possess. Attached Figure Description

[0036] Figure 1 This is a flowchart of an adaptive control method for a renewable resource production line based on reinforcement learning, according to the present invention.

[0037] Figure 2 This is a flowchart of the adaptive control method according to an embodiment of the present invention;

[0038] Figure 3 This is a schematic diagram of a weight self-tuning loop based on reinforcement learning, according to an embodiment of the present invention. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Other embodiments obtained by those skilled in the art based on the ideas in this specification without creative effort all fall within the protection scope of this invention.

[0040] Reference Figures 1 to 3 This invention provides an adaptive control method for a renewable resource production line based on reinforcement learning.

[0041] In this embodiment, the material reaction sensitivity parameter, in a general sense, refers to the dynamic process gain that characterizes the degree of influence of changes in key control inputs on core controlled output variables during the production process. The selection principle for this parameter is as follows: first, identify the core input-output relationship pair on the production line that best reflects the fluctuations in production results caused by uncertainties such as changes in raw material composition, catalyst activity decay, or operating point drift; then, use the dynamic gain of this relationship as the key parameter that needs to be identified online in real time. For example, in a chemical reaction process, it could be the gain of the neutralizer addition rate on pH value; in a smelting process, it could be the gain of the heating power on furnace temperature.

[0042] Example 1:

[0043] Reference Figure 1 , Figure 2 An adaptive control method for a renewable resource production line based on reinforcement learning, comprising:

[0044] The process variables of the copper salt product production line are acquired in real time and constructed into a state vector representing the real-time process state.

[0045] By using an online identifier, the historical sequence of the state vector is analyzed using the gating unit of a recurrent neural network, the material reaction sensitivity parameter that characterizes the material conversion characteristics under the current working condition is identified online, and the uncertainty information of the sensitivity parameter is estimated.

[0046] Based on a strategy network composed of multi-layer sensing mechanisms, nonlinear mapping is performed on uncertain information to output control gain modulation coefficients for compensating for changes in material conversion characteristics. Using the modulation coefficients, the preset performance benchmark controller parameter set and robust benchmark controller parameter set are dynamically weighted and combined to adaptively generate real-time controller parameters.

[0047] The controller output is adjusted based on real-time controller parameters; a value network composed of multi-layer sensing mechanisms is used to evaluate the state value based on the real-time process state, and a timing differential error signal is formed by combining the controller output with the expected state setpoint.

[0048] By utilizing the temporal difference error signal, the control update gradient is determined through a reinforcement learning-based control gradient algorithm, and the network weights of the online self-tuning policy network and value network are synchronously determined.

[0049] Furthermore, the process of constructing the state vector representing the real-time process state includes: acquiring material characteristic process variables reflecting the chemical composition of the material, furnace temperature process variables representing the core reaction conditions, and feed rate process variables through an online spectrometer, temperature sensor, and mass flow meter; filtering and synchronizing the acquired multiple sets of process variables; calibrating the processed multiple sets of process variables, and constructing the state vector.

[0050] Specifically, in an application scenario targeting a copper salt production line, multiple sets of sensors deployed at key process nodes acquire process variables in real time. An online spectrometer acquires material characteristic process variables reflecting the chemical composition of the materials, used to characterize real-time fluctuations in raw material composition. Temperature sensors acquire furnace temperature process variables characterizing core reaction conditions. Mass flow meters acquire feed rate process variables. Because the original acquired signals may contain noise interference and the sampling times of each variable are inconsistent, filtering and synchronization processing of the acquired process variables is required. Second-order Butterworth filters can be used for this filtering process. A low-pass filter is used, with its cutoff frequency preset according to the noise characteristics of the process, to smooth data and remove high-frequency noise. The method for determining the cutoff frequency is as follows: First, perform spectrum analysis on the raw process signals collected from the production line under various typical operating conditions to identify the signal frequency bands that can represent the main dynamic changes of the process, as well as the high-frequency noise frequency bands caused by measurement noise or electrical interference. Set the cutoff frequency at the boundary between the signal frequency band and the noise frequency band to effectively filter out noise while retaining the process dynamic information useful for control to the maximum extent. In the application scenario of this embodiment, the cutoff frequency can preferably be set to 0.2 Hz.

[0051] Synchronization processing employs a timestamp-based linear interpolation method to align the measured values ​​of all process variables to the start time of each control cycle, forming a temporally consistent data snapshot. The processed multiple sets of process variables are further calibrated using a max-min normalization method to linearly map sensor readings of different physical units and dimensions to a unified numerical range. For example, after calibration, these quantified process variables are arranged and combined in a preset fixed order to ultimately form a state vector that can comprehensively and in real-time characterize the current production process status.

[0052] The predetermined fixed order is determined by ensuring the consistency of the data structure for stable processing by the subsequent neural network model. Specifically, once selected, this order must remain constant throughout the offline model training and online identification and control processes. A preferred sorting method is to arrange the variables according to their logical relationship in the process flow. For example, variables representing raw material properties are arranged first, followed by variables representing core reaction conditions, and finally variables representing operational control quantities. For instance, in a typical copper salt production line application scenario, the acquired process variables may include material characteristic variables reflecting raw material composition, furnace temperature variables representing core reaction conditions, and feed rate variables representing operational settings. Accordingly, a reasonable fixed order can be determined as follows: the first element of the vector is the material characteristic variable, the second element is the furnace temperature variable, and the third element is the feed rate variable.

[0053] In a preferred embodiment, the process of constructing a state vector representing the real-time process state includes, after calibrating multiple sets of process variables and before constructing the state vector, weighting the calibrated multiple sets of process variables through an attention network; the attention network dynamically calculates a weight coefficient for each process variable based on the current values ​​of all process variables, and uses the weight coefficient to scale the corresponding process variables, and finally the scaled multiple sets of process variables constitute the state vector.

[0054] Specifically, after performing max-min normalization calibration on multiple sets of process variables after filtering and synchronization, these calibration values ​​are not directly combined into a state vector. Instead, they are first input into an attention network based on a self-attention mechanism. This network is a small, fully connected neural network that receives the set of all calibration values ​​as input and outputs a weight coefficient between 0 and 1 for each value, with the sum of all weight coefficients being one. The network is pre-trained offline to identify the relative importance of each process variable under the current operating condition based on the combination pattern of the input values. Subsequently, each calibration value is multiplied by its corresponding weight coefficient to obtain a set of dynamically weighted values. Finally, these weighted values ​​are arranged in a preset order to form the final state vector. The preset order is exactly the same as the fixed order used to combine the initial process variables, and its purpose is to ensure that the physical meaning represented by each element in the state vector remains stable and consistent before and after being weighted by the attention mechanism.

[0055] In a preferred embodiment, the attention network is pre-structured to include an input layer, a hidden layer with eight neurons, and an output layer with the same number of input variables. The hidden layer uses a modified linear unit as the activation function for nonlinear transformation, while the output layer uses a normalized exponential function to ensure that the sum of all output weight coefficients is exactly one. Further, the offline pre-training method includes the following steps: First, historical operating data of the production line is collected, and senior process engineers manually label the importance weight distribution of each process variable for each data sample according to different operating conditions, thus forming the labels required for training. Then, by optimizing a loss function designed to measure the difference between the network output weights and the engineer's labeled weights, supervised learning is used to train the network until it can accurately reproduce the engineer's judgment based on the input process variables.

[0056] By introducing an attention network, the state vector construction process can intelligently assign higher weights to the most critical process variables, suppress the interference of secondary information, provide inputs with higher information density for subsequent identification and control stages, and improve the overall accuracy of the algorithm.

[0057] The state vector construction method in this embodiment transforms heterogeneous raw measurement data from multiple sensors into a standardized mathematical expression that can be directly used by subsequent algorithms, providing a comprehensive and reliable data foundation for achieving accurate online identification and adaptive control.

[0058] Furthermore, the online identifier is specifically a hybrid structure cascaded with a recurrent neural network and an extended Kalman filter. The process of identifying the material reaction sensitivity parameter includes: in each control cycle, the recurrent neural network receives the historical state vector sequence as input, models the temporal dynamics of the process through its internal gating unit structure, and outputs a priori estimate of the material reaction sensitivity parameter at the current moment; the extended Kalman filter receives the prior estimate and simultaneously acquires the current state vector as measurement input, corrects the prior estimate through internal state transition and measurement update calculations, and outputs a posterior estimate of the material reaction sensitivity parameter.

[0059] Specifically, the online identifier adopts a cascaded structure with feedforward connections. Its front end is a recurrent neural network and its back end is an extended Kalman filter. In this embodiment, the recurrent neural network is specifically a gated recurrent unit network. At the beginning of each control cycle, the gated recurrent unit network receives a historical sequence composed of state vectors from the past several control cycles as input. The gated unit structure inside the network selectively transmits and forgets the sequence information through its update gate and reset gate, thereby modeling the temporal dynamics of the production process.

[0060] After the calculation is completed, the network outputs a scalar value, which is the prior estimate of the material reaction sensitivity parameter at the current moment. The extended Kalman filter then receives this prior estimate as input to its prediction stage and simultaneously acquires the state vector generated in the previous process step, using it as the measurement input. Internally, the filter performs state transition and measurement update calculations: the state transition step directly uses the prior estimate output by the gated loop unit network as the predicted state; the measurement update step calculates a correction amount based on the deviation between the current measurement input (i.e., the state vector) and the predicted state, and uses this correction amount to correct the prior estimate. The final result is a posterior estimate of the material reaction sensitivity parameter, which integrates trend predictions from historical data and real-time corrections from current data, serving as the most accurate parameter identification value within the current control cycle. The extended Kalman filter's process... The process model is preset as a random walk model, and its measurement model is a linear model representing the relationship between the sensitivity parameter and the state vector, obtained through offline analysis of historical data. The recurrent neural network is specifically a gated recurrent unit network with 64 units. This measurement model clearly defines the mathematical relationship between the material reaction sensitivity parameter (a scalar value) as a state variable and the state vector (a multidimensional vector) as a measurement value. Structurally, the model is a linear mapping relationship. The predicted value of the current state vector is calculated by multiplying the current material reaction sensitivity parameter value by a constant column vector with the same dimension as the state vector, which is determined in advance through offline calculation. Each element of the constant column vector represents the linear influence coefficient of the sensitivity parameter on the corresponding process variable, and its value is obtained by fitting a large amount of historical operating data using regression analysis methods such as least squares.

[0061] The hybrid structure identifier used in this embodiment combines the powerful nonlinear time-series data modeling capability of recurrent neural networks with the advantages of extended Kalman filters in performing optimal state estimation in noisy environments, thereby achieving fast and accurate online tracking and identification of key process parameters.

[0062] Furthermore, the estimation of the uncertainty information of the sensitivity parameter is accomplished by parsing the posterior state covariance matrix synchronously generated inside the extended Kalman filter. Specifically, this includes: identifying and locating the position index of the material reaction sensitivity parameter in the state vector inside the extended Kalman filter, extracting the corresponding numerical element from the main diagonal of the posterior state covariance matrix according to the position index, and using the element as the variance estimate of the sensitivity parameter and as the uncertainty information.

[0063] Specifically, in each iteration of the extended Kalman filter, while outputting the posterior estimate of the material reaction sensitivity parameter, a posterior state covariance matrix is ​​simultaneously updated and output. The dimension of this matrix corresponds to the dimension of the internal state vector of the filter, and the elements on its main diagonal represent the degree of uncertainty of the estimated value of each element in the state vector. In this embodiment, the internal state vector of the extended Kalman filter is set to contain only the material reaction sensitivity parameter as an element, so its position index is determined. The process of estimating the uncertainty information is to extract a unique numerical element from the main diagonal of the posterior state covariance matrix of this element. This numerical element is numerically equal to the variance of the posterior estimate of the material reaction sensitivity parameter. Finally, this variance estimate is directly used as the uncertainty information representing the credibility of the identification result and output to the subsequent control stage.

[0064] The extended Kalman filter has a clearly defined core observation model, which is a function pre-identified offline. This model describes, in both text and numerical terms, the mathematical relationship between the material reaction sensitivity parameter (a scalar) as a state variable and the state vector (a multidimensional vector) as a measurement value. In a preferred embodiment, the observation model is defined as a linear mapping relationship, i.e., the predicted state vector is obtained by multiplying the sensitivity parameter by a pre-defined column vector with the same dimension as the state vector. Each element of this column vector is obtained by analyzing a large amount of historical data and fitting it using the least squares method; they represent the linear influence coefficient of the sensitivity parameter on each process variable. The order of these influence coefficients in the column vector strictly adheres to the pre-defined fixed order of their corresponding process variables in the state vector. In the measurement update step, this column vector is used to calculate the predicted measurement value and, consequently, the deviation between the predicted and actual measurement values.

[0065] This embodiment quantifies the inherent uncertainty in the parameter identification process by analyzing the posterior state covariance matrix synchronously generated by the extended Kalman filter. This provides a key, quantifiable decision basis for the dynamic adjustment of the control strategy, rather than relying solely on point estimates of the parameters.

[0066] Further, the generation process of the real-time controller parameters includes: generating weighted performance control components and weighted robust control components; wherein, the weighted performance control components are obtained by multiplying each parameter in the performance reference controller parameter set with the control gain modulation coefficient; the weighted robust control components are obtained by multiplying each parameter in the robust reference controller parameter set with the complementary coefficient of the control gain modulation coefficient, and the sum of the modulation coefficient and the complementary coefficient is always one; the parameters at corresponding positions in the weighted performance control components and the weighted robust control components are added item by item to synthesize the final real-time controller parameters.

[0067] Specifically, the generation process is completed in a parameter synthesis module, which receives three inputs: a preset performance reference controller parameter set, a preset robust reference controller parameter set, and control gain modulation coefficients output by the policy network. First, a weighted performance control component is generated by multiplying each parameter (e.g., proportional, integral, and derivative terms) in the performance reference controller parameter set with the control gain modulation coefficient. A complementary coefficient is obtained by subtracting the control gain modulation coefficient from the numerical value. A weighted robust control component is then generated by multiplying each parameter in the robust reference controller parameter set with the complementary coefficient. The weighted performance control component and the weighted robust control component are then synthesized by adding them item by item, that is, adding the parameter values ​​at corresponding positions in the two components. The sum constitutes the final real-time controller parameter set, which will be used to adjust the controller output in the current control cycle.

[0068] The parameter generation method in this embodiment achieves a smooth transition between high performance and high robustness for the controller by dynamically weighting two sets of reference parameters. This method enables the controller to automatically adjust its response characteristics based on the quantitative information of process uncertainty, pursuing optimal performance when the operating conditions are stable and ensuring control stability when the operating conditions fluctuate.

[0069] In a preferred embodiment, the process of nonlinearly mapping the uncertainty information further includes calculating the time change rate of the uncertainty information within a continuous control period, and using the uncertainty information and its time change rate together as input, and having the policy network perform nonlinear mapping on the combined input to output the control gain modulation coefficient.

[0070] Specifically, after the parameter synthesis module receives the uncertainty information (i.e., variance estimate) of the current cycle output by the extended Kalman filter, the module also reads the uncertainty information value of the previous control cycle from the memory. By subtracting the current value from the value of the previous cycle and dividing by the control cycle duration, the time change rate of the uncertainty information is calculated. Subsequently, the current uncertainty information value and the calculated time change rate are combined to form a two-dimensional input vector. The policy network based on the multilayer sensing mechanism receives this two-dimensional input vector and, through its internal nonlinear mapping, finally outputs the control gain modulation coefficients used for parameter weighting.

[0071] By integrating information on the rate of change of uncertainty, the policy network can not only perceive the magnitude of process fluctuations but also judge their changing trends, thereby making more predictive decisions and improving the timeliness and accuracy of the controller's response in dynamic processes.

[0072] Furthermore, the performance benchmark controller parameter set is predetermined through this offline tuning method: a process model of the production line under nominal operating conditions is established, and the parameter set is calculated by analyzing the characteristic parameters of the open-loop response curve of the model using the Ziegler-Nichols step response method based on the model.

[0073] Specifically, the tuning of this parameter set is an offline preparation step completed before the adaptive control method is put into online operation. First, a process model characterizing the dynamic characteristics of the production line under nominal operating conditions is established. Here, nominal operating conditions refer to the ideal state where production is stable and material characteristics are within the standard range. This process model can be a first-order plus pure time delay transfer function model, and its parameters are identified by analyzing historical input and output data collected under nominal operating conditions. After the model is established, the Ziegler-Nichols step response method is used for parameter tuning. This method simulates a step input on the model, such as simulating a step change in the feed rate, and records the open-loop response curve of key output variables (such as furnace temperature). Three characteristic parameters are extracted from this response curve: process gain, time constant, and pure time delay. These three characteristic parameters are substituted into the standard tuning table of the Ziegler-Nichols method to calculate a set of controller parameters containing proportional, integral, and derivative terms. This set of parameters constitutes the performance benchmark controller parameter set.

[0074] This embodiment obtains a performance benchmark controller parameter set through offline tuning, which provides the optimal performance target for the controller to operate under ideal conditions. This parameter set is designed with the goal of pursuing fast response and minimum steady-state error, and constitutes an endpoint of the adaptive parameter weighted combination.

[0075] Furthermore, the robust reference controller parameter set is predetermined through an offline tuning method: a set of production line uncertainty process models including parameter perturbations and unmodeled dynamics is established, and based on the model set, H-infinity control theory is used to solve for controller parameters that can provide stable closed-loop performance for all models in the model set. This set of parameters is the robust reference controller parameter set.

[0076] Specifically, the tuning of this parameter set is also an offline step completed before the adaptive control method is put into online operation. Based on the nominal operating condition process model, a production line uncertainty process model set is established by introducing parameter perturbations and unmodeled dynamics. Parameter perturbations refer to setting upper and lower bounds for the key parameters in the model (such as process gain, time constant, etc.) to cover the fluctuation range that these parameters may experience in actual operation. Unmodeled dynamics are described by introducing a high-frequency gain-limited multiplicative uncertainty weight function to characterize the higher-order or nonlinear dynamic characteristics that the model fails to capture. Based on this uncertainty process model set, the controller is designed using H-infinity control theory. The goal of this theory is to solve for a controller that ensures the stability of the closed loop formed by any model in the model set, and that the impact of external disturbances on key output variables is suppressed to a preset minimum. By solving the corresponding optimization problem, a robust controller is obtained, and the corresponding proportional, integral, and derivative parameters are extracted from the controller. This set of parameters constitutes the robust benchmark controller parameter set. The model error boundary of this uncertain process model set is defined by a preset multiplicative uncertainty weight function (high-pass filter). The solution of the H-infinity control theory is based on a preset performance weight function (low-pass filter, used to penalize tracking error) and a control weight function (constant, used to limit control output).

[0077] The specific design method of the multiplicative uncertainty weighting function is as follows: by comparing the response of the nominal model at different frequencies with the frequency response data of the actual production process, the difference between the two is analyzed; a stable and minimum phase high-pass filter is designed so that its amplitude-frequency characteristic curve can completely enclose the relative error between the model and the actual process at all frequency points. The cutoff frequency of the filter is usually selected at the frequency point where the model begins to distort, and its high-frequency gain needs to be greater than the maximum possible relative error of the model in all high-frequency bands.

[0078] The design principle of the performance weighting function (low-pass filter) is that its bandwidth should match the desired closed-loop system response speed. For example, if the system is expected to complete the response within 1 second, the cutoff frequency of the filter can be set to around 1 radian / second. The selection of the control weighting function (constant) needs to take into account the limitations of the physical actuator. A larger constant value means a heavier penalty to the control output, which is suitable for situations where the actuator is prone to saturation or where it is desirable to save control energy. Its value is usually adjusted through design iterations to obtain a smooth and feasible control output while meeting performance requirements.

[0079] This embodiment obtains a robust reference controller parameter set through H-infinity control theory, which can ensure that the control loop maintains the stability of closed-loop operation when faced with significant changes in process characteristics and model mismatch. This parameter set sacrifices some control performance in exchange for the highest stability guarantee, and constitutes another endpoint of the adaptive parameter weighted combination.

[0080] Furthermore, the formation of the time-series differential error signal includes: the value network calculating the current state value estimate based on the current state vector, and obtaining the instantaneous feedback signal generated by the production line operation and the next-time state value estimate after discount factor attenuation; combining the current state value estimate with the instantaneous feedback signal and the attenuated next-time state value estimate to obtain the time-series differential error signal.

[0081] Specifically, the signal formation process is executed in each control cycle. First, the value network based on the multilayer sensing mechanism receives the current state vector as input. After nonlinear mapping calculation within the network, a scalar value is output, which is the current state value estimate.

[0082] Simultaneously, an immediate feedback signal is calculated based on the key performance indicators of the production line. To achieve optimal overall performance, the calculation of this signal requires a comprehensive evaluation of control performance and control costs, rather than focusing solely on the error of a single variable. A preferred calculation method is to calculate the error between the actual measured value and the expected setpoint of each key process variable, then square these error values ​​and multiply them by a preset performance weighting coefficient. Simultaneously, the final control output of the controller is also squared and multiplied by its corresponding cost weighting coefficient. Finally, all weighted error squares are added to the weighted control output squares. The sum is then negative, serving as the final immediate feedback signal. By appropriately setting the values ​​of each weight coefficient, the system's different emphases on tracking accuracy and economic cost can be flexibly balanced, thereby guiding the reinforcement learning algorithm to find a high-performance and efficient control strategy. Furthermore, the determination of these weight coefficients is based on production process knowledge and economic analysis: core process variables that affect the quality of the final product (such as pH value, furnace temperature, etc.) should be assigned relatively high performance weights; while control outputs that are costly or affect equipment lifespan (such as expensive catalyst addition, violent valve actions, etc.) should be assigned relatively high cost weights.

[0083] The process also requires obtaining the state vector at the next moment and inputting it into a target value network with the same structure but slower parameter updates to obtain the state value estimate at the next moment. Then, the estimate is multiplied by a preset discount factor less than one to obtain the state value estimate at the next moment after the discount factor decay.

[0084] Finally, a combination operation is performed to form an error signal: the instantaneous reward signal obtained from the above calculation is added to the decayed state value estimate of the next time step to obtain a target value; then, the current state value estimate output by the value network is subtracted from this target value, and the difference between the two is the final time-series differential error signal.

[0085] This embodiment provides a key, quantitative learning-driving signal for reinforcement learning algorithms by generating a temporal difference error signal. This signal accurately measures the accuracy of the value network's prediction of future returns, and its value directly guides the self-tuning process of the subsequent policy network and value network weights, which is the core link in realizing online optimization.

[0086] Furthermore, the reinforcement learning-based control gradient algorithm is a deep deterministic policy gradient algorithm. The online self-tuning of network weights specifically includes two distinct but synchronous update processes: the network weight update of the value network is to construct a value loss function by calculating the mean square error of the temporal difference error signal, and then adjust the weights along the negative gradient direction of the loss function; the network weight update of the policy network is to use the output of the value network as an evaluation of the policy performance, and then adjust the weights along the gradient direction of the evaluation value relative to the policy network parameters.

[0087] Specifically, this online self-tuning process is executed at the end of each control cycle and includes two parallel network weight update processes. The first process updates the value network weights, aiming to minimize the prediction error. It constructs a value loss function by calculating the mean square error of the time-series difference error signal generated in the previous process step; it then uses a gradient descent algorithm to calculate the gradient of this value loss function relative to all network weights of the value network, and makes small adjustments to the weights along the negative direction of this gradient. The second process updates the policy network weights, aiming to maximize long-term returns. It uses the output value of the value network as an evaluation of the current policy performance, and determines the direction of weight adjustment that can improve the evaluation value by calculating the gradient of this evaluation value relative to all network weights of the policy network. The weights of the policy network are then adjusted along this gradient direction. These two update processes are completed synchronously, constituting a single online self-tuning operation. The complete process is as follows: Figure 3 As shown.

[0088] In the online self-tuning process, the key network structure and training hyperparameters are preset as follows: Regarding the network structure, both the policy network and the value network employ a multilayer perceptron with two hidden layers. The number of hidden layer neurons in the policy network is 32 and 16, respectively, while the number of hidden layer neurons in the value network is 64 and 32, respectively. All hidden layers use modified linear units as activation functions. Regarding the key hyperparameters, the discount factor can be set to 0.99. The parameters of the target value network are obtained by soft updating the parameters of the main network, with a soft update coefficient of 0.5%. When executing the gradient descent algorithm, the learning rate of the value network can be set to 0.1%, and the learning rate of the policy network can be set to 0.01%.

[0089] In a preferred embodiment, the online self-tuning process of the network weights further includes storing an empirical data tuple containing historical state vectors, controller outputs, instantaneous feedback signals, and the state vector at the next time step into an empirical replay pool; assigning a priority to each empirical data tuple based on the absolute value of the time-series difference-division error signal corresponding to it; and, during the self-tuning of the network weights, performing weighted sampling from the empirical replay pool according to the priority to extract one or more empirical data tuples for calculating the control update gradient.

[0090] Specifically, at the end of each control cycle, a data set containing the current state vector, controller output, immediate feedback signal, and the state vector for the next time step is stored as an empirical data tuple in a fixed-capacity empirical replay pool. Simultaneously, the absolute value of the calculated timing differential error signal is also stored and used as the priority of this tuple. When performing network weight updates, the current data is not used directly; instead, a weighted sampling is performed on all tuples in the empirical replay pool according to their priority, extracting a small batch of data. Tuples with higher priority have a greater probability of being extracted. Finally, the calculation is based on this sampled small batch of data. The average control update gradient is calculated and used to synchronously adjust the network weights of the policy network and the value network. In a preferred embodiment, the capacity of the fixed-capacity experience replay pool can be set to 50,000 tuples, and the batch size for extracting a small batch of data can be set to 128. These values ​​are selected as a trade-off between learning efficiency and computational resources: a larger experience replay pool capacity helps to ensure the diversity of learning samples and break the temporal correlation between data, but it will consume more memory resources; while a moderate batch size can maintain high training efficiency while ensuring the stability of gradient estimation.

[0091] This embodiment introduces a priority experience replay mechanism, which enables the algorithm to learn more frequently from those more valuable and inspiring historical experiences, significantly accelerating the adaptation process to rare but critical operating conditions and greatly improving the overall learning efficiency and robustness of the control method.

[0092] The dual-network synchronous self-tuning mechanism adopted in this embodiment enables the control method to continuously learn and self-optimize from operational experience. This mechanism can fine-tune the nonlinear mapping relationship from uncertainty information to control gain online, thereby continuously improving the long-term comprehensive performance of adaptive control.

[0093] This invention utilizes an online identifier based on a cascaded recurrent neural network and an extended Kalman filter to quickly and accurately identify sensitivity parameters characterizing material conversion properties. By leveraging the uncertainty of these parameters, the strategy network dynamically adjusts the control gain, achieving a smooth and intelligent trade-off between high performance and high robustness. Simultaneously, an online self-tuning mechanism based on a deep deterministic strategy gradient algorithm is introduced, enabling the control strategy to continuously optimize itself based on real-time operational results. This significantly enhances the production line's adaptability to raw material fluctuations, ensuring the long-term stability of the control process and the consistency of the final product's quality.

[0094] Example 2:

[0095] This embodiment provides a specific process for applying the above-mentioned adaptive control method for a renewable resource production line based on reinforcement learning in a production line that uses copper-containing etching waste liquid to produce basic copper carbonate. The core control objective of this production line is to maintain the pH value in the reactor at a stable set point by precisely adjusting the addition rate of the neutralizing agent in order to cope with the interference of the continuous fluctuation of copper ion concentration and acidity in the feed waste liquid.

[0096] Within a specific control cycle, the implementation steps of this method are as follows:

[0097] First, a state vector representing the real-time process state is constructed. A process variable reflecting the material's chemical composition is obtained using an online ion concentration meter deployed in the feed pipeline; the current copper ion concentration in the waste liquid is measured to be 130 g / L. A pH meter and temperature sensor installed inside the reactor are used to obtain the pH and temperature process variables representing the core reaction conditions; the current pH is measured to be 5.2 and the temperature to be 55°C. The feed rate is measured to be 2.5 cubic meters per hour using a mass flow meter. These raw measurements are filtered by a third-order Butterworth low-pass filter and timestamped based on a 15-second control cycle. Subsequently, a maximum-minimum normalization method is used to calibrate the copper ion concentration, pH, temperature, and feed rate to their respective numerical ranges, resulting in calibration values ​​[0.85, 0.60, 0.75, 0.50]. These four values ​​together constitute the state vector for the current moment.

[0098] Next, an online identifier identifies the material reaction sensitivity parameter, defined as the "effective pH boosting coefficient per unit of neutralizer," which characterizes the material's buffering properties under the current operating conditions. A recurrent neural network containing 64 gated loop units receives the state vector sequence of the past 20 control cycles as input, models the temporal dynamics of the process, and outputs a prior estimate of the effective pH boosting coefficient, which is 0.12. Subsequently, an extended Kalman filter receives this prior estimate of 0.12 and obtains the current state vector [0.85, 0.60, 0.75, 0.50] as the measurement input. Through internal state transition and measurement update calculations, the prior estimate is corrected, and finally, a posterior estimate of the boosting coefficient is output, which is 0.11.

[0099] At the same time, to estimate the uncertainty information of the sensitivity parameter, the extended Kalman filter simultaneously generates a posterior state covariance matrix while outputting a posterior estimate of 0.11. The numerical elements corresponding to the boost coefficients are extracted from the main diagonal of this matrix to obtain its variance estimate of 0.008, which is used as the uncertainty information.

[0100] Subsequently, real-time controller parameters are generated based on this uncertainty information. A policy network composed of multilayer sensing mechanisms receives the uncertainty information value of 0.008 as input. This network outputs a control gain modulation coefficient with a value of 0.65 through internal nonlinear mapping. The preset performance benchmark controller parameter set (P=3.0, I=2.2, D=1.1) and robust benchmark controller parameter set (P=1.2, I=0.8, D=0.3) are used for calculation. By multiplying the performance parameter set with the modulation coefficient 0.65 and the robust parameter set with the complementary coefficient 0.35 (1-0.65), two weighted components are obtained. These two components are added one by one, and the real-time controller parameters are finally determined to be P=2.37, I=1.71, D=0.82. This set of parameters is immediately used to adjust the output power of the neutralizer addition pump.

[0101] Finally, a temporal differential error signal is generated and the network weights are self-tuned online. A multilayer perceptron-structured value network receives the current state vector [0.85, 0.60, 0.75, 0.50] and calculates the current state value estimate as -1.5. Based on the deviation between the pH setpoint of 5.0 and the measured value of 5.2, the instantaneous reward signal is calculated as -0.04. After 15 seconds, the state vector of the next moment is obtained and input into the target value network, resulting in a next moment state value estimate of -1.35 after attenuation by a discount factor of 0.98. The instantaneous reward signal is added to the attenuated next moment value estimate, and then the current state value estimate is subtracted to obtain a temporal differential error signal of 0.11. Using this error signal, the network weights of the value network and the policy network are updated synchronously through a deep deterministic policy gradient algorithm, completing one learning and optimization iteration.

[0102] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the scope of protection defined in the claims.

Claims

1. An adaptive control method for a renewable resource production line based on reinforcement learning, characterized in that, include: Real-time acquisition of process variables from the copper salt product production line; calibration and weighting of process variables; The attention network calculates a weight coefficient for each process variable based on the current values ​​of all process variables, and uses the weight coefficient to scale the corresponding process variables. The scaled multiple sets of process variables are then used to construct a state vector representing the real-time process state. The online identifier uses a recurrent neural network gating unit to analyze the historical sequence of the state vector, identifies the material reaction sensitivity parameter that characterizes the material conversion characteristics under the current working condition, and estimates the uncertainty information of the sensitivity parameter. The material reaction sensitivity parameter refers to the dynamic process gain that characterizes the degree of influence of changes in key control inputs on the core controlled output variable during the production process. Based on a strategy network composed of multi-layer sensing mechanisms, the time change rate of the uncertainty information within a continuous control cycle is calculated, and a nonlinear mapping is performed between the uncertainty information and the time change rate to output a control gain modulation coefficient for compensating for changes in material conversion characteristics. Using modulation coefficients, a preset performance reference controller parameter set and a robust reference controller parameter set are calculated to generate a weighted performance control component and a weighted robust control component. The weighted performance control component is obtained by multiplying each parameter in the performance reference controller parameter set with the control gain modulation coefficient. The weighted robust control component is obtained by multiplying each parameter in the robust reference controller parameter set with the complementary coefficient of the control gain modulation coefficient, and the sum of the modulation coefficient and the complementary coefficient is always one. The parameters at corresponding positions in the weighted performance control component and the weighted robust control component are added item by item to adaptively generate real-time controller parameters. The performance benchmark controller parameter set is predetermined using this offline tuning method: A process model of the production line under nominal operating conditions is established, and based on the process model, the Ziegler-Nichols step response method is used to calculate the parameter set of the performance benchmark controller by analyzing the characteristic parameters of the open-loop response curve of the model. The robust reference controller parameter set is predetermined using an offline tuning method. A set of production line uncertainty process models including parameter perturbations and unmodeled dynamics is established. Based on the set of production line uncertainty process models, H-infinity control theory is used to solve for controller parameters that can provide stable closed-loop performance for all models in the model set. This set of parameters is the robust benchmark controller parameter set. The controller output is adjusted based on real-time controller parameters; a value network composed of multi-layer sensing mechanisms is used to evaluate the state value based on the real-time process state, and a timing differential error signal is formed by combining the controller output with the expected state setpoint. By utilizing the temporal difference error signal, the control update gradient is determined through a reinforcement learning-based control gradient algorithm, and the network weights of the online self-tuning policy network and value network are synchronously determined.

2. The adaptive control method for a renewable resource production line based on reinforcement learning according to claim 1, characterized in that, The process of constructing the state vector representing the real-time process state includes: Using an online spectrometer, temperature sensor, and mass flow meter, process variables reflecting the chemical composition of the material, furnace temperature, and feed rate are acquired. The acquired process variables are then filtered and synchronized. The processed process variables are calibrated and used to form the state vector.

3. The adaptive control method for a renewable resource production line based on reinforcement learning according to claim 1, characterized in that, The online identifier is specifically a hybrid structure cascaded with a recurrent neural network and an extended Kalman filter. The process of identifying the material reaction sensitivity parameters includes: In each control cycle, the recurrent neural network receives the historical state vector sequence as input, models the temporal dynamics of the process through its internal gating unit structure, and outputs a prior estimate of the material reaction sensitivity parameter at the current moment. The extended Kalman filter receives the prior estimate and simultaneously acquires the current state vector as measurement input. Through internal state transition and measurement update calculations, it corrects the prior estimate and outputs a posterior estimate of the material reaction sensitivity parameter.

4. The adaptive control method for a renewable resource production line based on reinforcement learning according to claim 1, characterized in that, The estimation of the uncertainty information of the sensitivity parameter is accomplished by analyzing the posterior state covariance matrix synchronously generated inside the extended Kalman filter, specifically including: Identify and locate the position index of the material reaction sensitivity parameter in the internal state vector of the extended Kalman filter, extract the corresponding numerical element from the main diagonal of the posterior state covariance matrix according to the position index, and use the element as the variance estimate of the sensitivity parameter and as the uncertainty information.

5. The adaptive control method for a renewable resource production line based on reinforcement learning according to claim 1, characterized in that, The formation of the timing differential error signal includes: The value network calculates the current state value estimate based on the current state vector, and obtains the instantaneous feedback signal generated by the production line operation and the next moment state value estimate after discount factor decay; the current state value estimate, the instantaneous feedback signal, and the decayed next moment state value estimate are combined to obtain the time series difference error signal.

6. The adaptive control method for a renewable resource production line based on reinforcement learning according to claim 1, characterized in that, The control gradient algorithm based on reinforcement learning is a deep deterministic policy gradient algorithm. The online self-tuning of network weights specifically includes two distinct but synchronous update processes: The value network weight update is achieved by constructing a value loss function by calculating the mean square error of the temporal difference error signal and adjusting the weights along the negative gradient direction of the loss function. The policy network weight update is achieved by using the output of the value network as an evaluation of the policy performance and adjusting the weights along the gradient direction of the evaluation value relative to the policy network parameters.

Citation Information

Patent Citations

  • Adaptive prediction control method based on Hammerstein model

    CN103268069A

  • Factory-level real-time scheduling method for wafer circle in semiconductor manufacturing based on deep reinforcement learning

    CN119398463A

  • Igniter robot production line control method and system based on deep reinforcement learning

    CN119579120A

  • Speed planning optimization method and system for repeated path operation of industrial robot

    CN120921392A

  • Friction stir welding control method and equipment

    CN120962090A