Event-triggered self-tuning PID temperature control system and method

By using an event-triggered self-tuning PID control system, combined with dynamic thresholds and reinforcement learning, the PID parameters and utility function weights are adjusted in real time. This solves the problems of response delay and computational efficiency in nonlinear environments of traditional PID control, and achieves high robustness and precise temperature control.

CN121559843APending Publication Date: 2026-02-24NINGBO ZHISHENG OVEN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511615704.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Traditional PID control suffers from response delays and instability in nonlinear and uncertain environments. Event-triggered control has inaccurate triggering timing and low computational efficiency. Basic game theory methods have low resource utilization and lack adaptive optimization capabilities.

Method used

By combining dynamic thresholds, state potential game theory, and reinforcement learning, an event-triggered self-tuning PID control system is used to adjust PID parameters and utility function weights in real time, thereby optimizing resource consumption and performance indicators.

Benefits of technology

Achieve high robustness and precise temperature control in complex environments, thereby improving system stability and operating efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121559843A_ABST
    Figure CN121559843A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial automation control, and discloses an event-triggered self-tuning PID temperature control system and method. According to the system, a parameter updating loop is activated when specific conditions are met through an event triggering mechanism, and game theory self-learning and a preset performance control method are integrated into an extended PID controller structure through the loop, so that dynamic adjustment of parameters is achieved. The system can respond to temperature deviation and external disturbance in real time, and therefore multiple performance indexes such as overshoot, stabilization time and energy consumption are optimized at the same time. All the modules jointly optimize a control strategy through data interaction and function cooperation so as to reduce the computing resource consumption of the system. In a resource limited scene, the mechanism improves the robustness of the system. Compared with a traditional PID algorithm, the method is higher in adaptability in an uncertain environment. Through innovative combination of event triggering and reinforcement learning, the method aims at overcoming the limitations of response delay, remarkable overshoot, resource waste and the like of the existing algorithm in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of industrial automation control technology, specifically relating to an event-triggered self-tuning PID temperature control system and method, which is suitable for precise temperature control in nonlinear and uncertain environments such as chemical processes, printing presses, or reactors.

[0002] Definition: In this invention, "extended PID controller" refers to a proportional-integral-derivative controller that supports higher-order affine nonlinear systems. By adding higher-order terms, it overcomes the limitation of traditional PID controllers, which are only applicable to second-order systems. Its control law is defined as follows: in This is the output of the controller, i.e., the control signal. The gain coefficient of the integral term. (t) represents the adjustment error. This indicates the deviation between the system output and the desired set value. This is the gain coefficient of the proportional term. Here are the gain coefficients for each order of differential terms. Higher-order differential terms enable the controller not only to reflect the current state, cumulative effects, and trends, but also to sense more advanced dynamic characteristics (such as changing acceleration, jerk, etc.), thus providing the ability to stabilize and control complex high-order nonlinear systems.

[0003] State-based potential game is a game theory framework in which players (PID parameters) collaborate to optimize the overall state function, ultimately converging to Nash equilibrium. The "game theory self-learning module" refers to a self-learning algorithm based on state-based potential game theory, used to dynamically adjust PID parameters. Background Technology

[0004] Existing temperature control algorithms mainly rely on traditional PID control, event-triggered control, or basic game theory methods. Traditional PID controls perform well under stable conditions, but their response is delayed when faced with disturbances or changes in the operating point, leading to overshoot or instability and making them unsuitable for complex situations. Traditional event-triggered control often uses fixed thresholds, which can result in inaccurate triggering timing in nonlinear systems and also has low computational efficiency. While basic game theory methods can achieve partial cooperation, they suffer from low resource utilization and, in particular, lack the ability to perform adaptive optimization under multi-objective constraints.

[0005] With advancements in artificial intelligence algorithms, lightweight frameworks and reinforcement learning algorithms such as Q-learning can efficiently process dynamic data, providing a technological foundation for the system's self-learning process. Event-triggered mechanisms support dynamic parameter adjustment, updating only when necessary, reducing computational burden. The potential of these technologies in temperature control has not yet been fully explored. This invention innovatively integrates these technologies, providing an efficient solution: by using event-triggered mechanisms for local monitoring and parameter optimization, the risk of ineffective updates is avoided, while maintaining multi-objective optimization effects. Summary of the Invention

[0006] This invention relates to the field of industrial automation control technology, specifically an event-triggered self-tuning PID temperature control system and method. This scheme combines dynamic thresholding, state potential game theory, and reinforcement learning, and is designed for applications such as chemical processes, printing equipment, and reactors. It overcomes the shortcomings of traditional PID in terms of adaptability, computation, and efficiency, providing stable temperature control. The core of the "multi-objective optimization" described in this invention is the automatic adjustment of utility function weights through a reinforcement learning framework. This mechanism can dynamically weigh various performance indicators during operation, ultimately optimizing resource consumption while ensuring temperature control accuracy. The core solution of this invention is as follows: Step S1: Initialize PID parameters and utility function weights. Configure initial PID coefficients to reconcile control error with resource consumption targets; the weights of the utility function are dynamically adjusted according to preset scenarios or control modes (such as fast response mode or energy-saving mode).

[0007] Step S2: Real-time acquisition of temperature data and calculation of control error. This step uses a processing module to accurately calculate the deviation and employs a periodic sampling strategy to adapt to environmental changes.

[0008] Step S3: Determine the event trigger condition. If satisfied, activate the game theory self-learning module to update parameters. This step is responsible for the core function of online parameter optimization. It makes judgments by combining dynamic thresholds and error change rates, and introduces a game theory collaboration mechanism to improve the system's robustness in heterogeneous scenarios while reducing the computational load on resource-constrained devices. The self-tuning PID controller module includes a parameter update unit that receives the PID parameters output from the game theory self-learning module and reconstructs the control law, while also summarizing system status and performance indicators as self-learning input.

[0009] Step S4: Reconstruct the control law using a preset performance function and output the control signal. The preset performance function quickly constrains the system error within predetermined boundaries and is adjusted in conjunction with the actuator's actions to ensure efficient and accurate output of the control signal.

[0010] Step S5: Optimize the utility function weights through reinforcement learning feedback. Fine-tune the weights based on performance data, and coordinate adjustments with cloud computing resources to achieve continuous adaptive optimization of the system weights.

[0011] The technical solution of this invention addresses the shortcomings of traditional temperature control by providing a control method that combines high robustness and accuracy in uncertain environments, thereby effectively improving the stability and operating efficiency of the system. Attached Figure Description

[0012] Figure 1 This is an overall block diagram of the event-triggered self-tuning PID method and system of the present invention. The system collects the temperature of the controlled object in real time through sensors, and the error signal is obtained through the acquisition layer. At the same time, the set temperature and the error signal are sent to the trigger layer (event trigger module and game theory self-learning). The former judges the conditions, and the latter realizes parameter updates. The two together drive the output layer to generate signals to the actuator, which then acts on the controlled object to form a closed-loop control process; Figure 2 This is a flowchart illustrating the event triggering process of this invention. The process begins with "Start" and proceeds to data acquisition and error calculation e(t). Subsequently, based on the dynamic threshold δ(t) and the error change rate de / dt, triggering conditions are determined: when e(t) > δ(t) or de / dt exceeds the threshold, the game theory self-learning module is activated to update the PID parameters; if the system is in a stable state, the current parameters are maintained; if a disturbance is detected, the optimization program is immediately activated for rapid adjustment. After processing each branch, a preset performance function is called to reconstruct the control law and output the control signal, then the process returns to "Start" for continuous monitoring, forming a closed-loop operation. Figure 3 This is a temporal interaction diagram of the game theory self-learning mechanism of this invention. Within each optimization cycle, the parameter update unit of the self-tuning PID controller module first summarizes the system state / performance indicators and sends them to the game theory self-learning module. The game theory self-learning module performs high-frequency adjustments to the PID parameters based on the state potential game and outputs the updated parameters. The system utility function receives the parameters and calculates the global utility using the state. The evaluation result, along with the updated parameters, is fed back to the self-tuning PID controller module (parameter update unit) for control law reconstruction, forming a closed loop. This process coordinates with the weight update of reinforcement learning: the utility function weights are optimized at a lower frequency by Q-learning feedback based on historical performance data, thereby achieving continuous optimization of "high-frequency parameter adjustment—low-frequency weight optimization—closed-loop feedback." Figure 4This diagram illustrates the internal structure of the reinforcement learning module of this invention. The module takes the error signal e(t), the error rate of change de / dt, and the current weight vector as input, and establishes a state-action value table through the core of the Q-learning algorithm. The algorithm quantifies and evaluates control performance based on the reward function. Through a weight update mechanism, the weight parameters of the utility function are adjusted based on reward feedback in each training cycle to achieve multi-objective balance optimization. The output layer generates a new weight vector, which is used by the utility function optimization unit to calculate the control decision. This decision, combined with the output of the event triggering module, forms the final control signal, achieving continuous adaptive adjustment and closed-loop optimization. This module enhances the robustness and generalization ability of the system under dynamic operating conditions. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] It should be noted that in this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0015] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

[0016] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to an embodiment of an event-triggered self-tuning PID temperature control system and method in this application, which includes: Step S1: Initialize PID parameters and utility function weights In this step, the present invention achieves multi-objective initialization by setting a utility function. First, the system is initialized, including using an extended PID framework as the control model to handle temperature state information (current temperature, set temperature, disturbance level, and rate of change). Temperature sensors collect the current temperature of the device in real time and make a preliminary comparison with the set value to calculate the initial deviation. Simultaneously, a reinforcement learning component is introduced offline for pre-definition, using historical data from multiple different processes, loads, and device models to construct a scenario set. Each scenario is constructed by selecting a continuous temperature sequence (50-300 data points) under a specific process and load and injecting simulated disturbances. This sequence is then divided into a support set (for state-action training in Q-learning) and a query set (for evaluating configuration effectiveness and updating weights).

[0017] Based on the processed temperature state, the utility function is set as overshoot minimization and energy consumption minimization, and transformed into a multi-objective function through a weighted sum. Through the above initialization process, the system obtains an initial state that balances response speed and resource regulation during the startup phase, providing the necessary conditions for subsequent online adaptive optimization.

[0018] Step S2: Collect temperature data in real time and calculate the control error e(t). This invention achieves data optimization through a processing module. The system operates a data sampling module integrating filtering algorithms (such as Kalman filtering) to suppress noise and thus accurately calculate control errors. Simultaneously, the system dynamically adjusts the event triggering threshold based on the current load. The system dynamically divides the control process into multiple stages based on the error and its rate of change: in the initial stage, due to a large deviation, the system meets the event triggering conditions, thereby activating parameter updates; in the stable stage, the error is small, and the system maintains the current parameters; in the disturbance stage, the error rate of change is high, and the system immediately triggers the optimization program. The error for each stage is generated by the module, and the specific calculation basis includes the temperature change rate, disturbance level, and noise indicators, etc.

[0019] To ensure compatibility with online characteristics, a periodic sampling mechanism is adopted: every certain number of steps (e.g., every M steps, where M can be set to 5-20 depending on the equipment capacity), the error is recalculated using new data; during the interval, the old error continues to be used for monitoring. This mechanism effectively reduces the system's computational overhead while ensuring control accuracy. The calculation results are temporarily stored and passed as feedback to the event triggering module, thus forming a closed-loop collaboration of "real-time high-frequency calculation guiding optimization judgment".

[0020] Step S3: Determine the event triggering condition; if it is met, activate the game theory self-learning module to update parameters. In this step, parameter optimization is achieved through an event-triggered module. Each device makes operational decisions based on error data evaluation conditions. Subsequently, a dynamic threshold formula is used to share signals only when parameters need updating, rather than continuously transmitting raw data, to address differences in uneven system load. To handle variations in computing power among different devices, the system supports an asynchronous decision-making mechanism, prioritizing nodes with sufficient computing resources to perform parameter optimization tasks.

[0021] This mechanism introduces an error change rate and balances performance and resource consumption by finely controlling a threshold δ (typically set to 0.1-1): in high-response scenarios, δ is reduced to improve system sensitivity; in low-response scenarios, δ is increased to conserve resources. This keeps the performance impact of triggered behavior within a reasonable range, effectively avoiding over-triggering. Specifically, the threshold is adjusted using a moving average, such as... in A smoothing factor (typically 0.9) is used, and ε is a very small positive number (e.g., 0.001) to prevent division by zero errors. Additional measures include encrypted transmission of updates and random perturbations, with log verification to block invalid activations and avoid resource waste. The judgment period is fixed at periodicity to ensure that parameters are updated consistently with the threshold.

[0022] When the triggering condition is met, the game theory self-learning module is activated. This module uses a state-potential game framework, where the PID parameters Kp, Ki, and Kd act as players. Through collaboration, they optimize the global utility function U to update the parameters. In specific applications, the utility function of each player (e.g., Kp) is... in For the system's settling time, This represents the maximum overshoot or undershoot. and As auxiliary weight parameters, For the barrier function, Control state variables The 1-norm represents a cumulative measure of control deviation during the event. Measurable disturbance state variables The 1-norm quantifies the extent to which external disturbances affect the system during an event. It is a small positive number to prevent the denominator in the formula from being zero. Global Situation Function Ensure convergence to Nash equilibrium. The action set is a continuous range. ( The update rule uses a gradient method: in This is the step size. After triggering, the error and status information are summarized by the parameter update unit of the self-tuning PID controller module and sent to the game theory self-learning module for parameter learning and updating. Step S4: Reconstruct the control law using a preset performance function and output the control signal. The parameter update unit receives updated Kp, Ki, and Kd to reconstruct the control law and output control signals, including PID coefficients and error constraints, to achieve temperature adaptation and optimization. A preset performance function is introduced to enable the algorithm to quickly constrain errors: the function component is predefined offline and then invoked when the device is running online. Specifically, when the event triggering condition is met, the system automatically invokes the predefined performance function to quickly reconstruct the control law. The output process is quickly configured based on the support set and evaluated using the query set. The collaborative workflow between the function and the event triggering module is as follows: First, the event triggering module provides an activation signal and the current error state as input to the preset performance function. Second, offline predefined historical multi-scenario data is used to enhance the function's constraint capability. Finally, combined with real-time calculated data and parameter updates, rapid control response is achieved.

[0023] The specific form of the preset performance function is as follows: in This is the initial performance bound (typically 1). This is the final performance limit (typical value 0.1). This represents the decay rate (typically 0.05). This function is used to reconstruct the control law. in To correct the coefficients and ensure that the tracking error remains within a preset neighborhood, a support set is used for state-action training in Q-learning, while a query set is used to verify and adjust the reconstruction effect. The temperature is adjusted via actuators, and key performance indicators such as temperature fluctuation range, resource consumption, and response time are evaluated online. This adaptive framework for dynamic reconstruction effectively improves the speed, stability, and robustness of the temperature control system, maintaining good temperature control accuracy even under dynamic and uncertain environments.

[0024] Step S5: Optimize the utility function weights through reinforcement learning feedback. Weights are optimized based on historical performance data, including utility weighting and multi-objective balancing, to achieve adaptive and continuous optimization. The Q-learning component is deployed using a hybrid offline-online model: it is first pre-trained offline locally or in the cloud, and then the trained model is deployed on the device for online feedback optimization. The advantage of this model is that it front-loads high-computation tasks, thus avoiding the burden of real-time model training on resource-constrained devices. Specifically, when the system detects performance changes, it automatically calls upon the model weights obtained from cloud pre-training and efficiently adapts them to the system. The feedback process is based on rapid adaptation using the support set, and the effect is evaluated using the query set.

[0025] The collaborative logic between the game theory self-learning module and Q-learning is as follows: the game theory self-learning module is responsible for the real-time optimization of PID parameters, ensuring convergence through state potential game theory; Q-learning adjusts the utility function weights based on historical data, forming a feedback loop. Specifically, the Q-learning framework is designed as follows: State ,action =Weight adjustment direction and step size (e.g.) ),award The collaborative mechanism for learning and optimization comprises three stages: First, the parameter optimization module transmits global performance metrics to the Q-learning algorithm as the basis for updating weights; second, the algorithm pre-trains using historical multi-task data to enhance generalization ability; and third, the feedback stage combines local data and optimization results to achieve rapid weight adaptation. A support set is used for state-action training of Q-learning, while a query set is used to evaluate the optimization effect.

[0026] In one specific embodiment, the process of performing step S1 may specifically include the following steps: (1) The initial temperature of the controlled object is collected by a temperature sensor and compared with the set value to calculate the initial deviation. Then, a filtering algorithm (such as Kalman filtering) is called to denoise the raw data and normalize the temperature data to provide standardized input for subsequent control calculations.

[0027] (2) In the controller, the weights of the extended PID parameters and utility function are initialized according to the selected control mode (such as fast response or energy saving mode) to achieve the initial balance among multiple objectives.

[0028] (3) Based on the above initialization parameters, generate initial control signals and send them to the execution module. At the same time, initialize the reinforcement learning component to prepare for online optimization after the system is put into operation.

[0029] In one specific embodiment, the process of performing step S2 may specifically include the following steps: (1) The system periodically calls the data sampling algorithm to calculate the control error e(t) and its rate of change in real time based on the current temperature and the set value. This algorithm has real-time scheduling capability and can prioritize the processing of key data points to ensure control accuracy.

[0030] (2) The system can dynamically adjust the monitoring strategy according to the operating status. For example, when the detected disturbance amplitude exceeds the preset threshold, the system automatically increases the calculation and monitoring frequency of the error change rate in order to capture the rapid changes in the dynamic process.

[0031] (3) To reduce computational overhead, during non-core sampling periods, the system continues to use the error value from the previous cycle for status monitoring. Newly acquired data is sent to the event triggering module as the basis for judgment. This mechanism effectively saves computational resources while ensuring control continuity.

[0032] In one specific embodiment, the process of performing step S3 may specifically include the following steps: (1) The event triggering module receives the control error e(t) and its rate of change from S2. Based on this, the triggering parameters are initialized or updated according to the dynamic threshold formula, for example, the smoothing factor β is set to 0.9 to limit the impact of the instantaneous deviation of the error on the system.

[0033] (2) To address the possibility of multiple control loops in the system, this mechanism supports asynchronous judgment. When the error characteristics of different loops show significant differences, nodes with stronger computing power are prioritized to perform parameter optimization tasks, thereby improving the overall system response efficiency. The error change rate is incorporated into the judgment logic as one of the key triggering conditions.

[0034] (3) When the triggering condition is met, the game theory self-learning module is activated, and the PID parameters are updated based on the state potential game model. The system synchronously tracks the control performance improvement and computational resource consumption brought about by the parameter update, thereby ensuring control robustness in high-disturbance environments.

[0035] In one specific embodiment, the process of performing step S4 may specifically include the following steps: (1) The self-tuning PID controller calls a preset performance function to reconstruct the control law and constrain the tracking error boundary. The system loads the corresponding offline predefined parameters (such as initial performance boundary, decay rate, etc.) according to the current operating scenario to initialize the function.

[0036] (2) During operation, once the event triggering module issues an activation signal, the controller immediately uses this performance function to correct the control signal online. This function integrates knowledge trained from historical data, enabling fast and effective error constraints even based on limited local real-time data.

[0037] (3) Finally, the reconstructed control signal is output to the actuator (such as a heater or cooling device) to directly drive temperature regulation. The system monitors and records key indicators such as temperature fluctuation range and response time in real time to verify the effectiveness of the preset performance control.

[0038] In one specific embodiment, the process of performing step S5 may specifically include the following steps: (1) The reinforcement learning module receives historical performance data (such as overshoot, settling time, and energy consumption) and adjusts the weights of each objective in the utility function accordingly. This process first loads the parameters of the pre-trained Q-learning model in the cloud for initialization to ensure a high quality starting point for optimization.

[0039] (2) The system continuously monitors the control performance determined by the current weights. When the evaluation results indicate that there is room for performance improvement, an online adaptation mechanism based on a small amount of new data is triggered to combine local real-time data with system updates and quickly fine-tune the weights so that the multi-objective optimization strategy closely follows the dynamic operating conditions.

[0040] (3) The updated weights are applied to the calculation of the utility function, thereby affecting the optimization direction of the game theory self-learning in the next cycle. The system comprehensively evaluates the combined impact of this weight adjustment on multiple competitive indicators (such as control accuracy and energy consumption), thereby driving the overall system utility to continuously approach the optimal equilibrium.

Claims

1. An event-triggered self-tuning PID temperature control system and method, characterized in that, The system includes a temperature acquisition module, a self-tuning PID controller module, an event triggering module, a game theory self-learning module, a preset performance module, and an execution module. The method includes the following steps: S1: Initialize PID parameters and utility function weights. Set initial proportional coefficient Kp, integral coefficient Ki, and derivative coefficient Kd, and define the utility function in a weighted sum form to balance overshoot, settling time, and energy consumption. The weights can be adjusted according to the application scenario (such as fast response mode or energy-saving mode) to ensure good adaptability of the system in the initial stage; S2: Real-time temperature data acquisition and control error calculation. The sensor collects the difference between the actual temperature and the set temperature, and calculates the error rate of change and environmental disturbances. S3: Determine the event trigger condition. If it is met, activate the game theory self-learning module to update the parameters. The trigger condition is determined based on a dynamic threshold and the error rate of change, and the parameters are collaboratively optimized using a state potential game framework. S4: Reconstruct the control law using a preset performance function and output the control signal. A monotonically decreasing function ensures the error is within a preset neighborhood, and the final control output is generated and sent to the actuator. S5: Optimize utility function weights through reinforcement learning feedback. The Q-learning algorithm is used to automatically adjust weights based on historical performance data, achieving continuous multi-objective optimization.

2. The system and method according to claim 1, characterized in that: The utility function in S1 is based on multi-objective computation, used to handle dynamic environments, and supports rapid startup through initial weights.

3. The system and method according to claim 1, characterized in that: The data acquisition in S2 employs a real-time processing module. This module uses a filtering algorithm for noise reduction by default, but edge computing integration can also be selected. It uses periodic sampling: data is collected at intervals to support subsequent judgments.

4. The system and method according to claim 1, characterized in that: The event triggering mechanism in S3 uses a dynamic threshold formula to handle nonlinear changes. The threshold is used in an exponential form to limit error deviation, while simultaneously updating the game parameters collaboratively.

5. The system and method according to claim 1, characterized in that: The control law reconstruction in S4 is achieved through a preset performance function. The function component is predefined offline, based on historical datasets from multiple scenarios, with each scenario divided into a support set and a query set.

6. The system and method according to claim 1, characterized in that: The utility function optimization in S5 employs a reinforcement learning module. This module uses the Q-learning algorithm to process historical data. It employs a feedback mechanism: after performance evaluation, the weights are adjusted to achieve continuous optimization.