Temperature dynamic control system for optical fiber rod processing
By integrating data from multi-macro camera arrays, sensor networks, and edge computing nodes, combined with dynamic mathematical prediction and reinforcement learning algorithms, the multi-variable collaborative adjustment problem of the temperature control system in optical fiber rod processing was solved, achieving high-precision and efficient temperature control.
Patent Information
- Application Number
- CN202510939839.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-09
AI Technical Summary
The existing temperature control system for optical fiber rod processing has significant limitations in multi-variable coordinated regulation, making it difficult to balance control stability and dynamic response performance. It also lacks the fusion analysis of multi-source heterogeneous data, resulting in insufficient temperature control accuracy and response efficiency.
The acquisition unit consists of a multi-macro camera array, a production line sensor network, and edge computing nodes, combined with a multi-input and multi-output dynamic mathematical prediction model and a reinforcement learning algorithm to achieve real-time perception, accurate prediction, and adaptive control of the multi-dimensional status of the production line.
It significantly improves the temperature control accuracy and response efficiency, ensures the consistency of the temperature control process and the closed-loop operation of the system control logic, adapts to complex environmental changes, and improves the adaptability and robustness of the system.
Smart Images

Figure CN120447660B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of control systems, and in particular to a temperature dynamic control system for optical fiber light rod processing. Background Art
[0002] In current industrial manufacturing scenarios such as optical fiber rod processing, temperature control systems usually rely on sensor networks deployed on the production line to obtain environmental parameters such as temperature, and adjust the temperature through empirical rules or traditional control algorithms.
[0003] In the manufacturing process of optical fiber rods, temperature control accuracy directly impacts the dimensional consistency and material properties of the product. This process involves multiple highly coupled temperature zones, with complex heat conduction and feedback relationships between them. Traditional methods based on manual adjustment or static threshold judgments have significant limitations in multivariable coordinated adjustment, often failing to balance control stability and dynamic response performance. Furthermore, due to the lack of integrated analysis of multi-source heterogeneous data (such as images, tension, and temperature), existing systems' information utilization at the perception level is far from ideal, making it difficult to support high-precision predictive modeling and strategy optimization.
[0004] In view of the above shortcomings, it is necessary to propose a temperature control system architecture with data fusion, intelligent prediction, adaptive optimization and control instruction integration capabilities to meet the actual needs of multi-temperature zone dynamic control scenarios such as optical fiber and light rod processing. Summary of the Invention
[0005] The present application provides a temperature dynamic control system for optical fiber light rod processing to improve the accuracy and response efficiency of temperature control in industrial production processes.
[0006] The present application provides a temperature dynamic control system for optical fiber light rod processing, comprising:
[0007] The acquisition unit includes a multi-macro camera array, a production line sensor network, and an edge computing node. The multi-macro camera array is used to capture imaging data of production line materials. The production line sensor network includes an environmental perception unit and an equipment status sensor array for acquiring temperature, tension, humidity, and speed data. The edge computing node is used to fuse and pre-process the imaging data and acquired data and upload them to the cloud to form comprehensive production status data.
[0008] a control unit configured to receive the comprehensive production status data and establish a dynamic mathematical prediction model for a multi-input multi-output system; predict future temperature states of multiple temperature zones of the production line based on the dynamic mathematical prediction model, and generate temperature zone optimization control input instructions based on the prediction results;
[0009] The learning unit is configured to receive the temperature zone optimization control input instruction and historical temperature data, and generate an adaptive control strategy based on the temperature zone optimization control input instruction and the historical temperature data through a reinforcement learning algorithm.
[0010] The integration unit is configured to receive the temperature zone optimization control input instruction and the adaptive control strategy, integrate the two to form a final control instruction, and send the final control instruction to a production line temperature control execution device to adjust the temperature of each temperature zone.
[0011] The present application has the following beneficial technical effects: (1) The acquisition unit composed of a multi-macro camera array and multiple types of sensors can realize real-time sensing of the multi-dimensional state of the production line (including image, temperature, tension, humidity, speed, etc.), and the fusion processing is performed through the edge computing node, which significantly improves the timeliness and integrity of the data. (2) The control unit adopts a multi-input and multi-output dynamic mathematical prediction model, which can comprehensively consider the coupling relationship between the temperature zones, realize accurate prediction of the future temperature state of multiple temperature zones and generation of coordinated control input, and improve the temperature control precision. (3) The learning unit introduces historical data and current instructions into joint modeling through a reinforcement learning algorithm, has adaptive ability, can continuously optimize the control strategy according to the working condition change, and improves the adaptability and robustness of the system to complex environment. (4) The integration unit fuses the control strategies generated by the control unit and the learning unit to form a unified final control instruction, thereby realizing the coordinated regulation of each temperature zone and ensuring the consistency of the temperature control process and the closed-loop operation of the system control logic. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 is a schematic diagram of a temperature dynamic control system for optical fiber rod processing provided by the first embodiment of the present application. DETAILED DESCRIPTION
[0013] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced in a variety of ways beyond the specific embodiments described herein without departing from the scope of the present application, and those skilled in the art can make similar modifications without departing from the spirit of the present application, so the present application is not limited to the specific implementations disclosed below.
[0014] The first embodiment of the present application provides a temperature dynamic control system for optical fiber rod processing. Please refer to Figure 1 , which is a schematic diagram of the first embodiment of the present application. The following will be described in detail Figure 1 The first embodiment of the present application provides a temperature dynamic control system for optical fiber rod processing. Please refer to
[0015] The temperature dynamic control system for optical fiber rod processing includes a collection unit 101 , a control unit 102 , a learning unit 103 and an integration unit 104 .
[0016] The acquisition unit 101 includes a multi-macro camera array, a production line sensor network and an edge computing node. The multi-macro camera array is used to capture imaging data of materials on the production line. The production line sensor network includes an environmental perception unit and an equipment status sensor array for obtaining temperature, tension, humidity and speed data. The edge computing node is used to fuse and pre-process the imaging data and the acquired data and upload them to the cloud to form comprehensive production status data.
[0017] Acquisition unit 101 is used to collect and initially process high-precision, multi-dimensional data on key process states in industrial manufacturing environments, providing fundamental data support for subsequent temperature control modeling and decision-making. Acquisition unit 101 specifically comprises a multi-macro camera array, a production line sensor network, and edge computing nodes, all interconnected via wired or wireless communication to form a complete edge sensing architecture.
[0018] The multi-macro camera array consists of several industrial-grade macro cameras installed at key points on the production line. They preferably utilize an optical system with a resolution of at least 5MP, a frame rate exceeding 30fps, and zoom and low-light performance. They are installed at a distance of 10–30cm to accommodate the imaging requirements of the microscopic surface structures of materials such as cables, optical fibers, or optical rods. Each camera is connected to an edge computing node via an industrial interface (such as GigE, USB 3.0, or MIPI). The cameras capture images at a set interval (e.g., every second) and perform basic pre-processing on the images, including grayscale balancing, edge sharpening, and object identification (e.g., determining the melting state and radial shrinkage of the optical rod surface).
[0019] The production line sensor network includes an environmental sensing unit and an equipment status sensor array. The environmental sensing unit is used to obtain real-time environmental parameters related to temperature control, such as temperature sensors (such as thermocouples or PT100s, with a measurement accuracy of ±0.2°C), humidity sensors (such as capacitive humidity sensors, with a measurement range of 10%–90% RH and an accuracy of ±2%), air velocity meters, and ambient light sensors. The equipment status sensor array is used to collect real-time process parameters related to temperature control in the production line. Typical examples include tension sensors (such as strain gauge tensiometers, with a range of 0–100N and an accuracy of ±1%FS), speed sensors (such as magnetoelectric tachometers or encoders, with a measurement range of 0.1–5m / s), and equipment temperature sensors (such as embedded thermocouple arrays placed on the inner walls or core areas of multiple temperature zones, with a sampling period of no more than 1 second).
[0020] The edge computing node is composed of an embedded processor with edge intelligence capability, and is configured with a local memory of not less than 4 GB and a storage capacity of not less than 64 GB. The edge computing node integrates a multi-channel data acquisition module, which can receive the camera image stream and the sensor data stream in parallel. The edge computing node is built-in with a fusion processing algorithm, which specifically includes: image and physical data time alignment processing, sensor data filtering and interpolation algorithm, data standardization (such as normalization or Z-score conversion), time sequence sliding window structure feature vector construction and other steps. After the above processing, the edge computing node forms a complete production state feature data (including image coding features, temperature vector, tension vector, speed scalar, etc.) in each sampling period, and then packs it in JSON or Protobuf format and uploads it to the cloud or the control unit 102 through an industrial network (such as Ethernet / IP, 5G or industrial Wi-Fi).
[0021] The design of the acquisition unit 101 can ensure that the millisecond-level data update frequency and edge processing response are realized under dynamic working conditions, and the state perception accuracy is improved through heterogeneous data fusion, a high-credibility input data basis for subsequent control modeling and strategy learning is constructed, so as to meet the dependence of the industrial temperature control system on multi-source, high-frequency and stable data stream.
[0022] Further, the acquisition unit is also used for:
[0023] The edge computing node is based on the imaging data captured by the multi-macro camera array and the temperature, tension, humidity and speed data obtained by the production line sensor network, and establishes a sliding time window structure according to a preset time step, which is used to construct a time sequence nested state vector; wherein the sliding time window structure is used to continuously store the imaging data features and the corresponding temperature, tension, humidity and speed data of the previous M time steps in each sampling period, the imaging data features are extracted into fixed-length numerical vectors by an image processing algorithm, and the corresponding sensor data are spliced to form a multi-dimensional state vector;
[0024] The edge computing node performs cross-time step fusion processing on the state vector in the sliding time window structure to form a nested time sequence feature representation, which is used to reflect the production state change trend of the target material in a continuous time period;
[0025] The edge computing node uploads the nested time sequence feature representation as part of the production state comprehensive data to the cloud, and provides the control unit for establishing a dynamic mathematical prediction model of the multi-input multi-output system.
[0026] The acquisition unit includes edge computing nodes, which further perform time series structure construction and cross-time fusion operations during the data acquisition and processing process. Specifically, the edge computing node receives raw imaging data from a multi-macro camera array and simultaneously receives temperature, tension, humidity, and speed data from the production line sensor network. This data is time-series, that is, it is collected once at each time point, forming a time series.
[0027] To capture the dynamic trends of variables in the production process, edge computing nodes construct a sliding time window structure with a fixed length M (for example, M = 5) within each sampling period (e.g., once per second or minute). This structure continuously retains M sets of data from the current time point and the previous M–1 time points, continuously sliding forward to maintain the latest status information. The data at each time point includes imaging features and multiple sensor data.
[0028] When processing imaging data, edge computing nodes use preset image processing algorithms (such as grayscale conversion, edge detection, principal component analysis, and convolutional feature extraction) to compress each image frame into a fixed-length numerical vector, such as a 128-byte floating-point array. This vector describes key visual features such as the target material's surface condition, brightness variations, and edge morphology. Sensor data, including current temperature, tension, humidity, and speed, is arranged in a predefined order to form a complete physical state vector.
[0029] The edge computing node then concatenates the imaging feature vector at each time point with its corresponding sensor vector to form a multidimensional state vector. For example, if the image features are 128-dimensional and the sensor data consists of four items, each state vector is 132-dimensional. The sliding time window structure thus stores M 132-dimensional state vectors, forming a two-dimensional matrix of input data, denoted as an M × 132 time series tensor.
[0030] This structure not only preserves the instantaneous state at each point in time but also, through structural storage, captures dynamic evolution information over continuous time periods. Edge computing nodes further perform cross-timestep fusion processing on this tensor structure. For example, through algorithms such as sliding average, recursive weighting, temporal convolution, or self-attention, they extract comprehensive nested temporal features that reflect dynamic temperature trends, tension patterns, or image feature drift. The fused result remains a fixed-length vector (e.g., 256 dimensions) but incorporates trend information along the temporal dimension and the interactions between multiple source variables.
[0031] Finally, the edge computing node uploads this nested time-series feature representation as a component of the "production state comprehensive data" to the cloud or control unit via an industrial network protocol (e.g., MQTT, OPC UA, or HTTP API). The control unit utilizes this feature data as input to a multi-input multi-output (MIMO) model to perform temperature state prediction and subsequent control instruction computation.
[0032] Through this mechanism, the edge computing node not only completes the fusion of static data, but also realizes the modeling input preparation of state changes in continuous time periods, effectively improving the perception ability of the temperature control model to process evolution, especially suitable for high-speed dynamic change industrial production scenarios.
[0033] The control unit 102 is configured to receive the production state comprehensive data and establish a dynamic mathematical prediction model of a multi-input multi-output system. The control unit 102 is further configured to predict future temperature states of multiple temperature zones of the production line based on the dynamic mathematical prediction model and generate temperature zone optimization control input instructions according to the prediction results.
[0034] The control unit 102 is mainly used to construct a mathematical model based on the multi-dimensional production state comprehensive data provided by the acquisition unit 101 to simulate and predict the thermal change trend of multiple temperature zones in the entire industrial production process, and generate temperature control input instructions accordingly. The control unit 102 internally integrates a dynamic modeling and prediction calculation framework for a multi-input multi-output system (MIMO). The basic logic includes three core steps of model establishment, state prediction, and control instruction generation, which are executed through a combination of hardware computing platforms and industrial control software in the specific implementation process.
[0035] In the aspect of model establishment, the control unit 102 first analyzes the production state comprehensive data uploaded by the acquisition unit 101. The data includes continuous variables such as temperature, tension, humidity, and speed at each time step, as well as image feature encoding results related to the optical rod processing. The control unit selects a nonlinear state space modeling method, such as extended Kalman filter (EKF), long short-term memory neural network (LSTM), or state-enhanced autoregressive model (ARX+), to model the thermal coupling relationship, thermal inertia, environmental disturbance, and other characteristics of each temperature zone. The model input is the historical state vector of each temperature zone at time t and the control variable input value corresponding to the acquisition data, and the model output is the temperature prediction value of each temperature zone at time t+Δt, where Δt is the control time window, usually set to 1 second to 5 seconds. Physical constraints such as the temperature change rate of each temperature zone not exceeding a certain range (e.g., ±5°C / s) are considered in the modeling process, and optimization goals such as minimizing the mean square value of the prediction error are set.
[0036] During the state prediction phase, the control unit solves the model for each temperature zone independently or jointly, predicting the short-term temperature trajectory based on the constructed multi-input, multi-output dynamic model. This prediction is based not only on current observations but also on temperature trends, device response delays, and the influence of heat conduction between adjacent temperature zones, as determined during the system identification phase, to improve prediction accuracy. For example, the future temperature prediction for the current temperature zone i will take into account the state evolution of zones i-1 and i+1, as well as the lag in the control command's impact on them.
[0037] Based on the above prediction results, the control unit generates temperature zone optimization control input instructions through a rolling optimization strategy. Specifically, the control unit calls the model predictive control (MPC) algorithm to search for the optimal control variable trajectory within the set prediction time domain (such as 10 time steps) so that the future temperature trajectory is as close as possible to the target set value (such as 600°C±2°C) while satisfying the system constraints. The optimization problem can be formalized as a constrained quadratic programming (QP) or nonlinear programming (NLP) problem. The variables are temperature control execution quantities such as control power, damper opening, and cooling flow. The objective function is the weighted sum of the sum of squared temperature errors plus the control smoothing term. The constraints include the control variable boundaries, the maximum power of the equipment, the safe temperature range, etc.
[0038] Ultimately, the control unit outputs a set of temperature zone optimization control input instructions—a set of quantitative control values for each temperature zone (for example, setting zone 1 to 45% heating power and zone 2 to 25% cooling flow). These instructions are then passed to the learning unit 103 for further adaptive strategy generation. These instructions also serve as important input for the final control decision. The control unit can implement online updates by adjusting model parameters and input structures in real time to address model deviations caused by changing production conditions.
[0039] The structure and algorithm logic of the control unit ensure the dynamic prediction capability of the thermal behavior of each temperature zone and the real-time generation of optimal control input, which is the key to achieving precise temperature control in the system.
[0040] In order to help those skilled in the art fully understand the implementation of the control unit in this specification, a simplified specific embodiment is provided below to explain in detail the construction and prediction process of the dynamic mathematical prediction model of the multi-input and multi-output system, and to explain how to generate optimized control input instructions based on this.
[0041] Consider a light rod heating system with three temperature zones: Zone A, Zone B, and Zone C. Each zone is controlled by an independent heater. However, due to heat conduction, the zones are coupled. Heating one zone affects not only its own temperature but also adjacent zones. Furthermore, for simplicity, assume that the input to each heater is heating power expressed as a percentage (0%–100%). The system collects temperature data for each zone every minute.
[0042] In this example, the control unit first builds a simplified prediction model of the multiple-input multiple-output (MIMO) system as follows:
[0043] Assumptions:
[0044] Respectively represent the temperature of each temperature zone at time k;
[0045] denote the input power of each heater at time k;
[0046] For the control period, set it to 1 minute;
[0047] The constructed model is a first-order linear difference model, which is as follows:
[0048]
[0049] in, etc. are the temperature state transfer coefficients, reflecting the degree of coupling between temperature zones. To control the gain coefficient, it can be obtained by fitting historical production data. For example, the following coefficients are obtained through experiments
[0050]
[0051] The control unit obtains comprehensive data on the current production status, such as:
[0052]
[0053] Assume the current heating input is
[0054] Substituting into the model we can predict the temperature at the next moment:
[0055]
[0056] The control unit compares the target temperature setpoint (e.g. Target = 600°C, Target = 600°C, Target = 600 °C), it is found that the current control input is not enough to reach the target, so control optimization is needed.
[0057] At this time, the control unit executes a simple optimization strategy, for example, adjusting the input power to minimize the sum of squared temperature errors while ensuring that the device's rated power is not exceeded. For example, using gradient descent or enumeration optimization, the control unit can try the following control input combination:
[0058] The adjusted control input is again substituted into the model to calculate the predicted temperature, and the control input combination with the smallest error is selected as the zone optimization control input instruction and sent to the next module.
[0059] The above examples illustrate how the control unit establishes a MIMO dynamic model, predicts future zone temperature states based on the model, and generates zone optimization control input instructions accordingly. This process can be implemented on an industrial control platform (such as a PLC or industrial edge computing gateway) through conventional control software.
[0060] Furthermore, the control unit is also used to:
[0061] When establishing the dynamic mathematical prediction model of the multi-input multi-output system, a covariance penalty term for suppressing the coupling effect between zones is introduced when constructing the objective function, and the objective function is composed of a prediction error square term and a covariance term of multiple zone control quantity changes, wherein the covariance term is used to measure the joint fluctuation degree between control inputs of each zone at different times;
[0062] When solving the objective function, the optimal zone optimization control input instruction is obtained through an optimization algorithm, which limits the mutual interference of control changes between zones while meeting the convergence of each zone target temperature, avoiding system oscillation caused by coupling gain amplification;
[0063] Within each preset control period interval, based on the prediction error of the production state comprehensive data and the actual response of the current zone, the state transition matrix parameters of the multi-input multi-output system are updated using an online least squares method or a gradient descent algorithm to improve the adaptability and robustness of the dynamic mathematical prediction model to changes in working conditions.
[0064] In the temperature dynamic control system for optical fiber rod processing, when establishing the dynamic mathematical prediction model of the multi-input multi-output system, not only the traditional prediction accuracy optimization problem is considered, but also a covariance penalty term for suppressing the thermal coupling interference between multiple zones is further introduced to form a more stable and adaptable control strategy. This control modeling scheme is particularly suitable for complex control scenarios where there is strong thermal conduction and disturbance coupling between zones in processes such as optical rod heating and continuous annealing.
[0065] Specifically, the control unit constructs a multiple-input multiple-output (MIMO) model based on the comprehensive production status data and sets the objective function as follows:
[0066]
[0067] Indicates the first The predicted temperature vector of each temperature zone at the moment; Indicates the first The predicted temperature of temperature zone 1 at the moment; Indicates the first Time temperature zone The predicted temperature; Indicates the first Time temperature zone Here, Indicates the total number of temperature zones, each of which is an independent temperature regulation area.
[0068] Indicates the length of the prediction horizon, that is, the number of future prediction time steps set in model predictive control.
[0069] is the target temperature vector of each temperature zone corresponding to the moment; the prediction error term represents the Euclidean norm squared, which is used to measure the deviation of the predicted temperature from the target setting; represents the time difference vector of the control input, It is Always keep the temperature in the zone Control input; Indicates the total number of temperature zones;
[0070] Indicates all the predictions within the time domain The covariance matrix of the samples is used to reflect the degree of joint fluctuation of control changes between temperature zones;
[0071] represents the trace operation of the matrix, that is, the sum of the main diagonal elements of the covariance matrix;
[0072] It is a weight adjustment factor used to balance the priority of prediction accuracy and control coupling stability. The recommended value range is generally 0.1 to 5. It is adjusted according to different working conditions to ensure a balance between control smoothness and response speed.
[0073] Through the above objective function, when solving the temperature control optimization problem, the control unit not only focuses on the temperature control accuracy of each temperature zone individually, but also explicitly suppresses the drastic synchronous control fluctuations between temperature zones. In particular, when the system has strong coupling between multiple temperature zones, it effectively prevents the control strategies from amplifying each other and causing system oscillation or overshoot.
[0074] The optimization process can be implemented by constrained quadratic programming (QP) or nonlinear optimization algorithms. The boundaries of the control variables can be set to the physical limits of the actuators such as heaters and cooling valves. For example, the input value is within the range.
[0075] Furthermore, to address model mismatches caused by factors such as production cycle changes, equipment aging, or material thermal capacity differences in actual operating conditions, the control unit implements a model adaptive update mechanism within each preset control cycle (e.g., every 100 control steps or every hour) based on the deviation between the predicted temperature and the actual temperature response collected during the current period. This mechanism uses the comprehensive production status data as input and updates the state transition matrix and control gain matrix parameters in the MIMO model through an online least squares method or a backpropagation-based gradient descent method, maintaining the model's robustness to environmental changes and its responsiveness.
[0076] For example, suppose the model state update relationship is:
[0077]
[0078] This formula describes the system from the current state Develop to the next state relationship.
[0079] is the temperature zone state vector at the current moment; , Indicates the Temperature zone in time step Temperature value; Indicates the total number of temperature zones; It is Always keep the temperature in the zone Control input; Indicates the total number of temperature zones;
[0080] in, is the state transition matrix, is the control gain matrix, which are parameters to be learned.
[0081] A Yes The state transfer matrix is used to reflect the internal relationship of the temperature state evolving over time. Its diagonal elements Indicates the impact of the current temperature zone state on the next state, non-diagonal elements Reflects the thermal coupling effect between temperature zones.
[0082] The initial value can be obtained by fitting from historical temperature control data through system identification methods (such as the least squares method), or it can be set by engineering experience and then optimized through subsequent iterations.
[0083] yes The control gain matrix of represents the degree of direct effect of the control input on the state. It indicates the direct impact of the control signal on the temperature in this temperature zone. The non-diagonal elements are generally set to small values or zero unless there is a cross-zone control feedback path.
[0084] The update rule can take the following form:
[0085]
[0086] in The learning rate, which is the step size of each update, controls the speed and amplitude of parameter updates. The typical value can be adjusted between 0.001 and 0.05.
[0087] and The objective functions are Regarding the gradient of the model parameters, the objective function here can be the objective function provided in Formula 1.
[0088] If the least squares method is used, the recursive formula is used to directly update the parameters to minimize the sum of squares of the error residuals.
[0089] By introducing covariance constraints and adaptive update mechanisms, the control unit has multi-temperature zone coordinated control capabilities that are significantly superior to traditional MPC, and exhibits higher control stability and accuracy in dealing with complex thermal coupling processes and dynamically changing environments, ensuring the adjustability and reliability of the system in long-term operation.
[0090] The learning unit 103 is configured to receive the temperature zone optimization control input instruction and the historical temperature data, and generate an adaptive control strategy through a reinforcement learning algorithm based on the temperature zone optimization control input instruction and the historical temperature data.
[0091] The core function of learning unit 103 is to continuously train and update the adaptive control strategy using a reinforcement learning algorithm based on the temperature zone optimization control input instructions and historical temperature data provided by control unit 102, thereby improving the adaptability and robustness of the entire temperature control system in long-term operation. Its essence lies in the introduction of an environmental feedback mechanism, which enables the system to not only rely on physical modeling results but also automatically learn the optimal control method through feedback from the actual temperature control process. This is particularly suitable for complex industrial scenarios with dynamic disturbances or incomplete model inaccuracies.
[0092] In a specific implementation, the learning unit 103 first continuously receives the temperature zone optimization control input instructions generated by the control unit 102 in each control cycle, and records the environmental response at the corresponding time point, that is, the change trend of the actual temperature of each temperature zone. These data form an experience data pair in each control cycle, which is recorded as ,in Indicates the current temperature status. Indicates the current control input, It represents the reward value formed by the deviation between the actual temperature and the target temperature under the control input. Represents the temperature state at the next moment. This four-tuple constitutes the basic sample in the reinforcement learning process.
[0093] In terms of specific algorithm selection, learning unit 103 can employ a combination of offline pre-training and online updating, typically implemented using a reinforcement learning architecture based on a Deep Q-Network (DQN). Initially, historically collected temperature control data can be constructed into an experience replay buffer to initialize the DQN model's Q-function approximator. In actual operation, the system updates the Q-function by sampling a certain number of experience samples at set intervals (e.g., every 10 control cycles). Gradient descent training is performed by minimizing the time-delay error (TD) to ensure that the network outputs a Q-value function that accurately estimates the expected benefit of each control action under a given state.
[0094] In terms of status representation, It can be defined as a vector composed of the temperature, tension, humidity and speed of each temperature zone at the current moment, and the states of the latest frames can be used to form a time series to increase the model's sensitivity to changing trends; in the definition of action space, The temperature control input can be a discrete quantitative value (such as 0%, 10%, 20%...100%) or a continuous real control value. In actual deployment, different strategy output structures are selected according to the type of control actuator.
[0095] Learning unit according to current status The Q value estimated by the DQN network selects the action As a strategy recommendation, we can specifically adopt the ε-greedy strategy, which strikes a balance between exploring new strategies and leveraging the current optimal strategy. For example, we can execute the currently evaluated optimal control input 90% of the time and randomly sample 10% of the time to encourage strategy updates.
[0096] In addition, the learning unit also needs to set the reward function , which is used to guide the optimization direction of the reinforcement learning model. It is generally set to the negative square of the temperature control accuracy, for example .
[0097] This is the negative of the sum of the deviations between the actual temperature in each temperature zone and the target temperature, thus encouraging the strategy to stay as close to the target temperature as possible in the long run. It can also introduce control smoothness constraints, such as penalizing overly frequent or drastic control changes.
[0098] After several cycles of iterative training, the learning unit gradually acquires an adaptive control strategy that performs stably and reliably under a variety of production conditions. This strategy serves as a supplementary input, along with the model-based optimization input generated by the control unit, to participate in the strategy fusion of the integration unit 104, thereby improving the overall temperature control effect.
[0099] Through the above-mentioned design and training mechanism, the learning unit 103 not only has the ability to perceive historical information and extract control experience, but can also automatically adapt and adjust control behavior under new working conditions, realizing the evolution of temperature control strategy from rule-based to data-driven, ensuring that the system has long-lasting intelligent adjustment capabilities.
[0100] Furthermore, the learning unit is specifically used to:
[0101] In the process of generating an adaptive control strategy based on the temperature zone optimization control input instructions and historical temperature data, a composite reward function for reinforcement learning training is constructed. The composite reward function includes feedback items in multiple dimensions, including at least a temperature control deviation item, a control energy consumption item, a device action frequency item, and a device state stability item, and is used to simultaneously constrain temperature control accuracy, energy utilization efficiency, device usage load, and long-term system operation stability.
[0102] When performing a policy update, an immediate reward is calculated based on the weighted combination value of each feedback item in the current state, which is used to guide the adaptive control strategy to balance between multiple optimization objectives, and the state-action value function in the reinforcement learning model is updated based on the immediate reward;
[0103] It is used to dynamically evaluate the estimated confidence of the state-action value function and automatically adjust the strategy selection method according to the confidence evaluation result. When the confidence is high, the historical learning strategy is given priority, and when the confidence is low, the random exploration ratio is increased, thereby optimizing the balance between exploration and utilization during long-term training, and improving the convergence efficiency and temperature control stability of the adaptive control strategy.
[0104] In the temperature control system of this invention, the core task of the learning unit is to train a control strategy that can adapt to environmental changes based on existing temperature control instructions and historical temperature data. To achieve this goal, the system designs a multi-dimensional reward function, which guides the reinforcement learning model to optimize its control strategy.
[0105] First, the learning unit evaluates the current temperature control results in each control cycle and calculates an immediate reward value based on this. This reward value is not determined by a single factor, but rather takes into account four key dimensions.
[0106] The first dimension is temperature control deviation, which is the difference between the current actual temperature of each temperature-controlled zone and the target temperature for that zone. The system takes the temperature difference for each zone and calculates the overall average. A larger average temperature difference indicates poorer temperature control accuracy, resulting in lower rewards or even penalties. Conversely, a smaller average temperature difference indicates better control effectiveness and higher rewards.
[0107] The second dimension is controlling energy consumption. The system monitors the current power settings of the heating or cooling equipment in each temperature zone, such as the heater output percentage. If overall energy consumption is high, the system reduces the reward to encourage the model to learn a more energy-efficient control strategy. If energy consumption is low, the contribution is positively rewarded.
[0108] The third dimension is frequency of operation. The system compares the magnitude of change between the current control command and the previous cycle's command. Frequent and significant fluctuations in the control variable of a heater or cooler indicate excessive starts and stops, which can lead to mechanical fatigue or energy waste. The system identifies this behavior and penalizes it, encouraging the model to gradually develop a stable, continuous control style.
[0109] The fourth dimension is device state stability. The system reviews temperature fluctuations in each temperature zone over time, particularly the magnitude of fluctuations. Drastic temperature fluctuations in a zone, even if the average value is close to the target, indicate a problem with the control strategy, and the system will reduce the reward based on this dimension. Conversely, stable temperature fluctuations indicate a reliable strategy, and the system will increase the score for this dimension.
[0110] Each of these four dimensions is weighted and summed to form a comprehensive immediate reward value. This value is fed into the reinforcement learning model to guide it in adjusting the relationship between state and action, thereby gradually optimizing the control strategy.
[0111] During the execution strategy selection process, the system also needs to decide whether to directly select the optimal solution from the learned strategies or make a new attempt. To do this, the system evaluates the confidence in the current strategy model, which is also known as the confidence level. If the system finds that the past control strategies in the current state have repeatedly performed well and the scores of each action are stable, the confidence level is considered high, indicating that no exploration is required and the previously learned optimal control method can be used directly. If the system finds that the control effect fluctuates greatly and the action scores are unevenly distributed or unstable, the confidence level is considered low. At this time, the system will actively increase exploratory behavior, that is, randomly select a portion of multiple control solutions to execute in order to discover new potential strategies.
[0112] Through this confidence adjustment mechanism, the system achieves a dynamic balance between favoring utilization under stable working conditions and favoring exploration under uncertain working conditions, which not only improves learning efficiency but also enhances the system's adaptability to complex working conditions.
[0113] In summary, in each control cycle, the learning unit automatically calculates rewards based on multiple actual measurement data, dynamically adjusts the control strategy, and flexibly switches the strategy selection method, thereby continuously optimizing the overall temperature control performance.
[0114] The following is a reference implementation of the learning unit in the temperature control system of this invention. This code shows how to calculate rewards, update the reinforcement learning model, and adjust the policy selection method based on confidence in each control cycle.
[0115] import numpy as np
[0116] import random
[0117] #Simulation environment parameters
[0118] NUM_ZONES = 4 # Number of temperature control zones
[0119] TEMP_TARGET = np.array([120.0, 130.0, 125.0, 128.0]) # Target temperature for each zone
[0120] P_MAX = 100.0 # Maximum power setting percentage, used for energy consumption normalization
[0121] WINDOW_SIZE = 10 # Time window used to calculate temperature fluctuations
[0122] # Reward function weight
[0123] W_TEMP_ERROR = 0.4
[0124] W_ENERGY = 0.2
[0125] W_FREQUENCY = 0.2
[0126] W_STABILITY = 0.2
[0127] # Learning strategy parameters
[0128] EPSILON_INITIAL = 0.1 # Initial exploration probability
[0129] EPSILON_MIN = 0.01
[0130] ALPHA = 0.05 # learning rate
[0131] GAMMA = 0.9 # discount factor
[0132] CONFIDENCE_THRESHOLD = 0.001 # Variance threshold for determining whether confidence is high
[0133] # Initialize the reinforcement learning table (state-action mapping Q table)
[0134] q_table = {}
[0135] # Simulate temperature recording and control recording for fluctuation analysis
[0136] temp_history = [np.copy(TEMP_TARGET) for _ in range(WINDOW_SIZE)]
[0137] prev_control = np.zeros(NUM_ZONES)
[0138] # Calculate the current instant reward function
[0139] def compute_reward(current_temp, control_input, prev_control, temp_history_window):
[0140] # 1. Temperature control deviation
[0141] temp_error = np.abs(current_temp - TEMP_TARGET).mean()
[0142] # 2. Control energy consumption (total power percentage divided by maximum value)
[0143] energy = control_input.sum() / (NUM_ZONES * P_MAX)
[0144] # 3. Action frequency term (control change rate)
[0145] frequency = np.abs(control_input - prev_control).mean()
[0146] # 4. State stability term (temperature fluctuation standard deviation)
[0147] temp_array = np.array(temp_history_window)
[0148] stds = temp_array.std(axis=0)
[0149] stability = stds.mean()
[0150] # Weighted combination
[0151] reward = (-W_TEMP_ERROR * temp_error
[0152] -W_ENERGY * energy
[0153] -W_FREQUENCY * frequency
[0154] -W_STABILITY * stability)
[0155] return reward
[0156] # Update Q table function
[0157] def update_q_table(state, action, reward, next_state, q_table):
[0158] key = (tuple(state), action)
[0159] next_qs = [q_table.get((tuple(next_state), a), 0) for a in range(NUM_ZONES)]
[0160] best_future_q = max(next_qs) if next_qs else 0
[0161] old_value = q_table.get(key, 0)
[0162] new_value = old_value + ALPHA * (reward + GAMMA * best_future_q -old_value)
[0163] q_table[key] = new_value
[0164] # Strategy selection (with confidence adjustment)
[0165] def select_action(state, q_table, confidence):
[0166] if confidence > CONFIDENCE_THRESHOLD:
[0167] # Low confidence, increase exploration
[0168] if random.random() < EPSILON_INITIAL:
[0169] return random.randint(0, NUM_ZONES - 1)
[0170] else:
[0171] # High confidence, greedy choice
[0172] q_values = [q_table.get((tuple(state), a), 0) for a in range(NUM_ZONES)]
[0173] return int(np.argmax(q_values))
[0174] return random.randint(0, NUM_ZONES - 1)
[0175] # Confidence estimation function (measured by Q value variance)
[0176] def estimate_confidence(state, q_table):
[0177] q_values = [q_table.get((tuple(state), a), 0) for a in range(NUM_ZONES)]
[0178] return np.var(q_values)
[0179] # Example: Execute a control cycle (this function needs to be called periodically by the main control process)
[0180] def control_cycle(current_temp, control_input, prev_control, temp_history, q_table):
[0181] # 1. Calculate rewards
[0182] reward = compute_reward(current_temp, control_input, prev_control, temp_history)
[0183] # 2. Record temperature history (for next cycle fluctuation analysis)
[0184] temp_history.pop(0)
[0185] temp_history.append(current_temp)
[0186] # 3. Build status (here simplified to current temperature + control amount)
[0187] state = np.concatenate([current_temp, control_input])
[0188] next_state = state # Simplify the process, the real system should obtain the next cycle state
[0189] # 4. Update the reinforcement learning model
[0190] action = select_action(state, q_table, estimate_confidence(state,q_table))
[0191] update_q_table(state, action, reward, next_state, q_table)
[0192] return reward, action, temp_history
[0193] The integration unit 104 is configured to receive the temperature zone optimization control input instruction and the adaptive control strategy, integrate the two to form a final control instruction, and send the final control instruction to the production line temperature control execution device to adjust the temperature of each temperature zone.
[0194] The integrated unit 104 assumes the core responsibility of control instruction fusion and execution scheduling in the entire temperature control system. Its main function is to effectively integrate the temperature zone optimization control input instructions generated by the control unit 102 and the adaptive control strategy generated by the learning unit 103, generate the execution instructions that are ultimately used to control the actual production temperature control equipment, and complete the issuance of instructions to the temperature control execution device to ensure that the temperature adjustment actions of each temperature zone are coordinated and the response is accurate.
[0195] In practice, the integration unit 104 first receives the optimized control input instructions from the control unit 102 and the adaptive control strategy output by the learning unit 103 via a standard industrial communication interface (such as Modbus TCP / IP, OPC UA, or EtherCAT). These two sets of control quantities may have numerical differences or even inconsistencies in control logic, so the integration unit must be capable of strategy fusion and conflict judgment. The system uses an integration algorithm based on weighted fusion, which performs a weighted average of the two sets of control quantities according to adjustable weight parameters. For example, if the control unit output is set as the dominant strategy and the adaptive control strategy is set as the auxiliary strategy, the fusion process can be expressed as:
[0196]
[0197] in represents the control input generated by the model predictive control, represents the policy recommendation generated by reinforcement learning. α∈[0,1] is a weighting factor, set to 0.7 by default and dynamically adjusted based on system operational feedback. When the system is stable or the model prediction accuracy is high, α is large, and model control is the primary strategy. When the model prediction error is persistently large or external disturbances cause control deviation, the system automatically reduces α to increase the influence of the adaptive policy, thereby achieving self-regulation between model control and learning control.
[0198] In a specific light rod heating production process, assume that the production line has three temperature control zones: Zone A, Zone B, and Zone C. The system's current goal is to maintain the temperatures of these three zones at 600°C, 620°C, and 610°C, respectively. The control unit analyzes the current sensor data using model predictive control and recommends a set of control inputs. The model predictive control outputs are 65% heating power for Zone A, 60% for Zone B, and 70% for Zone C. Simultaneously, the learning unit uses reinforcement learning to combine historical operating data with the current state to generate a slightly different set of adaptive control strategies. These recommendations are 60% for Zone A, 63% for Zone B, and 68% for Zone C.
[0199] At this point, the integrated unit begins executing policy fusion. Since the system is currently operating in a stable state and the prediction model's error has been below the threshold over the past ten cycles, the system maintains the default fusion weight α = 0.7 and prioritizes the model predictive control recommendation. According to the fusion formula, the final fusion control instruction is:
[0200] The control value for temperature zone A is 0.7 × 65% + 0.3 × 60% = 63.5%, for temperature zone B it is 0.7 × 60% + 0.3 × 63% = 60.9%, and for temperature zone C it is 0.7 × 70% + 0.3 × 68% = 69.4%.
[0201] These fused control values will be packaged into final control instructions and sent to each temperature control execution device through the industrial network to drive the corresponding heater to make actual power adjustments.
[0202] However, suppose that over the next few cycles, the actual temperature in zone B continues to deviate from the target value by more than 3°C, and the system detects a decrease in the model's prediction performance. The integrated unit automatically adjusts the α value to 0.5 to increase the influence of the reinforcement learning strategy on the final decision. At this time, the control value for zone B is recalculated:
[0203] 0.5 × 60% + 0.5 × 63% = 61.5%.
[0204] Through this dynamic fusion process, the system can adaptively adjust between model control and learning control based on real-time performance feedback, making the control more flexible and robust while taking into account stability and responsiveness.
[0205] To further enhance the reliability of fusion decisions, integration unit 104 incorporates a control quality evaluation mechanism. This mechanism evaluates the results of each temperature control cycle and provides feedback to the fusion module, which dynamically adjusts the fusion weight for the next cycle. For example, if the actual temperature in a particular temperature zone deviates from the target range for an extended period by more than a set threshold (e.g., ±3°C), the system increases the weight of the adaptive strategy to correct for potential model deviations.
[0206] The resulting control command is a set of precise values for multi-zone control variables, such as: 60% heating power for zone 1, 45% cooling damper opening for zone 2, and 70% heat shield shielding for zone 3. This command is packaged in JSON, Protobuf, or a PLC-compatible format and distributed over the industrial control network to temperature control actuators, including PID control modules, variable frequency drives, electric heating controllers, and cooling valve blocks. Each time a control command is issued, the integrated unit also records the current execution status and response feedback for closed-loop calibration and fault diagnosis.
[0207] To ensure execution consistency and minimize response delays, the integrated unit is deployed in the edge control platform or PLC body, adopts multi-threading or real-time task scheduling mechanism to execute control logic, and configures a communication priority guarantee mechanism to avoid queuing delays of control data in the industrial bus.
[0208] The design of the integrated unit 104 ensures the coordination and integration of control strategies from different sources, solves the control deviation or strategy conflict problems that may exist between model prediction and reinforcement learning, and enables the entire temperature control system to have both model-based global control capabilities and data-driven rapid adaptation capabilities, thereby realizing high-precision coordinated control of multiple temperature zones under dynamic working conditions.
[0209] Furthermore, the integrated unit is specifically used for:
[0210] After receiving the temperature zone optimization control input instruction and the adaptive control strategy, a nonlinear function mapping method is used to determine a fusion weight factor based on the real-time control error fluctuation range and the corresponding historical temperature control deviation change rate of each temperature zone in the current control cycle, wherein the nonlinear function includes a S-type function or a confidence interval estimation function, which is used to convert the confidence of the control effect into a weight factor to automatically adjust the relative weight of the temperature zone optimization control input instruction and the adaptive control strategy in the fusion process;
[0211] The fusion weight factor is used for weighting the temperature zone optimization control input instruction and the adaptive control strategy, a fusion instruction set is formed, and the fusion process is independently performed on each temperature zone to ensure that the control strategy of each temperature zone is adapted to the local characteristics;
[0212] After the fusion instruction set is generated, the fused instruction is verified for logical consistency by a rapid evaluation module, the verification process includes judging whether the control gradient difference of adjacent temperature zones exceeds a threshold value, whether the control fluctuation amplitude of a single temperature zone is abnormal, and the like, so that the execution instability caused by the conflict between the model prediction and the learning strategy is avoided.
[0213] On the basis of verifying the logical rationality of the fusion instruction set, a final control instruction is generated and sent to a production line temperature control execution device, so as to adjust the temperature control output of each temperature zone in real time.
[0214] The integration unit in the application is a key module for fusing the control instructions of different sources after the control unit and the learning unit generate the temperature zone optimization control input instruction and the adaptive control strategy. In the fusion process, in order to reasonably balance the dominant weight of the model prediction control and the reinforcement learning strategy under different working conditions, a dynamic weight adjustment mechanism is introduced, and the mapping calculation of the fusion weight factor is completed through a nonlinear function.
[0215] Specifically, the system first collects the real-time control error fluctuation range of each temperature control area in the current control period, that is, the change amplitude of the difference between the current set temperature and the actual feedback temperature with time, and also reads the temperature control deviation trend of the temperature zone in the past several control periods, for example, whether there is deviation amplification, convergence or oscillation and the like. Based on the two input quantities, the system uses a nonlinear function for processing, commonly including an S-shaped function (for example, a sigmoid function) or a confidence estimation function based on a Bayesian framework, for mapping the "predictability" or "stability" of the current working condition into a fusion weight value between 0 and 1. The fusion weight value is the trust allocation proportion between the model control output and the learning control output in the current period of the temperature zone.
[0216] After the fusion weight is determined, the system extracts the temperature zone optimization control input instruction and the adaptive control strategy corresponding to each temperature control area, respectively, and combines them according to the corresponding fusion weight, so as to complete the fusion at the instruction level. This process is independently performed on each temperature zone to ensure that each area adjusts the dominance of the control source according to its own control effect and uncertainty, instead of simply using a unified weight, so as to improve the local adaptability and global coordination of the control.
[0217] After completing the preliminary calculation of the fusion instruction set, the system will not send it directly to the execution device, but will enter the rapid evaluation module for logical consistency verification. This module mainly includes two aspects of detection logic: the first is to determine the control gradient change between adjacent temperature zones, that is, whether the difference in temperature control instructions between two adjacent temperature zones exceeds the preset physical or process threshold, so as to avoid thermal shock or stress concentration problems caused by excessive differences in local control strategies; the second is to monitor the control change amplitude of each temperature control area itself. If the current fusion instruction changes too much compared to the previous cycle, the system will regard it as a potential abnormal fluctuation risk, which may be caused by the inconsistency between model prediction and learning strategy. Therefore, when the module determines an abnormality, it will trigger protection logic, such as smoothing, instruction rollback, or fusion weight readjustment.
[0218] After the aforementioned logical consistency verification is complete and no risks are identified, the fused instruction set is confirmed to be reasonable. The system then generates final control instructions, which are accurately distributed to the temperature control actuators in each temperature zone, such as electric heaters, cooling water valves, and other specific hardware controllers, to achieve real-time temperature adjustment in each temperature-controlled area.
[0219] The following is the entire fusion logic process of the integrated unit in the present invention, including the nonlinear calculation of the fusion weight, the fusion of the instructions of each area, the logic consistency verification and the final instruction issuance.
[0220] import numpy as np
[0221] from scipy.special import expit # sigmoid function
[0222] # Assume the system has 4 temperature control zones
[0223] NUM_ZONES = 4
[0224] # Control command input example (from control unit and learning unit)
[0225] model_outputs = np.array([65.0, 70.0, 72.5, 68.0]) # Output instructions from model predictive control
[0226] rl_outputs = np.array([64.0, 71.0, 73.0, 67.0]) # Output instructions from reinforcement learning policy
[0227] # The control error fluctuation range of the current cycle (the variation between the actual temperature value and the target temperature)
[0228] current_fluctuations = np.array([2.5, 1.0, 4.2, 0.8]) # Unit: °C
[0229] # Historical deviation change rate (the rate of change of temperature control deviation per unit time)
[0230] history_drift_rate = np.array([0.4, -0.1, 1.2, -0.3]) # Unit: °C / cycle
[0231] # Define fusion threshold parameters
[0232] GRADIENT_THRESHOLD = 5.0 # Maximum allowed control gradient difference between adjacent regions
[0233] MAX_CONTROL_DELTA = 10.0 # Maximum control variation allowed for a single region
[0234] # Final instruction record of the previous cycle (for fluctuation detection)
[0235] prev_final_outputs = np.array([65.0, 70.0, 72.0, 68.0])
[0236] def compute_fusion_weight(fluctuation, drift):
[0237] #Use nonlinear function to generate fusion weights.
[0238] # The sigmoid function is used as an example here. The more stable the input, the greater the output weight.
[0239] #sigmoid(x) = 1 / (1 + e^(-x))
[0240] # Combine the fluctuation range and drift rate into a fusion sensitivity factor
[0241] confidence_score = -1.0 * (fluctuation + abs(drift)) # Smaller means higher confidence
[0242] weight = expit(confidence_score) # sigmoid output is between 0 and 1
[0243] return weight
[0244] def fuse_controls(model_outputs, rl_outputs, fluctuation, drift):
[0245] #Perform strategy fusion for each temperature control area and return the fused control instruction array.
[0246] fused_outputs = np.zeros(NUM_ZONES)
[0247] for i in range(NUM_ZONES):
[0248] alpha = compute_fusion_weight(fluctuation[i], drift[i]) # Get fusion weight
[0249] # Use weighted average fusion model control and learning control strategy
[0250] fused_outputs[i] = alpha * model_outputs[i] + (1 - alpha) *rl_outputs[i]
[0251] return fused_outputs
[0252] def validate_fused_outputs(fused_outputs, prev_outputs):
[0253] #Verify the logical consistency of the fused instruction set.
[0254] #Includes: 1) Adjacent temperature zone control gradient detection; 2) Single zone control fluctuation detection.
[0255] # Check if the adjacent area control difference exceeds the threshold
[0256] for i in range(NUM_ZONES - 1):
[0257] if abs(fused_outputs[i] - fused_outputs[i + 1]) > GRADIENT_THRESHOLD:
[0258] print(f"Fusion failed: The control difference between regions {i} and {i+1} is too large")
[0259] return False
[0260] # Check if the control change of a single area is too large
[0261] for i in range(NUM_ZONES):
[0262] if abs(fused_outputs[i] - prev_outputs[i]) > MAX_CONTROL_DELTA:
[0263] print(f"Fusion failed: Region {i} control changes too much")
[0264] return False
[0265] return True
[0266] def control_dispatch(fused_outputs):
[0267] #Send the final instruction to the temperature control execution device (simulate printing here).
[0268] for i in range(NUM_ZONES):
[0269] print(f"Send control instructions to area {i}: {fused_outputs[i]:.2f}℃")
[0270] # === Main fusion process ===
[0271] # 1. Perform fusion
[0272] fused = fuse_controls(model_outputs, rl_outputs, current_fluctuations, history_drift_rate)
[0273] # 2. Verify the legitimacy of the fused control instructions
[0274] if validate_fused_outputs(fused, prev_final_outputs):
[0275] # 3. Issue control instructions
[0276] control_dispatch(fused)
[0277] else:
[0278] print("Fusion failed, use backup strategy or adjust fusion weight.")
[0279] Further, the integration unit comprises:
[0280] a heterogeneous knowledge distillation module configured to receive the warm zone optimization control input instruction and the adaptive control strategy, and to construct a bidirectional knowledge extraction network comprising a time-sensitive feature alignment layer configured to convert the warm zone optimization control input instruction and the adaptive control strategy to a unified representation space; the heterogeneous knowledge distillation module further comprises a complementary enhancement mechanism configured to dynamically identify the advantage distribution area of model predictive control and reinforcement learning control in different temperature control regions based on the unified representation space;
[0281] a warm zone state migration predictor configured to receive the fusion strategy in the unified representation space, to construct a warm zone state discretization dictionary, and to establish a Markov decision process model; the warm zone state migration predictor is configured to simulate and predict the state evolution path of the fusion strategy in future control cycles based on the state information of each temperature control region, and to output a prediction reliability evaluation index of each candidate control strategy guiding the system to reach the target temperature state;
[0282] a multi-objective resolver configured to receive the prediction reliability evaluation index, and to construct a Pareto optimization evaluation function based on four performance dimensions of temperature accuracy, energy efficiency, equipment life and system stability, to sort the advantages and disadvantages of multiple candidate fusion strategies, and to determine the optimal parameter combination of the final control instruction.
[0283] In the temperature dynamic control system for optical fiber rod processing according to the present application, the integration unit further comprises a heterogeneous knowledge distillation module, a warm zone state migration predictor and a multi-objective resolver, which cooperatively constitute a decision framework for fusion and optimization of control strategies, and the specific implementation is as follows.
[0284] First, the heterogeneous knowledge distillation module processes two types of control strategy data from the control unit and the learning unit: temperature zone optimization control input commands and adaptive control strategies. These two types of control logic differ significantly in their source structure and representation, making direct fusion difficult. To address this, the heterogeneous knowledge distillation module constructs a bidirectional knowledge extraction network consisting of a set of parallel input channels, each receiving the two types of strategies and normalizing them using a time-sensitive feature alignment layer. Specifically, the system nests and expands the inputs of each strategy across multiple time steps, extracts its temporal statistical features, such as amplitude, directionality, and stability indicators, and maps them into a vector representation of unified dimensions, resulting in a set of candidate fusion strategies in a unified representation space. To further enhance the policy representation capabilities of this unified space, the module embeds a set of complementary enhancement mechanisms, such as attention weighting or confidence weighting, to automatically identify the relative performance of model predictive control or reinforcement learning control in each temperature control zone. This dynamically assigns higher representation weights to the superior strategies, thereby constructing a preliminary fusion strategy with significant structural discrimination.
[0285] Next, the temperature zone state transition predictor receives the aforementioned fusion strategy characterization results and, combined with the production state data for each temperature zone collected during the current cycle, simulates the possible evolution of these strategies within future control cycles. To achieve traceable state prediction, the system first constructs a discretized dictionary of temperature zone states, converting continuous sensor data such as temperature and speed into a finite number of state labels, such as "undercooled," "slightly cold," "moderate," "slightly hot," and "overheated." Based on these discrete states, a Markov decision process model is then established, where state transition probabilities are estimated from historical operating data or initialized based on expert experience and updated through online learning. During the simulation prediction, the system inputs a candidate control strategy and iteratively simulates the state evolution path of each temperature control zone under its influence, predicting whether it can converge to the set target temperature state within a finite cycle. The output is expressed as a probability confidence value or prediction trajectory score, which characterizes the reliability of each fusion strategy in achieving its target.
[0286] Finally, after obtaining the predictive reliability evaluation indicators corresponding to all candidate fusion strategies, the multi-objective arbiter conducts a multi-dimensional performance analysis and ranking screening. Specifically, the system first evaluates four key performance indicators for each candidate strategy: temperature control accuracy, which is the mean square error between the current actual temperature and the target temperature; energy efficiency, which is the energy consumption level corresponding to the control execution; equipment life contribution, which determines the degree of equipment fatigue by analyzing the frequency and amplitude of instruction fluctuations; and system stability, which is structural stability factors such as temperature fluctuation range and control input consistency. The above four performance indicators together constitute a multi-objective optimization scenario. The arbiter uses the Pareto optimization principle to construct a multi-dimensional evaluation model to compare the dominant relationship of different strategies in multiple dimensions. Finally, the system selects a set of control parameter combinations that have balanced performance in all dimensions or the best overall performance as the final control instruction and sends it to the temperature control execution device.
[0287] The following code fully demonstrates how the integrated unit in the present invention collaborates and integrates and optimizes the final control strategy through three sub-modules: heterogeneous knowledge distillation, temperature zone state migration prediction, and multi-objective judgment.
[0288] import numpy as np
[0289] # ================== Heterogeneous Knowledge Distillation Module ===================
[0290] class HeterogeneousKnowledgeDistillation:
[0291] def __init__(self, input_dim, aligned_dim):
[0292] self.input_dim = input_dim # Original control strategy vector dimension
[0293] self.aligned_dim = aligned_dim # Unify vector dimensions after alignment
[0294] # Initialize the weight matrix for feature alignment (trainable parameters)
[0295] self.align_weight_model = np.random.rand(input_dim, aligned_dim)
[0296] self.align_weight_rl = np.random.rand(input_dim, aligned_dim)
[0297] def temporal_feature_align(self, control_seq_model, control_seq_rl):
[0298] # Receive policy sequences from model predictive control and reinforcement learning control, extract timing features and align them.
[0299] control_seq_model / control_seq_rl: shape = (time_steps, input_dim)
[0300] #Perform time series statistics (such as mean) and multiply by the mapping matrix
[0301] aligned_model = np.mean(control_seq_model, axis=0).dot(self.align_weight_model)
[0302] aligned_rl = np.mean(control_seq_rl, axis=0).dot(self.align_weight_rl)
[0303] return aligned_model, aligned_rl
[0304] def complementary_enhancement(self, aligned_model, aligned_rl,confidence_model, confidence_rl):
[0305] #Assign weights based on confidence and generate fusion strategy vectors
[0306] total_conf = confidence_model + confidence_rl + 1e-6
[0307] w_model = confidence_model / total_conf
[0308] w_rl = confidence_rl / total_conf
[0309] fused = w_model * aligned_model + w_rl * aligned_rl
[0310] return fused
[0311] # ================== State Transition Predictor ==================
[0312] class StateTransitionPredictor:
[0313] def __init__(self, num_states, transition_matrix=None):
[0314] self.num_states = num_states
[0315] if transition_matrix is None:
[0316] self.transition_matrix = np.full((num_states, num_states), 1.0 / num_states)
[0317] else:
[0318] self.transition_matrix = transition_matrix # shape =(num_states, num_states)
[0319] def predict_trajectory(self, initial_state, strategy_index, steps=5):
[0320] #Use the Markov model to simulate the state trajectory. initial_state is the current state index.
[0321] # strategy_index: The transition matrix used to distinguish different strategies (simplified to an index)
[0322] current = initial_state
[0323] trajectory = [current]
[0324] for _ in range(steps):
[0325] prob = self.transition_matrix[current]
[0326] current = np.random.choice(self.num_states, p=prob)
[0327] trajectory.append(current)
[0328] return trajectory
[0329] def evaluate_reliability(self, trajectory, target_state):
[0330] #Calculate the frequency of occurrence of the target state in the trajectory as a strategy reliability indicator
[0331] score = trajectory.count(target_state) / len(trajectory)
[0332] return score
[0333] # ================== Multi-target Arbitrator ==================
[0334] class MultiObjectiveArbiter:
[0335] def __init__(self, weights):
[0336] # weights: dict, representing the weight coefficient of each target dimension
[0337] self.weights = weights
[0338] def evaluate(self, strategies_metrics):
[0339] # Input a set of performance indicators for multiple strategies, each element is a dict:
[0340] #{'accuracy': float, 'energy': float, 'lifetime': float, 'stability': float}
[0341] #Return the optimal strategy index
[0342] scores = []
[0343] for metrics in strategies_metrics:
[0344] # Sort based on weighted linear combination (for more complex sorting, use Pareto sorting)
[0345] score = sum([self.weights[k] * metrics[k] for k inself.weights])
[0346] scores.append(score)
[0347] return int(np.argmax(scores))
[0348] Although the present application is disclosed as above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
Claims
1. A temperature dynamic control system for optical fiber rod processing, characterized in that: include: The acquisition unit includes a multi-macro camera array, a production line sensor network, and an edge computing node. The multi-macro camera array is used to capture imaging data of production line materials. The production line sensor network includes an environmental perception unit and an equipment status sensor array for acquiring temperature, tension, humidity, and speed data. The edge computing node is used to fuse and pre-process the imaging data and acquired data and upload them to the cloud to form comprehensive production status data. a control unit configured to receive the comprehensive production status data and establish a dynamic mathematical prediction model for a multi-input multi-output system; predict future temperature states of multiple temperature zones of the production line based on the dynamic mathematical prediction model, and generate temperature zone optimization control input instructions based on the prediction results; a learning unit, configured to receive the temperature zone optimization control input instruction and the historical temperature data, and generate an adaptive control strategy based on the temperature zone optimization control input instruction and the historical temperature data through a reinforcement learning algorithm; An integration unit is configured to receive the temperature zone optimization control input instruction and the adaptive control strategy, integrate the two, and form a final control instruction; and send the final control instruction to the production line temperature control execution device to adjust the temperature of each temperature zone; Wherein, the control unit is further used for: When establishing the dynamic mathematical prediction model of the multi-input multi-output system, a covariance penalty term for suppressing the coupling effect between temperature zones is introduced when constructing the objective function. The objective function is composed of a squared prediction error term and a covariance term of changes in control quantities of multiple temperature zones, wherein the covariance term is used to measure the degree of joint fluctuation between the control inputs of each temperature zone at different times. When solving the objective function, an optimal temperature zone optimization control input instruction is obtained through an optimization algorithm. The temperature zone optimization control input instruction satisfies the convergence of the target temperature of each temperature zone while limiting the mutual interference of control changes between the temperature zones to avoid system oscillation caused by coupling gain amplification; In each preset control cycle interval, based on the prediction error between the comprehensive production status data and the actual response of the current temperature zone, the state transfer matrix parameters of the multi-input multi-output system are updated using an online least squares method or a gradient descent algorithm to improve the adaptability and robustness of the dynamic mathematical prediction model to changes in operating conditions; The integrated unit is specifically used for: After receiving the temperature zone optimization control input instruction and the adaptive control strategy, a nonlinear function mapping method is used to determine a fusion weight factor based on the real-time control error fluctuation range and the corresponding historical temperature control deviation change rate of each temperature zone in the current control cycle, wherein the nonlinear function includes a S-type function or a confidence interval estimation function, which is used to convert the confidence of the control effect into a weight factor to automatically adjust the relative weight of the temperature zone optimization control input instruction and the adaptive control strategy in the fusion process; The temperature zone optimization control input instructions and the adaptive control strategy are weighted according to the fusion weight factor to form a fusion instruction set. The fusion process is performed independently in each temperature zone to ensure that the control strategy of each temperature zone adapts to its local characteristics; After generating the fused instruction set, the fast evaluation module verifies the logical consistency of the fused instructions. This verification process includes determining whether the control gradient difference between adjacent temperature zones exceeds a threshold and whether the control fluctuation amplitude of a single temperature zone is abnormal, thereby avoiding execution instability caused by conflicts between model prediction and learning strategies. On the basis of verifying the logical rationality of the fusion instruction set, a final control instruction is generated and sent to the temperature control execution device of the production line for real-time adjustment of the temperature control output of each temperature zone.
2. The temperature dynamic control system for optical fiber rod processing according to claim 1 is characterized in that: The acquisition unit is further used for: Based on the imaging data captured by the multi-macro camera array and the temperature, tension, humidity, and speed data acquired by the production line sensor network, the edge computing node establishes a sliding time window structure at preset time steps to construct a time-series nested state vector. The sliding time window structure is used to continuously store the imaging data features and corresponding temperature, tension, humidity, and speed data of the first M time steps within each sampling period. The imaging data features are extracted as fixed-length numerical vectors through an image processing algorithm and spliced with the corresponding sensor data to form a multidimensional state vector. The edge computing node performs cross-time step fusion processing on the state vector in the sliding time window structure to form a nested time series feature representation for reflecting the production state change trend of the target material in a continuous time period; The edge computing node uploads the nested time series feature representation as part of the comprehensive production status data to the cloud, and is used by the control unit to establish a dynamic mathematical prediction model of the multi-input multi-output system.
3. The temperature dynamic control system for optical fiber rod processing according to claim 1, characterized in that: The learning unit is specifically used for: In the process of generating an adaptive control strategy based on the temperature zone optimization control input instructions and historical temperature data, a composite reward function for reinforcement learning training is constructed. The composite reward function includes feedback items in multiple dimensions, including at least a temperature control deviation item, a control energy consumption item, a device action frequency item, and a device state stability item, and is used to simultaneously constrain temperature control accuracy, energy utilization efficiency, device usage load, and long-term system operation stability. When performing a policy update, an immediate reward is calculated based on the weighted combination value of each feedback item in the current state, which is used to guide the adaptive control strategy to balance between multiple optimization objectives, and the state-action value function in the reinforcement learning model is updated based on the immediate reward; It is used to dynamically evaluate the estimated confidence of the state-action value function and automatically adjust the strategy selection method according to the confidence evaluation result. When the confidence is high, the historical learning strategy is given priority, and when the confidence is low, the random exploration ratio is increased, thereby optimizing the balance between exploration and utilization during long-term training, and improving the convergence efficiency and temperature control stability of the adaptive control strategy.
4. The temperature dynamic control system for optical fiber rod processing according to claim 1, characterized in that: The integrated unit comprises: a heterogeneous knowledge distillation module for receiving the temperature zone optimization control input instructions and the adaptive control strategy and constructing a bidirectional knowledge extraction network, wherein the bidirectional knowledge extraction network includes a time-sensitive feature alignment layer for converting the temperature zone optimization control input instructions and the adaptive control strategy into a unified representation space; the heterogeneous knowledge distillation module also includes a complementary enhancement mechanism for dynamically identifying advantageous distribution areas of model predictive control and reinforcement learning control in different temperature control zones based on the unified representation space; a temperature zone state transition predictor, configured to receive the fusion strategy in the unified representation space, construct a temperature zone state discretization dictionary, and establish a Markov decision process model; the temperature zone state transition predictor is configured to simulate and predict the state evolution path of the fusion strategy over multiple future control cycles based on the current state information of each temperature control zone, and output a prediction reliability evaluation index for each candidate control strategy to guide the system to the target temperature state; A multi-objective arbiter is used to receive the predicted reliability evaluation index and construct a Pareto optimization evaluation function based on four performance dimensions: temperature accuracy, energy efficiency, equipment life and system stability, to rank multiple candidate fusion strategies and determine the optimal parameter combination for the final control instruction.
Citation Information
Patent Citations
Thermocouple fault diagnosis and processing method and system of semiconductor heat treatment device
CN103759860A
Environment self-sensing vehicle window fusion control method, device and equipment and storage medium
CN118462017A