Online heating control method and system based on dynamic model prediction and multi-modal feedback

By employing an online heating control method based on dynamic model prediction and multimodal feedback, the problems of hysteresis and measurement accuracy in heating systems are solved, achieving ultra-high precision temperature control and adaptive capability. This adapts to the production needs of multiple product categories, improving production efficiency and product quality.

CN122018603BActive Publication Date: 2026-08-04上海神众智能科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
上海神众智能科技有限公司
Filing Date
2026-04-14
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing heating control systems suffer from lag, insufficient measurement accuracy, and poor adaptability to environmental interference, making it difficult to achieve ultra-high precision temperature control of ±0.1℃ and unable to meet the needs of multi-category, flexible production.

Method used

An online heating control method based on dynamic model prediction and multimodal feedback is adopted. Through multi-dimensional temperature data acquisition and preprocessing, an online updated dynamic thermodynamic model is constructed. Combined with multimodal feedback and reinforcement learning, the heating power can be pre-controlled and environmental disturbances can be compensated, and the thermal characteristics of different heating objects can be adapted.

Benefits of technology

It achieves ultra-high precision temperature control within ±0.1℃, has strong adaptability, adapts to multi-category production, reduces production line switchover and debugging costs and time, and improves production efficiency and product yield.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018603B_ABST
    Figure CN122018603B_ABST
Patent Text Reader

Abstract

The application discloses an online heating control method and system based on dynamic model prediction and multi-modal feedback, and belongs to the technical field of industrial precise heating control.The application collects multi-modal sensing data in real time, constructs an online updated dynamic thermodynamic model, predicts future temperature change trends and generates heating power pre-control instructions, solves the hysteresis overshoot / oscillation problem of traditional PID control, eliminates single sensor error and random interference through multi-source sensor data fusion and environmental interference dynamic compensation, improves control accuracy to within + / -0.1 DEG C, and optimizes control parameters online through a reinforcement learning algorithm, and adapts to the thermal characteristic differences of different heating objects.The application can be widely applied to semiconductor manufacturing, precise mold heat treatment and other industrial scenes with extremely high requirements for temperature accuracy and uniformity, and has high response speed, ultra-high control accuracy and strong versatility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of semiconductor manufacturing technology, and more specifically, to an online heating control method and system based on dynamic model prediction and multimodal feedback. Background Technology

[0002] In the fields of semiconductor manufacturing and precision mold heat treatment, the temperature uniformity and temperature control accuracy of the heating process directly determine the product molding quality and performance stability. This is a key factor restricting the yield rate of high-end manufacturing and the implementation of core processes. Especially for high-precision workpieces such as semiconductor lithography molds, wafer carriers, and micro-structural components, even small temperature fluctuations and temperature differences can lead to workpiece deformation, dimensional deviations, and performance degradation, making it difficult to meet the stringent standards of chip manufacturing and precision molding.

[0003] Most existing industrial heating control solutions employ traditional PID control algorithms, which adjust based on the deviation between the current temperature and the target temperature. While simple in principle and easy to implement, they suffer from inherent drawbacks: First, heating systems are typically characterized by large inertia and large lag. PID control, being a reactive error adjustment, cannot predict temperature change trends in advance, making it prone to temperature overshoot and oscillation, and failing to meet the ultra-high precision control requirement of ±0.1℃. Second, traditional PID control parameters are fixed. When the heating object changes or environmental parameters dynamically change, control performance deteriorates significantly, requiring manual parameter retuning by professionals. This results in poor adaptability and fails to meet the needs of multi-category, flexible production.

[0004] To address the shortcomings of PID control, some existing technologies have introduced model predictive control schemes. However, most of these use thermodynamic models with fixed parameters, which cannot adapt to real-time changes in the thermal characteristics of the heated object online. They also fail to consider the impact of dynamic environmental disturbances such as airflow, humidity, and power fluctuations, resulting in insufficient model prediction accuracy and limiting the breakthrough in control precision. Furthermore, most existing technologies employ a single thermocouple temperature measurement scheme, which suffers from measurement delays, inability to reflect the overall temperature of the heated object from a single point measurement, and susceptibility to temperature drift caused by environmental interference. The measurement error of a single sensor directly limits the upper limit of control precision, making it impossible to stably achieve a control precision of ±0.1℃. In addition, existing technologies mostly compensate for environmental disturbances with fixed parameters, failing to dynamically adapt to real-time changing disturbance factors, resulting in poor compensation effects and an inability to cope with complex industrial environments.

[0005] Therefore, developing an ultra-high precision online heating control scheme that can fundamentally solve the problem of heating system hysteresis, eliminate multi-source measurement errors, and dynamically adapt to environmental interference and the characteristics of the heated object has become an urgent technical problem to be solved in this field. Summary of the Invention

[0006] Technical problems to be solved

[0007] To address the aforementioned shortcomings of existing technologies, the present invention aims to provide an online heating control method and system based on dynamic model prediction and multimodal feedback. This method solves the technical problems of hysteresis overshoot, insufficient measurement accuracy, poor adaptability to environmental interference, and weak versatility of heating objects in traditional control schemes. It stably achieves ultra-high precision temperature control within ±0.1℃, while possessing strong adaptability and high reliability, meeting the stringent control requirements of high-end industrial scenarios.

[0008] Technical solution

[0009] The online heating control method based on dynamic model prediction and multimodal feedback includes the following steps:

[0010] S1 Multimodal sensing data acquisition and preprocessing: Real-time acquisition of multi-dimensional temperature data of the heated object, operating status data of the heating equipment, and environmental parameter data of the heating cavity. The acquired raw data is preprocessed by filtering and denoising, time alignment and outlier removal to obtain a standardized sensing dataset.

[0011] S2 Online Update Dynamic Thermodynamic Prediction Model Construction: Based on the law of conservation of energy and the heat transfer characteristics of the heating system, combined with preprocessed sensing data, a lumped parameter dynamic thermodynamic model is constructed. The thermophysical parameters and heat loss coefficient of the model are corrected online through real-time collected sensing data. Based on the corrected dynamic thermodynamic model, the temperature change trend of the heated object in the next N control cycles is predicted.

[0012] S3 Model-based heating power pre-control command generation: Using the preset target temperature curve as the control target and the predicted temperature change trend as the basis, the constrained rolling optimization objective function is solved in each control cycle to obtain the optimal heating power pre-control command, thereby realizing the advanced pre-control of temperature.

[0013] S4 Multimodal Feedback Fusion and Real-time Compensation Correction: Adaptive weighted fusion of multi-dimensional temperature data and operating status data is performed to obtain a high-precision actual temperature feedback value; an environmental interference compensation model is constructed, and the real-time interference compensation amount is calculated based on environmental parameter data. Combined with the deviation between the temperature feedback value and the target temperature, the heating power pre-control command is closed-loop corrected to generate the final heating power control command.

[0014] S5 Reinforcement Learning-Based Adaptive Online Optimization of Control Parameters: Constructs a parameter optimization reinforcement learning agent, with temperature control deviation, overshoot, and control stability as core reward indicators, and dynamic thermodynamic model parameters, rolling optimization weight coefficients, and environmental compensation gain coefficients as optimization objects, to iteratively optimize the control parameters of the entire process online, adapting to the differences in thermal characteristics of different heating objects.

[0015] According to one or more embodiments of the present invention, in step S1, the multi-dimensional temperature data includes at least two contact temperature measurement data and at least one non-contact temperature measurement data. The contact temperature measurement data is temperature data collected by thermocouples and platinum resistance thermometers, and the non-contact temperature measurement data is temperature data collected by infrared temperature sensors and infrared thermal imagers. The operating status data includes the real-time output power, input voltage, input current, and on / off status of the heating module. The environmental parameter data includes the airflow velocity, ambient humidity, ambient temperature, and cavity wall temperature inside the heating cavity.

[0016] According to one or more embodiments of the present invention, in step S2, the construction and updating of the dynamic thermodynamic model specifically includes:

[0017] S21 Based on the law of conservation of energy, an initial lumped-parameter thermodynamic model is constructed, comprehensively considering the heat input of the heating module, the convective heat loss between the heated object and the environment, and the radiative heat loss. The model expression is as follows:

[0018]

[0019] in, The total heat capacity of the object being heated and the heating platform. For the real-time temperature of the object being heated, Heating power-to-thermal energy conversion efficiency For heating power, The convective heat transfer coefficient is... For heat exchange area, For ambient temperature, For radiative emissivity, It is the Stefan-Boltzmann constant;

[0020] S22 uses the recursive least squares method to identify and correct the heat capacity in the model online based on preprocessed real-time sensing data. Conversion efficiency convective heat transfer coefficient This allows for the acquisition of a dynamic thermodynamic model that is updated in real time, adapting to changes in thermophysical parameters during the heating process.

[0021] S23 is based on a dynamic thermodynamic model, using the current state of the control cycle as the initial value, to predict the temperature sequence for the next N control cycles. The value of N ranges from 5 to 50, and the value of the control cycle ranges from 1 ms to 100 ms, balancing prediction accuracy and real-time calculation.

[0022] Furthermore, in S3, the rolling optimization objective function focuses on minimizing temperature prediction deviation and power fluctuation, and its expression is:

[0023]

[0024] The constraints are:

[0025]

[0026] in, The predicted temperature for the i-th control cycle. The target temperature for the i-th control cycle is... This represents the power change between adjacent control cycles. These are the weighting coefficients. These are the upper and lower limits of the heating power. To maximize the allowable rate of temperature change and avoid thermal stress damage to the heated object due to sudden temperature changes.

[0027] According to one or more embodiments of the present invention, in step S4, adaptive weighted fusion is performed on the measurement data of the M temperature sensors, and the measurement variance of each sensor within the sliding time window is calculated respectively. The weights of each sensor are adaptively assigned based on the variance. Sensors with smaller variances are assigned higher weights to maximize the elimination of random errors. The specific expression is:

[0028]

[0029] The final fusion temperature value is:

[0030]

[0031] in, The temperature measurement value is the preprocessed value of the i-th sensor.

[0032] According to one or more embodiments of the present invention, in step S4, the environmental interference compensation model includes an airflow heat loss compensation module, a power fluctuation compensation module, and a sensor temperature drift compensation module.

[0033] The airflow heat loss compensation module calculates the change in convective heat transfer coefficient based on the real-time collected airflow velocity to obtain convective heat loss compensation power, thereby offsetting the heat loss fluctuations caused by airflow.

[0034] The power fluctuation compensation module calculates the output deviation of heating power based on the deviation between the real-time collected input voltage and the rated voltage, obtains the power fluctuation compensation amount, and eliminates the power output error caused by grid voltage fluctuation.

[0035] The sensor temperature drift compensation module corrects the drift of the temperature sensor's measured values ​​based on the difference between the ambient temperature and the calibration temperature, thereby eliminating sensor system errors caused by changes in ambient temperature.

[0036] According to one or more embodiments of the present invention, in S5, the reinforcement learning agent adopts the proximal policy optimization PPO algorithm, and its state space includes the current temperature deviation, temperature change rate, model prediction error, environmental parameters, and thermal characteristic parameters of the heated object; the action space includes the identification weights of the dynamic thermodynamic model, the weight coefficients of the rolling optimization, the gain coefficients of the environmental compensation, and the weight allocation coefficients of the data fusion.

[0037] The reward function focuses on controlling accuracy and stability, and its expression is:

[0038]

[0039] in, The deviation between the current fusion temperature and the target temperature. Let Variance be the temperature deviation within the sliding window. This represents the change in the power control command. For weighting coefficients; when hour, It is a positive reward value, otherwise it is 0, guiding the agent to converge to the optimal parameters for ultra-high precision control.

[0040] According to one or more embodiments of the present invention, S4 further includes a sensor fault diagnosis and fault tolerance step: real-time monitoring of the deviation between the measured values ​​of each sensor and the fused temperature value; when the deviation of a certain sensor continues to exceed a preset fault threshold, the sensor is determined to be faulty, the faulty sensor data is automatically removed, the fusion weight of the remaining sensors is reallocated, the continuous and stable operation of the control closed loop is ensured, and the system reliability is improved.

[0041] According to one or more embodiments of the present invention, the method further includes a pre-calibration step: after power-on, a preset heating-heating-cooling calibration process is executed, full-temperature-range thermal response data of the heated object is collected, the initial thermodynamic model parameters and initial values ​​of control parameters are identified, and after pre-calibration is completed, the online control mode is entered to shorten the convergence time of the initial control.

[0042] Corresponding to the above method, the present invention also provides an online heating control system based on dynamic model prediction and multimodal feedback, comprising:

[0043] The multimodal sensing unit is used to collect multi-dimensional temperature data of the heated object, operating status data of the heating equipment, and environmental parameter data of the heating cavity in real time, and output a pre-processed standardized sensing dataset.

[0044] The dynamic model prediction unit, which is connected in communication with the multimodal sensing unit, is used to build and update the dynamic thermodynamic model online to predict the temperature change trend of the heated object in the next N control cycles.

[0045] The power pre-control generation unit communicates with the dynamic model prediction unit and is used to solve the rolling optimization objective function based on the predicted temperature trend and the target temperature curve to generate the optimal heating power pre-control command.

[0046] The multimodal feedback compensation unit is communicatively connected to the multimodal sensing unit and the power pre-control generation unit, respectively. It is used to fuse multi-source data to obtain high-precision temperature feedback values, calculate real-time compensation for environmental interference, perform closed-loop correction of power pre-control commands, and generate final heating power control commands.

[0047] The adaptive parameter optimization unit is connected to the dynamic model prediction unit, the power pre-control generation unit, and the multimodal feedback compensation unit, respectively. It has a built-in reinforcement learning agent for online iterative optimization of the control parameters of the entire process and adaptive thermal characteristics of different heating objects.

[0048] The heating execution unit is communicatively connected to the multimodal feedback compensation unit. It is used to receive heating power control commands, output the corresponding power to the heating module, and complete the heating control.

[0049] Beneficial effects

[0050] Compared with existing technologies, this invention provides an online heating control method and system based on dynamic model prediction and multimodal feedback, which has the following beneficial effects:

[0051] 1. This invention fundamentally solves the problem of lag in traditional heating control. By predicting future temperature change trends through an online-updated dynamic thermodynamic model and combining it with rolling optimization model predictive control, it achieves advanced control of heating power, completely avoiding temperature overshoot and oscillation caused by traditional PID post-adjustment, and significantly improving the response speed and stability of temperature control.

[0052] 2. This invention breaks through the accuracy limit of a single sensor. Through adaptive weighted fusion of multi-source temperature sensors, it balances the long-term stability of contact temperature measurement with the rapid response of non-contact temperature measurement, eliminating the random and systematic errors of a single sensor. At the same time, through a dynamic compensation model for multi-dimensional environmental interference, it offsets the effects of common industrial interferences such as airflow, power fluctuations, and temperature drift, and can stably achieve ultra-high precision temperature control within ±0.1℃, meeting the stringent requirements of high-end industrial scenarios.

[0053] 3. This invention has strong adaptability and versatility. By using a reinforcement learning agent to optimize the control parameters of the entire process online, it can adapt to the differences in thermal characteristics of different heating objects such as metals, semiconductors, and ceramics without manual parameter tuning. At the same time, it can adapt to the dynamic changes in environmental parameters. It can be directly applied to multi-category, flexible industrial production scenarios, which greatly reduces the debugging cost and time of production line switching.

[0054] 4. This invention has high reliability and industrial applicability. The entire process algorithm runs online in real time, and the control cycle can be as low as milliseconds, which is suitable for the needs of continuous online production in industry. At the same time, it has built-in sensor fault diagnosis and fault tolerance mechanism, which can maintain the stable operation of the control closed loop when a single sensor fails, avoid production accidents caused by sensor failure, and has extremely high practical value in industrial field. Attached Figure Description

[0055] Figure 1 The flowchart shows the steps of an online heating control method based on dynamic model prediction and multimodal feedback.

[0056] Figure 2 A schematic diagram of the framework for adaptive parameter optimization in reinforcement learning. Detailed Implementation

[0057] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0058] It should be noted that, unless otherwise specified, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0059] In this invention, unless otherwise stated, the directional terms such as "up" and "down" generally refer to the directions shown in the accompanying drawings, or to the vertical, perpendicular, or gravitational direction; similarly, for ease of understanding and description, "left" and "right" generally refer to the left and right shown in the accompanying drawings; "inner" and "outer" refer to the inner and outer contours of each component itself, but the above directional terms are not intended to limit this invention.

[0060] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings.

[0061] This embodiment is applied to the high-precision heating control scenario of semiconductor wafer annealing process. The wafer annealing process requires temperature control accuracy to be stable within ±0.1℃, with controllable heating rate, no overshoot, no oscillation, and can adapt to the thermal characteristics differences of wafers of different sizes and materials.

[0062] The online heating control system based on dynamic model prediction and multimodal feedback described in this embodiment has the following hardware architecture:

[0063] Multimodal sensing unit: including 2 K-type armored thermocouples (contact temperature measurement, range 0-600℃, accuracy ±0.2℃), 1 high-speed infrared temperature sensor (non-contact temperature measurement, range 0-500℃, response time 1ms, accuracy ±0.3℃), 1 power sensor, 1 wind speed sensor, 1 temperature and humidity sensor, and 1 voltage and current sensor.

[0064] Controller: An industrial-grade FPGA controller is used, with a control cycle of 10ms to ensure the real-time performance of the algorithm;

[0065] Heating actuator: includes a silicon controlled rectifier power regulator and an infrared heating tube array, rated power 5kW, power regulation resolution 0.1%;

[0066] Human-machine interaction unit: Industrial touch screen, used to set target temperature curves, display real-time temperature data and system status.

[0067] The online heating control method based on dynamic model prediction and multimodal feedback described in this embodiment is implemented in the following steps:

[0068] S0: Power-on precalibration

[0069] No-load calibration: Under no-load conditions, the heating platform is subjected to a complete calibration process of linear heating from 25℃ to 200℃ (heating rate 2℃ / s) → holding at 200℃ for 300s → natural cooling at 25℃. The time-series data of temperature, power, and environmental parameters throughout the process are collected by the multi-modal sensing unit. Based on the recursive least squares method, the core thermodynamic model parameters of the heating platform, such as initial heat capacity C, heat conversion efficiency η, convective heat transfer coefficient h, and radiative emissivity ε, are identified.

[0070] Load calibration: After loading the 8-inch silicon wafer to be processed, perform a small temperature rise test from 25℃ to 100℃ (heating rate 1℃ / s), collect thermal response data during the heating process, identify the total thermophysical parameters of the wafer and heating platform combination, complete the construction of the initial dynamic thermodynamic model, and initialize control parameters such as rolling optimization weight coefficient, environmental compensation gain coefficient, and data fusion weight. After the pre-calibration is completed, the system enters the online closed-loop control mode.

[0071] S1: Multimodal sensing data acquisition and preprocessing

[0072] Within each 10ms control cycle, multi-modal sensing units collect real-time sensing data across all dimensions. The specific data collection content and preprocessing process are as follows:

[0073] Data collection scope:

[0074] Multi-dimensional temperature data: Contact temperature data collected by two thermocouples at the center and edge of the wafer carrier area of ​​the heating platform, and non-contact temperature data collected by one infrared temperature sensor at the center of the upper surface of the wafer, forming a multi-source temperature measurement redundancy architecture with at least two contact and one non-contact sources.

[0075] Heating equipment operating status data: real-time output power, input voltage, input current, and thyristor on / off status of the heating module;

[0076] Environmental parameters of the heating chamber: airflow velocity, ambient temperature, ambient humidity, and chamber wall temperature;

[0077] Data preprocessing: The raw data collected above are preprocessed in sequence as follows: a 5-point moving average filter is used to remove high-frequency random noise, linear interpolation is used to achieve 10ms-level time alignment of multi-sensor data, and the 3σ criterion is used to remove outliers that exceed the normal fluctuation range. Finally, a standardized sensing dataset is obtained and synchronously output to the dynamic model prediction unit and the multimodal feedback compensation unit.

[0078] S2: Dynamic Thermodynamic Model Construction and Temperature Prediction

[0079] After receiving the standardized sensing dataset, the dynamic model prediction unit performs model building, online correction, and temperature prediction operations, achieving a perfect match. Figure 1 The logical flow of this step is as follows:

[0080] Initial Model Construction: Based on the law of conservation of energy, a lumped-parameter thermodynamic model is constructed, comprehensively considering the heat input of the heating module, convective heat loss between the heated object and the environment, and radiative heat loss. The model expression is as follows:

[0081]

[0082] in, The total heat capacity of the object being heated and the heating platform. For the real-time temperature of the object being heated, Heating power-to-thermal energy conversion efficiency For heating power, The convective heat transfer coefficient is... For heat exchange area, For ambient temperature, For radiative emissivity, It is the Stefan-Boltzmann constant;

[0083] Online correction of model parameters: Based on the preprocessed real-time sensing data, the recursive least squares method is used to identify and correct the total heat capacity C, heat conversion efficiency η, and convective heat transfer coefficient h in the model online, adapting to the drift of thermal property parameters caused by temperature changes during the heating process, and obtaining a dynamic thermodynamic model that is updated in real time.

[0084] Multi-step temperature prediction: Using the temperature, power, and environmental parameters of the current control cycle as initial values, the wafer temperature change sequence for the next 20 control cycles (i.e., the next 200ms) is predicted through a dynamic thermodynamic model that is updated in real time, providing a basis for advance control and balancing prediction accuracy and real-time calculation.

[0085] S3: Generation of heating power pre-control command

[0086] Using a preset wafer annealing target temperature curve (25℃→300℃, heating rate 5℃ / s, holding time 30min, cooling rate 2℃ / s) as the control objective, the constrained rolling optimization objective function is solved in each control cycle:

[0087]

[0088] The constraints are:

[0089]

[0090] By solving the above optimization problem through quadratic programming, the optimal heating power pre-control command for the current control cycle is obtained, realizing the advanced pre-control of temperature and avoiding overshoot caused by lag at the source.

[0091] S4: Multimodal Feedback Fusion and Real-time Compensation Correction

[0092] After receiving the standardized sensing dataset and heating power pre-control command, the multimodal feedback compensation unit performs multi-source data fusion, environmental interference compensation, closed-loop correction, and fault tolerance operations, achieving a perfect match. Figure 1 The feedback compensation logic for this step is as follows:

[0093] Adaptive weighted fusion of multi-source temperature data: For three temperature measurement data points (two thermocouples and one infrared temperature sensor), the measurement variance of each sensor within a 500ms sliding time window is calculated. The weights of each sensor are adaptively assigned based on the variance. Sensors with smaller variances are assigned higher weights. The weighting formula is as follows:

[0094]

[0095] The final fusion temperature value is:

[0096]

[0097] in, The temperature measurement value after preprocessing of the i-th sensor is obtained through this fusion algorithm, which takes into account both the long-term stability of contact temperature measurement and the fast response of non-contact temperature measurement, eliminates the random error and systematic error of a single sensor, and obtains a high-precision actual temperature feedback value.

[0098] Dynamic compensation for environmental disturbances: Construct a multi-dimensional environmental disturbance compensation model and calculate the real-time compensation amount for each.

[0099] Airflow heat loss compensation: Based on the real-time collected airflow velocity, the real-time change of the convective heat transfer coefficient is calculated, and then the convective heat loss compensation power is obtained to offset the heat loss change caused by airflow fluctuations in the cavity.

[0100] Power fluctuation compensation: Based on the deviation between the real-time collected input voltage and the rated 220V, the output deviation of heating power is calculated to obtain the power fluctuation compensation amount, thereby eliminating the power output error caused by grid voltage fluctuations.

[0101] Sensor temperature drift compensation: Based on the difference between the real-time ambient temperature and the 25℃ calibration temperature, the drift of the measured values ​​of each temperature sensor is corrected to eliminate the sensor system error caused by changes in ambient temperature.

[0102] Power command closed-loop correction: Combining the real-time deviation between the fused temperature feedback value and the target temperature, as well as the total environmental interference compensation calculated above, the heating power pre-control command is closed-loop corrected to generate the final heating power control command, which is output to the heating execution unit to drive the thyristor power regulator to output the corresponding power to the infrared heating tube array, thus completing the heating control.

[0103] Sensor fault diagnosis and fault tolerance: Real-time monitoring of the deviation between the measured values ​​of each sensor and the fused temperature value. When the deviation of a certain sensor continuously exceeds the preset fault threshold of 0.5℃ and the duration exceeds 5 control cycles, the sensor is determined to be faulty. The measurement data of the faulty sensor is automatically removed, and the fusion weight of the remaining normal sensors is reallocated to ensure the continuous and stable operation of the control closed loop and avoid production interruption and product scrapping caused by sensor failure.

[0104] S5: Adaptive Online Optimization of Control Parameters

[0105] A reinforcement learning agent based on the PPO algorithm is constructed. Its state space includes: current temperature deviation, temperature change rate, model prediction error, airflow speed, ambient temperature, and wafer thermal properties; its action space includes: model identification weights, rolling optimization weight coefficients, environmental compensation gain coefficients, and data fusion weight coefficients.

[0106] The reward function is set as follows:

[0107]

[0108] Among them, when hour, Otherwise, it is 0;

[0109] Appendix Figure 2 The hardware foundation of the reinforcement learning adaptive parameter optimization implementation method is an industrial-grade FPGA + ARM heterogeneous controller. The FPGA is responsible for real-time data acquisition, timing control, and heating power command output for steps S1-S4, with a fixed control cycle of 10ms. The ARM core deploys the reinforcement learning agent algorithm, running strictly in sync with the main control flow. Supporting hardware includes multiple contact and non-contact temperature sensors, environmental / equipment status sensors, a thyristor power regulator, and a heating module. The interactive object is the entire heating control system environment shown on the left side of the attached diagram, covering the entire chain of dynamic model prediction, power pre-control generation, multimodal feedback compensation, and heating execution, forming a complete closed loop of "environmental perception - intelligent decision-making - action execution - effect feedback - iterative optimization." The specific implementation of each link is as follows:

[0110] 1) Implementation of state space input

[0111] This section corresponds to the attached document. Figure 2 The state space input module is the perception entry point for the intelligent agent. Its specific hardware components and functions are as follows:

[0112] Multimodal sensing unit: corresponding Figure 2 The leftmost multimodal sensing unit module serves as the system's sensing input, used to collect real-time multi-dimensional temperature data of the heated object, operating status data of the heating equipment, and environmental parameter data of the heating cavity, and outputs a pre-processed standardized sensing dataset. In this embodiment, this unit specifically includes two K-type armored thermocouples (contact temperature measurement, range 0-600℃, accuracy ±0.2℃, installed at the center and edge of the wafer carrier area of ​​the heating platform, respectively), one high-speed infrared temperature sensor (non-contact temperature measurement, range 0-500℃, response time 1ms, accuracy ±0.3℃, vertically aligned with the center area of ​​the upper surface of the wafer), one power sensor (connected in series in the heating main circuit, used to collect the real-time output power of the heating module), one wind speed sensor (installed at the air inlet of the heating cavity, collected the airflow speed inside the cavity), one temperature and humidity sensor (installed on the side wall of the heating cavity, collected the ambient temperature and humidity inside the cavity), and one voltage and current sensor (connected in parallel to the power input terminal of the heating circuit, collected the input voltage and input current).

[0113] Dynamic model prediction unit: corresponding Figure 2The dynamic model prediction unit module is connected to the signal output of the multimodal sensing unit via its signal input terminal. It receives standardized sensing datasets, constructs and updates a dynamic thermodynamic model online, and predicts the temperature change trend of the heated object over the next N control cycles. In this embodiment, the core computation of this unit is deployed on the first computing core of the industrial-grade FPGA controller, with a control cycle set to 10ms to ensure real-time model identification and temperature prediction.

[0114] Power pre-control generation unit: corresponding Figure 2 The power pre-control generation unit module in the system has its signal input terminal communicatively connected to the signal output terminal of the dynamic model prediction unit. It receives the predicted temperature change trend sequence, combines it with a preset target temperature curve, solves the rolling optimization objective function, and generates the optimal heating power pre-control command. In this embodiment, the core computation of this unit is deployed on the second computation core of the FPGA controller, enabling parallel processing with the dynamic model prediction unit and reducing control latency.

[0115] Multimodal feedback compensation unit: corresponding Figure 2 The multimodal feedback compensation unit module has its first signal input terminal communicatively connected to the signal output terminal of the multimodal sensing unit, and its second signal input terminal communicatively connected to the signal output terminal of the power pre-control generation unit. It receives multi-source sensing data and heating power pre-control commands, performs adaptive weighted fusion of the multi-source temperature data to obtain a high-precision temperature feedback value, calculates real-time compensation for environmental interference, and performs closed-loop correction of the power pre-control commands based on temperature deviations to generate the final heating power control command. In this embodiment, the unit also integrates a sensor fault diagnosis and fault tolerance module to achieve closed-loop stable operation under abnormal conditions.

[0116] Adaptive parameter optimization unit: corresponding Figure 2 The adaptive parameter optimization unit module has its signal input terminals connected to the signal output terminals of the dynamic model prediction unit, power pre-control generation unit, and multimodal feedback compensation unit, respectively. It incorporates a reinforcement learning agent based on the PPO algorithm to receive control state data throughout the entire process, iteratively optimize the control parameters online, and adapt to the differences in thermal characteristics of different heating objects. In this embodiment, the operation of this unit is deployed on the third processing core of the FPGA controller, achieving decoupled operation of parameter optimization and real-time control, thus avoiding impact on the real-time performance of the main control closed loop.

[0117] Heating execution unit: corresponding Figure 2The rightmost heating execution unit module has its signal input terminal communicatively connected to the signal output terminal of the multimodal feedback compensation unit. It receives the final heating power control command and outputs the corresponding power to the heating module to complete the heating control. In this embodiment, the unit includes a silicon controlled rectifier (SCR) power regulator and an infrared heating tube array, with a rated power of 5kW and a power regulation resolution of 0.1%, enabling linear and precise power regulation.

[0118] 2) Implementation of dual network architecture

[0119] This section corresponds to the attached document. Figure 2 The value network, policy network, and state value evaluation modules adopt the classic actor-critic dual network architecture of the PPO algorithm, balancing decision rationality and training stability.

[0120] Value network and state value assessment: Taking the state space vector as input, it outputs the long-term value assessment result of the current system state, judges the degree of benefit of the current operating state to achieve the final control goal, provides an objective benchmark for subsequent policy updates, and avoids blind policy adjustments;

[0121] Policy network: synchronously receives state space input and outputs action decisions with corresponding optimized parameters. It is the core execution decision module of the agent.

[0122] 3) Implementation of reward function calculation

[0123] This section corresponds to the attached document. Figure 2 The reward function calculation module is the core guide for agent optimization. It is fully anchored to the core control objectives of ±0.1℃ temperature control accuracy, no overshoot, and high stability, and calculates real-time rewards based on the control effect feedback of the main process.

[0124] The core design of the reward function is as follows: a gradient penalty term is set for temperature control deviation, continuous temperature fluctuation, and drastic changes in heating power, while a high positive reward is set—positive incentive is triggered when the temperature deviation is ≤0.1℃, otherwise there is no reward, which accurately guides the agent to converge to the ultra-high precision control range first, which is in line with the core technical goal of this invention.

[0125] 4) Calculation of Advantage Function and Implementation of Strategy Update

[0126] This section corresponds to the attached document. Figure 2 The advantage function calculation and strategy update modules are the core guarantees for the PPO algorithm to adapt to industrial scenarios:

[0127] Advantage function calculation: Combining the state value assessment results with real-time rewards, the relative advantage of the current action is calculated to determine whether the current parameter adjustment can bring about a control effect improvement that exceeds the baseline, providing a reliable basis for judging the merits of the strategy update and avoiding ineffective adjustments;

[0128] This implementation method uses the Generalized Dominance Estimation (GAE) method to calculate the dominance function. The specific calculation formula is as follows:

[0129]

[0130] Among them, time difference error Discount factor A value of 0.99 represents the degree of importance attached to future rewards; decay coefficient The value is 0.95, representing the bias and variance of the balance advantage estimate.

[0131] In practice, a trajectory sequence of length 10 is constructed every 10 control cycles. The advantage function value at each moment is calculated based on this sequence, avoiding noise interference from single-step advantage estimation, improving the stability of strategy updates, and fully adapting to the stringent stability requirements of industrial control scenarios.

[0132] Strategy update implementation: Utilizing the pruning mechanism at the core of the PPO algorithm, the update magnitude is strictly limited, ensuring the deviation between the old and new strategies remains within a controllable range. This completely resolves the pain points of traditional reinforcement learning, such as easy divergence and system oscillation, and adapts to the stability requirements of continuous online industrial production. This implementation method completes a strategy iteration every 10 control cycles, balancing optimization efficiency with the real-time performance of the main control process.

[0133] 5) Motion space output and closed-loop execution implementation

[0134] This section corresponds to the attached document. Figure 2 The action space output module, with its optimized actions output by the strategy network, directly affects the core parameters of the entire heating control process, covering: online identification weights of thermodynamic models, weight coefficients of rolling optimization objective functions, environmental interference compensation gains, multi-source temperature data fusion weights, and sensor fault diagnosis thresholds. This enables collaborative optimization of parameters across the entire chain from perception, modeling, prediction to compensation, rather than local adjustment of a single parameter as in existing technologies.

[0135] After the action is output, the main control process executes the heating control of the next cycle based on the updated parameters. The resulting new state and control effect are fed back to the reinforcement learning agent, forming a continuous closed-loop iterative optimization until the heating control process ends.

[0136] The intelligent agent iterates and optimizes in each control cycle to maximize the cumulative reward. It adjusts the control parameters of the entire process online to adapt to the thermal characteristics of wafers of different sizes and materials. Without the need for manual parameter tuning, it can ensure that the control accuracy is stable within ±0.1℃.

[0137] This embodiment has been tested in practice. In the entire process of semiconductor wafer annealing, there is no overshoot during the heating stage, and the temperature fluctuation during the holding stage is stably controlled within ±0.08℃, which fully meets the requirements of high-end semiconductor manufacturing processes. When changing to wafers of different materials, the system can complete parameter adaptive optimization within 10 seconds without manual adjustment, which greatly improves the production efficiency and product yield of the production line.

Claims

1. An online heating control method based on dynamic model prediction and multi-modal feedback, characterized in that: Includes the following steps: S1 Multimodal Sensing Data Acquisition and Preprocessing: Real-time acquisition of multi-dimensional temperature data of the heated object, operating status data of the heating equipment, and environmental parameter data of the heating cavity. The acquired raw data undergoes filtering, noise reduction, time-series alignment, and outlier removal preprocessing to obtain a standardized sensing dataset. The multi-dimensional temperature data includes at least two contact temperature measurement data streams and at least one non-contact temperature measurement data stream. The contact temperature measurement data is collected by thermocouples and platinum resistance thermometers, while the non-contact temperature measurement data is collected by infrared temperature sensors and infrared thermal imagers. The operating status data includes the real-time output power, input voltage, input current, and on / off status of the heating module. The environmental parameter data includes the airflow velocity, ambient humidity, ambient temperature, and cavity wall temperature within the heating cavity. S2 Online Update Dynamic Thermodynamic Prediction Model Construction: Based on the law of conservation of energy and the heat transfer characteristics of the heating system, combined with preprocessed sensing data, a lumped-parameter dynamic thermodynamic model is constructed. The model's thermophysical parameters and heat loss coefficient are corrected online using real-time collected sensing data. Based on the corrected dynamic thermodynamic model, the temperature change trend of the heated object over the next N control cycles is predicted. The construction and updating of the dynamic thermodynamic model specifically includes: S21 Based on the law of conservation of energy, an initial lumped parameter thermodynamic model is constructed: Where C is the total heat capacity of the object being heated and the heating platform, T(t) is the real-time temperature of the object being heated, and η is the heating power-to-thermal energy conversion efficiency. Where is the heating power, h is the convective heat transfer coefficient, and A is the heat transfer area. Where ε is the ambient temperature, ε is the emissivity, and σ is the Stefan-Boltzmann constant; Based on the preprocessed real-time sensing data, S22 uses the recursive least squares method to identify and correct the heat capacity C, conversion efficiency η, and convective heat transfer coefficient h in the model online, and obtains a dynamic thermodynamic model that is updated in real time. S23 is based on a dynamic thermodynamic model, using the current control cycle state as the initial value, to predict the temperature sequence for the next N control cycles. The value of N ranges from 5 to 50, and the value of the control cycle ranges from 1 ms to 100 ms. S3 Model-Based Heating Power Pre-Control Command Generation: Using a preset target temperature curve as the control objective and the predicted temperature change trend as the basis, a constrained rolling optimization objective function is solved in each control cycle to obtain the optimal heating power pre-control command, achieving proactive temperature control. The rolling optimization objective function is: The constraints are: in, The predicted temperature for the i-th control cycle. The target temperature for the i-th control cycle is... This represents the power change between adjacent control cycles. These are the weighting coefficients. These are the upper and lower limits of the heating power. This represents the maximum permissible rate of temperature change. S4 Multimodal Feedback Fusion and Real-time Compensation Correction: Adaptive weighted fusion of multi-dimensional temperature data and operating status data is performed to obtain a high-precision actual temperature feedback value; an environmental interference compensation model is constructed, and the real-time interference compensation amount is calculated based on environmental parameter data. Combined with the deviation between the temperature feedback value and the target temperature, the heating power pre-control command is corrected in a closed loop to generate the final heating power control command. S5 adaptive online optimization of control parameters based on reinforcement learning: Construct a parameter optimization reinforcement learning agent, with temperature control deviation, overshoot, and control stability as the core reward indicators, and dynamic thermodynamic model parameters, rolling optimization weight coefficients, and environmental compensation gain coefficients as optimization objects, to iteratively optimize the control parameters of the entire process online and adapt to the differences in thermal characteristics of different heating objects.

2. The online heating control method based on dynamic model prediction and multimodal feedback according to claim 1, characterized in that: In S4, the specific method of adaptive weighted fusion is as follows: For the measurement data from the M temperature sensors, calculate the measurement variance of each sensor within the sliding time window. The weights of each sensor are adaptively assigned based on the variance. : The final fusion temperature value is: in, The temperature measurement value is the preprocessed value of the i-th sensor.

3. The online heating control method based on dynamic model prediction and multimodal feedback according to claim 1, characterized in that: In S4, the environmental interference compensation model includes an airflow heat loss compensation module, a power fluctuation compensation module, and a sensor temperature drift compensation module. The airflow heat loss compensation module calculates the change in the convective heat transfer coefficient based on the real-time collected airflow velocity to obtain the convective heat loss compensation power. The power fluctuation compensation module calculates the output deviation of heating power based on the deviation between the real-time collected input voltage and the rated voltage, and obtains the power fluctuation compensation amount. The sensor temperature drift compensation module corrects the drift of the temperature sensor's measured values ​​based on the difference between the ambient temperature and the calibration temperature, thereby eliminating system errors caused by ambient temperature.

4. The online heating control method based on dynamic model prediction and multimodal feedback according to claim 1, characterized in that: In S5, the reinforcement learning agent uses the proximal policy optimization PPO algorithm. Its state space includes the current temperature deviation, temperature change rate, model prediction error, environmental parameters, and thermal characteristic parameters of the heated object; its action space includes the identification weights of the dynamic thermodynamic model, the weight coefficients of the rolling optimization, the gain coefficients of the environmental compensation, and the weight allocation coefficients of the data fusion. The reward function is: in, The deviation between the current fusion temperature and the target temperature. Let Variance be the temperature deviation within the sliding window. This represents the change in the power control command. For weighting coefficients; when hour, If it is a positive reward value, then it is 0 otherwise.

5. The online heating control method based on dynamic model prediction and multimodal feedback according to claim 1, characterized in that: The S4 also includes a sensor fault diagnosis and fault tolerance step: real-time monitoring of the deviation between the measured values ​​of each sensor and the fused temperature value; when the deviation of a certain sensor continues to exceed the preset fault threshold, the sensor is determined to be faulty, the faulty sensor data is automatically removed, and the fusion weights of the remaining sensors are reallocated to ensure the continuous and stable operation of the control closed loop.

6. The online heating control method based on dynamic model prediction and multimodal feedback according to claim 1, characterized in that: The method also includes a pre-calibration step: after power-on, a preset heating-heating-cooling calibration process is executed, the full-temperature-range thermal response data of the heated object is collected, the initial thermodynamic model parameters and the initial values ​​of the control parameters are identified, and after the pre-calibration is completed, the online control mode is entered.

7. An online heating control system based on dynamic model prediction and multimodal feedback, characterized in that: include: The multimodal sensing unit is used to collect multi-dimensional temperature data of the heated object, operating status data of the heating equipment, and environmental parameter data of the heating cavity in real time, and output a pre-processed standardized sensing dataset. The dynamic model prediction unit, which is connected in communication with the multimodal sensing unit, is used to build and update the dynamic thermodynamic model online to predict the temperature change trend of the heated object in the next N control cycles. The power pre-control generation unit communicates with the dynamic model prediction unit and is used to solve the rolling optimization objective function based on the predicted temperature trend and the target temperature curve to generate the optimal heating power pre-control command. The multimodal feedback compensation unit is communicatively connected to the multimodal sensing unit and the power pre-control generation unit, respectively. It is used to fuse multi-source data to obtain high-precision temperature feedback values, calculate real-time compensation for environmental interference, perform closed-loop correction of power pre-control commands, and generate final heating power control commands. The adaptive parameter optimization unit is connected to the dynamic model prediction unit, the power pre-control generation unit, and the multimodal feedback compensation unit, respectively. It has a built-in reinforcement learning agent for online iterative optimization of the control parameters of the entire process and adaptive thermal characteristics of different heating objects. The heating execution unit is communicatively connected to the multimodal feedback compensation unit. It is used to receive heating power control commands, output the corresponding power to the heating module, and complete the heating control.