Power regulation and detection method and system for temperature controller based on AI intelligence

By using an AI-based intelligent temperature controller power adjustment method, and leveraging digital twin models and multimodal sensing data, combined with convolutional neural networks and graph neural networks, precise personalized control of materials is achieved. This solves the problems of unstable product yield and poor performance consistency in traditional heating furnace control methods, thereby improving production efficiency and product quality.

CN121557754BActive Publication Date: 2026-03-31HUNAN SANSUO INTERNET OF THINGS INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-23
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional heating furnace control methods cannot achieve personalized and optimized processing of materials, resulting in unstable product yield and poor performance consistency, and failing to effectively utilize the potential of multi-zone heating units.

Method used

An AI-based intelligent temperature controller power regulation method is adopted. By acquiring the digital twin model of the material and multimodal sensing data, combined with convolutional neural networks and graph neural networks, a personalized physical profile is constructed, and hierarchical reinforcement learning is performed to generate a refined power regulation scheme.

Benefits of technology

It enables precise insight into individual material differences, proactively avoids the risk of thermal failure, improves product yield and performance consistency, ensures a dynamic optimal balance between production efficiency and safety, and has high robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121557754B_ABST
    Figure CN121557754B_ABST
Patent Text Reader

Abstract

The application provides an AI intelligent-based temperature controller power regulation and detection method and system, relates to the technical field of intelligent manufacturing, and obtains online multi-modal sensing data of a material to be processed, identifies micro defect distribution through a physical information neural network and a graph neural network, generates individualized time-varying evolution field data representing the thermodynamic evolution law of the material in a heating process, and constructs a digital twin model thereof. Through interaction between an agent and the digital twin model in a hierarchical reinforcement learning framework, a high-level strategy selects a macro process curve, and a low-level strategy generates a partitioned candidate power regulation scheme capable of actively avoiding thermal failure risks on this basis. Through Monte Carlo simulation, the expected time integral performance value of the scheme under various material disturbance scenarios is evaluated, and a composite reward signal is formed to iteratively optimize the agent strategy, thereby improving product yield and performance consistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent manufacturing technology, and particularly to a method and system for power regulation and detection of a thermostat based on AI intelligence. Background Technique

[0002] In the manufacturing processes of photovoltaic cells and semiconductor wafers, heat treatment (such as rapid thermal annealing, diffusion, oxidation, etc.) is a crucial process step. By precisely controlling the temperature curve in a high-temperature furnace body, purposes such as doping, activation, defect repair, and thin film growth of materials can be achieved. The uniformity, stability, and precision of the heat treatment process directly determine the electrical properties, mechanical strength, and yield of the final product. Therefore, efficient and intelligent power regulation and control of the heating furnace body are the core technical challenges for improving product quality and reducing production costs.

[0003] Traditional control methods usually rely on a set of fixed standard operating procedures formulated based on experience. The standard operating procedures set a unified power output and heating-up curve for all materials to be processed, completely ignoring the microscopic differences between individual materials. For example, silicon wafers in different batches or even within the same batch may have different types, sizes, and distributions of microscopic defects (such as microcracks, impurity clusters, cutting damages, etc.) inside. For materials with high-risk defects, using the standard rapid heating-up curve is extremely likely to cause stress concentration, resulting in fragmentation during the processing; while for materials with excellent quality, using an overly conservative heating-up curve will reduce production efficiency. This unified and lagging control method cannot achieve personalized and optimal processing of materials, which is the fundamental reason for unstable product yield and poor performance consistency. Moreover, modern heating furnace bodies are usually equipped with an array of heating units that are independently controllable in multiple zones, which provides a physical basis for realizing refined spatial temperature field control. However, traditional control methods are difficult to effectively utilize this ability. Simultaneously optimizing the power output of dozens of heating units at hundreds of time points to suppress local risks while meeting the global heating-up target is an extremely complex high-dimensional, non-linear optimization problem, far beyond the solving ability of traditional control algorithms and manual experience. Therefore, actual operations often simplify to unified or simple linkage control of each zone, resulting in rough control means and being unable to fully utilize the potential of multi-zone heating, far from the true optimal control. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for power regulation and detection of a thermostat based on AI intelligence to solve the problems raised in the above background technique.

[0005] To achieve the above purpose, the present invention provides the following technical solutions:

[0006] The AI-based intelligent temperature controller power adjustment and detection method is applied to the heating furnace body of photovoltaic production. The specific steps include: S1: Obtain the time-varying evolution field data of the material and thermodynamics of the digital twin model of the material to be processed within the preset processing cycle. The time-varying evolution field data is configured to characterize the law of evolution of the physical and chemical properties of the material at each position over time during the heating process.

[0007] S2: In a hierarchical reinforcement learning framework, obtain the candidate power adjustment scheme generated by the agent for the current thermodynamic state of the digital twin model;

[0008] S3: Based on time-varying evolution field data, the digital twin model with candidate power adjustment schemes is subjected to accelerated time-domain simulation to calculate the time integral performance value of the candidate power adjustment schemes within the preset processing cycle.

[0009] S4: Generate a reward signal based on the time integral performance value. The reward signal is configured to characterize the overall performance of the candidate power adjustment scheme throughout the process and is used to update the agent's policy network.

[0010] Furthermore, S1 includes: acquiring multimodal sensing data online, acquiring multimodal sensing data at the instant before the material enters the furnace body via the conveyor belt; the multimodal sensing data includes visual image data and hyperspectral image data;

[0011] The steps for generating time-varying evolution field data using a physical information neural network model specifically include: using a convolutional neural network to perform pixel-level segmentation and feature extraction on hyperspectral and machine vision images in online multimodal sensing data, outputting classification confidence vectors representing the defect types of each micro-region; constructing a spatial adjacency graph of the material surface, where nodes represent each micro-region and edges represent their thermodynamic transfer relationships; using a graph neural network, with the classification confidence vector as the initial feature of the nodes, and utilizing the evolution mechanism of how different defects interact in the thermal field extracted from the process knowledge graph to construct constraints for information propagation, iteratively reasoning on the spatial adjacency graph, and outputting an interactive evolution parameter set that can characterize the warping and stress evolution rate of each micro-region during heating; and using the interactive evolution parameter set as the dynamic boundary condition of the physical information neural network model to generate time-varying evolution field data.

[0012] Furthermore, the feature extraction steps include: constructing the average gray value, texture complexity index, spectral impurity index, and spectral slope of each micro-region, and combining them in a predetermined order to generate a fused feature vector;

[0013] The average gray value is calculated as follows: obtain the gray value of each pixel in the micro-region, and perform an arithmetic average operation on all the obtained gray values. The average value is the average gray value.

[0014] The texture complexity index is calculated as follows: The center pixel of the micro-region is taken as the reference pixel, and multiple neighboring pixels around the reference pixel are obtained. The grayscale value of each neighboring pixel is compared with the grayscale value of the reference pixel. When the grayscale value of the neighboring pixel is greater than or equal to the grayscale value of the reference pixel, a binary value "1" is generated; otherwise, a binary value "0" is generated. All generated binary values ​​are combined in a preset order to form a binary sequence. The binary sequence is converted into a corresponding decimal value, which is the texture complexity index.

[0015] The spectral impurity index is calculated as follows: the spectral reflectance intensity value at a first preset wavelength is obtained as the first intensity value, and the spectral reflectance intensity value at a second preset wavelength is obtained as the second intensity value; the first intensity value and the second intensity value are divided, and the quotient is the spectral impurity index.

[0016] The spectral slope is calculated as follows: obtain the spectral reflection intensity value at the third preset wavelength as the third intensity value, and obtain the spectral reflection intensity value at the fourth preset wavelength as the fourth intensity value; calculate the difference between the third intensity value and the fourth intensity value to obtain the intensity difference; calculate the difference between the third preset wavelength and the fourth preset wavelength to obtain the wavelength difference; divide the intensity difference by the wavelength difference, and the quotient is the spectral slope.

[0017] Furthermore, the step of inputting the fused feature vector into the convolutional neural network model and outputting the final classification confidence vector specifically includes: S10001, for each preset defect type, obtaining the weight vector and bias term that uniquely correspond to the current defect type from the convolutional neural network model;

[0018] S10002. Perform a weighted summation operation on the fused feature vector and weight vector to obtain the weighted sum;

[0019] S10003. Add the weighted sum to the bias term, and the sum is the original score for the current defect type.

[0020] S10004. Traverse all preset defect types and repeat steps S10001 to S10003 to generate an original score vector containing the original scores of each defect type.

[0021] S1232. For each original score in the original score vector, perform an exponential operation with the base e of the natural logarithm and the current original score as the exponent to obtain a corresponding exponential score; sum all the generated exponential scores to obtain a total value; divide each exponential score by the total value, and the quotient is the final classification confidence vector for the corresponding defect type; based on the final classification confidence vector for the corresponding defect type, infer the warping evolution rate and stress evolution rate.

[0022] Furthermore, based on the warp evolution rate and stress evolution rate, time-varying evolution field data are generated. Specific steps include:

[0023] S141. Obtain a time and temperature sequence describing the material heating process. The sequence includes a series of discrete time points and the corresponding measured or preset temperature value at each time point.

[0024] S142. Set the initial state of all micro-regions at time point t=0, where the initial transverse stress is 0 MPa and the initial vertical displacement is 0 micrometers, and obtain the initial temperature from the time and temperature sequence in step S141, denoted as Tinitial.

[0025] S143. Define a calculation method for calculating the current temperature at any target time point t:

[0026] S1431. In the time and temperature sequences obtained in step S141, find the two closest consecutive time points (denoted as t) that include the target time point t. a and t b ) and their corresponding temperature values ​​(denoted as T) a and T b );

[0027] S1432. Based on two time points and their temperature values, the precise current temperature of the target time point t is calculated using linear interpolation.

[0028] S144. Steps for calculating the vertical displacement at any time point:

[0029] S1441. Obtain the warping evolution rate of a specified micro-region;

[0030] S1442. Subtract the current temperature T0 calculated in step S1432 from the initial temperature T0 set in step S142 to obtain the temperature difference value.

[0031] S1443. Multiply the warping evolution rate obtained in step S1441, the temperature difference obtained in step S1442, and the target time t together. The resulting product is the total vertical displacement at time t.

[0032] S145. Steps for calculating transverse stress at any time point:

[0033] S1451, Obtain the stress evolution rate of a specified micro-region;

[0034] S1452. Multiply the stress evolution rate obtained in step S1451, the temperature difference obtained in step S1442, and the target time t together. The resulting product is the total transverse stress at time point t.

[0035] Furthermore, the step of calculating the time integral performance value in S3 specifically includes: within the time integral interval defined by the preset processing cycle, performing an integral operation on the preset instantaneous performance health value, which characterizes the comprehensive performance of the digital twin model at any point in time; the instantaneous performance health value takes at least the temperature uniformity index and product yield prediction value of the digital twin model at the corresponding point in time as positive inputs, and the instantaneous energy consumption index associated with the candidate power adjustment scheme as negative inputs.

[0036] Furthermore, the integration operation is a stochastic integration process based on the Monte Carlo method, which specifically includes: performing multiple weighted random samplings from a pre-set material disturbance scenario library containing various batch-to-batch differences of materials to be processed and their occurrence probabilities to generate a set of virtual material processing paths; independently calculating the corresponding time integral performance value for each virtual material processing path; and weighting the time integral performance values ​​calculated for all virtual material processing paths according to their corresponding path occurrence probabilities to obtain the final risk-adjusted expected time integral performance value.

[0037] Furthermore, in S4, the step of generating the reward signal specifically includes: monitoring the temperature uniformity index in real time during the simulation of calculating the time integral performance value; when the temperature uniformity index is lower than the preset lower limit of the process window, generating a process penalty value proportional to the index deviation value; after the simulation process ends, decomposing the final expected time integral performance value into a cumulative yield component characterizing the overall product quality and a cost penalty component characterizing the total energy consumption, forming a multi-dimensional terminal reward vector; combining the process penalty value with the multi-dimensional terminal reward vector to form a composite reward signal, which is used to update the policy network.

[0038] Furthermore, it also includes step S5: after the strategy network converges, by performing perturbation analysis on the power output of each heating unit in the candidate power adjustment scheme at different time points, the sensitivity of the expected time integral performance value to the change of each power output is calculated, and the set of key control points that have the greatest impact on the overall process performance is identified based on the sensitivity.

[0039] A temperature controller power regulation and detection system based on AI intelligence, comprising:

[0040] The time-varying evolution field generation module is used to acquire the time-varying evolution field data of the material and thermodynamics of the digital twin model of the material to be processed within a preset processing cycle;

[0041] The hierarchical policy generation module is used to obtain candidate power adjustment schemes generated by the agent for the current thermodynamic state of the digital twin model in a hierarchical reinforcement learning framework.

[0042] The simulation and performance evaluation module is used to perform accelerated time-domain simulations of digital twin models that apply candidate power regulation schemes based on time-varying evolution field data, thereby calculating the time integral performance value of the candidate power regulation schemes.

[0043] The network update module is used to generate a reward signal based on the time integral performance value, and to update the agent's policy network using the reward signal.

[0044] Compared with the prior art, the beneficial effects of the present invention are:

[0045] This invention acquires multimodal sensing data, including visual and hyperspectral data, of the materials to be processed online. By combining convolutional neural networks, graph neural networks, and physical information neural networks, it constructs a high-precision, personalized digital twin model capable of proactively predicting the evolution path of defect interactions. This approach no longer ignores individual material differences but creates a unique physical profile for each piece of material, revealing potential failure risks caused by microscopic defects at their source. This proactive risk insight for each individual material provides accurate and reliable physical evidence for subsequent intelligent decision-making, reducing the passivity and blindness of traditional control methods.

[0046] This invention also employs a hierarchical reinforcement learning framework to decompose complex control tasks into high-level macroscopic process curve selection and low-level fine-grained regional power adjustment. This achieves the beneficial effect of proactively avoiding thermal failure risks while ensuring overall processing efficiency, thereby improving product yield and performance consistency. The high-level strategy of this invention ensures the strategic rationality of the global heating rhythm, while the low-level strategy performs power fine-tuning based on real-time predictions from a digital twin model to suppress local stress and warpage. This strategy achieves a dynamic optimal balance between production efficiency and product safety, ensuring high quality of the final product.

[0047] This invention introduces a Monte Carlo-based stochastic integral process to conduct a comprehensive, risk-adjusted performance evaluation of candidate power adjustment schemes within a disturbance scenario library containing various batch-to-batch material variations. This results in a highly robust control strategy that effectively withstands the impact of raw material fluctuations in real production. Instead of optimizing based on a single ideal condition, this invention ensures that the power scheme remains robust under various potential non-ideal conditions. By combining information-rich composite reward signals, the strategy learned by the agent is more closely aligned with actual operating conditions, enhancing the stability and reliability of the process. Attached Figure Description

[0048] Figure 1 This is a schematic diagram illustrating the overall process of the method of the present invention;

[0049] Figure 2 This is a schematic diagram of the system flow of the present invention. Detailed Implementation

[0050] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0051] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0052] Example 1:

[0053] Please see Figure 1 This invention provides a technical solution: an AI-based intelligent temperature controller power adjustment and detection method, applied to a photovoltaic production heating furnace, the specific steps of which include:

[0054] S1: Obtain the time-varying evolution field data of the material and thermodynamics of the digital twin model of the material to be processed within the preset processing cycle. The time-varying evolution field data is configured to characterize the evolution of the physical and chemical properties of the material at each location over time during the heating process.

[0055] Based on online multimodal sensing data of the material to be processed before it enters the furnace, the distribution pattern of micro-defects on the surface of the material is identified; and combined with a pre-set process knowledge graph that stores the correspondence between material defects and thermal failure mechanisms, time-varying evolution field data is generated through a physical information neural network model.

[0056] The steps for generating time-varying evolution field data using a physical information neural network model specifically include: using a convolutional neural network to perform pixel-level segmentation and feature extraction on hyperspectral and machine vision images in online multimodal sensing data, outputting classification confidence vectors representing the defect types of each micro-region; constructing a spatial adjacency graph of the material surface, where nodes represent each micro-region and edges represent their thermodynamic transfer relationships; using a graph neural network, with the classification confidence vector as the initial feature of the nodes, and utilizing the evolution mechanism of how different defects interact in the thermal field extracted from the process knowledge graph to construct constraints for information propagation, iteratively reasoning on the spatial adjacency graph, and outputting an interactive evolution parameter set that can characterize the warping and stress evolution rate of each micro-region during heating; and using the interactive evolution parameter set as the dynamic boundary condition of the physical information neural network model to generate time-varying evolution field data.

[0057] The application scenario of this embodiment is the intelligent control of the rapid thermal processing (RTP) process for a monocrystalline silicon wafer (hereinafter referred to as "material") with a size of 158.75mm×158.75mm and a thickness of 180μm in the photovoltaic industry.

[0058] S1: Example of generating time-varying evolution field data: The core of this step is to generate a high-precision digital twin model of the current single crystal silicon wafer before it enters the heating furnace, and predict its thermodynamic evolution law within a preset processing cycle (e.g., heating to 1000°C and holding for 30 seconds).

[0059] S11. Online acquisition of multimodal sensing data: The online detection system performs a non-contact scan on the material just before it enters the furnace body via the conveyor belt, acquiring a set of multimodal sensing data.

[0060] Multimodal sensing data includes visual image data and hyperspectral image data;

[0061] The visual image data, acquired via a linear charge-coupled device (CCD) camera, has a resolution of 8192×8192 pixels and a grayscale of 16 bits. Currently, this visual image data is primarily used to identify macroscopic and microscopic morphological defects, such as edge cracks, surface scratches, and pits.

[0062] Hyperspectral image data, acquired via a pushbroom hyperspectral imager, has a spatial resolution aligned with visual images, covering a spectral range of 400 nm to 1700 nm and containing 256 spectral channels. This hyperspectral image data is currently used to identify minute differences in material composition, such as uneven doping concentrations, or microscopic regions of aggregated oxygen or carbon impurities.

[0063] Data structure example: The collected data is integrated into a data packet, the structure of which is shown in Table 1.

[0064] Table 1: Structure of Multimodal Sensing Data Packets in a Single Acquisition

[0065]

[0066] S12: Microscopic Defect Identification and Classification Based on Convolutional Neural Network Model: The acquired multimodal data is input into a pre-trained convolutional neural network model for pixel-level segmentation and feature extraction. The pre-trained convolutional neural network model is obtained through the following steps: First, a training dataset containing a large amount of historical material multimodal sensing data is constructed. Each sample in this dataset is manually annotated at the pixel level by process experts to clarify the true defect type in each microscopic region. Then, this annotated dataset is input into the convolutional neural network model architecture. Using the cross-entropy loss function as the optimization objective, the backpropagation algorithm and gradient descent optimizer (set to Adam optimizer) are used to iteratively train the network weights until the model reaches a preset classification accuracy (e.g., 98%) on an independent validation set. Finally, the network weights are saved, forming the final usable pre-trained model.

[0067] The convolutional neural network (CNN) model acquires grayscale values, texture features, and spectral curve features from aligned visual and hyperspectral data using image processing techniques. The CNN model analyzes these features pixel-by-pixel (or in units of 3×3 pixels) of microscopic regions. For each microscopic region on the material surface (e.g., a 10 μm × 10 μm region), the CNN outputs a classification confidence vector. This vector represents the probability that the current microscopic region belongs to different preset defect types. The specific steps include:

[0068] S121. Convert visual image data into grayscale image data, and extract the spectral reflectance intensity value of each key wavelength of hyperspectral data;

[0069] S122. Construct the average gray value, texture complexity index, spectral impurity index and spectral slope of each micro-region, and combine them in a predetermined order to generate a fused feature vector.

[0070] The average gray value is calculated as follows: obtain the gray value of each pixel in the micro-region, and perform an arithmetic average operation on all the obtained gray values. The average value is the average gray value.

[0071] The texture complexity index is calculated as follows: The center pixel of the micro-region is taken as the reference pixel, and multiple neighboring pixels around the reference pixel are obtained. The grayscale value of each neighboring pixel is compared with the grayscale value of the reference pixel. When the grayscale value of the neighboring pixel is greater than or equal to the grayscale value of the reference pixel, a binary value "1" is generated; otherwise, a binary value "0" is generated. All generated binary values ​​are combined in a preset order to form a binary sequence. The binary sequence is converted into a corresponding decimal value, which is the texture complexity index.

[0072] The spectral impurity index is calculated as follows: the spectral reflectance intensity value at a first preset wavelength (set to 850 nm) is obtained as the first intensity value, and the spectral reflectance intensity value at a second preset wavelength (set to 950 nm) is obtained as the second intensity value; the first intensity value and the second intensity value are divided, and the quotient is the spectral impurity index.

[0073] The spectral slope is calculated as follows: obtain the spectral reflectance intensity value at the third preset wavelength (set to 1200nm) as the third intensity value, and obtain the spectral reflectance intensity value at the fourth preset wavelength (set to 1100nm) as the fourth intensity value; calculate the difference between the third intensity value and the fourth intensity value to obtain the intensity difference; calculate the difference between the third preset wavelength and the fourth preset wavelength to obtain the wavelength difference; divide the intensity difference by the wavelength difference, and the quotient is the spectral slope.

[0074] S123, The step of inputting the fused feature vector into the convolutional neural network model and outputting the final classification confidence vector specifically includes: S1231, the original score generation step:

[0075] S10001. For each preset defect type, obtain the weight vector and bias term that uniquely correspond to the current defect type from the convolutional neural network model;

[0076] S10002. Perform a weighted summation operation on the fused feature vector and the weight vector (that is, multiply the elements at corresponding positions in the vectors and then sum them) to obtain the weighted sum;

[0077] S10003. Add the weighted sum to the bias term, and the sum is the original score for the current defect type.

[0078] S10004. Traverse all preset defect types and repeat steps S10001 to S10003 to generate an original score vector containing the original scores of each defect type.

[0079] S1232, Classification Confidence Vector Normalization Steps: For each original score in the original score vector, perform an exponential operation with the base e of the natural logarithm and the current original score as the exponent to obtain a corresponding exponential score; sum all generated exponential scores to obtain a total value; divide each exponential score by the total value, and the quotient is the classification confidence vector for the corresponding defect type; combine all final confidence scores in the order corresponding to the defect types to generate the final classification confidence vector.

[0080] Data structure example: Four key defect types are preset: normal micro-regions, microcracks, impurity clusters, and cutting damage. The output of the current neural network can be organized into the structure shown in Table 2.

[0081] Table 2: Confidence vector for micro-region defect classification output by the convolutional neural network

[0082]

[0083] Set a confidence threshold of 0.7. When the classification confidence vector of the corresponding defect type exceeds 0.7, it is identified as the corresponding defect type.

[0084] S13. Interactive Evolution Parameter Inference Based on Graph Neural Networks; This step aims to simulate how different defects interact when heated, thereby predicting key thermodynamic evolution rates. This includes: constructing a spatial adjacency graph, abstracting the material surface as a graph, which contains a set of nodes and edges. Each node represents a micro-region, and its initial feature is the classification confidence vector corresponding to Table 2. Edges specifically connect adjacent nodes in space. The edge weights are set based on the material's inherent thermal conductivity and the distance between nodes, representing the strength of the thermodynamic transfer relationship between nodes. Constraints are set for the spatial adjacency graph, extracting inference rules from a pre-defined process knowledge graph as physical constraints for the graph neural network information propagation process, including:

[0085] Rule 1 (Stress Concentration Effect): If the type of the removed node is "microcrack", then during information propagation, the information it transmits to the adjacent "normal micro-region" nodes will have an additional stress growth factor greater than 1.

[0086] Rule 2 (Differential Heat Absorption Effect): If the node type is "impurity cluster," then in subsequent calculations, the current heat absorption rate is set to 1.2 times that of the normal micro-region. The graph neural network model performs multiple iterations on the constructed spatial adjacency graph. In each iteration, each node updates its own features based on the state of its neighboring nodes and the constraints of the S131 spatial adjacency graph. After a set number of iterations, the network converges, outputting a stable set of parameters that characterizes the dynamic evolution trend of each micro-region.

[0087] Interactive evolutionary parameter inference based on graph neural networks specifically includes:

[0088] S131, Steps for generating the neighborhood influence vector:

[0089] S1311. For the target micro-region (e.g. (3015, 4096)), first identify all its spatially adjacent neighborhood micro-regions.

[0090] S1312. Obtain the classification confidence vector of each neighborhood micro-region (from Table 2), and obtain the weight values ​​of the edges connecting the target micro-region and the current neighborhood region;

[0091] S1313. Multiply the classification confidence vector of the neighborhood region with the weight value of the edge to obtain the weighted neighborhood vector.

[0092] S1314. Repeat steps S1311 and S1313 for all neighboring regions, and sum all the generated weighted neighborhood vectors to obtain a neighborhood influence vector that integrates the influence of all neighbors.

[0093] S132. Steps for calculating the warpage evolution rate:

[0094] S1321. Obtain the classification confidence vector of the target micro-region itself, and the neighborhood influence vector calculated in S1314.

[0095] S1322. The classification confidence vector of the target micro-region itself is concatenated with the neighborhood influence vector to form a combined feature vector;

[0096] S1323. Application Rule 2: Check the portion of the combined feature vector contributed by neighbors of the "impurity cluster" type. If it exists, multiply the value of this portion by the preset thermal absorptivity factor (set to 1.2).

[0097] S1324. Obtain the first weight vector preset for calculating the warp evolution rate;

[0098] S1325. Perform a weighted summation operation (i.e., dot product operation) on the combined feature vector (which may have been modified by rules) obtained in step S1323 and the first weight vector. The resulting scalar value is the warping evolution rate of the target micro-region.

[0099] S133, Calculation steps for stress evolution rate:

[0100] S1331, Similarly, obtain the combined feature vector generated in S1322);

[0101] S1332. Application Rule 1: Check whether the target micro-region itself is a "microcrack". If so, multiply the part of the combined feature vector contributed by the target micro-region itself by a preset stress growth factor (a scalar greater than 1).

[0102] S1333, Obtain a second weight vector that is different from the first weight vector and is preset for calculating the stress evolution rate;

[0103] S1334. Perform a weighted summation (i.e., dot product) on the combined feature vector (which may have been modified by rules) obtained in step S1332 and the second weight vector. The resulting scalar value is the stress evolution rate of the target micro-region. The "first weight vector" and "second weight vector" mentioned in steps S1325 and S1334 are not fixed parameters, but belong to two independent, pre-trained linear regression models or small multilayer perceptrons. The training process for these two models is as follows: A training dataset is constructed, whose input is a combined feature vector containing neighborhood influences obtained through graph neural network inference, and whose output label is the actual warpage and stress evolution rate of each micro-region under a standard heating curve, simulated and calculated using high-precision offline finite element analysis software. Supervised learning training is performed on this dataset, enabling the model to learn the precise mapping relationship from the features output by the GNN to the physical evolution rate.

[0104] Data structure example: The final output of a graph neural network is a set of interactive evolutionary parameters, as shown in Table 3.

[0105] Table 3: Set of interactive evolutionary parameters output by the graph neural network

[0106]

[0107] Note: The micro-regions (3015, 4096) were identified as microcracks, and their predicted warpage and stress evolution rates were higher than those of the normal micro-regions.

[0108] S14: Based on the warp evolution rate and stress evolution rate, generate time-varying evolution field data. Specific steps include:

[0109] S141. Obtain a time and temperature sequence describing the material heating process. The sequence includes a series of discrete time points and the corresponding measured or preset temperature value at each time point.

[0110] S142. Set the initial state of all micro-regions at time point t=0, where the initial transverse stress is 0 MPa and the initial vertical displacement is 0 micrometers, and obtain the initial temperature from the time and temperature sequence in step S141, denoted as Tinitial.

[0111] S143. Define a calculation method for calculating the current temperature at any target time point t:

[0112] S1431. In the time and temperature sequences obtained in step S141, find the two closest consecutive time points (denoted as t) that include the target time point t. a and t b ) and their corresponding temperature values ​​(denoted as T) a and T b );

[0113] S1432. Based on two time points and their temperature values, the precise current temperature at the target time point t is calculated using linear interpolation, denoted as T0. The specific formula is: T0 = T a +(T b -T a )×(t×t a ) / (t b ×t a );

[0114] S144. Steps for calculating the vertical displacement at any time point:

[0115] S1441. Obtain the warping evolution rate of a specified micro-region;

[0116] S1442. Subtract the current temperature T0 calculated in step S1432 from the initial temperature T0 set in step S142 to obtain the temperature difference value.

[0117] S1443. Multiply the warping evolution rate obtained in step S1441, the temperature difference obtained in step S1442, and the target time t together. The resulting product is the total vertical displacement at time point t.

[0118] S145. Steps for calculating transverse stress at any time point:

[0119] S1451, Obtain the stress evolution rate of a specified micro-region;

[0120] S1452. Multiply the stress evolution rate obtained in step S1451, the temperature difference obtained in step S1442, and the target time t together. The resulting product is the total transverse stress at time point t, as shown in Table 4.

[0121] Table 4: Generated time-varying evolution field data (excerpt of some micro-regions and time points)

[0122]

[0123] Table 4 clearly shows that, due to the presence of initial defects and the inference of interactions, the micro-regions (3015, 4096) accumulated stress far exceeding that of normal micro-regions during the heating process, resulting in greater warping displacement. Through the detailed explanation of the above embodiments, this technical solution successfully integrates online sensing data, domain knowledge, and physical simulation. It generates a unique "physical evolution script" for each specific material to be processed, considering the interactive effects of micro-defects. This high-precision time-varying evolution field data provides a solid data foundation for subsequent feedforward, refined, and zoned adjustment of the temperature controller power, thereby improving the yield and performance consistency of photovoltaic products. This method has a clear technical path and feasibility, and improves upon existing technologies in terms of prediction accuracy and physical insight.

[0124] Example 2:

[0125] S2: In a hierarchical reinforcement learning framework, obtain the candidate power adjustment scheme generated by the agent for the current thermodynamic state of the digital twin model;

[0126] In the hierarchical reinforcement learning framework, an agent interacts with a digital twin model. The agent's action space is defined as the power applied to each heating unit and the time series matrix within a preset processing cycle. Based on the reward signal, the agent continuously optimizes its policy network to generate candidate power adjustment schemes.

[0127] The agent is a hierarchical reinforcement learning agent, which includes: a high-level policy network, configured to select a macroscopic process curve prototype as the target from a preset process curve prototype library based on the initial characteristics of the material to be processed and the global process objective; and a low-level policy network, configured to receive the macroscopic process curve prototype as an instruction and perform a fine-grained parameter search for power and time series matrices within the temperature and time envelope defined by the macroscopic process curve prototype to generate candidate power adjustment schemes.

[0128] The core of this step lies in using a hierarchical reinforcement learning framework to transform the digital twin model and its time-varying evolution field data generated in Example 1 for a specific material into an optimal, zoned power adjustment scheme for the heating furnace. This scheme aims to precisely execute heating tasks while proactively avoiding the risk of thermal failure caused by microscopic defects in the material. The physical environment of this example is set as follows: the heating source inside the photovoltaic production heating furnace consists of a 4×4 independently controllable infrared lamp array, with each lamp array unit covering a specific quadrant of the material surface, and its output power can be continuously adjusted between 0% and 100%.

[0129] S21: Selection of macroscopic process curves for high-level strategy networks; The role of high-level strategy networks is to select the most suitable macroscopic heating strategy for the material to be processed from a global perspective.

[0130] S211. Constructing a global defect feature vector: Aggregate the defect classification confidence data of each micro-region generated in Example 1 (S12) (as shown in Table 2) to generate a low-dimensional global defect feature vector. This vector is configured to characterize the overall defect status of the material. The specific construction steps include: (1) Calculating the total confidence integral value of the "microcrack" type: Sum the confidence of "microcracks" in all micro-regions to obtain a scalar characterizing the severity of cracks. (2) Calculating the maximum connected area of ​​the "impurity cluster" type: Identify spatially adjacent micro-regions whose confidence of "impurity clusters" is higher than the preset confidence threshold (preferred value is 0.7), and calculate the number of micro-regions contained in the maximum connected area. (3) Calculating the edge distribution density of the "cutting damage" type: Only count the sum of the confidence of "cutting damage" within a 100-pixel width band of the material edge. Combine the three scalar values ​​calculated above to form a three-dimensional global defect feature vector. For the material in this embodiment, since there are high-confidence microcracks in the micro-region (3015, 4096), the total confidence integral value of its "microcracks" will be significantly higher.

[0131] S212. Definition of the Macroscopic Process Curve Prototype Library: The system has a pre-stored library of various macroscopic process curve prototypes. Each prototype represents a different global heating strategy, which defines the envelope of the ideal average temperature of the material over time throughout the entire processing cycle. Examples are shown in Table 5 below.

[0132] Table 5: Macroscopic Process Curve Prototype Library

[0133]

[0134] S213, Selection Decision of Macroscopic Process Curve: The global defect feature vector generated in S211 is input into the pre-trained high-level strategy network (set as a small multilayer perceptron). Based on its knowledge learned from a large amount of historical processing data, the high-level strategy network outputs a selection instruction pointing to the prototype of the optimal macroscopic process curve. In this embodiment, since the global defect feature vector shows a high "microcrack" index, the high-level strategy network determines that the material has a high risk of local fracture, and therefore selects P03: the stepped stress release curve as the target instruction issued to the low-level strategy network.

[0135] S22: Refined power sequence generation of the low-level policy network: After receiving the high-level instruction (i.e., the P03 curve), the task of the low-level policy network is to interact with the digital twin model constructed in Example 1 at high frequency under the guidance of the macro-objective to generate specific power and time series matrices for each heating unit.

[0136] S221, Definition of State Space, Action Space and Reward Function: (1) State: At any time point t, the state input of the low-level agent is a composite data structure, which includes: The thermodynamic state of the current digital twin model: that is, the complete field data of temperature, lateral stress and vertical displacement of each micro-region at time point t, as shown in Table 4 of Example 1. Macroscopic target deviation: the difference between the average temperature of the current material and the target temperature of the P03 curve at time point t. (2) Action: The action of the low-level agent is a 16-dimensional floating-point vector, where each element corresponds to the power setting value (range [0,1]) of the heating unit at the next time point (e.g., the time interval is 0.1 seconds). (3) Reward function: In order to guide the agent to learn the optimal strategy, the calculation formula of the reward function Rt is set as: Rt=wtemp×Rtemp-wstress×Rstress-wwarp×Rwarp; where Rtemp represents the temperature tracking reward: when the error between the average temperature of the material and the target temperature of the P03 curve decreases, a positive reward is given. The temperature tracking reward Rtemp is calculated as follows: This reward is configured to measure the degree to which the average temperature of the material matches the target temperature of the macroscopic process curve (e.g., the PO3 curve) at the current time point t. Preferably, a Gaussian function is used for calculation, which can provide a smooth reward when the error is small and a rapidly decaying reward when the error is large. Rtemp = exp[-α×(Tavgactual-Ttarget)] 2 Where: Tavgactual is the arithmetic mean of all micro-region temperature values ​​obtained from the digital twin model at time t. Ttarget is the target temperature value retrieved from the selected macro-process curve prototype (P03) at time t. α is a positive constant used to adjust the reward sensitivity. Preferably, the value of α is set to 0.01. This setting ensures that the reward value remains at approximately 0.37 when the temperature deviation is ±10 degrees Celsius, and quickly approaches 1 when the deviation is smaller, thereby encouraging the agent to perform precise temperature control while having a certain tolerance for small, unavoidable fluctuations.

[0137] The stress over-limit penalty Rstress is calculated as follows: it is configured to apply a significant negative excitation when the lateral stress in any micro-region predicted by the digital twin model exceeds a preset safety threshold. Preferably, the penalty value is quadratic with the degree of exceeding the threshold to impose a more severe penalty for serious over-limit behavior. Rstress = [max(0, σmax-σthreshold)] 2 Where: σmax is the maximum value of the "lateral stress" in all micro-regions obtained from the digital twin model at time point t. σthreshold is a preset safety threshold for lateral stress in the material. Preferably, based on the mechanical properties of single-crystal silicon, the value of σthreshold is set to 150 MPa. When σmax does not exceed σthreshold, Rstress is 0, and no penalty is incurred.

[0138] The calculation method for the warpage penalty Rwarp is as follows: It is configured to apply a penalty when the predicted maximum warpage displacement of the material exceeds the upper limit of the process-allowed flatness. Rwarp = [max(0, dmax - dthreshold)] 2 Where: dmax is the maximum absolute value of the "vertical displacement" of all micro-regions obtained from the digital twin model at time point t. dthreshold is a preset upper limit for process flatness. Preferably, for a silicon wafer with a thickness of 180μm in this embodiment, the value of dthreshold is set to 50 micrometers (μm). When dmax does not exceed dthreshold, Rwarp is 0, and no penalty is incurred.

[0139] Preferred values ​​for the weighting coefficients: The weighting coefficients wtemp, wstress, and wwarp are used to balance the importance of different objectives. In photovoltaic production, avoiding material breakage or irreversible deformation is the primary task. Therefore, the weight of the penalty item should be significantly higher than the weight of the reward item. Preferably, the weighting coefficients are set as follows: wtemp = 1.0, wstress = 5.0, wwarp = 5.0. This weighting configuration explicitly conveys the following policy priority to the reinforcement learning agent: avoiding the risk of stress or warping exceeding limits, which is five times more important than simply tracking the temperature curve. Setting wstress and wwarp to the same order of magnitude is because both are related to the physical integrity of the material and should be given equal importance. This setting ensures that when exploring the optimal power scheme, the agent will prioritize generating actions that ensure material safety, even if this may cause a small deviation in tracking the target temperature curve.

[0140] S222, Policy Optimization and Candidate Solution Generation: The low-level policy network (e.g., using the Proximal Policy Optimization (PPO) algorithm) iterates the policy in the following loop: (1) Obtain the state: At time t (e.g., t=15 seconds), obtain the complete state of the digital twin model. (2) Generate actions: The low-level policy network outputs actions containing the power values ​​of 16 heating units based on the current state. (3) Interact with the environment: Apply this power action to the digital twin model. The model advances one step according to the physical evolution law defined in S14, calculating the new thermodynamic state (new temperature, stress, displacement field) at time t+Δt. (4) Calculate the reward: Calculate the reward value Rt using the reward function defined in S221 based on the new state and the P03 objective. (5) Network update: Update the parameters of the low-level policy network using the obtained (state, action, reward, new state) tuple. Repeat the above steps until the entire processing cycle is covered. After sufficient training, given initial materials, the low-level network can generate a complete power and time series matrix that maximizes the cumulative reward. This matrix is ​​the final candidate power adjustment scheme.

[0141] Data structure example: Continuing the scenario of Example 1, at time point t=15 seconds, the power adjustment scheme fragment generated by the low-level agent is shown in Table 6.

[0142] Table 6: Candidate Power Adjustment Schemes (Excerpt from time point t=15 seconds)

[0143]

[0144] Through the detailed explanation of the above embodiments, this technical solution demonstrates how to generate a forward-looking, zoned, and refined power regulation scheme based on a high-precision material digital twin model and a hierarchical reinforcement learning framework. This scheme moves beyond blind or uniform heating, opting instead for adaptive intelligent control tailored to the unique defect distribution of each material. High-level strategies ensure the rationality of the overall heating rhythm, while low-level strategies are finely adjusted during execution, proactively mitigating failure risks predicted by time-varying evolution field data, thereby improving the yield and performance consistency of the final product.

[0145] Example 3:

[0146] S3: Based on time-varying evolution field data, the digital twin model with candidate power adjustment schemes is subjected to accelerated time-domain simulation to calculate the time integral performance value of the candidate power adjustment schemes within the preset processing cycle.

[0147] The steps for calculating the time integral performance value in S3 specifically include: within the time integral interval defined by the preset processing cycle, performing an integral operation on the preset instantaneous performance health value, which characterizes the comprehensive performance of the digital twin model at any point in time; the instantaneous performance health value takes at least the temperature uniformity index and product yield prediction value of the digital twin model at the corresponding point in time as positive inputs, and the instantaneous energy consumption index associated with the candidate power adjustment scheme as negative inputs.

[0148] The integration operation is a stochastic integration process based on the Monte Carlo method, which specifically includes: multiple weighted random samplings from a pre-set material disturbance scenario library containing batch-to-batch differences of various materials to be processed and their occurrence probabilities to generate a set of virtual material processing paths; for each virtual material processing path, its corresponding time integration performance value is calculated independently; the time integration performance values ​​calculated for all virtual material processing paths are weighted and averaged according to their corresponding path occurrence probabilities to obtain the final risk-adjusted expected time integration performance value.

[0149] S4: Generate a reward signal based on the time integral performance value. The reward signal is configured to characterize the overall performance of the candidate power adjustment scheme throughout the process and is used to update the agent's policy network.

[0150] In S4, the steps for generating the reward signal specifically include: monitoring the temperature uniformity index in real time during the simulation of time integral performance value calculation; generating a process penalty value proportional to the index deviation value when the temperature uniformity index is lower than the preset lower limit of the process window; after the simulation process ends, decomposing the final expected time integral performance value into a cumulative yield component representing the overall product quality and a cost penalty component representing the total energy consumption, forming a multi-dimensional terminal reward vector; and combining the process penalty value with the multi-dimensional terminal reward vector to form a composite reward signal, which is used to update the policy network.

[0151] The immediate reward signal (Rt, defined in S221): This signal is used for the "inner loop" training of the low-level policy network. In each interaction with the digital twin model (at a single time point), the agent provides immediate feedback and learns based on Rt, rapidly exploring and optimizing fine-grained control actions within a single time point. This is the direct input to standard reinforcement learning algorithms such as PPO.

[0152] The composite reward signal (defined in S4) is used for the "outer loop" evaluation and calibration of the entire policy network (including high-level and low-level systems). Once a complete candidate power adjustment scheme (i.e., a complete processing cycle) is generated, the system initiates the Monte Carlo evaluation in S3 to obtain a comprehensive evaluation of the scheme throughout the entire process and after risk adjustment, forming the composite reward signal in S4. This signal is richer and more macroscopic, used to perform a round of overall performance updates on the agent's policy to ensure that the policy converges towards the global optimum. In short, Rt guides the agent "how to take each step well," while the composite reward signal tells the agent "whether this path is good or not, and how to choose a better path in the future."

[0153] It also includes step S5: after the strategy network converges, by performing perturbation analysis on the power output of each heating unit in the candidate power adjustment scheme at different time points, the sensitivity of the expected time integral performance value to the change of each power output is calculated, and the set of key control points that have the greatest impact on the overall process performance is identified based on the sensitivity.

[0154] The core objective of this embodiment is to conduct a comprehensive performance evaluation of the candidate power adjustment schemes generated by the hierarchical reinforcement learning agent in Embodiment 2, taking into account the uncertainties of real-world production. Subsequently, based on the evaluation results, an information-rich composite reward signal is generated to guide further optimization by the agent. Finally, after the agent's policy converges, sensitivity analysis is performed on the optimal scheme to identify the key control points that have the greatest impact on the process results.

[0155] S3: Calculation of time integral performance value based on Monte Carlo simulation: This step aims to quantify the overall performance of the power regulation scheme generated in Example 2 over the entire processing cycle, rather than relying solely on immediate rewards in the reinforcement learning environment.

[0156] S31, Construction of instantaneous performance health;

[0157] The instantaneous performance health P(t) is calculated to evaluate the material state at any time point t during simulation. Instantaneous performance health integrates three dimensions: product quality, uniformity, and energy consumption. The preferred calculation formula is: P(t) = wu × U(t) + wy × Y(t) – we × E(t);

[0158] The preferred values ​​and calculation methods for each component and its weighting coefficient are as follows:

[0159] Temperature uniformity index U(t): U(t) = 1 - (Tmax(t) - Tmin(t)) / Tavg(t). This index reflects the temperature uniformity of the material surface; the closer to 1, the better. The weight wu is set to 2.0.

[0160] Predicted product yield Y(t): Y(t) = (1 - (σmax(t) / σthreshold)) 4 )×(1-(dmax(t) / dthreshold) 4 This function, based on the time-varying evolution field data generated in Example 1, predicts the potential failure rate at time t caused by stress (σmax) and warpage (dmax). The weight wy is set to 10.0 to highlight its decisive impact on the final product quality.

[0161] Instantaneous energy consumption index E(t): E(t) is the total power output (normalized value) of the 16 heating units at time t, which is directly related to production costs. The weight we is set to 0.5.

[0162] S32. Define a material disturbance scenario library; to make the evaluation results more closely reflect the batch-to-batch differences that exist in actual production, the system presets a scenario library containing various material disturbances and their probabilities of occurrence, as shown in Table 7.

[0163] Table 7: Material Disturbance Scenario Library (Example)

[0164]

[0165] S33. Monte Carlo Simulation and Calculation of Expected Time Integral Performance Value: The system applies the power regulation scheme generated in Example 2 and performs accelerated time-domain simulation on the digital twin model of Example 1. This process is not a single simulation, but rather multiple (set to 1000 times) weighted random samplings based on Table 8 to simulate different virtual material processing paths.

[0166] For each path, calculate its time integral performance value Vpath: Vpath=∫P(t)dt (integration interval is 0 to 30 seconds), as shown in Table 8.

[0167] Table 8: Monte Carlo Simulation Path Sampling and Performance Value Calculation (Excerpt from the first 4 times)

[0168]

[0169] After completing 1000 simulations, the performance values ​​of all paths are weighted and averaged according to the probability of occurrence of each scenario to obtain the final risk-adjusted expected time integral performance value E[V].

[0170] Assume the average performance values ​​for each scenario are: VC0 = 285.8, VC1 = 271.5, VC2 = 250.9, VC3 = 278.0. Then E[V] = 285.8 × 0.70 + 271.5 × 0.15 + 250.9 × 0.10 + 278.0 × 0.05 = 279.78. In the Monte Carlo simulation of S33, a total of 1000 independent accelerated time-domain simulations were performed. Each simulation randomly selected a scenario (C0, C1, C2, or C3) based on the probability in Table 7 (Material Disturbance Scenario Library). Therefore, these 1000 simulations can be divided into four groups: Group 1: All simulations that selected scenario C0 (baseline scenario). Based on a 70% probability, this group contains approximately 700 simulations. Group 2: All simulations that selected scenario C1 (high doping concentration). Based on a 15% probability, this group has approximately 150 simulations. The third group: all simulations drawn for scenario C2 (initial microcrack size is relatively large). Based on a 10% probability, this group has approximately 100 simulations. The fourth group: all simulations drawn for scenario C3 (wafer thickness is slightly thinner). Based on a 5% probability, this group has approximately 50 simulations. VC0 to VC3 are the arithmetic mean performance values ​​of these four simulation groups.

[0171] S4: Generation of composite reward signal: This step transforms the macroscopic performance evaluation value calculated in S3 into a composite reward signal that can more precisely guide the agent's policy network update.

[0172] S41. Generation of Process Penalty Values: During the simulation in S3, the system monitors key process indicators in real time. It is assumed that the process window requires the temperature uniformity index U(t) to never fall below 0.98.

[0173] Data Example: In simulation sequence 2 (scenario C2) in Table 8, at t=18.2 seconds, the system detects U(18.2)=0.975. At this time, an instantaneous process penalty value is generated: Penaltyproc=k×(Uthreshold-U(t))=100×(0.98-0.975)=-0.5. (k is the penalty coefficient) This instantaneous negative signal -0.5 will be recorded and used together with the final terminal reward for policy updates. Wherein, Uthreshold represents the preset lower limit of the process window for the temperature uniformity index;

[0174] S42. Generation of Multidimensional Terminal Reward Vector: After the simulation, the final expected time integral performance value E[V] is decomposed into a multidimensional terminal reward vector that better reflects the specific contribution. The cumulative yield component, denoted as Ryield, is the expected value of the weighted integral of the yield prediction value Y(t). For example, Ryield = 255.4 is calculated. The cost penalty component, denoted as Cenergy, is the expected value of the weighted integral of the instantaneous energy consumption E(t). For example, Cenergy = -24.3 is calculated. The resulting multidimensional terminal reward vector is: Rterminal = [255.4, -24.3].

[0175] S43. Combination of Composite Reward Signals: The process penalty value is combined with the multi-dimensional terminal reward vector to form the final composite reward signal used to update the agent's policy network. Composite reward signal = {process penalty value: [-0.5], Rterminal: [255.4, -24.3]} This composite signal conveys richer information to the agent: "The final yield of this solution is good, and the energy consumption is well controlled, but uneven temperature occurred in the middle of the process, which needs to be improved."

[0176] S5: Identification of the set of key control points: After the policy network converges after multiple rounds of training, post-processing analysis is performed on the current optimal power regulation scheme to mine process knowledge.

[0177] S51. Disturbance Analysis and Sensitivity Calculation: Select several representative control points in the optimal power regulation scheme for disturbance analysis. Apply a small disturbance (e.g., increase the absolute power by 1%) to the power value of each point, and then re-execute the full set of Monte Carlo simulations in S3 to calculate the new expected time integral performance value E[V'].

[0178] The formula for calculating sensitivity S is: S = (E[V'] - E[V]) / ΔPower; where ΔPower represents the amount of power disturbance applied by the heating unit at the control point (e.g., an increase of 1% in absolute power). As shown in Table 9.

[0179] Table 9: Identification of Key Control Points (Example of Sensitivity Analysis)

[0180]

[0181] Through this analysis, the system automatically identified (heating unit (2,3), time 5.0 seconds) as the primary critical control point with the greatest impact on the overall process performance. This data-driven set of critical control points provides process engineers with valuable, quantitative decision-making support for optimizing SOPs and setting critical monitoring windows.

[0182] Figure 1The isometric view on the left side of the figure illustrates the physical scenario in which this invention is applied: a multi-zone heating furnace, a conveyor belt transporting materials to be processed (such as silicon wafers), and a multimodal sensing device in front of the furnace inlet. This part visually presents the industrial environment that this invention transforms and optimizes. The technical roadmap on the right side of the figure breaks down the intelligent control closed loop of this invention in detail, with each flowchart corresponding to a core step: The first flowchart, "Online Sensing and Digital Twin Construction," corresponds to step S1. It describes the process by which the system acquires the initial state data of the material through the multimodal sensing device, constructs its personalized digital twin model, and generates time-varying evolution field data. The second flowchart, "Hierarchical Reinforcement Learning and Scheme Generation," corresponds to step S2. It indicates that, under the hierarchical reinforcement learning framework, the agent generates partitioned candidate power adjustment schemes aimed at avoiding risks based on the state of the digital twin model. The third flowchart, "Monte Carlo Simulation and Performance Evaluation," corresponds to step S3. It describes the process by which the system performs accelerated time-domain simulation of candidate schemes and calculates their expected time integral performance values ​​under various perturbation scenarios based on the Monte Carlo method. The fourth flowchart, "Compound Reward Feedback and Policy Optimization," corresponds to step S4. It demonstrates the closed-loop optimization process by which the system generates a compound reward signal based on the evaluation results and uses iterative updates to the agent's policy network, ultimately achieving intelligent adjustment and detection of the thermostat's power.

[0183] The core technical principle of this invention lies in constructing a full-process intelligent control system that goes from personalized physical insight to hierarchical intelligent decision-making, and then to risk quantification assessment and closed-loop feedback optimization.

[0184] In terms of technical principles, this invention first constructs a unique, high-precision digital twin model for each piece of material entering the furnace, encompassing the distribution of microscopic defects and their interactions, through online multimodal perception and deep learning models (CNN, GNN). The core innovation of this model lies in its ability to proactively predict the time-varying evolution path of thermally induced failure risks such as stress and warping caused by initial defects during the heating process, utilizing Physical Information Neural Network (PINN) and process knowledge graph.

[0185] Subsequently, in the control strategy generation stage, this invention innovatively employs a hierarchical reinforcement learning (HRL) framework for task decomposition. The high-level policy network, based on a global assessment of material defects, macroscopically selects the optimal process curve prototype ("stepped stress release curve") to achieve strategic risk avoidance. The low-level policy network, guided by this macroscopic instruction, interacts frequently with the digital twin model to execute specific, zoned power adjustment actions, closely tracking the target temperature while performing tactical, refined local risk avoidance.

[0186] In the strategy evaluation and optimization phase, this invention introduces a risk assessment mechanism based on the Monte Carlo method. Instead of evaluating solutions under a single ideal condition, it performs thousands of random simulations on candidate solutions within a perturbation scenario library containing variations between batches of various materials, thereby calculating the risk-adjusted expected time integral performance value. Furthermore, this invention designs a composite reward signal that includes process penalties and multi-dimensional terminal rewards. This signal not only evaluates the final result but also focuses on the stability of the process, providing the agent with richer and more precise optimization directions, guiding the policy network towards global optimum and high robustness.

[0187] Example 4:

[0188] A temperature controller power regulation and detection system based on AI intelligence, please refer to... Figure 2 ,include:

[0189] The time-varying evolution field generation module is used to acquire the time-varying evolution field data of the material and thermodynamics of the digital twin model of the material to be processed within a preset processing cycle;

[0190] The hierarchical policy generation module is used to obtain candidate power adjustment schemes generated by the agent for the current thermodynamic state of the digital twin model in a hierarchical reinforcement learning framework.

[0191] The simulation and performance evaluation module is used to perform accelerated time-domain simulations of digital twin models that apply candidate power regulation schemes based on time-varying evolution field data, thereby calculating the time integral performance value of the candidate power regulation schemes.

[0192] The network update module is used to generate a reward signal based on the time integral performance value, and to update the agent's policy network using the reward signal.

[0193] In this embodiment, the system constructs a dedicated failure prediction model for each material through the time-varying evolution field generation module, achieving forward-looking risk insight. Based on this, the hierarchical strategy generation module can formulate personalized optimal power adjustment schemes, mitigating thermally induced failures caused by microscopic defects in materials at the source. Combining Monte Carlo risk assessment from the simulation and performance evaluation module with closed-loop optimization based on composite reward signals from the network update module ensures that the final generated control strategy has high robustness in the face of real production fluctuations.

[0194] It should be noted that all calculation formulas in this application employ regression analysis, including but not limited to machine learning algorithms, to deeply analyze the collected parameters and identify their natural trends and interrelationships. Specialized software, such as Python's Scikit-learn library or the R language, is used to automatically generate mathematical models that match the data. Then, cross-validation and other methods are used to objectively evaluate the model performance, and continuous feedback and optimization are combined to ensure that the created formulas truly reflect the inherent laws of the data, thereby guaranteeing their effectiveness and accuracy. In all calculation formulas in this application, the parameters in each formula undergo dimensionless processing within a consistent range to ensure that different physical quantities are compared on the same scale; dimensionless processing techniques include, but are not limited to, min-max-normalization and Z-score standardization.

[0195] The algorithm of this invention is implemented as a Python script. Before executing the core logic, the program first executes a data loading module (e.g., using the widely used pandas library in Python), configured to read the aforementioned spreadsheet file and load its contents into the program's working memory (e.g., a DataFrame data structure). Subsequent algorithm steps will directly query and retrieve the required configuration parameters from this in-memory data structure.

[0196] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. The AI intelligent-based temperature controller power regulation and detection method is applied to a photovoltaic production heating furnace body, characterized in that, The specific steps include: S1: Obtain the time-varying evolution field data of the material and thermodynamics of the digital twin model of the material to be processed within the preset processing period, and the time-varying evolution field data is configured to represent the law of the physical and chemical properties of the material at each position evolving with time during heating; online acquisition of multi-modal perception data, at the moment when the material enters the furnace body through the conveying belt, multi-modal perception data is acquired; multi-modal perception data includes visual image data and hyperspectral image data; S2: In a hierarchical reinforcement learning framework, obtain the candidate power adjustment scheme generated by the agent for the current thermodynamic state of the digital twin model; the agent is a hierarchical reinforcement learning agent, which includes: a high-level policy network configured to select a macro process curve prototype as a target from a preset process curve prototype library according to the initial characteristics of the material to be processed and the global process target; a low-level policy network configured to receive the macro process curve prototype as an instruction and perform fine parameter search on the power and time sequence matrix within the temperature and time envelope defined by the macro process curve prototype to generate a candidate power adjustment scheme; S3: Based on the time-varying evolution field data, perform time-domain simulation on the digital twin model to which the candidate power adjustment scheme is applied, so as to calculate the time integral performance value of the candidate power adjustment scheme within the preset processing period; The step of calculating the time integral performance value in S3 includes: integrating the preset instantaneous performance health degree representing the comprehensive performance of the digital twin model at any time point within the time integral interval defined by the preset processing period; the instantaneous performance health degree at least takes the temperature uniformity index and product yield prediction value of the digital twin model at the corresponding time point as positive input, and takes the instantaneous energy consumption index associated with the candidate power adjustment scheme as negative input; S4: Generate a reward signal according to the time integral performance value, and the reward signal is configured to represent the overall performance of the candidate power adjustment scheme and is used to update the policy network of the agent.

2. The AI intelligence-based temperature controller power adjustment and detection method according to claim 1, characterized in that: S1 includes: The step of generating time-varying evolution field data through a physical information neural network model includes: performing pixel-level segmentation and feature extraction on hyperspectral and machine vision images in online multi-modal perception data through a convolutional neural network, and outputting a classification confidence vector representing the defect type of each micro area; constructing a spatial adjacency graph of the material surface, wherein the nodes represent each micro area and the edges represent the thermodynamic transfer relationship therebetween; Using the classification confidence vector as the initial feature of the node, and using the evolution mechanism extracted from the process knowledge graph about how different defects affect each other in the thermal field, the constraint of information propagation is constructed to perform iterative reasoning on the spatial adjacency graph, and an interactive evolution parameter set representing the warping and stress evolution rate of each micro area during the heating process is output; the interactive evolution parameter set is used as the dynamic boundary condition of the physical information neural network model to generate the time-varying evolution field data. 3.The AI intelligence-based temperature controller power adjustment and detection method of claim 2, wherein: The feature extraction step includes: constructing the average gray value, texture complexity index, spectral impurity index and spectral slope of each micro area, and combining them in a predetermined order to generate a fusion feature vector; The average gray value is calculated by: obtaining the gray value of each pixel in the micro area, and performing an arithmetic average operation on all obtained gray values, and the obtained average value is the average gray value; The texture complexity index is calculated by: taking the center pixel of the micro area as the reference pixel, and obtaining a plurality of neighborhood pixels around the reference pixel; comparing the gray value of each neighborhood pixel with the gray value of the reference pixel, and when the gray value of the neighborhood pixel is greater than or equal to the gray value of the reference pixel, a binary value "1" is generated, otherwise a binary value "0" is generated; combining all generated binary values in a predetermined order to form a binary sequence; converting the binary sequence into a corresponding decimal value, which is the texture complexity index; The spectral impurity index is calculated by: obtaining the spectral reflectance intensity value at the first preset wavelength as the first intensity value, and obtaining the spectral reflectance intensity value at the second preset wavelength as the second intensity value; dividing the first intensity value by the second intensity value, and the quotient is the spectral impurity index; The spectral slope is calculated by: obtaining the spectral reflectance intensity value at the third preset wavelength as the third intensity value, and obtaining the spectral reflectance intensity value at the fourth preset wavelength as the fourth intensity value; calculating the difference between the third intensity value and the fourth intensity value to obtain an intensity difference; calculating the difference between the third preset wavelength and the fourth preset wavelength to obtain a wavelength difference; dividing the intensity difference by the wavelength difference, and the quotient is the spectral slope.

4. The AI intelligence-based temperature controller power adjustment and detection method of claim 3, wherein: The step of inputting the fusion feature vector into the convolutional neural network model and outputting the final classification confidence vector includes: S10001, for each preset defect type, obtaining a weight vector and a bias term corresponding to the current defect type from the convolutional neural network model; S10002, performing weighted sum operation on the fusion feature vector and the weight vector to obtain a weighted sum; S10003, adding the weighted sum and the bias term, and the sum is the original score for the current defect type; S10004, repeating steps S10001 to S10003 to generate an original score vector containing the original scores of each defect type; S1232, for each original score in the original score vector, taking the natural logarithm base e as the base and the current original score as the exponent to perform exponential operation to obtain a corresponding exponential score; summing all generated exponential scores to obtain a total value; dividing each exponential score by the total value to obtain the final classification confidence vector for the corresponding defect type; based on the final classification confidence vector of the corresponding defect type, inferring to obtain the warping evolution rate and the stress evolution rate.

5. The AI intelligence-based temperature controller power adjustment and detection method of claim 4, wherein: Based on the warping evolution rate and the stress evolution rate, time-varying evolution field data is generated, and the specific steps include: S141, acquire a time and temperature sequence describing the material heating process, the sequence containing a series of discrete time points and corresponding measured or preset temperature values at each time point; S142, set the initial state of all micro-regions at time point t=0, where the initial lateral stress is 0 MPa and the initial vertical displacement is 0 microns, and obtain the initial temperature from the time and temperature sequence of step S141, denoted as Tinitial; S143, define a calculation method for calculating the current temperature at any target time point t: S1431, in the time and temperature sequence acquired in step S141, find the two consecutive time points containing the target time point t and their corresponding temperature values; S1432, based on the two time points and their temperature values, calculate the accurate current temperature at the target time point t using linear interpolation; S144, the step of calculating the vertical displacement at any time point: S1441, acquire the warping evolution rate of the specified micro-region; S1442, subtract the initial temperature Tinitial set in step S142 from the current temperature T0 calculated in step S1432 to obtain a temperature difference value; S1443, multiply the warping evolution rate acquired in step S1441, the temperature difference value obtained in step S1442, and the target time t to obtain the total vertical displacement at time point t; S145, the step of calculating the lateral stress at any time point: S1451, acquire the stress evolution rate of the specified micro-region; S1452, multiply the stress evolution rate acquired in step S1451, the temperature difference value obtained in step S1442, and the target time t to obtain the total lateral stress at time point t.

6. The AI intelligence-based temperature controller power adjustment and detection method of claim 1, wherein: The integral operation is a random integral process based on the Monte Carlo method, which specifically includes: generating a set of virtual material processing paths by multiple weighted random sampling from a preset material disturbance scenario library containing differences between various batches of materials to be processed and their occurrence probabilities; for each virtual material processing path, independently calculating its corresponding time integral performance value; and weighting and averaging all time integral performance values calculated for the virtual material processing paths according to their corresponding path occurrence probabilities to obtain the final risk-adjusted expected time integral performance value.

7. The AI intelligence-based temperature controller power adjustment and detection method of claim 6, wherein: In S4, the step of generating a reward signal, specifically including: monitoring the temperature uniformity index in real time during the calculation and simulation of the time integral performance value; generating a process penalty value proportional to the index deviation value when the temperature uniformity index is lower than the lower limit of the preset process window; after the end of the simulation process, decomposing the final expected time integral performance value into a cumulative yield component representing the product quality of the whole process and a cost penalty component representing the total energy consumption, forming a multi-dimensional terminal reward vector; combining the process penalty value with the multi-dimensional terminal reward vector to form a composite reward signal for updating the policy network.

8. The AI intelligence-based temperature controller power adjustment and detection method of claim 7, wherein: The method further comprises the following step S5: after the policy network converges, performing perturbation analysis on the power output of each heating unit in the candidate power adjustment scheme at different time points, calculating the sensitivity of the expected time integral performance value to the change of each power output, and identifying the key regulation point set with the greatest impact on the overall process performance according to the sensitivity.

9. An AI intelligent-based temperature controller power regulation and detection system using the AI intelligent-based temperature controller power regulation and detection method of any one of claims 1-8. The method comprises the following steps: a time-varying evolution field generation module, configured to obtain time-varying evolution field data of materials and thermodynamics of a digital twin model of a material to be processed within a preset processing period; a hierarchical strategy generation module, configured to obtain a candidate power adjustment scheme generated by an agent for a current thermodynamic state of the digital twin model in a hierarchical reinforcement learning framework; a simulation and performance evaluation module, configured to perform accelerated time-domain simulation on the digital twin model to which the candidate power adjustment scheme is applied based on the time-varying evolution field data, so as to calculate a time integral performance value of the candidate power adjustment scheme; a network updating module, configured to generate a reward signal according to the time integral performance value, and update the policy network of the agent by using the reward signal.

Citation Information

Patent Citations

  • Intelligent monitoring method and system for whole hot working process and storage medium

    CN120746406A

  • Metal plate flexible production line real-time scheduling method and system based on digital twinning

    CN120975517A