Smt production line active energy carbon perception production scheduling method and system based on primitive reinforcement learning

CN122713664APending Publication Date: 2026-09-08FUJIAN FUYAO UNIVERSITY OF SCIENCE & TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610870418.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

[0015]本发明的目的旨在解决现有 SMT 产线能源 - 碳协同调度中的耦合建模难、碳追溯弱、实时性差、安全约束不足、多智能体失衡等技术瓶颈,结合半马尔可夫决策过程、图神经网络、元学习与 Transformer 注意力机制,提供一种基于图元强化学习的SMT产线主动式能源碳感知生产调度方法及系统

Benefits of technology

[0049] This invention constructs a closed-loop management and control system covering the entire process, integrating SMDP, graph neural networks, meta-learning, and Transformer technologies to solve the core challenge of energy-carbon coordinated scheduling in SMT production lines, achieving significant results in terms of technology, economy, environment, and application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122713664A_ABST
    Figure CN122713664A_ABST
Patent Text Reader

Abstract

This invention relates to a proactive energy and carbon-sensing production scheduling method and system for SMT production lines based on primitive reinforcement learning, belonging to the field of intelligent manufacturing and low-carbon management technology. Addressing the challenges of coupling discrete scheduling with the continuous thermodynamics of reflow soldering in SMT production lines, large dynamic electricity price / carbon factor disturbances, and imbalanced credit allocation among multiple agents, this invention constructs a dynamic coupling model of production-thermal-electricity-carbon flow, reconstructing scheduling as a semi-Markov decision process, and proposes a GMeta-MATD3 primitive reinforcement learning architecture. The spatiotemporal topology of production and power grid is extracted through a multi-flow graph convolutional network, and a meta-learning module is embedded to infer the physical context and predict peer actions. A Transformer self-attention evaluator is used to solve credit allocation, achieving thermal recipe clustering, proactive waiting, and peak avoidance. This invention balances production efficiency, economy, and low carbon emissions, is suitable for highly mixed, small-batch SMT scenarios, and can be widely applied to energy management and production scheduling systems in SMT workshops for consumer electronics, automotive electronics, and semiconductor packaging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of intelligent manufacturing, low-carbon electronic manufacturing, energy-carbon collaborative management and control, multi-agent reinforcement learning, and flexible interactive cross-technology of power distribution networks. Specifically, it relates to an active energy carbon sensing production scheduling method and system for SMT production lines based on primitive reinforcement learning. Background Technology

[0002] Driven by the global "dual-carbon" strategy and the upgrading of the electronics manufacturing industry, SMT production lines, as core high-energy-consuming processes in electronics manufacturing, account for over 70% of the total energy consumption of thermal equipment such as reflow ovens, making them a key control unit for industrial carbon emission reduction. Modern SMT workshops have formed a deeply coupled cyber-physical system of "production-energy-grid," integrating processes such as solder paste printing, component mounting, multi-area reflow soldering, and AOI inspection, and supported by external conditions such as microgrids, time-of-use electricity pricing, and dynamic carbon emission factors, playing a core role in promoting the green transformation of electronics manufacturing. However, current low-carbon scheduling and energy management of SMT production lines face four major technical bottlenecks, and traditional methods are no longer suitable for complex operating conditions:

[0003] 1. The lack of a discrete-continuous coupled scheduling mechanism leads to low energy utilization efficiency.

[0004] SMT production lines exhibit a heterogeneous coupling between discrete production logic and continuous thermodynamics: production scheduling is driven by discrete events (work order assignment, equipment routing, batch switching), while reflow ovens are continuous thermal processes (multi-zone heating / cooling, asymmetric thermal inertia, and sequence-dependent time settings).

[0005] Traditional scheduling only optimizes time indicators such as manufacturing cycle and delay rate, treats reflow oven setting time as a static constant, ignores the dynamic energy consumption of temperature transition in multiple regions, and frequent recipe switching leads to high transition energy consumption and serious waste of thermal inertia; without establishing a coordination mechanism between production sequence and thermal parameters, "slow thermal response" and "fast production scheduling" are mismatched, resulting in low renewable energy absorption rate, large peak-valley difference in grid load, and extremely low energy resource allocation efficiency.

[0006] 2. The power-carbon flow coupling management system is imperfect, making carbon emission traceability and control difficult.

[0007] Carbon emissions from SMT production lines are deeply intertwined with power flow and thermal energy consumption. The distribution of carbon emissions is influenced by multiple factors, including reflow soldering power, power grid topology, carbon emission factors, and work order sequences. Current technologies can only provide macroscopic statistics on total carbon emissions from the production line. They have not established a dynamic coupling model of production-thermal-power-carbon flow, making it impossible to quantify the carbon footprint of each work order, each piece of equipment, and each time period. This makes it difficult to achieve accurate carbon emission traceability and dynamic management.

[0008] Meanwhile, existing research is based on the assumption of static electricity price / carbon factor, and does not take into account the real-time fluctuation of microgrid carbon emission factor and the impact of ambient temperature on heat dissipation rate. Carbon emission reduction measures lack specificity and cannot achieve efficient utilization of carbon quotas while ensuring production delivery. It is difficult to balance production efficiency, energy costs and carbon emission reduction targets.

[0009] 3. Insufficient real-time performance and security of scheduling under dynamic operating conditions.

[0010] SMT production line operations face multiple dynamic disturbances: emergency work order insertion, dynamic recipe changes, time-of-use electricity price fluctuations, real-time changes in carbon emission factors, and reflow oven heat dissipation rate drifting with the environment, placing extremely high demands on the real-time performance of scheduling decisions. Traditional mixed-integer programming and metaheuristic algorithms have high computational complexity and are time-consuming, failing to meet real-time scheduling requirements.

[0011] In recent years, multi-agent reinforcement learning (MARL) has been gradually applied to production line scheduling, but standard MARL has three major drawbacks: ① Environmental non-stationarity leads to training oscillations and policy collapse; ② Imbalance in credit allocation among multiple agents, with aggressive heating of local equipment triggering grid peak violations; ③ The lack of integration of thermodynamic and physical constraints results in a lack of physical interpretability in decision-making, which can easily lead to risks to grid safety and production delivery.

[0012] 4. Existing technological research has limitations and lacks a comprehensive collaborative optimization solution.

[0013] Current research focuses primarily on optimizing single aspects: some optimize only production scheduling without incorporating energy-carbon management; some optimize only energy consumption without coupling it with production constraints; some employ basic reinforcement learning but fail to address challenges related to topological perception, non-stationarity, and credit allocation; and no comprehensive energy-carbon coordinated scheduling system has been established, encompassing "data perception - coupled modeling - intelligent decision-making - execution feedback."

[0014] In summary, existing technologies cannot solve the four core problems of discrete-continuous coupling, precise carbon flow control, dynamic safety scheduling, and multi-agent collaboration in SMT production lines. There is an urgent need to develop an intelligent control method that takes into account production scheduling, thermal coupling, energy-carbon synergy, and grid security, and to build a systematic proactive energy-carbon sensing and scheduling system for SMT production lines. Summary of the Invention

[0015] The purpose of this invention is to address the technical bottlenecks in existing SMT production line energy-carbon collaborative scheduling, such as difficulty in coupled modeling, weak carbon traceability, poor real-time performance, insufficient safety constraints, and imbalance among multiple agents. By combining semi-Markov decision processes, graph neural networks, meta-learning, and Transformer attention mechanisms, this invention provides an active energy and carbon sensing production scheduling method and system for SMT production lines based on graph primitive reinforcement learning.

[0016] The core objective of this invention is:

[0017] 1. Construct a dynamic coupling model of production-thermal-electricity-carbon flow to quantify the correlation between thermodynamic transition, work order sequence, and power grid load in multiple regions of reflow soldering, so as to achieve accurate traceability and dynamic calculation of carbon emissions for each batch, each piece of equipment, and each time period;

[0018] 2. Reconstruct the SMT scheduling into a semi-Markov decision process (SMDP) to adapt to variable thermal transition time and solve the problem of deep coupling between discrete production logic and continuous thermodynamics;

[0019] 3. Propose the GMeta-MATD3 primitive reinforcement learning architecture, which integrates multi-flow graph convolutional network (MF-GCN), meta-learning module and Transformer self-attention evaluator to solve the problems of environmental non-stationarity and multi-agent credit allocation;

[0020] 4. Establish a multi-objective optimization system for production, energy, and carbon to achieve the shortest manufacturing cycle, the lowest energy cost, and the smallest carbon emissions, while simultaneously meeting grid peak load, delivery time, and thermal process constraints.

[0021] 5. Construct a closed-loop control system covering the entire process of "perception-modeling-decision-execution-feedback" to achieve real-time acquisition of multi-source data, dynamic updating of coupled models, intelligent generation of scheduling decisions, precise execution of control commands, and online optimization of strategies.

[0022] To achieve the above objectives, the technical solution of this invention is: a proactive energy and carbon sensing production scheduling method for SMT production lines based on primitive reinforcement learning, comprising:

[0023] Collect and preprocess production data from surface mount technology production lines and microgrid data.

[0024] Construct a mixed integer nonlinear programming model that couples production, thermal engineering, electricity, and carbon flow, and define the optimization objectives of manufacturing cycle, energy cost, and carbon emissions, as well as constraints on production, thermal engineering, and power grid.

[0025] The scheduling problem is refactored into a semi-Markov decision process driven by discrete events, defining the state space, action space, reward function, and variable time transition rules.

[0026] A multi-flow graph convolutional network is constructed to extract the spatiotemporal topology embedding of the production-power grid, and the temperature of multiple regions of the reflow oven, conveyor belt speed, thermal formula, buffer queue, dynamic electricity price, carbon emission factor and power margin are represented as graph node features;

[0027] Embed a meta-learning module in a multi-agent system to infer the physical context and predict the actions of fellow agents by using historical trajectories;

[0028] A centralized evaluator based on the Transformer self-attention mechanism is adopted to dynamically decompose the joint value and calculate the individual advantage function;

[0029] Based on the GMeta-MATD3 architecture, a multi-agent strategy is trained under the centralized training and distributed execution paradigm to generate scheduling decisions for work order assignment, equipment routing, and hot recipe selection.

[0030] Execute scheduling decisions and issue control commands, collect execution feedback data to update primitive reinforcement learning models and multi-objective scheduling strategies.

[0031] Furthermore, the multi-flow graph convolutional network models the SMT production line as a directed multi-relationship graph, with nodes including SMT equipment, buffers, and microgrid nodes, and edges including production material flow edges and grid power flow edges; the hidden representation update formula for the (l+1)th layer of the multi-flow graph convolutional network is:

[0032]

[0033] in, To incorporate the normalized adjacency matrix of self-loops, To generate the edge matrix, For the power grid side matrix, It is the identity matrix. For degree matrix, Let l be the trainable weight matrix of the l-th layer. For activation function, This represents the embedding representation of the l-th layer node.

[0034] Furthermore, the state space of the semi-Markov decision process includes a job feature matrix, a machine feature matrix, and a dynamic environment matrix; the action space is a set of macro actions including work order assignment actions, equipment routing actions, and thermal recipe control actions; and the reward function is a weighted negative penalty for manufacturing cycle, energy cost, carbon emissions, delay, and grid peak violation.

[0035] Furthermore, the scheduling decision includes clustering work orders with the same or similar thermal curves, so that the work orders pass through the reflow oven continuously with the minimum feed gap to reduce thermal transition energy consumption; when the peak of time-of-use electricity price or high carbon emission window is predicted, the agent performs an active waiting action to transfer the high energy consumption thermal transition to the low time-of-use electricity price period or low carbon emission period.

[0036] Furthermore, the meta-learning module includes a historical network, a state network, and a prediction network; the historical network encodes features of past trajectories to generate historical embedding representations, the state network infers the current physical context based on recent trajectories, and the prediction network integrates local observations, historical embeddings, and state embeddings to deterministically estimate the next action of the companion agent.

[0037] Furthermore, the Transformer-enhanced centralized evaluator concatenates each agent's local observations, current actions, historical embeddings, state embeddings, and predicted peer actions into a feature vector, and calculates the marginal contribution of each agent by stacking self-attention layers; the temporal difference objective in the target network uses deterministic actions predicted by the meta-learning module to replace randomly sampled actions in order to reduce variance.

[0038] Furthermore, the thermodynamic transition time of the multi-zone reflow oven is the maximum value of the transition time of each independent temperature control zone, and the specific calculation formula is as follows:

[0039]

[0040] in, The transition time of the z-th temperature control zone when the reflow oven l switches from the i-th operation to the j-th operation. Let z be the heating rate of the z-th region. Let z be the cooling rate of the z-th region. The target temperature required for operation i in the z-th zone of reflow oven l. The target temperature required for operation j in the z-th zone of reflow oven l.

[0041] Furthermore, the reward function includes an energy cost penalty, a carbon emission penalty, a delay penalty, and a grid peak penalty; the energy cost and carbon emission are obtained by integrating the operating power, transition power, and standby power over the scheduling period.

[0042] Furthermore, the decision-making time of the semi-Markov decision process is triggered by discrete events, including the completion of machine tasks and the transition to idle time, the arrival of urgent orders, and the reaching of the target set value in the slowest temperature zone of the reflow oven.

[0043] This invention also provides an active energy and carbon sensing production scheduling system for SMT production lines based on primitive reinforcement learning, comprising:

[0044] The multi-source information sensing module is configured to collect production data, thermal data, microgrid data, and external environmental data from the SMT production line, and perform data preprocessing.

[0045] The production-thermal-electric carbon coupling modeling module is configured to construct a multi-objective scheduling model that includes discrete production logic and continuous thermodynamic constraints, and to reconstruct the scheduling problem as a semi-Markov decision process.

[0046] The primitive reinforcement learning decision module is configured to extract spatiotemporal topological features based on a multi-flow graph convolutional network, infer the physical context and predict peer actions through the meta-learning module, and use a Transformer-enhanced centralized evaluator for multi-agent credit allocation and policy optimization.

[0047] The execution and feedback closed-loop module is configured to send scheduling decisions to the SMT production line equipment controller, monitor the operating status, and provide feedback to optimize the scheduling strategy.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] This invention constructs a closed-loop management and control system covering the entire process, integrating SMDP, graph neural networks, meta-learning, and Transformer technologies to solve the core challenge of energy-carbon coordinated scheduling in SMT production lines, achieving significant results in terms of technology, economy, environment, and application.

[0050] 1. Technical effects: Enables coupled scheduling and security management, improving the level of intelligence.

[0051] 1.1 For the first time, SMDP was used to achieve deep coupling between discrete SMT scheduling and continuous thermodynamics of reflow soldering, reducing thermal transition energy consumption from 12.5 kWh / batch to 3.2 kWh / batch;

[0052] 1.2 The GMeta-MATD3 architecture solves the problems of environmental non-stationarity and credit allocation, improving convergence speed by 40% and eliminating policy crashes;

[0053] 1.3 Multi-flow graph convolutional networks achieve topology awareness, reducing the peak violation rate of the power grid to 0 and the decision response time to ≤500ms;

[0054] 1.4 The closed-loop system enables online updates of the model and strategy, adapting to dynamic electricity prices, carbon factors, and thermal parameter drift, significantly improving robustness.

[0055] 2. Economic benefits: Reduces energy costs and improves production efficiency.

[0056] 2.1 Active energy arbitrage reduced energy costs by 34.5%, with a total energy cost of only 1020.5 yuan for 100 batches;

[0057] 2.2 Formula clustering optimization of the production sequence shortened the manufacturing cycle to 13.1 hours and reduced the delay rate to 1.5%;

[0058] 2.3 The load peak-valley difference is reduced by 34.5%, reducing grid peak penalty and improving the economic efficiency of the production line.

[0059] 3. Environmental impact: Significantly reduces carbon emissions and promotes low-carbon manufacturing.

[0060] 3.1 Total carbon emissions were reduced by 36.8%, with carbon emissions from 100 batches amounting to only 810.2 kg;

[0061] 3.2 The daily carbon intensity remained stable below 0.55 kg CO2 / kWh, meeting the requirements for low-carbon management;

[0062] 3.2 Proactive peak avoidance and thermal stepping enable green production and grid-friendly interaction.

[0063] 4. Application effect: Highly scalable and with high engineering implementation value.

[0064] 4.1 It is compatible with IEEE 33 / 141 node power distribution systems, completes 500 batches of dispatch within 500 milliseconds, and can be deployed at scale;

[0065] 4.2 Modular design, expandable to high-energy-consuming manufacturing scenarios such as semiconductor packaging and power battery assembly lines;

[0066] 4.3 Seamlessly integrates with existing SMT MES and EMS systems, requiring no large-scale equipment modifications and resulting in low application costs. Attached Figure Description

[0067] Figure 1 This is a flowchart of the core processes in an SMT production line.

[0068] Figure 2 This is a system architecture diagram for a hybrid production line for surface mount technology (SMT).

[0069] Figure 3 This is an interactive framework diagram of a semi-Markov decision process (SMDP).

[0070] Figure 4 This is the overall architecture diagram of GMeta-MATD3.

[0071] Figure 5 This is a comparison chart of the training convergence of various algorithms in the IEEE 33-node system.

[0072] Figure 6 Comparison of FIFO with the scheduling Gantt chart of this invention.

[0073] Figure 7 This is a topology diagram of the IEEE 141 node power distribution system.

[0074] Figure 8 This is a comparison chart of the training convergence of various algorithms for the IEEE 141-node system.

[0075] Figure 9 This is a comparison chart of the thermodynamic state and power curves of a reflow oven.

[0076] Figure 10This is a graph showing the relationship between production line load and carbon emission fluctuations.

[0077] Figure 11 This is a sensitivity analysis plot of the time decay discount factor. Detailed Implementation

[0078] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0079] This invention provides a proactive energy and carbon sensing production scheduling method for SMT production lines based on primitive reinforcement learning, comprising:

[0080] Collect and preprocess production data from surface mount technology production lines and microgrid data.

[0081] Construct a mixed integer nonlinear programming model that couples production, thermal engineering, electricity, and carbon flow, and define the optimization objectives of manufacturing cycle, energy cost, and carbon emissions, as well as constraints on production, thermal engineering, and power grid.

[0082] The scheduling problem is refactored into a semi-Markov decision process driven by discrete events, defining the state space, action space, reward function, and variable time transition rules.

[0083] A multi-flow graph convolutional network is constructed to extract the spatiotemporal topology embedding of the production-power grid, and the temperature of multiple regions of the reflow oven, conveyor belt speed, thermal formula, buffer queue, dynamic electricity price, carbon emission factor and power margin are represented as graph node features;

[0084] Embed a meta-learning module in a multi-agent system to infer the physical context and predict the actions of fellow agents by using historical trajectories;

[0085] A centralized evaluator based on the Transformer self-attention mechanism is adopted to dynamically decompose the joint value and calculate the individual advantage function;

[0086] Based on the GMeta-MATD3 architecture, a multi-agent strategy is trained under the centralized training and distributed execution paradigm to generate scheduling decisions for work order assignment, equipment routing, and hot recipe selection.

[0087] Execute scheduling decisions and issue control commands, collect execution feedback data to update primitive reinforcement learning models and multi-objective scheduling strategies.

[0088] This invention also provides an active energy and carbon sensing production scheduling system for SMT production lines based on primitive reinforcement learning, comprising:

[0089] The multi-source information sensing module is configured to collect production data, thermal data, microgrid data, and external environmental data from the SMT production line, and perform data preprocessing.

[0090] The production-thermal-electric carbon coupling modeling module is configured to construct a multi-objective scheduling model that includes discrete production logic and continuous thermodynamic constraints, and to reconstruct the scheduling problem as a semi-Markov decision process.

[0091] The primitive reinforcement learning decision module is configured to extract spatiotemporal topological features based on a multi-flow graph convolutional network, infer the physical context and predict peer actions through the meta-learning module, and use a Transformer-enhanced centralized evaluator for multi-agent credit allocation and policy optimization.

[0092] The execution and feedback closed-loop module is configured to send scheduling decisions to the SMT production line equipment controller, monitor the operating status, and provide feedback to optimize the scheduling strategy.

[0093] The following is a detailed implementation process of the present invention.

[0094] like Figure 1 As shown, this invention presents an active energy carbon sensing production scheduling method and system for SMT production lines based on primitive reinforcement learning. By constructing a four-level core architecture consisting of a "multi-source information sensing module, a production-thermal-electric carbon coupling modeling module, a primitive reinforcement learning decision-making module, and an execution and feedback closed-loop module," it achieves intelligent and integrated management and control of the entire process.

[0095] The system of this invention is a closed-loop management and control system. Each module interacts through a standardized data interface. The process is as follows: the multi-source information sensing module collects and preprocesses production, thermal, power, and carbon factor data from the SMT production line; the production-thermal-electricity-carbon coupling modeling module constructs the MINLP model and SMDP model to calculate thermal transition energy consumption, carbon emissions, and grid load; the primitive reinforcement learning decision module generates proactive scheduling decisions based on the GMeta-MATD3 architecture; and the execution and feedback closed-loop module issues instructions, monitors the operating status, and provides feedback to optimize the model and strategy, forming a complete closed loop.

[0096] The functions of each module in the system are as follows:

[0097] 1. Multi-source information sensing module, serving as the system's data foundation, is deployed in various SMT devices, microgrid nodes, and environmental monitoring units to achieve real-time acquisition, time synchronization, and standardized preprocessing of multi-source heterogeneous data; specific implementation is as follows:

[0098] (1) Data collection dimensions

[0099] 1) Production data: work order batches, bill of materials, hot recipe ID, delivery date, buffer queue, and processing time for each process;

[0100] 2) Thermal data: Real-time temperature of multiple zones in the reflow oven, heating / cooling power, conveyor belt speed, thermal transition time, and heat dissipation rate;

[0101] 3) Power data: Real-time power of each device, microgrid node voltage, branch current, peak power limit, and time-of-use electricity price;

[0102] 4) Carbon and environmental data: real-time carbon emission factors of the power grid, ambient temperature, humidity, and total carbon allowance.

[0103] (2) Data preprocessing methods

[0104] 1) Time synchronization: Based on BeiDou / GPS, timestamps are aligned and unified to the second-level scale;

[0105] 2) Outlier removal: 3σ criterion + Isolation Forest algorithm to remove sensor fault data;

[0106] 3) Missing value completion: Linear interpolation + LSTM time series prediction to complete missing data;

[0107] 4) Data standardization: Normalize to the [0,1] interval to eliminate the influence of units;

[0108] 5) Categorized storage: Stored on the industrial big data platform, supporting quick access from modules.

[0109] 2. The production-thermal-electric carbon coupling modeling module, as the core analysis unit, constructs a multi-objective MINLP model and a SMDP model, coupling production logic, thermal characteristics, and electricity-carbon flow constraints to provide quantitative basis for the decision-making module; the specific implementation is as follows:

[0110] (1) Multi-objective MINLP model

[0111] 1) Optimization objectives: Minimize manufacturing cycle time, energy costs, and total carbon emissions;

[0112] 2) Core constraints:

[0113] Routing and production flow constraints: uniqueness of work order allocation, process sequence constraints, and no overlapping processing;

[0114] Multi-region thermal coupling constraints: Asymmetric heating / cooling rates and bottleneck regions determine the total setup time;

[0115] Peak power constraint: Total power ≤ maximum allowable power of microgrid;

[0116] Delivery time constraint: Work order completion time ≤ delivery time.

[0117] (2) Reconstruction of Semi-Markov Decision Process (SMDP)

[0118] 1) State space: s k =[J k M k Ek This includes operational characteristics, equipment thermodynamic state, and dynamic power grid environment;

[0119] 2) Motion Space: Macro Actions This includes work order assignment, route allocation, and hot recipe selection.

[0120] 3) Reward function: Negatively penalizes energy costs, carbon emissions, delays, and peak violations to guide the optimal strategy;

[0121] 4) Transfer dynamics: Determining the variable transition time τ using the multi-region bottleneck effect. k It is adapted to non-stationary thermophysical properties.

[0122] (3) Electricity-carbon flow coupling calculation

[0123] 1) Energy cost: The time-of-use electricity price is calculated by integrating the operating power, transient power, and standby power.

[0124] 2) Carbon emissions: The dynamic carbon emission factor emissions of each power stage are calculated by integral calculation;

[0125] 3) Thermal transition energy consumption: Calculate the sequence-dependent transition energy consumption of multi-region heating / cooling.

[0126] 3. The primitive reinforcement learning decision module, as the core decision-making unit, implements centralized training and distributed execution (CTDE) based on the GMeta-MATD3 architecture, solving the challenges of non-stationarity, credit allocation, and topology awareness; the specific implementation is as follows:

[0127] (1) Multi-flow graph convolutional network (MF-GCN)

[0128] Modeling the SMT production line as a production-grid multi-relationship graph, extracting material flow and embedding microgrid spatiotemporal topology, enables the agent to perceive buffer congestion, global peak power, and carbon emission constraints.

[0129] (2) Meta-learning module

[0130] Each agent embeds a meta-learning module, which includes a history network, a state network, and a prediction network:

[0131] 1) Historical network encoding of historical trajectories to generate contextual embeddings;

[0132] 2) State network infers real-time physical context (heat dissipation rate, environmental parameters);

[0133] 3) Predict the network to deterministically estimate the next action of the peer agent, giving the agent the ability to proactively wait for energy arbitrage.

[0134] (3) Transformer Enhanced Centralized Evaluator

[0135] Replace the fully connected evaluator with a stacked self-attention Transformer:

[0136] 1) Dynamically decompose joint value to solve the credit allocation problem among multiple agents;

[0137] 2) Replace random sampling with meta-predictive actions to reduce TD error variance and stabilize training;

[0138] 3) Calculate the individual advantage function to achieve fair policy updates.

[0139] (4) Security constraint embedding mechanism

[0140] By embedding grid peak value, thermal process, and delivery time constraints into the reward function, high penalties are imposed for constraint violations and rewards are given for constraint satisfaction, ensuring that the decision-making is safe and feasible.

[0141] 4. Execution and feedback closed-loop module, which enables precise issuance of scheduling instructions, operation monitoring, effect evaluation, and model feedback optimization; the specific implementation is as follows:

[0142] (1) Execution unit

[0143] Scheduling commands are sent to the SMT equipment controller via Modbus TCP / IP and OPC UA industrial protocols.

[0144] 1) Smart agent of the placement machine: performs work order assignment and route allocation;

[0145] 2) Reflow oven intelligent system: performs functions such as heat recipe selection, conveyor belt speed adjustment, and deep sleep mode;

[0146] 3) Collaborative execution: recipe clustering, proactive waiting, and peak avoidance.

[0147] (2) Feedback unit

[0148] 1) Operation monitoring: Real-time collection of production progress, thermal parameters, power load, and carbon emission data;

[0149] 2) Performance assessment: Calculate manufacturing cycle time, energy costs, carbon emissions, delay rate, and peak violation rate;

[0150] 3) Feedback optimization: Update the parameters of the coupled model online, add real-time experience to the replay buffer, and iteratively optimize the reinforcement learning strategy.

[0151] The following is a detailed description of the implementation functions of the production-thermal-electric carbon coupling modeling module.

[0152] 1. System Architecture and Hybrid Flow Process

[0153] The SMT manufacturing system operates in a multi-variety, small-batch production mode, as shown in Figure 2. It processes a set of jobs (PCB batches) J={1,2,…,N}, where each job has a different bill of materials (BOM), PCB substrate and solder paste type. Production strictly follows the sequence of K stages defined by K={1,2,…,K}. The system has typical HFS characteristics, integrating continuous flow operations in upstream stages (solder paste printing, component placement) with recipe-driven pipeline batch operations in the final stage (multi-zone reflow soldering, k=K).

[0154] To accommodate parallel processing, each stage k is equipped with a set of heterogeneous machines M k . Specifically, reflow oven l∈M K varies in physical dimensions and heating configuration, and includes Z l independent temperature control zones.

[0155] For scheduling logic, binary decision variables are defined: X i,k,l ∈{0,1} equals 1 if job i is allocated to reflow oven l at stage k, and 0 otherwise. Y i,j,k,l ∈{0,1} equals 1 if job i is processed immediately before job j on reflow oven l at stage k, and 0 otherwise. Continuous time variables S i,k ≥0 and C i,k ≥0 represent the start time and completion time, respectively. For upstream stages (k<K), the standard processing time is p i,k,l . However, for reflow ovens, the processing time is jointly determined by the physical length L of the oven l and the dynamically adjustable conveyor speed determined by the thermal recipe of the job :

[0156]

[0157] 2. Multi-zone thermodynamics and recipe-based pipeline

[0158] The key bottleneck and main energy consumer is the reflow soldering stage. Each job i requires a strictly customized thermal recipe, which is defined by the desired conveyor speed and the temperature vector of Z l zones corresponding to the allocated reflow oven l .

[0159] Switching the recipe of reflow oven l from that of job i to that of job j generates a thermodynamic state transition time SU i,j,lBecause each region operates independently, the total setup time is dominated by a "bottleneck effect"—the furnace is only ready when the slowest region has completed its transition. In the scheduling model, this complex multi-region thermal inertia is expressed as a highly constrained, sequentially dependent setup time. Furthermore, the transition is asymmetric: rapid active heating ( ) and slow passive cooling ( The transition time for a specific region z is:

[0160] (1)

[0161] in, The transition time of the z-th temperature control zone when the reflow oven l switches from the i-th operation to the j-th operation. Let z be the heating rate of the z-th region. Let z be the cooling rate of the z-th region. The target temperature required for operation i in the z-th zone of reflow oven l. The target temperature required for operation j in the z-th zone of reflow oven l.

[0162] The total setup time for the welding furnace is the maximum value across all zones:

[0163] (2)

[0164] Recipe-based production lines: To minimize dry drying and maximize energy efficiency, operations with identical or highly similar thermal recipes can continuously share the same welding furnace. As a binary similarity indicator variable, when jobs i and j share the same recipe (T) i,l =T j,l and = The value is 0 if it is true, and 1 otherwise. =0, operation j can enter the welding furnace immediately after the minimum feed gap Δtgap, realizing continuous production line flow without waiting for operation i to completely exit the welding furnace.

[0165] Transition power in all Z l The dynamic integral over the region is:

[0166] (3)

[0167] Where P z (t) Output active heating power P according to the transition direction of the region heat,z or standby power P cool,z .

[0168] 3. Multi-objective optimization formulation

[0169] The system operates under a dynamic microgrid, constrained by real-time time-of-use (TOU) electricity prices c(t) and fluctuating grid carbon emission factors e(t). The MINLP model minimizes... To determine the Pareto optimal schedule.

[0170] Minimize manufacturing cycle:

[0171] (4)

[0172] Total energy cost minimization: Energy cost accurately captures the economic impact of time-based load shifts and multi-regional transitions.

[0173] (5)

[0174] Minimize total carbon emissions:

[0175] (6)

[0176] 4. System constraints

[0177] Let B represent a sufficiently large positive real number (Big-M).

[0178] Routing and upstream flow constraints:

[0179] (7)

[0180] (8)

[0181] (9)

[0182] (10)

[0183] Constraint (10) stipulates that for discrete upstream stages, operations cannot overlap.

[0184] Reflux pipeline with multi-region thermal coupling: For the reflux stage (k=K), disjunction constraints are fundamentally upgraded to support recipe-based pipelines. If the recipe does not match (I... i,j =1), job j must wait for job i to complete and the multi-region transition to end. If the recipe matches (I i,j =0), job j directly follows job i, separated only by the feed gap Δtgap:

[0185] (11)

[0186] Logical sorting must be mutually exclusive:

[0187] (12)

[0188] (13)

[0189] (14)

[0190] Constraints (12) and (13) serve as validity prerequisites. They ensure that the ordering relationship between job i and job j can only exist if both jobs are actively assigned to the same reflow oven l in stage k, where Y=1 or Y=1. If either job is routed to a different oven, these ordering variables are forced to zero.

[0191] Constraint (14) enforces mutual exclusion and completeness. If job i and job j are both assigned to the same reflow oven l (i.e., X + X = 2), the right side is evaluated to 1. This strictly requires the left side to be at least 1, ensuring that either job i must precede job j, or job j must precede job i. This prevents scheduling conflicts and ensures a clear, non-overlapping processing order.

[0192] Spatiotemporal coupling of peak demand: Let Φ be a binary indicator variable, which is 1 when t is in [S,C]. Let Ψ be a binary indicator variable, which is 1 when t is in [C,C+SU] and Y=1.

[0193] (15)

[0194] Delivery time constraints:

[0195] (16)

[0196] 5. Semi-Markov Decision Process (SMDP) Representation for DRL Implementation

[0197] Transforming complex MINLP models into semi-Markov decision processes (SMDPs) represents a crucial leap from traditional analytical solvers to modern data-driven control. Markov decision processes, such as... Figure 3 As shown, the agent interacts with the environment to learn its policy, and each agent has a meta-learning module for predicting the actions of other agents.

[0198] SMDP is significantly superior to standard Markov decision processes because performing actions in an SMT production line, particularly multi-zone thermodynamic transitions and variable conveyor speeds, consumes a continuous, variable time span τ, rather than a fixed discrete time step.

[0199] The following definition is a state-action-reward tuple specifically tailored for handling meta-reinforcement learning (Meta-RL) algorithms in mixed assembly line workshops.<S,A,P,R> A rigorous mathematical framework.

[0200] In the proposed SMDP, the agent does not need to make a decision every consecutive microsecond. We define the discrete decision time t. k When a specific discrete event occurs, the agent is awakened to make a decision:

[0201] (1): The machine completes its previous task and transitions to an idle state (or reaches the pipeline feed gap Δt). gap ).

[0202] (2): New high-priority emergency order arrival system.

[0203] (3): The slowest temperature zone of the target reflow oven has successfully reached the set value, marking the completion of the "bottleneck" transition and the readiness for processing.

[0204] 5.1. State Space (S)

[0205] At any decision time t k System status s k The logical state of the manufacturing queue, the multi-regional thermodynamic state of the equipment, and the dynamic environment must be encapsulated holistically. We decouple the state vector into three characteristic matrices: s k =[J k M k E k ].

[0206] Job Feature Matrix (J) k ): Describes the current task attributes in the waiting queue (buffer). Since routing and conveyor speed are dynamically allocated, processing time is not static. For each pending job i, let T req,i This represents the specific thermal formula vector required for the operation in different regions, v req,i D represents the standard conveyor speed required to satisfy the duration of the thermal profile. i -t k This indicates the remaining slack time until the delivery date.

[0207] Machine Feature Matrix (M) k ): Describes the continuous and discrete physical states of heterogeneous equipment. For each reflow oven l, T cur,l (t k ) represents all Z of welding furnace l l Real-time physical temperature vectors for each independent region. cur,l (t k This indicates the current operating speed of the conveyor belt in the welding furnace. (Recipe) I D l This indicates the encoding of the currently executed hot recipe, which the agent uses to determine pipeline similarity (I). i,j P mach,l (t k) represents the aggregated instantaneous physical power consumption across all regions.

[0208] Dynamic Environment Matrix (E k ): Describes the non-stationary time series information of the external power grid, where c(t) k ),c(t k +Δt) represents the TOU electricity price at the current moment and within the future forecast window, e(t) k ),e(t k +Δt) represents the predicted grid carbon emission factor curve, Q acc (t k This represents the cumulative instantaneous power demand of the workshop (used to prevent exceeding the peak limit Q). max ).

[0209] 5.2. Action Space (A)

[0210] Given the complex characteristics of the High-Speed ​​Fire Array (HFS) and multi-region thermal coupling, directly defining the action space as a continuous multi-region temperature would trigger the "curse of dimensionality" and prevent convergence. Therefore, the action 'a' generated by the agent... k Designed as a route-recipe macro action vector: .

[0211] Assign actions ( ): Select the next job i to process from the queue. The inclusion of virtual Wait actions enables proactive idle time for energy arbitrage.

[0212] Route allocation ( ): Determine the specific heterogeneous reflow oven l to which the selected job will be assigned, and form a parallel queue based on the oven capacity.

[0213] Thermodynamics and Kinematic Control This indicates that the agent does not output the temperature of a single region, but rather draws from the predefined library R. lib Select a standard thermal recipe. This single action sets all Z-recipes simultaneously. l The target temperature vector T of each region target And actively adjust the conveyor belt speed v conv This determines the processing time (p=L). l / v conv The Sleep command forces the furnace into deep sleep mode.

[0214] Driven by this action space, the agent autonomously learns to route jobs requiring similar recipes to the same welding furnace, continuously feeding them into the conveyor belt to maximize the production line effect.

[0215] 5.3. Transfer Dynamics and Nonstationarity

[0216] In SMDP, execute action a k Transition the system to state s k +1 and generate a random dwell time span τ k .

[0217] Time evolution: t k+1 =t k +τ k

[0218] τ k Physical mechanism: Transition time is strictly determined by pipelined logic (I i,j The decision is determined by the multi-regional bottleneck effect. If the agent clusters similar tasks (matching recipes, I=0), the next decision time is during the brief feed interval Δt. gap This occurs later. If a forced recipe switch is performed (I=1), τ k Maximum area setting time must be bridged:

[0219] (17)

[0220] Non-stationarity of transfer: In real-world scenarios, the passive cooling rate of each region It is highly sensitive to changes in the factory environment. Therefore, the transition probability P exhibits Hidden Markov characteristics. This provides an optimal theoretical entry point for incorporating meta-reinforcement learning: the meta-learner can extract the latent variable z from historical trajectories. k Infer the real, potential physical thermal environment online.

[0221] 5.4. Reward Function (R)

[0222] The design of the reward function determines the agent's ability to autonomously learn complex strategies, such as "recipe clustering," "energy-carbon arbitrage," and "conveyor belt speed synchronization." The MINLP objective function is directly mapped to the integral-based instantaneous reward r of SMDP. k .

[0223] In the interval [t] k ,t k +τ k Within [the context], the immediate reward r received by the agent. k Represented as the negative sum of multiple penalty terms:

[0224] (18)

[0225] , , , For their respective weights, the exact calculus of these terms is defined as:

[0226] Dynamic economic cost penalty (R)cost ):

[0227] (19)

[0228] Dynamic carbon emission penalty (R carbon ):

[0229] (20)

[0230] In these two spans of all Z l Driven by the components of the regional integral, the agent self-penalizes unnecessary recipe switching and autonomously learns to group tasks to avoid large-scale multi-regional transition power P. trans .

[0231] Delay penalty (R) delay ):

[0232] (twenty one)

[0233] Completion time depends on the selected conveyor belt speed Make the decision proactively.

[0234] Peak demand hard penalty (R) penalty If the duration is τ k Within the system, the aggregated concurrent power exceeds the microgrid's allowable limit Q. max Then, a huge, precipitous punishment will be imposed:

[0235] (twenty two)

[0236] Where I is the indicator function and Θ is a very large penalty constant.

[0237] The following is a detailed description of the implementation functions of the primitive reinforcement learning decision module.

[0238] Primitive Reinforcement Learning Architecture for Proactive Scheduling

[0239] To address the aforementioned complex SMT production line, a novel multi-agent deep reinforcement learning framework, Grapheta-MATD3, is proposed. Since SMT production lines inherently operate as coupled HFS networks of material, power, and carbon flow, traditional vector-based state representations cannot capture spatial topology and dynamic parallel routing. Furthermore, environmental non-stationarity (such as fluctuating cooling rates) also presents challenges. The dynamic carbon factor e(t) and the necessity of proactive scheduling (such as performing "wait" actions or routing jobs to utilize the recipe pipeline) pose significant challenges to algorithm convergence. Figure 4 As shown.

[0240] To address these issues, GMeta-MATD3 integrates a Multi-Flow Graph Convolutional Network (MF-GCN) for topological state extraction, a meta-learning module for inferring dynamic hidden context, and a Transformer-based evaluator that accurately evaluates joint actions using multi-head self-attention. The algorithm operates under a centralized training distributed execution (CTDE) paradigm.

[0241] 1. Multi-Agent CTDE Framework and Agent Definition

[0242] By formulating the problem as a graph-enhanced SMDP, the foundation for the GMeta-MATD3 framework is laid. Similar to solving grid volatility, it involves a dynamic TOU electricity price c(t) and variable thermal inertia v. - The introduced nonstationarity is addressed by utilizing a historical parsing recurrent neural network that generates a latent context vector. The Actor-Critic network uses this vector along with the topological states extracted by the GCN. k The conditional output enables the agent to dynamically perform "time-sharing arbitrage" and "recipe clustering" without having to solve NP-Hard MINLP online.

[0243] The scheduling problem is decomposed into a cooperative multi-agent system (MAS). An agent is assigned to each critical machine stage. For example, the heterogeneous reflow oven is controlled by a dedicated agent responsible for recipe selection (dynamically indicating multi-zone setpoints and conveyor speeds), while the upstream placement machine is controlled by an agent responsible for dispatching and routing.

[0244] Each agent i relies only on its local observations during distributed execution. and inferred meta-context Through its Actor network Generate Routes - Recipe Macro Action During intensive training, the global Critic network uses global state s k Assess joint actions a k To calculate gradient updates.

[0245] 2. Multi-flow graph convolutional networks

[0246] To capture the spatial coupling of production and energy, the workshop is modeled as a directed multi-relationship graph G=(V,E) prod E grid ).

[0247] Node (V): V represents heterogeneous machines and buffers. Input node feature matrix X (0) Includes the decoupled state vector defined in SMDP: [J k M k E k [Including the current status of each oven] .

[0248] Production edge (E) prod ): E prod It represents the physical routing probability of PCB batches between sequential stages, capturing the parallel queues of the HFS architecture.

[0249] Electric grid edge (E grid ): E grid Connecting all machines to a virtual "microgrid node" represents a shared peak demand constraint (Q). max (and carbon emissions) are coupled.

[0250] MF-GCN extracts spatial embeddings through layer-by-layer propagation. The basic operations of GCN can be represented according to

[23] . In particular, the hidden representation of layer l is updated as follows:

[0251] (twenty three)

[0252] in, To incorporate the normalized adjacency matrix of self-loops, To generate the edge matrix, For the power grid side matrix, It is the identity matrix. For degree matrix, Let l be the trainable weight matrix of the l-th layer. It is the ReLU activation function. This represents the embedding of nodes in layer l. The final GCN layer... The output is used as the graph-aware observation embedding for each agent. .

[0253] 3. Meta-learning module for proactive behavior prediction

[0254] As SMT production lines scale up, MAS (Multi-Agent System) accelerates the achievement of task objectives. However, as each agent continuously learns and adjusts its routing and recipe selection strategies, the environment becomes highly non-stationary. This dynamic interaction and potential policy conflicts complicate the learning process and hinder algorithm convergence.

[0255] To address this non-stationarity and endow agents with proactive scheduling capabilities, an embedded meta-learning (ML) module is introduced into each agent. The fundamental goal of this ML module is to learn the learning process itself and use prior knowledge to predict upcoming macro-actions of other agents. We assume that agents can observe the states and actions of their peers but cannot access their internal parameters or precise policies.

[0256] To achieve efficient online inference, the meta-reinforcement learning architecture adopts a context-driven paradigm. Specifically, it consists of three specialized neural network modules: a history network (acting as a context encoder, extracting task-specific features from past transitions), a state network (a prediction network for real-time decision-making), and a context encoder.

[0257] Historical Network ( ):make This represents the j-th past historical trajectory observed by agent i. (Network) Each past segment is analyzed independently to generate a feature embedding: These embeddings are summed to form a comprehensive historical representation:

[0258] (twenty four)

[0259] State networks ( State network The immediate operational state context is inferred by parsing the most recent stage trajectory, and previous state representations are iteratively embedded to output the updated current state embedding:

[0260] (25)

[0261] Prediction Network ( Finally, given the state of the new observation. The predictive network output is used to deterministically estimate the next action of its peers. :

[0262] (26)

[0263] Integration in Active Pipelines: This ML prediction mechanism is transformative in the context of HFS-based SMT production lines. Through successful prediction... (For example, inferring that a specific downstream reflow oven will maintain a high-temperature formulation R) x The upstream placement machine agent can proactively adjust its routing strategy, akroute, to intentionally route items that require R... x The tasks are assigned to that specific welding furnace. This maximizes the pipeline indicator variable (I=0) and completely bypasses the multi-zone transition bottleneck. Furthermore, these predicted actions... It directly replaces the high-variance random sampling action in the time difference (TD) objective calculation, stabilizing the overall multi-agent training process.

[0264] 4. Enhance the centralized evaluator (TSA) by stacking self-attention Transformers.

[0265] In the standard CTDE paradigm, all agents typically share a joint global reward. In the highly coupled HFSSMT production line, this creates a severe "credit allocation" problem. For example, an upstream placement agent might route mismatched workpieces to the soldering furnace, forcing the furnace agent to switch recipes during peak electricity price windows. This triggers a large-scale multi-region transition penalty (Ptrans). If a global joint reward is used, innocent parallel agents will be unfairly penalized, hindering their ability to evaluate local policies.

[0266] To properly evaluate the individual contribution of each agent, the standard multilayer perceptron evaluator is abandoned, and a Transformer is introduced into a centralized Critic network via TSA. The TSA architecture has two key improvements: (i) it explicitly utilizes the output of the meta-learning module, and (ii) it stacks self-attention layers to dynamically adjust weights based on the importance of each agent's specific state and routing / recipe action.

[0267] 1) Meta-information self-attention mechanism: During intensive training, the evaluator receives the input matrix The feature vector of agent i is constructed by concatenating its local observations, macro actions, and meta-predicted future actions. .

[0268] The SA layer embeds and maps this metadata into a query (Q), key (K), and value (V) matrix:

[0269] (27) |

[0270] Attention weights are calculated using scaled dot products, enabling the evaluator to dynamically focus on bottleneck agents (e.g., specific welding furnaces undergoing multi-region thermal transitions):

[0271] (28) |

[0272] 2) Value decomposition and advantage assessment: Through the forward propagation of TSA, we obtain the corresponding Q value of each individual agent. The joint expected return is strictly decomposed:

[0273] (29)

[0274] To isolate the relative value of a specific agent's routing or recipe actions and to eliminate baseline noise caused by other agents, we compute the advantage function:

[0275] (30)

[0276] in This is achieved through the marginal contribution A based purely on agent i.i They updated their strategies to address the problem of imbalanced learning.

[0277] 3) Variance reduction in TD error calculation: Calculate the TD error δ k This typically involves randomly sampling the next action. Significant variance is introduced. By integrating the TSA evaluator with the meta-learning module, deterministic meta-predictive actions are achieved. Replace random sampling.

[0278] The improved TD error updated by SMDP becomes:

[0279] (31)

[0280] This physical information TD target stabilizes the training trajectory, enabling the SMT agent to execute highly coordinated time arbitrage and recipe pipeline strategies.

[0281] 5. Active TD Error Update Mechanism

[0282] To train the GMeta-MATD3 network, we utilize the dual-delay deep deterministic policy gradient (TD3) method. To leverage the proactive context inferred from the meta-learning module and reduce the high variance caused by scattered exploration, we fundamentally modify the standard TD objective computation.

[0283] For the HFS SMDP formulation, the improved TD objective y k Calculated as:

[0284] (32)

[0285] in This represents the continuous-time discount factor, taking into account the variable physical transition time τ of the SMT thermodynamic process. k The characteristics of Critic. The Critic parameter ω i Update by minimizing the mean squared error (MSE) loss on the sampled mini-batch B:

[0286] (33)

[0287] Finally, the Actor network parameters θ for each agent i i Deterministic policy gradient updates guided by a centralized evaluator enhanced by Transformer:

[0288] (34)

[0289] By explicitly integrating graph topology, meta-inference context, and attention-based value decomposition, GMeta-MATD3 provides a mathematically robust solution for performing proactive, recipe-aware scheduling in highly dynamic HFS environments. The execution logic of the proposed GMeta-MATD3 algorithm is summarized in Algorithm 1, which details the continuous interactions and proactive TD updates within SMDP.

[0290] Algorithm 1: GMeta-MATD3 Production Scheduling Algorithm

[0291] Required: SMDP environment, MF-GCN, meta-modules ( , ), Actor Network TSA-Critic Network

[0292] Require: Playback buffer B, batch size B, discount factor Soft update rate Delayed update frequency d.

[0293] 1: Initialization: Target Network and .

[0294] 2: for episode = 1 to M do

[0295] 3: Reset SMDP and receive the initial global state s0

[0296] 4: for decision time k=0,1,…do

[0297] 5: / / Decentralized proactive execution

[0298] 6: For each agent i in {1,..., N}, do

[0299] 7: From s via MF-GCN k Extracting spatial features

[0300] 8: Utilize the history encoder and state encoder Inferring active context

[0301] 9: Select macro actions (routing, recipe, dispatch)

[0302] 10: end for

[0303] 11: Execute joint action a kObserve the transition duration τ k Reward r k , and the next state s k+1

[0304] 12: Storage Transfer At B

[0305] 13: / / Centralized Meta-training

[0306] 14: if|B|>B then

[0307] 15: Mini-batch transfer from sample size B

[0308] To verify the superiority and physical feasibility of the proposed GMeta-MATD3 architecture in continuous-discrete coupled manufacturing systems, numerous numerical experiments were conducted. The experiments directly implemented the SMDP formulation and the multi-objective MINLP model. These were based on real industrial data and standard microgrid models, including IEEE 33-node and IEEE 141-node distribution systems.

[0309] 1. Experiment setup and two-layer network interaction

[0310] The simulation environment is constructed as a two-layer "production-energy-carbon" interaction framework, instantiating the multi-relationship graph topology defined above. The upper layer simulates an SMT production line with heterogeneous multi-zone reflow ovens. The lower energy-carbon layer employs the classic IEEE 33-node and extended IEEE 141-node power distribution systems for main body verification and scalability verification, respectively. In both configurations, the SMT facility acts as the primary industrial load, directly coupled to a heavy-load power grid node (e.g., node 18 in a 33-node network).

[0311] Two-way interaction mechanism:

[0312] (1) Bottom-up state awareness: IEEE distribution network calculates power and carbon flow in real time, combining the nodal marginal electricity price c(t), real-time carbon emission intensity e(t), and remaining microgrid capacity Q. acc (t) is fed back to the upper-level dynamic environment matrix E defined in the SMDP state space. k (Section 4.1).

[0313] (2) Top-down proactive response: The dispatching actions of the SMT production line, especially the activation, heating, and cooling of the laser reflow oven, generate the instantaneous power curve P total (t). The dynamic load curve is injected into the corresponding power distribution system as a disturbance, directly changing the global voltage distribution and carbon flow topology, while forcibly enforcing the spatiotemporal coupling constraint of peak demand (Equation (15)).

[0314] To comprehensively evaluate the proposed method, GMeta-MATD3 is benchmarked against six classic and state-of-the-art control and MARL algorithms. These baselines include Rolling Time-Domain Model Predictive Control (MPC), Independent PPO (IPPO), Multi-Agent PPO (MAPPO), Traditional MATD3, and MADDPG. Furthermore, it is compared with SQDDPG, a hybrid architecture combining DDPG with Sequential Quadratic Programming (SQP) for enforcing strict physical constraints.

[0315] The key physical parameters are derived from real SMT production data and strictly matched to the multi-zone thermodynamic model established above. For heterogeneous reflow ovens, the number of independent temperature control zones is set to Z. l =4 (preheating, soaking, reflux, and cooling). Active heating power P per zone heat,z It is 45kW, while the standby / passive cooling power P cool,z The power is 4kW. To capture the multi-region bottleneck effect, the asymmetric thermal inertia rate is calibrated to... =2.5degC / s (rapid heating) and =0.1 degC / s (passive cooling). The minimum feed gap for the production line based on the formulation is set to Δt. gap =0.5 minutes.

[0316] Regarding the processing tasks, the production workload was configured to simulate highly mixed, small-batch manufacturing. The evaluation covered two task scales: a primary scale of N=100 PCB batches for the IEEE 33-node system, and a large-scale expansion of N=500 batches for the IEEE 141-node system. To capture the inherent heterogeneity of the operations, standard processing times for discrete upstream stages (such as pick-and-place) were randomly generated according to a uniform distribution p∼U(0.5, 1.2) hours. Furthermore, the batch thermal requirements were uniformly distributed across three different formulation curves: R... A (220degC), R B (240degC) and R C (260°C). Delivery date D i Generates a strict benchmark for latency metrics by multiplying the standard tight-loose routing factor by the cumulative baseline processing time.

[0317] For dynamic microgrid environments, the peak price of the time-of-use electricity price c(t) is 1.2 yuan / kWh (e.g., 10:00-12:00), and the valley price is 0.3 yuan / kWh, while the grid carbon emission factor e(t) fluctuates between 0.4 and 0.8 kgCO2 / kWh. The local node capacity is limited to Q. max =150kW. To enforce strict peak avoidance, an overwhelming penalty constant Θ=104 is imposed in the SMDP reward function.

[0318] For the GMeta-MATD3 algorithm, the network architecture and hyperparameters were carefully tuned to ensure stable convergence. MF-GCN uses 3 hidden layers and 128-dimensional embeddings to extract the spatial topology, while the meta-learner employs LSTM for context inference, with a hidden state size of 64. A continuous-time decay discount factor of γ=0.99 and a target network soft update rate of ρ=0.005 were used, which validated maximum algorithm stability in the sensitivity analysis (below). The learning rates for the Actor and TSA-Critic networks were 1e−4 and 3e−4, respectively. The experience replay buffer capacity was set to 10. 6 Active TD error updates are performed using a batch size of B=256.

[0319] All experiments maintained the same hyperparameter and physical parameters in both IEEE 33-node and IEEE 141-node scenarios to ensure fair comparison. The multi-agent system followed a centralized training and distributed execution paradigm, with each agent using a meta-learning module for proactive peer action prediction.

[0320] The GMeta-MATD 3-meta reinforcement learning model was trained using the PyTorch deep learning framework on a high-performance workstation equipped with an NVIDIA Tesla V100 GPU (32GB VRAM) and a 2.1GHz 6-core Intel Xeon(R) Gold 6130 CPU. All training experiments, including the main IEEE 33-node case and the IEEE 141-node scalability extension, were conducted under the same hardware and software configurations to ensure reproducibility and fair performance comparisons.

[0321] 2. Algorithm convergence and overall performance

[0322] Figure 5 The training process of each method is shown, and Table 1 details the multidimensional scheduling metrics for 100 job batches.

[0323] Table 1. Comparison of overall performance of production, energy, and carbon indicators on the IEEE 33-node system (task size N=100)

[0324] MPC (Rolling Time Domain) -1,450.2 14.5 8.5 1,850.4 1,420.5 0 IPPO -1,020.6 14.3 8.2 1,350.8 1,080.2 10 MADDPG -1,000.4 13.9 5.5 1,280.5 1,010.4 7 MAPPO -980.2 13.7 4.5 1,220.2 970.8 4 SQDDPG -960.8 13.6 3.8 1,180.9 935.3 1 MATD3 -940.5 13.4 3.2 1,150.4 910.6 2 GMeta-MATD3 (This article) -850.3 13.1 1.5 1,020.5 810.2 0

[0325] Over a training span of 4,000 segments, GMeta-MATD3 achieved the fastest convergence and the highest global cumulative reward of -850.3. By leveraging meta-predictive actions to stabilize TD error updates, it established a robust, oscillation-free plateau as early as the 1,400th segment.

[0326] While state-of-the-art MARL baselines—such as MATD3 (-940.5) and Constrained Hybrid SQDDPG (-960.8)—narrowed the final performance gap to 10.6% and 13.0%, respectively, their training trajectories suffered from severe high-frequency variance and recurring catastrophic policy collapses. These abrupt, cliff-like drops in joint rewards, followed by slow recovery, typically occurred when dispersed agents performed conflicting exploratory actions that disrupted the coordination of previously learned strategies.

[0327] Furthermore, MADDPG (-1000.4) and MAPPO (-980.2) lack the topology awareness provided by MF-GCN, rendering them "blind" to the spatial capacity distribution of the IEEE 33-node microgrid. This leads to frequently overlapping heating tasks triggering large-scale peak demand violations (Q violations) at local nodes during peak TOU periods. max This results in dramatic oscillations between temporary gains and severe penalties. IPPO performs the worst among data-driven methods (-1,020.6) because its completely independent learning cannot handle non-stationary multi-region thermal coupling.

[0328] Finally, while conventional rolling temporal MPC mathematically guarantees that physical constraints are satisfied, its truncated prediction horizon traps it in severe local optima, resulting in a static reward of -1,450.2, a 70.5% decrease compared to GMeta-MATD3. This stark contrast demonstrates that purely passive constraint enforcement is far inferior to the proactive, farsighted temporal arbitrage and recipe pipeline achieved by our primitive reinforcement learning architecture.

[0329] To verify the physical feasibility of the learned strategy, a representative 10-batch scheduling window was analyzed. Figure 6 The placement time follows p~U(0.5, 1.2) hours, and the batch requires heat treatment of the formulation R. A (220degC), R B (240°C) or R C (260°C). Recipe switching produces significant transition power (Ptrans), while the same recipe shares a minimum 0.1-hour feed gap (Δt). gap ).

[0330] At the FIFO baseline ( Figure 6 a) Under these conditions, the reflow oven passively processes operations in chronological order (R) A →R C →R B This triggers frequent energy-intensive heating transitions. Crucially, blindly implementing large-scale heating transitions during peak time-of-use pricing and carbon windows (10:00-12:00) carries the risk of severe economic penalties and negative impacts on microgrid Q. max Capacity violation risk.

[0331] Conversely, GMeta-MATD3 ( Figure 6 b) Using meta-prediction to autonomously cluster identical formulas (R A →R B →R C This establishes a seamless unidirectional thermodynamic pipeline, significantly reducing total setup time. Upon reaching peak demand, the upstream agent proactively executes a Wait macro, forcing the welding furnace into deep sleep (Pcool). High-temperature RC operations are intentionally postponed to the green electricity off-peak period after 12:00, perfectly executing spatiotemporal energy arbitrage while strictly meeting manufacturing cycle constraints.

[0332] 3. Scalability verification on a 141-node network

[0333] To rigorously evaluate the scalability and generalization capabilities of the proposed architecture, numerical experiments were extended from the main 33-node microgrid to a significantly more complex large-scale environment: the IEEE 141-node distribution network. In this extended scenario, the SMT manufacturing facility was scaled up to include eight parallel upstream placement lines and three independent multi-zone reflow ovens to meet significantly higher throughput requirements (N=500 job batches).

[0334] Expanded environments exacerbate the inherent challenges of proactive scheduling. Expanded graph topologies introduce deeper spatial energy-carbon couplings, while the increased number of parallel machines exacerbates the non-stationarity of multi-agent systems.

[0335] Figure 7 Training convergence for 4,000 segments in a 141-node environment is demonstrated. Despite perturbations from the extended state-action space, GMeta-MATD3 exhibits remarkable resilience. Utilizing MF-GCN for spatial embedding and meta-learner prediction of thermodynamic transitions, it achieves a stable Pareto-optimal plateau near segment 2,000, with a final joint reward of -2,850.5. In contrast, the baseline deteriorates severely under the "curse of dimensionality." MAPPO (-3,480.2) and MADDPG (-3,650.8), lacking graph-aware topological embedding, suffer from "topological blindness," leading to massive power spikes and unrecoverable policy collapse. IPPO (-3,910.4) exhibits extreme oscillations as the independent exploration of numerous agents makes the environment highly non-stationary.

[0336] Table 2 Scalability performance on IEEE 141 bus network (task size N=500, 3 parallel soldering furnaces)

[0337] MPC (Rolling Time Domain) -4,863.4 56.5 31.2 5,240.6 4,820.5 0 IPPO -3,910.4 48.5 18.2 4,120.5 3,950.4 35 MADDPG -3,650.8 46.2 14.5 3,850.2 3,710.6 28 MAPPO -3,480.2 45.4 11.2 3,680.5 3,520.3 22 SQDDPG -3,310.5 44.8 9.5 3,510.4 3,380.2 12 MATD3 -3,150.2 43.6 7.2 3,320.8 3,190.5 15 GMeta-MATD3 (This article) -2,850.5 41.2 3.5 2,950.6 2,850.5 2

[0338] Table 2 summarizes the metrics for the expanded N=500 task size. Crucially, traditional rolling time-domain MPC struggles to cope with dimensionality explosion. Forced to terminate within a 4-hour computation limit, the MINLP solver produces short-sighted, suboptimal solutions with a static reward of -4,863.4. While MPC guarantees zero physics violations, its truncated time domain results in significant time delays and staggering energy costs.

[0339] In comparison, the GMeta-MATD3 completed the online inference of the entire 500-batch scheduling in just 320 milliseconds. Furthermore, its efficient management of a complex pipeline of three independent welding furnaces limited carbon emissions to 2,850.5 kg CO2, demonstrating its absolute superiority in large-scale energy-sensing manufacturing.

[0340] 4. Production Scheduling and Active Energy-Carbon Response Analysis

[0341] To analyze the potential decision-making intelligence of the trained agent, a microscopic analysis was performed on the generated scheduling trajectory. Figure 8 High-resolution Gantt plots and corresponding thermodynamic and power curves of the reflow oven are shown.

[0342] Visualization shows that the traditional heuristic method of frequently relying on the first-in-first-out (FIFO) rule or purely time-centered allocation causes the target temperature of the reflow oven to oscillate unstablely between high and low setpoints. This "zigzag" thermal curve generates a large amount of wasted energy due to continuous active heating and air heating, which fundamentally violates the recipe-based pipeline mechanism established by formula (11).

[0343] Conversely, GMeta-MATD3 exhibits superior "proactive follow-and-arbitrage" behavior. By fully leveraging the meta-learning prediction module and graph-enhanced observation embeddings, the multi-agent system autonomously converges to two advanced cooperative strategies:

[0344] (1) Unidirectional thermodynamic stepping: Instead of passively following a time-sequential queue reaction, upstream bonding agents actively cluster thermal formulations with the same or similar targets (T). req The incoming batches. By maximizing the pipeline indicator variable (I=0), the job is only performed with the minimum feed gap (Δt). gap The continuous flow through the furnace, this strategic formulation grouping significantly reduces the average multi-zone transition energy consumption from 12.5 kWh / batch under FIFO to a highly concentrated 3.2 kWh / batch, effectively eliminating energy-intensive passive cooling and reheating cycles.

[0345] (2) Proactive "Wait" and Load Arbitrage: A true sign of production-energy-carbon coupling was observed between 10:00 and 12:00. During this window, the microgrid experienced a significant surge in time-of-use pricing and grid carbon emission factors. This was achieved through the latent context vector z. k Anticipating this environmental penalty, the placement agent intentionally executes a virtual Wait macro action defined in the SMDP action space. This deliberately halts material feeding, allowing the downstream reflow oven agent to transition to a deep sleep mode (P). sleep The delayed tasks were then processed in batches during the "Green Power Valley" at 13:00. This perfect execution of time arbitrage confirms that the system can proactively shape its production rhythm according to the dynamic energy-carbon flow, sacrificing marginal time slack in exchange for significant economic and environmental benefits.

[0346] 5. Long-term fluctuation analysis of the IEEE 33-node energy-carbon network

[0347] To verify the system's robustness over extended timescales, the load and carbon emission fluctuation curves of the 33-node system were continuously extracted over a one-month (30-day) period, such as... Figure 9 As shown.

[0348] In the baseline scenario without proactive scheduling, the random arrival of SMT orders triggers unpredictable and severe load spikes at node 18 (peak values ​​up to 180kW, see below). Figure 9 a). These severe transient surges are the main trigger for harmful voltage drops in the distribution network.

[0349] After implementing GMeta-MATD3, the system's power load was smoothly smoothed out, even when handling the same total number of manufacturing batches. The peak-to-valley difference was mathematically reduced by 34.5%.

[0350] In addition, such as Figure 9 As shown in (b), under the proposed strategy, the daily average carbon intensity of the microgrid is strictly maintained within a safety margin of less than 0.55 kg CO2 / kWh, completely eliminating the severe carbon emission violations observed in the baseline. This conclusively demonstrates that a multi-agent swarm controlled by the TSA evaluator value decomposition (Section 5.4) can ensure the physical and environmental robustness of the grid while fully competing for production efficiency.

[0351] 6. Sensitivity analysis of core SMDP hyperparameters

[0352] The novelty of the proposed SMDP lies in the continuous physical transition time τ k (Controlled by the multi-region bottleneck effect in Equation (2)) and coupled with the reinforcement learning temporal difference (TD) update mechanism. To systematically evaluate this, sensitivity analyses were performed on two key hyperparameters: the time decay discount factor. and target network soft update rate .

[0353] Figure 10 This indicates that the continuous time discount mechanism This profoundly impacts the long-term perspective of scheduling strategies. When When the energy density is less than 0.95, time penalties dominate; agents become excessively short-sighted, frantically processing batches to minimize manufacturing cycles while refusing to perform wait actions in exchange for future low-carbon windows, causing energy costs to soar to over 1,800 yuan. Conversely, when... At a value of 0.99, the algorithm achieves the optimal Pareto tradeoff, balancing the manufacturing cycle constraint (13.1 hours) with the minimum energy-carbon cost. However, [the following text appears to be incomplete and requires further context: "will..."] Pushing it to 0.995 will induce "excessive waiting", resulting in severe delay penalties and worsened manufacturing cycles.

[0354] Meanwhile, Figure 111 highlights the different soft update rates ( The vulnerability of multi-agent coordination in networks. The algorithm exhibits maximum stability and the highest joint reward (-850.3) when the value is 0.005. Beyond 0.01, the target network updates become overly aggressive. This volatility introduces a severe bias in the behavior prediction of the meta-learner (evidenced by soaring MSE on the secondary axis). This prediction failure subsequently poisons the value decomposition of the TSA evaluator, undermining the trust and cooperation pipeline in MAS, manifesting as huge reward variance and eventual policy collapse.

[0355] This invention addresses the proactive energy-carbon-aware scheduling problem in SMT production lines by reformulating the deep coupling of discrete workpiece assignment with continuous multi-region thermodynamics as a semi-Markov decision process. To solve this problem, the invention proposes the GMeta-MATD3 framework, which integrates a graph neural network for spatiotemporal topology embedding, a meta-learning module for latent context inference and proactive time arbitrage, and an attention-based centralized evaluator for solving multi-agent credit allocation. Extensive experiments using real SMT data and IEEE power distribution systems demonstrate its significant superiority over state-of-the-art baselines. GMeta-MATD3 establishes a dominant Pareto front—achieving a manufacturing cycle time of 13.1 hours, energy costs of RMB 1,020.5, CO2 emissions of 810.2 kg, a delay rate of 1.5%, and zero grid violations. Specifically, autonomous thermodynamic stepping reduces transition energy consumption from 12.5 kWh / batch to 3.2 kWh / batch, while proactive peak avoidance reduces load peak-to-valley difference by 34.5%, strictly limiting daily carbon intensity to below 0.55 kgCO2 / kWh. Scalability testing validated its potential for real-world industrial deployment, demonstrating online inference of 500 batches scheduled within 500 milliseconds.

[0356] The above are preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention shall fall within the protection scope of the present invention.

Claims

1. A proactive energy and carbon sensing production scheduling method for SMT production lines based on primitive reinforcement learning, characterized in that, include: Collect and preprocess production data from surface mount technology production lines and microgrid data. Construct a mixed integer nonlinear programming model that couples production, thermal engineering, electricity, and carbon flow, and define the optimization objectives of manufacturing cycle, energy cost, and carbon emissions, as well as constraints on production, thermal engineering, and power grid. The scheduling problem is refactored into a semi-Markov decision process driven by discrete events, defining the state space, action space, reward function, and variable time transition rules. A multi-flow graph convolutional network is constructed to extract the spatiotemporal topology embedding of the production-power grid, and the temperature of multiple regions of the reflow oven, conveyor belt speed, thermal formula, buffer queue, dynamic electricity price, carbon emission factor and power margin are represented as graph node features; Embed a meta-learning module in a multi-agent system to infer the physical context and predict the actions of fellow agents by using historical trajectories; A centralized evaluator based on the Transformer self-attention mechanism is adopted to dynamically decompose the joint value and calculate the individual advantage function; Based on the GMeta-MATD3 architecture, a multi-agent strategy is trained under the centralized training and distributed execution paradigm to generate scheduling decisions for work order assignment, equipment routing, and hot recipe selection. Execute scheduling decisions and issue control commands, collect execution feedback data to update primitive reinforcement learning models and multi-objective scheduling strategies.

2. The active energy and carbon sensing production scheduling method for SMT production lines based on primitive reinforcement learning according to claim 1, characterized in that, The multi-flow graph convolutional network models the SMT production line as a directed multi-relationship graph, with nodes including SMT equipment, buffers, and microgrid nodes, and edges including production material flow edges and grid power flow edges; the hidden representation update formula of the (l+1)th layer of the multi-flow graph convolutional network is: in, To incorporate the normalized adjacency matrix of self-loops, To generate the edge matrix, For the power grid side matrix, It is the identity matrix. For degree matrix, Let l be the trainable weight matrix of the l-th layer. For activation function, This represents the embedding representation of the l-th layer node.

3. The active energy and carbon sensing production scheduling method for SMT production lines based on primitive reinforcement learning according to claim 1, characterized in that, The state space of the semi-Markov decision process includes a job feature matrix, a machine feature matrix, and a dynamic environment matrix. The action space is a set of macro actions including work order assignment actions, equipment routing actions, and thermal recipe control actions. The reward function is a weighted negative penalty for manufacturing cycle, energy cost, carbon emissions, delay, and grid peak violation.

4. The active energy and carbon sensing production scheduling method for SMT production lines based on primitive reinforcement learning according to claim 1, characterized in that, The scheduling decision includes clustering work orders with the same or similar thermal curves, so that the work orders pass through the reflow oven continuously with the minimum feed gap to reduce heat transfer energy consumption; when the peak of time-of-use electricity price or high carbon emission window is predicted, the agent performs an active waiting action to transfer the high energy consumption heat transfer to the low time-of-use electricity price or low carbon emission period.

5. The active energy and carbon sensing production scheduling method for SMT production lines based on primitive reinforcement learning according to claim 1, characterized in that, The meta-learning module includes a historical network, a state network, and a prediction network. The historical network encodes features of past trajectories to generate historical embedding representations. The state network infers the current physical context based on recent trajectories. The prediction network integrates local observations, historical embeddings, and state embeddings to deterministically estimate the next action of the companion agent.

6. The active energy and carbon sensing production scheduling method for SMT production lines based on primitive reinforcement learning according to claim 1, characterized in that, The Transformer-enhanced centralized evaluator concatenates each agent's local observations, current actions, historical embeddings, state embeddings, and predicted peer actions into a feature vector, and calculates the marginal contribution of each agent by stacking self-attention layers; the temporal difference objective in the target network uses deterministic actions predicted by the meta-learning module to replace randomly sampled actions in order to reduce variance.

7. The active energy and carbon sensing production scheduling method for SMT production lines based on primitive reinforcement learning according to claim 1, characterized in that, The thermodynamic transition time of the reflow oven in multiple zones is the maximum value of the transition time of each independent temperature control zone, and the specific calculation formula is as follows: in, The transition time of the z-th temperature control zone when the reflow oven l switches from the i-th operation to the j-th operation. Let z be the heating rate of the z-th region. Let z be the cooling rate of the z-th region. The target temperature required for operation i in the z-th zone of reflow oven l. The target temperature required for operation j in the z-th zone of reflow oven l.

8. The active energy and carbon sensing production scheduling method for SMT production lines based on primitive reinforcement learning according to claim 1, characterized in that, The reward function includes an energy cost penalty, a carbon emission penalty, a delay penalty, and a grid peak penalty; the energy cost and carbon emission are obtained by integrating the operating power, transition power, and standby power over the scheduling period.

9. The active energy and carbon sensing production scheduling method for SMT production lines based on primitive reinforcement learning according to claim 1, characterized in that, The decision-making time of the semi-Markov decision process is triggered by discrete events, including the completion of machine tasks and the transition to idle time, the arrival of urgent orders, and the reaching of the target set value in the slowest temperature zone of the reflow oven.

10. A proactive energy and carbon sensing production scheduling system for SMT production lines based on primitive reinforcement learning, characterized in that, include: The multi-source information sensing module is configured to collect production data, thermal data, microgrid data, and external environmental data from the SMT production line, and perform data preprocessing. The production-thermal-electric carbon coupling modeling module is configured to construct a multi-objective scheduling model that includes discrete production logic and continuous thermodynamic constraints, and to reconstruct the scheduling problem as a semi-Markov decision process. The primitive reinforcement learning decision module is configured to extract spatiotemporal topological features based on a multi-flow graph convolutional network, infer the physical context and predict peer actions through the meta-learning module, and use a Transformer-enhanced centralized evaluator for multi-agent credit allocation and policy optimization. The execution and feedback closed-loop module is configured to send scheduling decisions to the SMT production line equipment controller, monitor the operating status, and provide feedback to optimize the scheduling strategy.