Intelligent decision system for energy consumption optimization and safety production accident prevention in smelting industry

CN122549948APending Publication Date: 2026-08-11FUJIAN METALLURGICAL IND DESIGN INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0002]冶炼行业是能源密集型产业,能源介质种类多、管网拓扑复杂、工序耦合性强,能源放散与短缺现象并存,能源利用效率长期偏低

Benefits of technology

[0015]本发明通过构建能流-风险流双向耦合网络,首次将能源流动与风险传播在同一图结构下统一建模,打破了传统能源管理与安全管理割裂的局面,为能效与安全的协同分析提供了结构化数据基础;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122549948A_ABST
    Figure CN122549948A_ABST
Patent Text Reader

Abstract

This invention relates to the field of energy efficiency optimization and accident prevention technology in the smelting industry, and particularly to an intelligent decision-making system for energy consumption optimization and safety production accident prevention in the smelting industry. The system includes a multi-source production and risk data perception module, an energy flow-risk flow bidirectional coupling network construction module, an energy efficiency-safety joint evolution prediction and management cost accounting module, and a multi-objective resource scheduling decision-making module. The system constructs a bidirectional coupling network of energy flow and risk propagation, uses a spatiotemporal graph attention network to predict the probability of energy supply-demand imbalance and the risk diffusion index, and uses this to construct a comprehensive objective function that integrates energy costs and safety accident losses. A reinforcement learning method incorporating a dynamic shrinkage mechanism of the safety boundary is employed to generate a safety-energy efficiency collaborative decision-making scheme. This invention achieves collaborative decision-making for energy efficiency optimization and accident prevention, significantly reducing energy waste and accident risks, and improving the scientific nature and robustness of decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of energy efficiency optimization and accident prevention technology in the metallurgical industry, specifically to an intelligent decision-making system for energy consumption optimization and safety production accident prevention in the metallurgical industry. Background Technology

[0002] The smelting industry is an energy-intensive industry with a wide variety of energy media, complex pipeline topology, and strong process coupling. Energy release and shortage coexist, and energy utilization efficiency has been low for a long time. At the same time, the smelting production environment has inherent hazards such as high-temperature melting, flammable and explosive materials, and toxic gases. Equipment deterioration and fluctuations in the concentration of environmental hazards can easily lead to safety accidents.

[0003] Currently, energy management systems and safety management systems in smelting enterprises have long operated independently: energy dispatch is driven by economic costs and lacks dynamic consideration of safety risks; safety control is mainly based on threshold alarms and manual responses, making it difficult to predict the trend of risk spread and its feedback constraints on energy dispatch.

[0004] In addition, existing scheduling decision-making methods mostly rely on static rules or single-objective optimization, which cannot achieve a dynamic balance between energy efficiency improvement and accident prevention. They also lack a closed-loop self-optimization mechanism based on actual execution feedback, resulting in poor on-site adaptability of decision-making schemes and difficulty in coping with complex and ever-changing production conditions.

[0005] Therefore, there is an urgent need for an intelligent decision-making system that can deeply integrate energy flow and risk flow information and achieve synergistic optimization of energy efficiency and safety. Summary of the Invention

[0006] The purpose of this invention is to provide an intelligent decision-making system for optimizing energy consumption and preventing safety accidents in the smelting industry, so as to solve the problems mentioned in the background art.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] An intelligent decision-making system for energy consumption optimization and safety accident prevention in the smelting industry, including:

[0009] The multi-source production and risk data perception module is used to acquire real-time energy medium consumption data, material flow data, equipment operating status data, and environmental safety monitoring data for the smelting process.

[0010] The energy flow-risk flow bidirectional coupling network construction module constructs an energy flow directed graph containing energy nodes and material nodes based on real-time energy medium consumption data and material flow data; it constructs a risk propagation directed graph containing hazard source nodes and risk-bearing body nodes based on equipment operation status data and environmental safety monitoring data; and it establishes cross-graph coupling edges by identifying the physical spatial mapping relationship and energy release association between energy nodes in the energy flow directed graph and hazard source nodes in the risk propagation directed graph, thereby generating an energy flow-risk flow bidirectional coupling network.

[0011] The energy efficiency-safety joint evolution prediction and management cost accounting module inputs the bidirectional coupling network of energy flow and risk flow in the current time window into a pre-trained spatiotemporal graph attention network to predict the probability of medium supply and demand imbalance of each energy node and the risk diffusion index of each hazard source node in the future scheduling cycle; calculates the energy release and shortage penalty cost based on the medium supply and demand imbalance probability, and calculates the expected loss cost of safety accidents based on the risk diffusion index, and constructs a comprehensive scheduling management objective function that includes the energy release and shortage penalty cost and the expected loss cost of safety accidents;

[0012] The multi-objective resource scheduling decision module aims to minimize the comprehensive scheduling management objective function. It uses the pipeline safety pressure threshold and the permissible concentration of hazardous sources for each energy medium as initial constraints, constructing a Markov decision process model. During the solution process, a dynamic shrinkage mechanism for the safety boundary is introduced: when the predicted risk diffusion index exceeds a preset warning threshold, the scheduling action space of the corresponding energy medium is dynamically shrunk according to the risk diffusion gradient, and the weight coefficient of the expected loss cost of safety accidents in the comprehensive scheduling management objective function is increased. A multi-objective reinforcement learning algorithm is used to iteratively optimize within the dynamically shrunk action space, generating a safety-energy efficiency collaborative decision scheme that includes cross-process allocation instructions for energy mediums and capacity downgrading scheduling instructions for high-risk processes.

[0013] The decision-making scheme distribution and closed-loop feedback module is used to output the safety-energy efficiency collaborative decision-making scheme to the smelting manufacturing execution system or energy management system for scheduling and execution, and to obtain the actual energy efficiency indicators and safety status indicators after execution, which are used to update the model parameters of the spatiotemporal graph attention network and the multi-objective reinforcement learning algorithm.

[0014] As can be seen from the technical solution provided by the present invention above, the intelligent decision-making system for energy consumption optimization and safety accident prevention in the smelting industry provided by the present invention has the following beneficial effects:

[0015] This invention constructs a bidirectional coupled network of energy flow and risk flow, and for the first time models energy flow and risk propagation in a unified manner under the same graph structure, breaking the traditional separation between energy management and safety management, and providing a structured data foundation for the collaborative analysis of energy efficiency and safety;

[0016] This invention utilizes a spatiotemporal graph attention network to capture the joint spatiotemporal evolution of energy flow and risk flow, which can accurately predict the probability of medium supply and demand imbalance and risk diffusion index within future scheduling cycles, thereby improving the foresight and accuracy of the prediction and providing a reliable basis for advanced scheduling decisions.

[0017] This invention integrates the costs of energy release and shortage penalties and the expected losses from safety accidents into the overall scheduling management objective function. Through a dynamic shrinkage mechanism of the safety boundary, it automatically tightens the scheduling action space and increases the weight of safety costs when the risk increases, thereby achieving a dynamic balance between energy efficiency optimization and accident prevention, and reducing overall operating costs and safety risks.

[0018] This invention employs multi-objective reinforcement learning to solve Markov decision process models, balancing short-term gains and long-term cumulative rewards in continuous decision-making. The generated scheduling scheme has both energy-saving and efficiency-enhancing objectives and safety assurance, exhibiting stronger adaptability and optimization depth compared to traditional rule-based scheduling methods.

[0019] This invention introduces a decision-making scheme distribution and closed-loop feedback module, which backpropagates the deviation between the actual execution result and the predicted value to the prediction model and decision-making strategy for online updates, forming a complete closed loop of perception-prediction-decision-execution-feedback, enabling the model to continuously adapt to changes in working conditions and continuously improve decision-making accuracy and on-site executability. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the intelligent decision-making system for optimizing energy consumption and preventing safety accidents in the smelting industry. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0022] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific embodiments.

[0023] like Figure 1 As shown, this embodiment of the invention provides an intelligent decision-making system for energy consumption optimization and safety accident prevention in the smelting industry, including:

[0024] The multi-source production and risk data perception module is used to acquire real-time energy medium consumption data, material flow data, equipment operating status data, and environmental safety monitoring data for the smelting process.

[0025] The energy flow-risk flow bidirectional coupling network construction module constructs an energy flow directed graph containing energy nodes and material nodes based on real-time energy medium consumption data and material flow data; it constructs a risk propagation directed graph containing hazard source nodes and risk-bearing body nodes based on equipment operation status data and environmental safety monitoring data; and it establishes cross-graph coupling edges by identifying the physical spatial mapping relationship and energy release association between energy nodes in the energy flow directed graph and hazard source nodes in the risk propagation directed graph, thereby generating an energy flow-risk flow bidirectional coupling network.

[0026] The energy efficiency-safety joint evolution prediction and management cost accounting module inputs the bidirectional coupling network of energy flow and risk flow in the current time window into a pre-trained spatiotemporal graph attention network to predict the probability of medium supply and demand imbalance of each energy node and the risk diffusion index of each hazard source node in the future scheduling cycle; calculates the energy release and shortage penalty cost based on the medium supply and demand imbalance probability, and calculates the expected loss cost of safety accidents based on the risk diffusion index, and constructs a comprehensive scheduling management objective function that includes the energy release and shortage penalty cost and the expected loss cost of safety accidents;

[0027] The multi-objective resource scheduling decision module aims to minimize the overall scheduling management objective function. It uses the pipeline safety pressure threshold and the permissible concentration of hazardous sources for each energy medium as initial constraints, constructing a Markov decision process model. During the solution process, a dynamic shrinkage mechanism for the safety boundary is introduced: when the predicted risk diffusion index exceeds a preset warning threshold, the scheduling action space of the corresponding energy medium is dynamically shrunk according to the risk diffusion gradient, and the weight coefficient of the expected loss cost of safety accidents in the overall scheduling management objective function is increased. A multi-objective reinforcement learning algorithm is used to iteratively optimize within the dynamically shrunk action space, generating a safety-energy efficiency collaborative decision scheme that includes cross-process allocation instructions for energy mediums and capacity downgrading scheduling instructions for high-risk processes.

[0028] In this embodiment, the multi-source production and risk data perception module serves as the data perception foundation layer of the intelligent decision-making system for energy consumption optimization and safety production accident prevention in the smelting industry. This module deploys multiple types of sensor groups at key energy flow nodes, material flow nodes, equipment operating positions, and environmentally sensitive areas in the smelting process to achieve comprehensive collection, preprocessing, feature extraction, and standardized encapsulation of four types of heterogeneous data: energy medium consumption, material flow, equipment operating status, and environmental safety monitoring. The module normalizes the sampling rate and suppresses cross-channel crosstalk in the raw signals, eliminating electromagnetic coupling interference and completing time-series alignment. It then performs multi-parameter joint analysis and dynamic calibration on various raw sequences to extract structured features. Finally, through multi-source cross-validation and spatiotemporal consistency checks, it removes abnormal feature values ​​and outputs high-quality, uniformly formatted real-time energy medium consumption data, material flow data, equipment operating status data, and environmental safety monitoring data. This provides reliable data support for the construction of the energy flow-risk flow bidirectional coupling network and subsequent intelligent decision-making, specifically including:

[0029] Multi-sensor signal acquisition and synchronization unit: Energy metering sensor groups, including flow meters, pressure transmitters, and temperature sensors, are deployed at key energy flow nodes in the smelting process to collect flow, pressure, and temperature signals of the energy medium, respectively; material tracking sensor groups, including weighing sensors, RFID tag readers, and material composition analyzers, are deployed at material flow nodes to collect material quality, batch identification, and composition information, respectively; operating condition monitoring sensor groups, including vibration sensors, current transformers, and temperature sensors, are deployed at equipment operating stations to collect vibration, current, and temperature signals of the equipment, respectively; environmental monitoring sensor groups, including combustible gas detectors, toxic gas detectors, and dust concentration sensors, are deployed in environmentally sensitive areas to collect parameters such as hazardous gas concentration and dust concentration, respectively; each sensor group synchronously acquires raw signals at a fixed sampling frequency, and a high-precision timestamp is added to each sampling point to obtain a heterogeneous multi-source raw signal set with timestamps;

[0030] Since different types of sensors have different sampling frequencies, the sampling rate of the heterogeneous multi-source raw signal set needs to be normalized to be unified to a preset standard sampling frequency, which facilitates subsequent time-series alignment and fusion analysis. A resampling algorithm based on cubic spline interpolation is adopted, the mathematical expression of which is: ,in, For the resampled signal in time The value at; For the original sampled signal at the 1st The value at that location, This is the original sampling period; This represents the original number of sampling points; For cubic spline interpolation kernel function; The target sampling period;

[0031] After sampling rate normalization, a blind source separation algorithm based on independent component analysis is used to suppress crosstalk between multi-channel signals and eliminate electromagnetic coupling interference between sensor cables; the separation model for crosstalk suppression is expressed as: ,in, The observed signal matrix consists of the normalized signals from each channel; The separation matrix is ​​estimated using a fixed-point iterative algorithm that maximizes non-Gaussianity; The result is the matrix of independent source signals after separation; the separated signals are remapped to the original channel positions to obtain the signals after crosstalk removal.

[0032] Using the standard clock provided by the central time service system as a reference, the time alignment of each signal after crosstalk removal is performed; a two-way time comparison method is used to calculate the time offset of each channel signal, and time shift correction is used to align each signal to a unified time axis; for time jitter caused by transmission delay or buffering, a dynamic time warping algorithm is used to minimize the alignment error; finally, the original energy time sequence, the original material time sequence, the original operating condition time sequence, and the original environmental time sequence under a unified time sequence reference are obtained.

[0033] Multi-parameter joint analysis and feature extraction unit: Performs multi-parameter joint analysis of flow rate, pressure, and temperature on the original energy time series; since the density of the energy medium is significantly affected by temperature and pressure, dynamic measurement calibration is required through the equation of state to achieve accurate conversion from volume to mass; for gaseous media, the following calibration formula is used: ,in, For mass flow rate; Volumetric flow rate; The absolute pressure of the medium; The molar mass of the medium; It is the compression factor; This is the universal gas constant; The absolute temperature of the medium is used; for liquid media, a temperature compensation formula is used for density correction; through the above calibration, the original energy time series is transformed into an energy structured characteristic series containing flow rate, pressure, temperature and calibrated mass flow rate;

[0034] Batch traceability processing is performed on the original time sequence of materials. By reading the batch identification information from the RFID tag, the material flow data is associated with the production batch information to establish a full-process traceability chain from raw materials to finished products. At the same time, the composition information collected by the material composition analyzer is used to map the composition of the materials and generate a structured feature sequence of materials that includes quality, batch identification and content of major components.

[0035] Operating mode identification is performed on the original time series of operating conditions; an energy feature extraction method based on wavelet packet decomposition is adopted to decompose the equipment vibration signal into multiple frequency bands, calculate the energy proportion of each frequency band, and construct the frequency domain feature vector of the equipment operating state; at the same time, based on the equipment current signal and temperature signal, the load characteristics and temperature rise characteristics of the equipment are extracted, and by comparing with the baseline characteristics under normal operating conditions, the degradation trend and abnormal mode of the equipment are identified, resulting in a structured feature sequence of operating conditions containing vibration characteristics, current characteristics and temperature characteristics;

[0036] Multi-factor concentration gradient analysis was performed on the original time series of the environment; based on the concentration data of combustible gas, toxic gas and dust at each monitoring point, the temporal gradient and spatial gradient of the concentration were calculated to identify the concentration change trend and diffusion direction; at the same time, combined with the meteorological data and ventilation conditions of the plant area, spatial interpolation was performed on the concentration distribution to generate environmental hazard concentration distribution information, and an environmental structured feature sequence containing concentration value, gradient value and spatial distribution was obtained.

[0037] Multi-source cross-validation and data encapsulation unit: Performs multi-source cross-validation on energy structured feature sequences, material structured feature sequences, operating condition structured feature sequences, and environmental structured feature sequences; verifies data consistency by establishing physical correlation models between different data sources; specifically including: verifying the matching degree between energy metering data and the theoretical heat required for material heating based on the energy consumption logic during material heating; verifying the consistency between operating condition data and energy data based on the relationship between equipment operating power and energy medium consumption; for data anomalies identified during validation, a multivariate anomaly detection algorithm based on Mahalanobis distance is used for location and removal, the calculation formula of which is: ,in, For sample vectors Mahalanobis distance; This is the mean vector of the normal data samples; This is the inverse of the covariance matrix of the normal data samples; when the Mahalanobis distance exceeds a preset threshold, the sample is marked as an abnormal feature value and removed.

[0038] The data that passed cross-validation were subjected to spatiotemporal consistency checks. In the time dimension, the continuity of timestamps and the uniformity of sampling intervals of each data sequence were checked, and missing time segments were marked and interpolated. In the spatial dimension, the matching relationship between the data collected by each sensor and the corresponding spatial location was checked to ensure the accuracy of the correspondence between data and physical location. A Kalman filter algorithm based on spatiotemporal constraints was used to smooth the data to eliminate measurement noise and random fluctuations.

[0039] The data, after cross-validation and spatiotemporal consistency verification, is standardized and encapsulated; the data fields of various structured feature sequences are uniformly defined, and serialization and encapsulation are performed using standard JSON or Protocol Buffers data formats, with the addition of data source identifiers, timestamp ranges, and quality level labels, generating real-time energy medium consumption data, material flow data, equipment operating status data, and environmental safety monitoring data that can be directly called by upper-level modules.

[0040] In this embodiment, the energy flow-risk flow bidirectional coupling network construction module is the core modeling layer of the intelligent decision-making system for energy consumption optimization and safety production accident prevention in the smelting industry. This module receives real-time energy medium consumption data, material flow data, equipment operating status data, and environmental safety monitoring data output by the multi-source production and risk data perception module, and constructs an energy flow directed graph representing the energy-material flow relationship and a risk propagation directed graph representing the influence relationship between the hazard source and the risk-bearing body, respectively. Then, by identifying the physical spatial mapping relationship and energy release association between nodes in the two types of graphs, cross-graph coupling edges are established, ultimately generating a unified energy flow-risk flow bidirectional coupling network. This network provides a complete graph structure data foundation for subsequent energy efficiency-safety joint evolution prediction and multi-objective resource scheduling decisions, enabling energy scheduling decisions to simultaneously consider energy efficiency improvement goals and safety risk constraints, achieving deep fusion modeling of the two physical processes, specifically including:

[0041] The directed energy flow graph construction unit classifies energy medium types based on the real-time energy medium consumption data output by the multi-source production and risk data perception module. According to medium attributes, energy media are categorized into types such as electricity, natural gas, coke oven gas, blast furnace gas, converter gas, steam, compressed air, and industrial water. Simultaneously, based on the smelter plant's pipeline network geographic information system and pipeline connection matrix, a breadth-first search algorithm is used to analyze the pipeline topology, identifying pipeline branch relationships, pipe diameters, and transmission paths from the energy supply inlet to the energy consumption terminals of each process. This generates an energy transmission topology containing pipeline connection relationships, flow direction, and pipeline attribute labels, and outputs the classification results for each type of energy medium.

[0042] Material batch identification is performed on material flow data. By parsing the batch code in the RFID tag data and combining it with the production plan information in the manufacturing execution system, the source process, target process, and material type of each batch of materials are determined. At the same time, based on the material flow sequence between each process in the smelting process, the transfer path of materials between processes such as raw material yard, sintering, ironmaking, steelmaking, continuous casting, and steel rolling is analyzed to generate a material transfer topology containing process node connection relationships and transfer direction information, and the material batch identification results are output.

[0043] Spatiotemporal feature aggregation processing is performed on the energy transmission topology and energy medium classification results. The spatial coordinates of each pipeline node or equipment node, the type of medium it carries, and the dynamic flow attributes obtained by the multi-source production and risk data perception module are spliced ​​and normalized to generate an energy node feature vector, whose dimensions include spatial location components, medium type coding components, and dynamic flow statistics components. Spatiotemporal feature aggregation processing is also performed on the material transfer topology and material batch identification results. The spatial coordinates of each material transfer node, batch identification code, and dynamic quality attributes are spliced ​​and normalized to generate a material node feature vector, whose dimensions include spatial location components, batch identification code components, and dynamic quality statistics components.

[0044] Based on the energy transmission topology and energy flow direction, directed edges for energy transmission are constructed between the feature vectors of energy nodes with direct pipeline connections, with the direction consistent with the medium flow direction. The edge weight of each directed edge for energy transmission is determined by a combination of transmission loss and delay time, and is calculated using the following formula: ,in, To start from energy nodes Pointing to energy nodes The edge weights of directed edges for energy transmission; This is the transmission loss weighting coefficient; For nodes With nodes The length of the pipe between them; This is the friction coefficient of the pipeline; The nominal diameter of the pipe; This is the weighting coefficient for delay time; The average flow velocity of the medium in the pipe;

[0045] Based on the material transfer topology and material flow direction, directed edges for material transfer are constructed between the feature vectors of material nodes with direct material transfer relationships, with the direction consistent with the material flow direction. The edge weight of each directed edge for material transfer is determined by a combination of two dimensions: mass loss and conversion time, and is calculated using the following formula: ,in, From the material node Pointing to material node The material transfer has the edge weight of the directed edge; This is the weighting coefficient for quality loss; For materials from nodes Transfer to node Quality loss rate during the process; This is the conversion time weighting coefficient; For nodes With nodes The transfer distance between them; This refers to the average speed of material transfer.

[0046] Based on the energy-material consumption correspondence between the feature vectors of energy nodes and the feature vectors of material nodes, energy-driven directed edges are constructed from the feature vectors of energy nodes to the feature vectors of material nodes. This correspondence is determined by the process mechanism of the smelting process: the material heating or melting process of a certain process requires the consumption of a specific type of energy medium. The edge weight of each energy-driven directed edge is the energy consumption intensity, calculated using the following formula: ,in, To start from energy nodes Pointing to material node The edge weight of the energy-driven directed edge; For material nodes The energy consumed by the corresponding process within the statistical period is determined by the energy nodes. The total amount of energy supplied; Material nodes within the same statistical period The total amount of materials processed in the corresponding process;

[0047] All energy node feature vectors, all material node feature vectors, all energy transmission directed edges, all material transfer directed edges, and all energy-material driven directed edges are aggregated into a unified graph structure, which is jointly stored using an adjacency matrix and a node attribute matrix to generate an energy flow directed graph.

[0048] Risk propagation directed graph construction unit: identifies equipment degradation patterns from equipment operating status data; inputs vibration features, current features, and temperature features from the structured feature sequence of operating conditions into a pre-trained convolutional autoencoder network, and identifies the degree to which the equipment status deviates from normal operating conditions through reconstruction error analysis;

[0049] The pre-training process of a convolutional autoencoder network is as follows:

[0050] Training data construction: During periods when the smelting equipment is in normal operation and no signs of deterioration are confirmed by professional inspectors, vibration, current, and temperature signals of the equipment are continuously collected using a set of operating condition monitoring sensors. Vibration signals are collected using an accelerometer at a sampling frequency of 10 kHz, with each acquisition window of one second serving as a vibration signal sample. Current and temperature signals are collected at the original sampling frequency of their respective sensors, and the average value within the same time window is taken as the current and temperature characteristic values ​​for that window. The collected vibration signal samples are then subjected to Fast Fourier Transform (FFT). The first 512 frequency components of the amplitude spectrum are taken as the vibration frequency domain feature vector by the Lie transform. The current feature value, temperature feature value and vibration frequency domain feature vector are concatenated to form a working condition feature vector with dimension 514. At least 10,000 working condition feature vectors under normal working conditions are collected to form a normal sample training set. In addition, at least 1,000 working condition feature vectors of the equipment under four known degradation modes, including bearing wear, rotor imbalance, gear pitting and winding insulation aging, are collected to form a validation set, which is used to determine the reconstruction error threshold and evaluate the detection sensitivity of the model for various degradation modes.

[0051] Network Structure: The encoder consists of three stacked one-dimensional convolutional layers and two fully connected layers. The first convolutional layer has a kernel size of 7, a stride of 2, and 32 output channels; the second convolutional layer has a kernel size of 5, a stride of 2, and 64 output channels; the third convolutional layer has a kernel size of 3, a stride of 1, and 128 output channels. Each convolutional layer is followed by a batch normalization layer and a modified linear unit activation function. The feature maps output from the convolutional layers are flattened and then input into two fully connected layers with 256 and 64 hidden units respectively. Each fully connected layer is followed by a modified linear unit activation function. The final output of the encoder is a one-dimensional... The latent space feature vector has a degree of 16. The decoder structure is mirror-symmetric to the encoder: first, two fully connected layers map the latent space feature vector back to the feature space of the same dimension as the output of the encoder's convolutional layer, with 64 and 256 hidden units in the fully connected layers, respectively; then, three transposed convolutional layers gradually restore the original input dimension, with the kernel size, stride, and number of output channels of each transposed convolutional layer being consistent with the corresponding convolutional layer in the encoder; followed by batch normalization layers and modified linear unit activation functions; the dimension of the decoder output layer is consistent with the dimension of the input condition feature vector, the activation function is linear activation, and the output is the reconstructed condition feature vector.

[0052] Loss function: The mean squared error between the input working condition feature vector and the reconstructed working condition feature vector is used as the loss function. ,in, To reconstruct the mean squared error loss; Batch size; The dimension of the working condition feature vector is set to 514. For the first The first sample Original eigenvalues; For the first The first sample 3D reconstruction of eigenvalues;

[0053] Training strategy: The Adam optimizer is used for training, with an initial learning rate of 0.001, a batch size of 128, and a maximum number of training epochs of 100. An early stopping strategy is employed during training: training is terminated if the reconstruction error on the validation set decreases for five consecutive training epochs. After training, the reconstruction error distribution of all normal training samples is calculated, and the upper 99th percentile of this distribution is taken. As a preset baseline threshold; during the equipment degradation pattern recognition stage, the operating condition feature vector of the equipment to be detected is input into the trained convolutional autoencoder network to calculate its reconstruction error; if the reconstruction error exceeds If the equipment deviates from normal operating conditions, the feature decomposition method will be used to further analyze the degradation mode type.

[0054] The specific operation of the feature decomposition method is as follows: the cosine similarity of the working condition feature vector of the equipment to be tested with the typical fault feature vectors of various known degradation modes in the normal sample training set is compared, and the degradation mode with the highest cosine similarity is selected as the degradation mode type of the equipment; if the highest cosine similarity is less than 0.5, the degradation mode is marked as an unknown abnormal type and the manual diagnosis process is triggered; the degradation degree score is obtained by mapping the ratio of reconstruction error to preset benchmark threshold to the zero to one interval through sigmoid normalization, and the closer the score is to one, the more severe the degradation degree.

[0055] For equipment whose reconstruction error exceeds a preset benchmark threshold, the feature decomposition method is further used to extract fault feature vectors and determine the type of degradation mode, including bearing wear, rotor imbalance, gear pitting, and winding insulation aging; finally, a sequence of equipment health status containing equipment identification, degradation mode type, and degradation degree score is generated.

[0056] Multi-factor concentration gradient analysis was performed on environmental safety monitoring data. Taking the concentrations of combustible gases, toxic gases, and dust as the analysis objects, the concentration-time change rate of each monitoring point within a continuous time window was calculated. Based on the spatial distribution of multiple monitoring points, the spatial distribution of concentrations in the entire plant area was estimated using the Kriging interpolation method. At the same time, combined with the wind speed, wind direction, and atmospheric stability information in the plant area's meteorological data, a Gaussian plume diffusion model was used to predict the diffusion trend of hazardous substances in the future period, generating an environmental hazard concentration distribution sequence containing a concentration distribution grid and a diffusion trend vector.

[0057] The equipment health status sequence is used to determine the equipment risk type; based on the degradation mode type and degradation degree score, equipment with degradation degree exceeding the high-risk threshold is identified as equipment-related hazard sources, which pose a risk of accidental energy release due to failure; other equipment, facilities or areas within the scope of the equipment-related hazard sources are identified as equipment-related risk bodies; the environmental hazard concentration distribution sequence is used to determine the environmental risk type, areas with flammable or toxic gas concentrations exceeding the safety threshold are identified as environmental hazard sources, and sensitive areas such as personnel activity areas, control rooms and evacuation routes are identified as personnel exposure risk bodies;

[0058] Hazard sources of equipment and environment are fused together; the degree of equipment deterioration, concentration of hazardous substances, and potential energy release intensity are quantified and spliced ​​to generate a unified hazard source feature vector, the dimensions of which include deterioration score component, concentration component, energy release potential component, and spatial location component; Hazard-bearing bodies of equipment and personnel exposure areas are fused together; entity type coding, vulnerability score, and exposure duration statistics are quantified and spliced ​​to generate a unified hazard-bearing body feature vector, the dimensions of which include entity type component, vulnerability score component, exposure duration component, and spatial location component.

[0059] Based on the spatial distance relationship between the hazard source feature vector and the hazard-bearing body feature vector, the airflow diffusion direction, and the energy release propagation path, a directed edge for risk propagation is constructed from the hazard source feature vector to the hazard-bearing body feature vector. The edge weight of each directed edge for risk propagation is determined by a combination of three dimensions: propagation delay time, attenuation coefficient, and exposure intensity, and is calculated using the following formula: ,in, To start from the hazardous source node Pointing to the risk-bearing node The risk propagation has the edge weight of the directed edge; This is the propagation delay time weighting coefficient; Hazardous source node With risk-bearing node Spatial distance between them; The speed of propagation in the medium of risk transmission is determined by the type of propagation, such as the speed of sound, airflow speed, or heat conduction speed. This is the attenuation coefficient weighting coefficient; This represents the energy attenuation ratio along the propagation path; Exposure intensity weighting coefficient; For risk-bearing nodes At the hazardous source node Evaluation value of exposure intensity under the influence;

[0060] All hazard source feature vectors, all risk-bearing body feature vectors, and all risk propagation directed edges are aggregated into a unified graph structure, which is jointly stored using an adjacency matrix and a node attribute matrix to generate a risk propagation directed graph.

[0061] Cross-graph coupling and bidirectional network generation unit: Extracts spatial coordinates of all energy nodes in the directed energy flow graph, parses spatial position components from the feature vectors of energy nodes, and generates a set of spatial coordinates of energy nodes; extracts spatial coordinates of all hazard source nodes in the directed risk propagation graph, parses spatial position components from the feature vectors of hazard sources, and generates a set of spatial coordinates of hazard source nodes; calculates the spatial distance between the two sets of spatial coordinates, and uses a KD-tree-based nearest neighbor search algorithm to identify pairs of energy nodes and hazard source nodes whose spatial distance is less than a preset co-location threshold, forming a candidate set of co-location node pairs;

[0062] The physical space mapping relationship of the candidate set of co-located nodes is verified; for each co-located node pair, the plant equipment and facility layout diagram and piping and instrumentation diagram are retrieved to confirm whether the pipeline or equipment corresponding to the energy node and the equipment or area corresponding to the hazard source node are indeed close to or co-located in physical space; the co-located node pairs that pass the verification are marked as valid physical space mapping relationships.

[0063] Based on the energy medium type and hazard source type of co-located node pairs, an energy release potential analysis is performed on the effective physical space mapping relationship to quantify the energy impact intensity that hazard source nodes may generate on energy nodes; the energy release potential is calculated using the following formula: ,in, Energy release intensity; The density of the medium transported by the energy node; The volume of the medium that may be ignited or detonated within the influence range of the hazardous source node; The calorific value of the medium. This is the ignition probability factor;

[0064] Simultaneously, a risk feedback impact analysis is conducted based on the influence of the risk propagation intensity of hazardous source nodes on energy transmission paths, quantifying the degree of feedback constraint on energy transmission after a hazardous event occurs; the risk feedback intensity is calculated using the following formula: ,in, As for the intensity of risk feedback; This is the feedback intensity scaling factor; Hazardous source node The set of risk-bearing nodes connected by the directed edges of risk propagation; To start from the hazardous source node Pointing to the risk-bearing node The risk propagation has the edge weight of the directed edge; For risk-bearing nodes Sensitivity coefficient to dependence on energy transmission routes;

[0065] Based on the effective physical space mapping relationship, a forward-coupled directed edge is constructed from the energy node to the hazard source node, with the energy release intensity as the edge weight. This represents the energy supply and risk amplification effect of the flammable and explosive medium carried by the energy node on the hazard source node. At the same time, a reverse-coupled directed edge is constructed from the hazard source node to the energy node, with the risk feedback intensity as the edge weight. This represents the safety threat and scheduling constraint effect of the hazard source node on the energy node after an accident. All forward-coupled and reverse-coupled directed edges are collectively referred to as cross-graph coupling edges.

[0066] The directed graphs for energy flow, risk propagation, and all cross-graph coupling edges are integrated into a unified network structure to generate a bidirectional coupled network for energy flow and risk flow. This network includes four types of nodes: energy nodes, material nodes, hazard source nodes, and risk-bearing body nodes, as well as five types of directed edges: directed edges for energy transmission, directed edges for material transfer, directed edges for energy and material drive, directed edges for risk propagation, and directed edges for cross-graph coupling. The unified network is stored in a heterogeneous graph data structure and supports indexing and querying by edge type and node type, providing structured graph data for the input of the subsequent spatiotemporal graph attention network.

[0067] In this embodiment, the energy efficiency-safety joint evolution prediction and management cost accounting module is the prediction and quantitative evaluation layer of the intelligent decision-making system for energy consumption optimization and safety production accident prevention in the smelting industry. This module receives the energy flow-risk flow bidirectional coupling network for the current time window generated by the energy flow-risk flow bidirectional coupling network construction module, inputs it into a pre-trained spatiotemporal graph attention network, and simultaneously captures the joint evolution law of energy flow state and risk flow state in the time and spatial dimensions. It predicts the probability of medium supply-demand imbalance at each energy node and the risk diffusion index of each hazardous source node within the future scheduling cycle. Based on the prediction results, the module calculates the energy release and shortage penalty costs from the energy dimension and the expected loss cost of safety accidents from the safety dimension, and weights and merges the two types of costs into a comprehensive scheduling management objective function. It also embeds the pipeline safety pressure threshold and the allowable concentration of hazardous sources as boundary constraints. This module provides a quantitative optimization objective function for subsequent multi-objective resource scheduling decisions, enabling scheduling decisions to achieve a dynamic balance between economic costs and safety risks, specifically including:

[0068] Spatiotemporal graph sequence construction and feature extraction unit:

[0069] For the bidirectional coupling network of energy flow and risk flow in the current time window and the bidirectional coupling network of energy flow and risk flow in multiple historical time windows, time-series slices are sampled according to a preset time step. The preset time step is set according to the scheduling cycle characteristics of the smelting process, usually taking one-tenth to one-fifth of the scheduling cycle. Each time-series slice contains the state characteristics of energy nodes, material nodes, hazard source nodes, risk-bearing body nodes, and cross-graph coupling characteristics at that moment, forming a complete graph structure snapshot. All time-series slices are arranged in chronological order to obtain a spatiotemporal evolution graph sequence, which is formally represented as: ,in, A sequence of spatiotemporal evolution diagrams; For the first A snapshot of the graph structure at each time step; This represents the total number of time steps in a time-series slice.

[0070] For each graph structure in the spatiotemporal evolution graph sequence, independent attention parameter matrices are assigned to five types of directed edges based on an edge type differentiation mechanism; specifically... Assign attention parameter matrices to the directed edges of energy transmission. Assign attention parameter matrices to the directed edges of material transfer. Assign attention parameter matrices to energy-driven directed edges. Assign attention parameter matrices to the directed edges of risk propagation. Attention parameter matrices are assigned to cross-graph coupled directed edges; each parameter matrix has the same dimension, but the parameter values ​​are learned independently to capture the semantic differences of different types of edges;

[0071] For each edge type, multi-head graph attention is calculated based on the edge weight, source node features, and target node features of that type; taking the directed edge for energy transmission as an example, the first... A focus of attention from energy nodes Pointing to energy nodes The formula for calculating attention weights is: ,in, For the first A focus of attention from energy nodes Pointing to energy nodes Intra-graph spatial attention weights; For the first The learnable attention vector parameters corresponding to the directed edges of energy transmission under each attention head; For the first Attention parameter matrix of directed edges for energy transmission under each attention head; For nodes The current feature vector; For nodes The current feature vector; To start from energy nodes Pointing to energy nodes The edge weights of directed edges for energy transmission; For energy nodes The set of all neighboring nodes connected by the directed edge of the energy transmission; This represents a vector concatenation operation;

[0072] The results from multiple attention heads are averaged or concatenated to obtain the spatial attention weight matrix within the graph. Based on this weight matrix, the features of neighboring nodes are weighted and aggregated to obtain the spatial dependency features of nodes at each time step. The update formula for this feature is as follows: ,in, For energy nodes Node spatial dependency characteristics after energy propagation edge type spatial aggregation; The in-graph spatial attention weights are calculated after averaging the long positions. The nonlinear activation function is used; the same computational paradigm is used for the remaining edge types to obtain the node spatial dependency features corresponding to each edge type.

[0073] Spatiotemporal co-evolution prediction unit: Performs temporal attention calculation on the spatial dependency features of nodes at each time step in the spatiotemporal evolution graph sequence; constructs a self-attention mechanism in the temporal dimension to capture the evolution patterns of energy flow state and risk flow state over time; for energy nodes... The formula for calculating the time attention weight across time steps is: ,in, For energy nodes At time step Time step Time attention weights; These are the learnable temporal attention vector parameters; and These are the query mapping matrix and the key mapping matrix, respectively. For energy nodes At time step The node spatial dependency feature; For energy nodes At time step The node spatial dependency features are obtained; the features of all time steps are weighted and aggregated based on the temporal attention weight to obtain the node temporal evolution features.

[0074] Cross-graph feature interaction calculations are performed on the temporal evolution characteristics of each energy node and each hazard source node through cross-graph coupling directed edges. Forward coupling directed edges transmit the feature information of energy nodes to hazard source nodes, quantifying the positive driving effect of energy flow on risk flow; backward coupling directed edges transmit the feature information of hazard source nodes to energy nodes, quantifying the negative constraint effect of risk flow on energy flow. The cross-graph feature interaction calculation formula is as follows: ,in, For nodes Cross-graph interaction features; The positive coupling feature fusion coefficient; To point to the node through positively coupled directed edges The set of all neighboring nodes; For the node Pointing to node The edge weight of the positively coupled directed edge, i.e., the energy release intensity; These are the reverse coupling feature fusion coefficients; To point to a node via a reverse-coupled directed edge The set of all neighboring nodes; For the node Pointing to node The edge weight of the reverse coupling directed edge, i.e., the risk feedback strength;

[0075] The node temporal evolution features and cross-graph interaction features are fused, and the fusion ratio of the two types of features is automatically learned through a gating fusion mechanism to obtain joint spatiotemporal evolution features: ,in, For nodes The joint spatiotemporal evolution characteristics; The gated weight vector is learned from node temporal evolution features and cross-graph interaction features through a fully connected network; For nodes The node time evolution characteristics; This is element-wise multiplication;

[0076] The pre-training process of the spatiotemporal graph attention network is as follows:

[0077] Training Data Construction: Multi-source production and risk data from the smelter area were collected over a continuous historical operating period. Following the processing flow of the energy flow-risk flow bidirectional coupling network construction module, a snapshot of the energy flow-risk flow bidirectional coupling network graph structure for each historical time window was generated. A set of training input samples was constructed from the graph structure snapshots of multiple consecutive historical time windows. The actual media supply-demand imbalance records and actual risk event records occurring within the future scheduling cycle of the last time window of this set of samples were used as monitoring labels. Specifically, for the monitoring labels of the media supply-demand imbalance prediction branch, binary labels were generated based on whether the actual emission and actual shortage of each energy node within the historical scheduling cycle exceeded the preset normal fluctuation range; if they exceeded, they were marked as 1, otherwise as 0. For the monitoring labels of the risk diffusion prediction branch, continuous value labels were generated after normalization based on the number and severity of accidents or near-miss events occurring within the actual impact range of each hazard source node within the historical scheduling cycle, with values ​​ranging from 0 to 1. All training samples were divided into a training set and a validation set according to time sequence, with a division ratio of 8:2.

[0078] Network Structure: The spatiotemporal graph attention network consists of a stacked spatial attention layer, a temporal attention layer, and a cross-graph feature interaction layer. In the spatial attention layer, each of the five types of directed edges is assigned an independent set of attention parameter matrices. Each set of parameter matrices contains four attention heads, each with a dimension of 64. The temporal attention layer uses a scaled dot product self-attention mechanism, with four attention heads, each with a dimension of 64. The cross-graph feature interaction layer performs bidirectional feature transfer through forward-coupled and reverse-coupled directed edges. The gating network in the gating fusion mechanism uses a single-hidden-layer fully connected network with a hidden layer dimension of 32. The medium supply and demand prediction branch consists of three fully connected layers with hidden layer dimensions of 128, 64, and 32 respectively. The last layer has an output dimension of one and is followed by a sigmoid activation function. The risk diffusion prediction branch has the same network structure as the medium supply and demand prediction branch but with independent parameters. The last layer is followed by an exponential activation function.

[0079] Loss function: A multi-task joint loss function is adopted, which is a weighted combination of the medium supply and demand forecasting loss and the risk diffusion forecasting loss; the medium supply and demand forecasting loss adopts a binary cross-entropy loss function. , in, Forecast losses for medium supply and demand; The total number of energy nodes; For energy nodes The true label of supply and demand imbalance; For energy nodes The probability of supply and demand imbalance in the predicted medium;

[0080] The loss for risk diffusion prediction uses the mean squared error loss function: ,in, Predict losses for risk diffusion; This represents the total number of hazardous source nodes. Hazardous source node The predicted risk diffusion index; Hazardous source node The true risk diffusion index label;

[0081] The joint loss function is: ,in, For the joint loss function, For risk prediction loss weighting coefficients, take... ; for Regularization intensity coefficient; These are all the trainable parameters of the spatiotemporal graph attention network;

[0082] Training strategy: The Adam optimizer is used for mini-batch gradient descent training, with an initial learning rate of 0.001 and a batch size of 32. A learning rate decay strategy is employed during training: the learning rate decreases to 0.8 times its current value every 20 training epochs. All training samples are iterated once per training epoch, with a maximum of 200 training epochs. To prevent overfitting, Dropout regularization with a ratio of 0.2 is applied after each spatial attention layer and temporal attention layer. The training termination condition is: if the joint loss on the validation set does not decrease for 15 consecutive training epochs, an early stopping mechanism is triggered, restoring the model parameters with the lowest validation set loss as the final pre-trained weights. After pre-training, the model parameters are saved as a pre-trained weight file for loading during the online inference stage.

[0083] Medium supply and demand imbalance prediction unit: The joint spatiotemporal evolution characteristics of each energy node are input into the medium supply and demand prediction branch; this branch is constructed by a multilayer perceptron, which contains multiple fully connected hidden layers, each followed by batch normalization and modified linear unit activation functions; the multilayer perceptron gradually maps the high-dimensional joint spatiotemporal evolution characteristics to a low-dimensional supply and demand state representation space, and extracts abstract features related to the supply and demand balance of energy medium; the output dimension of the last hidden layer is the preset supply and demand state dimension;

[0084] The output of the multilayer perceptron is mapped to a probability range of zero to one using a sigmoid activation function, outputting the probability of medium supply-demand imbalance at each energy node; the mathematical expression of sigmoid activation is: ,in, For energy nodes The probability of supply and demand imbalance of the medium; For multilayer sensing machines to energy nodes The output value; this probability value reflects the combined probability that the energy node will experience either oversupply leading to release or undersupply leading to shortage within the future scheduling cycle;

[0085] Risk diffusion prediction unit: The joint spatiotemporal evolution characteristics of each hazard source node are input into the risk diffusion prediction branch; this branch is also constructed by a multilayer perceptron, whose structure is symmetrical with the network architecture of the medium supply and demand prediction branch, but the parameters are trained independently; the multilayer perceptron gradually maps the joint spatiotemporal evolution characteristics to the risk diffusion state representation space and extracts abstract features related to risk propagation and diffusion trends;

[0086] The output of the multilayer perceptron is mapped using an exponential activation function to output the risk diffusion index of each hazard source node; the mathematical expression of the exponential activation function is: ,in, Hazardous source node The risk diffusion index; The baseline for the spread of basic risks; This is the diffusion scale adjustment coefficient; For multilayer sensing machines to detect hazard source nodes The output value of the risk diffusion index reflects the overall strength of the outward diffusion of dangerous energy or hazardous substances from the dangerous source node and its impact on surrounding risk-bearing bodies within the future scheduling cycle.

[0087] Energy release and shortage penalty cost accounting unit: The probability of supply-demand imbalance at each energy node is quantified by combining the rated supply capacity, real-time inventory level, and historical supply-demand fluctuation characteristics of each energy node; the formula for calculating the supply-demand deviation is: ,in, For energy nodes The supply-demand deviation during the future scheduling cycle, with positive values ​​indicating oversupply and negative values ​​indicating undersupply. For energy nodes The rated supply capacity; It is a capacity utilization adjustment factor; For energy nodes Real-time inventory levels; This is the inventory buffer coefficient; The historical volatility adjustment factor is calculated based on historical supply and demand volatility characteristics and is obtained by normalizing the standard deviation of historical supply and demand deviations.

[0088] Based on the sign of the supply-demand deviation, the deviation is separated into energy emission forecast and energy shortage forecast: , ,in, For energy nodes The predicted amount of energy emissions; For energy nodes The predicted amount of energy shortage;

[0089] The predicted energy emission is dynamically priced based on the time-of-use energy pricing coefficient, peak-valley adjustment factor, and energy medium type coefficient corresponding to the future scheduling cycle. The calculation formula is as follows: ,in, Penalties for energy emissions; The total number of energy nodes; For energy nodes Time-of-use energy pricing coefficient for the corresponding period in the future scheduling cycle; This is the peak-valley adjustment factor, which takes a value greater than one during peak periods and a value less than one during valley periods. This is the energy medium type coefficient, reflecting the differences in the economic and environmental costs caused by the release of different energy media;

[0090] The predicted energy shortage is quantified by calculating the production stoppage loss based on the process energy demand coefficient, unit energy output value, and shortage urgency coefficient. The calculation formula is as follows: ,in, The cost of punitive energy shortages; For energy nodes The process energy demand coefficient of the process being served reflects the sensitivity of that process to energy interruptions; For energy nodes The unit energy output value of the service process, that is, the economic value created by each unit of energy supply; The shortage urgency coefficient is dynamically determined based on the proportion of the shortage to the normal supply.

[0091] Summing the energy release penalty cost and the energy shortage penalty cost yields the total energy release and shortage penalty cost: ,in, The costs of energy waste and shortages;

[0092] Safety accident expected loss cost accounting unit:

[0093] The risk diffusion index of each hazard source node is used to classify the accident consequence level based on the vulnerability attributes of the affected entities within the influence range of each hazard source node, personnel exposure density, and emergency response capability. This classification employs a fuzzy inference system, using the risk diffusion index, vulnerability of the affected entities, personnel exposure density, and emergency response capability as input variables. These variables are mapped to the accident severity level through a pre-defined fuzzy rule base, with the output being an accident severity level index represented by continuous values. Simultaneously, based on historical accident statistics and the risk diffusion index, the probability of accident occurrence is estimated using a Poisson distribution probability model. ,in, Hazardous source node The probability of an accident occurring within the prediction period; The baseline accident incidence rate parameter;

[0094] The severity level of the accident is mapped to the direct economic loss using the asset vulnerability coefficient, asset replacement value coefficient, and environmental pollution control cost coefficient, and the expected value of the direct economic loss is calculated: ,in, Hazardous source node The expected value of direct economic losses; Hazardous source node The accident severity index; The asset vulnerability coefficient reflects the proportion of the insured's assets that would be lost under the impact of an accident; Hazardous source node The replacement value of assets within the affected area; This represents the cost coefficient for environmental pollution control. Hazardous source node The unit treatment cost of the corresponding pollutants;

[0095] The probability distribution of the accident is used to calculate indirect economic losses based on the process recovery time coefficient, output per unit time, and supply chain disruption loss coefficient. The expected value of the indirect economic loss is then calculated. ,in, Hazardous source node The expected value of indirect economic losses; This is the process recovery time coefficient; Hazardous source node The estimated time required for the process to return to normal operation after the accident; The output value per unit time for the affected process; This is the supply chain disruption loss coefficient. Additional costs incurred due to supply chain disruptions;

[0096] We obtain the expected cost of the safety accident by weighted summing of the expected values ​​of direct and indirect economic losses: ,in, The expected cost of losses due to a safety accident; This represents the total number of hazardous source nodes. This is the weighting coefficient for direct economic losses; This is the weighting coefficient for indirect economic losses;

[0097] The integrated scheduling management objective function construction unit is as follows: The weighted summation of the penalty costs for energy release and shortages and the expected loss costs from safety accidents is used to construct the integrated scheduling management objective function. ,in, The objective function for comprehensive scheduling and management; Energy cost weighting coefficient; This is the safety cost weighting coefficient; the weighting coefficient reflects the preference between energy efficiency optimization and safety risk control in scheduling decisions, and the weighting coefficient of expected loss cost in safety accidents will be dynamically adjusted in the safety boundary dynamic shrinkage mechanism of the multi-objective resource scheduling decision module;

[0098] The pipeline safety pressure threshold and the permissible concentration of hazardous sources are used as boundary constraints and embedded as penalty terms in the comprehensive scheduling management objective function. When the scheduling scheme causes the pipeline pressure to exceed the safety pressure threshold or the hazardous source concentration to exceed the permissible concentration, a constraint violation penalty term is added to the objective function. , in, The objective function for integrated scheduling and management after embedding constraints; To constrain the penalty coefficient for violations; Predicting operating pressure for the pipeline network; This refers to the safe pressure threshold for the pipeline network. For the predicted hazardous concentration of the hazard source; The permissible concentration of the hazardous source is defined; this objective function will serve as a core component of the immediate reward function in subsequent Markov decision process models, guiding reinforcement learning algorithms to find the optimal balance between safety and energy efficiency.

[0099] In this embodiment, the multi-objective resource scheduling decision module is the core decision layer of the intelligent decision-making system for energy consumption optimization and safety production accident prevention in the smelting industry. This module receives the comprehensive scheduling management objective function output by the energy efficiency-safety joint evolution prediction and management cost accounting module, as well as the network state characteristics output by the energy flow-risk flow bidirectional coupling network construction module. With minimizing the comprehensive scheduling management objective function as the optimization objective, and using the pipeline safety pressure threshold and permissible concentration of hazardous sources for each energy medium as initial constraints, a Markov decision process model is constructed, formalizing the energy scheduling and safety control problem into a sequential decision problem. During the model solving process, the module innovatively… A dynamic shrinkage mechanism for safety boundaries is introduced: when the predicted risk diffusion index exceeds the preset warning threshold, the scheduling action space of the corresponding energy medium is dynamically shrunk according to the risk diffusion gradient, and the weight coefficient of the expected loss cost of safety accidents in the comprehensive scheduling management objective function is increased, so that the scheduling strategy automatically shifts towards the safety priority direction; the module uses a multi-objective reinforcement learning algorithm to iteratively optimize within the dynamically shrunk action space, and finally generates a safety-energy efficiency collaborative decision-making scheme that includes cross-process allocation instructions for energy mediums and capacity downgrading scheduling instructions for high-risk processes, providing intelligent scheduling guidance for smelting production that takes into account both energy saving and efficiency improvement and accident prevention, specifically including:

[0100] The Markov decision process modeling unit extracts features from the state characteristics of energy nodes, material nodes, hazard source nodes, and risk-bearing body nodes in the bidirectional energy flow-risk flow coupled network within the current time window. It extracts state components directly related to scheduling decisions from the joint spatiotemporal evolution characteristics of each type of node. Simultaneously, it collects current operating pressure data of the pipeline network for each energy medium and current hazard concentration monitoring data of the hazard source. The node feature components are then time-series concatenated with pipeline operating parameters and hazard concentration parameters to generate the system's current observed state vector. The structure of the system's current observed state vector is as follows: , in, For a moment The system's current observation state vector; For energy nodes At any moment The joint spatiotemporal evolution characteristics; The total number of energy nodes; Hazardous source node At any moment The joint spatiotemporal evolution characteristics; This represents the total number of hazardous source nodes. For a moment The pressure vector composed of the operating pressures of each pipeline segment; For a moment A concentration vector composed of the hazard concentrations at each monitoring point; This represents a vector concatenation operation; it constructs the state space of a Markov decision process model from all possible current observed state vectors of the system. ;

[0101] Based on two main decision types—cross-process scheduling instructions for energy media and capacity downgrading scheduling instructions for high-risk processes—the pipeline valve opening adjustment, pipeline flow allocation ratio, and capacity load adjustment rate of high-risk processes for each energy media are discretized and mapped. The pipeline valve opening adjustment is discretized into several levels, each corresponding to a percentage change in opening. The pipeline flow allocation ratio is discretized into several preset allocation schemes, each satisfying the total flow conservation constraint. The capacity load adjustment rate is discretized into several levels, from full load to minimum safe load. All discretized scheduling actions are combined to generate an initial scheduling action set, in the following form: ,in, This is a scheduling action vector in the initial set of scheduling actions; For the first The valve opening adjustment amount; The number of adjustable valves; For the first The flow distribution ratio of the pipeline network; The quantity allocated to the pipeline flow; For the first Capacity load adjustment rate of high-risk processes; To determine the number of adjustable processes; construct the initial action space of the Markov decision process model from all initial scheduling action sets. ;

[0102] Obtain state transition records within historical scheduling cycles. Extract the system observation state vectors to the next time step after executing different scheduling actions under different system observation state vectors from the historical database, and construct a state transition sample dataset. Based on this sample dataset, train a smelting process state transition network to predict the system observation state vector of the next time step. The smelting process state transition network consists of a deep neural network with an encoder-decoder structure. The encoder receives the current state vector and action vector as input, and the decoder outputs the predicted mean and variance of the next state vector. Based on the prediction results and the Gaussian distribution assumption, calculate the transition probability distribution between adjacent time step states. ,in, In the state Execute action After transitioning to state The transition probability; This is the predicted mean vector output by the state transition network; The prediction covariance matrix is ​​given for the state transition network; Represent a multivariate Gaussian distribution; summarize the transition probability distributions corresponding to all state-action pairs into a state transition probability matrix. ;

[0103] The energy release and shortage penalty costs and the expected loss costs of safety accidents in the comprehensive scheduling and management objective function are negatively mapped, meaning that the lower the cost, the higher the reward, thus transforming the cost minimization objective into the cumulative reward maximization objective. Simultaneously, the violation penalty term is calculated based on the degree of violation of the pipeline safety pressure threshold and the permissible concentration of hazardous sources, combined with the current observed state vector of the system. The mathematical expression of the immediate reward function is: , in, In the state Execute action The instant reward value obtained afterward; For the comprehensive scheduling and management objective function in the state and actions The value to be taken below; For state The operating pressure of the pipeline network below; This refers to the safe pressure threshold for the pipeline network. For state The concentration of hazardous sources below; Permissible concentration for hazardous sources; This is the penalty coefficient for exceeding the limit; this immediate reward function ensures that scheduling decisions strictly adhere to safety constraints while pursuing the lowest economic cost.

[0104] Set the time step for future scheduling cycles and discount attenuation coefficient The time step is consistent with the sampling step of the spatiotemporal evolution diagram sequence in the energy efficiency-safety joint evolution prediction and management cost accounting module, and the discount decay coefficient is between zero and one, used to balance short-term rewards and long-term cumulative rewards; the state space is... Initial motion space State transition probability matrix Instant reward function and discount attenuation coefficient The model is encapsulated to generate a Markov decision process model for multi-objective reinforcement learning, whose core optimization objective is to maximize the expected cumulative discount reward. ,in, The optimal scheduling strategy; Any feasible strategy; For expectation operators; This represents the total number of time steps for future scheduling cycles.

[0105] Safety boundary dynamic contraction mechanism unit: During the solution of the Markov decision process model, the risk diffusion index of each hazard source node is monitored in real time; the updated risk diffusion index vector at each time step is obtained from the energy efficiency-safety joint evolution prediction and management cost accounting module. When the risk diffusion index of any identified hazardous source node exceeds the preset warning threshold... When the time comes, the dynamic shrinkage mechanism of the safety boundary will be immediately triggered; the warning triggering condition is: ,in, Hazardous source node At any moment The risk diffusion index; The preset warning threshold;

[0106] Once an early warning is triggered, the edge weights of the directed edges in the risk propagation graph are extracted. Combined with the spatial distribution of the risk diffusion index of each hazard source node, the risk diffusion gradient of the hazard source node is calculated. The risk diffusion gradient measures the direction and intensity of risk propagation from high-diffusion nodes to surrounding nodes, and its calculation formula is as follows: ,in, Hazardous source node The risk diffusion gradient; To propagate risk through directed edges and hazard source nodes The set of all connected neighboring nodes; Hazardous source node The risk diffusion index; Hazardous source node With the hazardous source node Spatial distance between them; To start from the hazardous source node Pointing to the dangerous source node The risk propagation has the edge weight of the directed edge; For the node Pointing to node The unit direction vector; For nodes The sum of the weights of all outgoing risk propagation edges is used to normalize the weights of individual edges; the risk diffusion gradients of all risk source nodes exceeding the threshold are combined to obtain the risk evolution trend vector. ;

[0107] Identify associated energy nodes affected by risk based on risk evolution trend vectors; for each associated energy node, extract the energy consumption intensity of its directed edge driven by material nodes in the energy flow-risk flow bidirectional coupling network, and calculate the maximum allowable load reduction rate for that associated energy node: ,in, For energy nodes Maximum allowable load reduction rate; For energy nodes The magnitude of the risk diffusion gradient of the nearest hazard source node; For energy nodes The average energy dissipation intensity of all energy-driven directed edges; The preset risk diffusion index warning threshold; the initial action space is determined using the maximum allowable load reduction rate. Energy nodes The upper limit of the value for capacity downgrade scheduling instructions for relevant high-risk processes and the lower limit of the value for cross-process allocation instructions for energy media are truncated; the capacity load adjustment fee rate is... Its upper bound shrinks to: ;

[0108] For the adjustment amount of valve opening in the pipeline network Its lower bound shrinks to: ;

[0109] After truncating all affected motion dimensions, the dynamically shrunken motion space is obtained. ;

[0110] Based on the deviation between the risk diffusion index and the preset warning threshold, the safety risk penalty gain coefficient is calculated using a nonlinear mapping function. The nonlinear mapping function employs an exponential growth curve to ensure that the penalty gain increases faster as the risk deviation increases. ,in, This is the gain coefficient for safety risk penalty; The gain adjustment parameter is used to amplify and update the initial weighting coefficients of the expected loss cost of safety accidents in the comprehensive scheduling and management objective function. ,in, The updated weighting coefficient for expected loss costs in safety incidents; The initial weighting coefficients are used to determine the expected losses from safety accidents; the updated weighting coefficients are then used to replace the corresponding terms in the comprehensive scheduling and management objective function to obtain the updated comprehensive scheduling and management objective function. This leads to the updated instant reward function. ;

[0111] Multi-objective reinforcement learning iterative optimization unit: A multi-objective reinforcement learning model with a dual-network architecture is constructed: the Actor network is responsible for approximating the policy function, and the Critic network is responsible for approximating the value function; the input to the Actor network is the system's current observed state vector. The output is the dynamically shrunk motion space. The sampling probability distribution of each scheduling action is determined using a Gaussian distribution as the action sampling strategy. ,in, Parameterize the Actor network as The policy function represents the state. Select action The probability density; This is the action mean vector output by the Actor network; This is the action variance vector output by the Actor network;

[0112] The input to the Critic network is the current observed state vector of the system. and action vectors The output is a value expectation estimate of the state-action pair. ,in, These are the parameters of the Critic network; both networks employ a multi-layer fully connected network structure, with the hidden layers using the modified linear unit activation function.

[0113] In each iteration, the Actor network expands its action space after dynamic shrinkage. Internal sampling generates exploration scheduling actions To ensure sufficient exploration, attenuated Gaussian exploration noise is added during the sampling process: , ,in, It is a standard normally distributed random noise vector; The identity matrix will be used; the scheduling action will be explored. Applying this to the state transition network of the smelting process yields the next state. And based on the updated instant reward function Calculate the immediate reward value; the Critic network calculates the expected value of exploratory scheduling actions based on the updated integrated scheduling management objective function. ;

[0114] The Actor network updates its parameters using the policy gradient method. The formula for calculating the policy gradient is: ,in, For the objective function Regarding Actor network parameters policy gradient, For the parameter vector Operators for finding the gradient; The Critic network serves as an experience replay buffer, storing state transition samples. The network parameters are updated using temporal difference error, calculated using the following formula: ,in, For a moment The timing difference error, The discount attenuation coefficient ranges from (0,1); the Critic network updates its parameters by minimizing the square of the temporal difference error. ,in, is the learning rate of the Critic network;

[0115] The process alternately executes action exploration, value assessment, Actor network parameter update, and Critic network parameter update steps. After a preset number of iterations or when the magnitude of the policy gradient change falls below a convergence threshold, the iteration terminates, yielding the optimal scheduling policy network. In each iteration, if the dynamic shrinkage mechanism of the safety boundary is triggered again, the shrinkage range and weight coefficient of the action space are recalculated to achieve adaptive dynamic adjustment of the safety constraints.

[0116] Scheduling scheme generation and encapsulation unit: generates and encapsulates the current observation state vector of the system at the current time step. Input to the optimal scheduling policy network The Actor network outputs the optimal action mean vector. This is used as the target scheduling action sequence: ,in, Schedule the action sequence for the target; For the first The optimal opening adjustment value for each valve; For the first The optimal flow distribution ratio for the pipeline network; For the first The optimal capacity load adjustment rate for each high-risk process;

[0117] The target scheduling action sequence is parsed, and numerical scheduling actions are converted into instruction formats executable by industrial control systems. Specifically, this includes: converting valve opening adjustment amounts into valve setpoint instructions for distributed control systems, including target valve identifier, current opening value, and target opening value; converting pipeline flow allocation ratios into flow adjustment instructions for energy management systems, including pipeline section identifier, current flow allocation, and target flow allocation; and converting capacity load adjustment rates into process capacity control instructions for manufacturing execution systems, including process identifier, current load rate, and target load rate. All instructions are sorted according to the time priority of scheduling execution and encapsulated to generate a safety-energy efficiency collaborative decision-making scheme that includes instructions for cross-process allocation of energy media and capacity degradation scheduling instructions for high-risk processes. Each instruction in the scheme includes four elements: instruction type, execution object, execution parameters, and expected execution time window, ensuring that downstream modules can accurately execute scheduling decisions.

[0118] Furthermore, the training process of the smelting process state transition network is as follows:

[0119] Training data construction: State transition records are extracted from the historical scheduling database of the smelter area; each record contains a triple, which is the system observation state vector at the current moment. The scheduling action vector to be executed at the current moment. and the system observation state vector actually observed at the next moment. The construction method of the system observation state vector and scheduling action vector is consistent with the definition in the Markov decision process modeling unit of this module; the components of each dimension of the state vector are Z-score standardized before being input into the network, so that the mean of the numerical distribution of each dimension is zero and the standard deviation is one; the components of each dimension of the action vector are normalized by maximum and minimum values, mapping the values ​​to the interval from negative one to one; all state transition triples are arranged in chronological order, and the first 80% are used as the training set, and the remaining 20% ​​are used as the test set;

[0120] Network Structure: The encoder receives the concatenated state vector and action vector as input, with the input dimension equal to the sum of the state vector dimension and the action vector dimension. The encoder consists of three fully connected hidden layers with 256, 128, and 64 hidden units respectively. Each layer is followed by a batch normalization layer and a modified linear unit activation function. The encoder output is a latent space feature vector with a dimension of 32. The decoder receives the latent space feature vector as input, and its network structure is mirror-symmetric to the encoder: the number of hidden units is 64, 128, and 256 respectively, with each layer followed by a batch normalization layer and a modified linear unit activation function. The decoder's output layer consists of two parallel output heads: the mean output head outputs the predicted mean of the next state vector. Its dimension is equal to the state vector dimension, and the activation function is linear activation; the variance output head outputs the diagonal elements of the predicted variance of the next state vector. Its dimensions are the same as the mean output head, and the activation function is softplus activation to ensure that the output is positive; the complete prediction covariance matrix. The state vector is represented by a diagonal matrix consisting of diagonal elements of variance, which assumes that each dimension of the state vector is conditionally independent given the current state and action.

[0121] Loss function: Based on the assumption that the state transition follows a Gaussian distribution, the negative log-likelihood loss function is adopted. ,in, The negative log-likelihood loss is used for training the state transition network; Batch size; For the first The determinant of the prediction covariance matrix for each sample; For the first The true next state vector of each sample; For the first The predicted mean vector of each sample; For the first The predicted covariance matrix of each sample; since the covariance matrix is ​​in diagonal form, the above loss function simplifies to: ,in, The dimension of the state vector; For the first The first sample Dimensional prediction variance; For the first The first sample The true next state value; For the first The first sample Dimensional prediction mean;

[0122] Training strategy: The AdamW optimizer is used for training, with an initial learning rate of 0.0005, a weight decay coefficient of 0.0001, and a batch size of 64. A cosine annealing learning rate scheduling strategy is employed during training, with a maximum training epoch of 150. To prevent drastic fluctuations in the predicted mean of the encoder-decoder structure during the initial training phase, a smaller learning rate (0.1 times the initial learning rate) is used for warm-up training of the variance output head weights during the first 10 training epochs. The training termination condition is triggered if the test set loss does not decrease for 10 consecutive training epochs, restoring the model parameters with the lowest test set loss as the final pre-trained weights. After pre-training, the model parameters are permanently saved for use in online Markov decision process model solving. During the solving process, the network parameters remain fixed and do not participate in gradient updates during reinforcement learning.

[0123] In this embodiment, the decision-making scheme issuance and closed-loop feedback module is the execution and feedback layer of the intelligent decision-making system for energy consumption optimization and safety production accident prevention in the smelting industry. This module receives safety-energy efficiency collaborative decision-making schemes generated by the multi-objective resource scheduling decision-making module, performs industrial protocol adaptation and multi-channel redundant issuance of energy medium cross-process allocation instructions and high-risk process capacity downgrade scheduling instructions included in the schemes. Simultaneously, it collects instruction execution status feedback, actual equipment operating status, and environmental safety monitoring update results through the real-time data interfaces of the distributed control system, manufacturing execution system, and energy management system. Based on the deviation between the execution status feedback and the predicted status, the module performs graded diagnosis and dynamic compensation processing on the deviation, feeding back the compensated and corrected execution results to the upstream prediction model and decision-making model, forming a complete closed loop of scheduling-execution-monitoring-feedback-correction. This module achieves a secure and reliable connection between the intelligent decision-making layer and the industrial field control layer, and is the key last-mile link in the system from data perception to intelligent decision-making and then to precise execution, specifically including:

[0124] The decision-making scheme protocol adaptation unit receives the safety-energy efficiency collaborative decision-making scheme generated by the multi-objective resource scheduling decision module, classifies the instruction set in the scheme according to instruction type, extracts four elements for each instruction: instruction type, execution object, execution parameters, and expected execution time window, and stores them in the instruction parsing buffer; for energy medium cross-process allocation instructions, it parses the target valve identifier, current opening value, and target opening value to generate a valve control parameter triplet; for high-risk process capacity downgrade scheduling instructions, it parses the process identifier, current load rate, and target load rate to generate a load control parameter triplet; simultaneously, it performs time slot alignment processing on the expected execution time window to ensure that the time slot granularity of all instructions is consistent with the control cycle of the industrial field control system, and the aligned time slot is denoted as... ;

[0125] A protocol adaptation knowledge base is constructed, storing data frame templates for commonly used communication protocols in distributed control systems. This includes three mainstream industrial control communication protocols: Modbus TCP, OPC UA, and Profinet, as well as interface call specifications for manufacturing execution systems and energy management systems. For valve control parameter triplets and load control parameter triplets, protocol matching is performed based on the control system type of the target device. A template mapping algorithm is used to fill the triplet parameters into the register address field or object node field of the corresponding protocol data frame. During the data frame encapsulation process, a cyclic redundancy check (CRC) code is calculated, and a timestamp and instruction sequence number are appended to generate a traceable communication data frame carrying unique identification information. The protocol adaptation mapping rules are formally expressed as follows: ,in, This refers to the encapsulated communication data frame. For template mapping functions corresponding to the communication protocol of the target device; It serves as a unique identifier for the target device within the control system; The current parameter value of the execution object; The target parameter value for the execution object;

[0126] For critical scheduling instructions that require operator confirmation or manual execution, the valve control parameter triplet and load control parameter triplet are converted into standardized human-machine interaction operation instruction text. The operation instruction text includes the target equipment name, equipment location, current operating parameters, target adjustment parameters, adjustment steps, and operation safety precautions. At the same time, an operation confirmation receipt code is generated to track the execution status of manual operation instructions.

[0127] All pending instructions in the instruction parsing buffer undergo instruction conflict detection. Conflict detection rules include: whether mutually exclusive opening adjustment instructions are assigned to the same target equipment in the same time slot; whether instructions in the same pipeline section in the same time slot cause the sum of flow allocation ratios to exceed the total constraint; and whether adjustment instructions in the same process in the same time slot exceed the maximum allowable load change rate of the process. When a conflict is detected, it is resolved according to preset priority rules, with priorities from high to low as follows: safety emergency shutdown instructions, capacity downgrade scheduling instructions, and energy medium cross-process allocation instructions. The instruction sequence after conflict resolution is sorted according to the increasing order of time slots and the decreasing order of priority to generate the final instruction issuance queue.

[0128] Multi-channel redundant command delivery and execution confirmation unit: For each communication data frame in the final command delivery queue, a dual-channel redundant communication delivery mechanism is established; the main channel is directly connected to the main controller of the distributed control system or the core server of the manufacturing execution system through an industrial Ethernet bus, and reliable transmission is achieved using a transmission control protocol; the backup channel is transmitted in parallel through an industrial wireless network or a ring redundant network independent of the main channel, using a user datagram protocol for low-latency transmission, and an acknowledgment and retransmission mechanism is added at the application layer to ensure reliability; the physical routes of the two channels are kept separate in the plant network topology to avoid command delivery interruption due to single point network failure;

[0129] The final instruction queue is grouped according to time slots. Each group of instructions is sent simultaneously on the primary and backup channels, with the sending time preceding the expected execution time window by a preset communication delay tolerance value. After sending, a two-way heartbeat confirmation mechanism is initiated: upon receiving the instruction, the control end immediately sends an instruction reception confirmation signal back to the decision end; after the instruction is executed, the control end sends an instruction execution completion confirmation signal back to the decision end, which includes the instruction sequence number, the actual execution timestamp, and the execution completion status code. The time interval of the heartbeat confirmation signals is determined by the following formula: ,in, The time interval for heartbeat confirmation signals; To determine the minimum heartbeat interval, the upper limit of the round-trip delay time of the communication link is taken; The number of times the acknowledgment signal is collected within each time slot; if no acknowledgment signal is received within the preset timeout period, the redundant channel's retransmission mechanism is automatically triggered, with the maximum number of retransmissions being the preset maximum number of retries;

[0130] When the dynamic shrinkage mechanism of the security boundary of the multi-objective resource scheduling decision module triggers the highest level security alarm, a safety emergency stop command is generated. The safety emergency stop command adopts a high-priority broadcast mode, bypassing the normal command queue, and is broadcast to all relevant control terminal devices simultaneously through the main channel and backup channel. An emergency flag is set in the broadcast data frame. After the control terminal recognizes the flag, it immediately interrupts the current scheduling command sequence, prioritizes the execution of the emergency stop operation, and returns the emergency stop execution result to the decision terminal through a dedicated interrupt channel.

[0131] The execution status monitoring and deviation calculation unit acquires feedback values ​​of actual valve opening, actual pipeline flow, actual pipeline pressure, and actual equipment power through the distributed control system data interface; it acquires actual process load rate, actual material handling volume, and actual material transfer status through the manufacturing execution system data interface; it acquires actual energy consumption, actual energy inventory level, and actual pipeline operating parameters through the energy management system data interface; and it acquires actual concentrations of hazardous gases, actual equipment vibration characteristics, and actual temperature characteristics through the safety monitoring data interface. All multi-source real-time data streams are synchronized and aligned according to a unified timestamp benchmark to generate an execution status feedback dataset, the time synchronization accuracy of which is consistent with the acquisition period of the system's observed status vector.

[0132] For each issued instruction, its target parameter value is compared item by item with the corresponding actual execution result in the execution status feedback dataset to calculate the instruction-level execution deviation; for valve opening adjustment instructions, the execution deviation calculation formula is: ,in, For the first Command-level execution deviation rate of each valve; For the first The actual opening feedback value of each valve; For the first The target opening value of each valve; For the first The full-range opening range of each valve; for capacity load adjustment commands, the execution deviation calculation formula is: ,in, For the first Instruction-level execution deviation rate for each process; For the first Feedback value of the actual load rate of each process; For the first Target load rate for each process; For the first The allowable load rate adjustment range for each process;

[0133] The actual observed values ​​corresponding to the state variables predicted by the energy efficiency-safety joint evolution prediction and management cost accounting module in the execution state feedback dataset are extracted, and the deviations are calculated with the predicted values ​​at the corresponding time steps to obtain the state-level prediction deviation. For the deviation of the predicted value of the energy node medium supply-demand imbalance probability, cross-entropy deviation is used as the metric. , in, It is a measure of the predicted deviation of the probability of supply and demand imbalance in the energy node medium. This is the probability of actual supply and demand imbalance calculated based on the actual supply and demand data in the execution status feedback dataset; The predicted supply-demand imbalance probability is output by the energy efficiency-safety joint evolution prediction and management cost accounting module; for the deviation of the predicted risk diffusion index value of the hazardous source node, a relative deviation metric is used. ,in, This represents the predicted deviation value of the risk diffusion index at the hazardous source node; This is the actual risk diffusion index calculated based on the actual risk monitoring data in the execution status feedback dataset; The predicted risk diffusion index is output by the energy efficiency-safety joint evolution prediction and management cost accounting module; It is a small smoothing constant;

[0134] Deviation Grading Diagnosis and Dynamic Compensation Unit: Based on a comprehensive judgment of the instruction-level execution deviation rate and the state-level prediction deviation metric, deviations are divided into three levels: If all instruction-level execution deviation rates are below the low deviation threshold and all state-level prediction deviation metrics are below the low prediction deviation threshold, it is judged as Level 1 mild deviation; if any instruction-level execution deviation rate exceeds the low deviation threshold but is below the high deviation threshold, or any state-level prediction deviation metric exceeds the low prediction deviation threshold but is below the high prediction deviation threshold, it is judged as Level 2 moderate deviation; if any instruction-level execution deviation rate exceeds the high deviation threshold, or any state-level prediction deviation metric exceeds the high prediction deviation threshold, it is judged as Level 3 severe deviation. A source tracing diagnosis program is initiated for Level 2 and Level 3 deviations, using a fault tree inference engine to analyze the root causes of the deviations. The input to the fault tree inference engine is the deviation type, the system observation state vector corresponding to the deviation occurrence time, and the execution state feedback dataset. The output is the classification result of the deviation root causes, including: actuator response hysteresis, actuator mechanical jamming, sensor measurement drift, abnormal changes in pipeline resistance, external uncontrollable disturbances, and prediction model parameter offset.

[0135] For a level 1 minor deviation, the execution of the current scheduling scheme is not terminated. Instead, a proportional-integral fine-tuning compensation algorithm is used to correct the scheduling instructions for subsequent time slots online. The correction magnitude is determined by the cumulative exponentially weighted moving average of the current deviation. ,in, For a moment The dynamic compensation correction amount; This is the proportional gain coefficient; The deviation metric at the current moment is the weighted average of the instruction-level execution deviation rate and the state-level prediction deviation metric. This is the integral gain coefficient; The length of the sliding window; The attenuation factor is used to add the dynamic compensation correction amount to the scheduling instruction parameters of the next time slot, generate the compensated scheduling instruction, and perform in-situ replacement in the instruction parsing buffer.

[0136] For a moderate deviation at level two, the execution of subsequent instructions in the affected scheduling section of the current scheduling scheme is suspended, and the root cause of the deviation is addressed. If the root cause is actuator response lag or actuator mechanical jamming, an actuator maintenance request is generated, and the actuator control parameters are adjusted as needed. If the root cause is sensor measurement drift, the sensor online calibration process is triggered, and redundant sensor data is used for cross-validation. If the root cause is abnormal changes in pipeline resistance, the flow allocation ratio of the affected pipeline is recalibrated. After the root cause is addressed, the Markov decision process in the multi-objective resource scheduling decision module is rerun based on the actual state feedback. A strategy evaluation is performed starting from the current system observation state vector, a corrected scheduling action sequence is generated, the instructions in the affected scheduling section of the original scheduling scheme are replaced, and scheduling execution is resumed after compensation is completed.

[0137] For Level 3 severe deviations, a dynamic shrinkage mechanism for the safety boundary is immediately triggered, forcibly downgrading the capacity load adjustment rate of the deviation-related section to the lowest safe load level to isolate the scope of the deviation's impact. Simultaneously, an emergency dispatch plan is activated, re-executing the complete Markov decision process modeling and multi-objective reinforcement learning iterative optimization process with the current system observation state vector as the initial state to generate a new emergency dispatch plan that overwrites the original plan. During the generation of the emergency dispatch plan, the weight coefficient of the expected loss cost of the safety accident is set to the preset emergency maximum value to ensure that the emergency plan prioritizes safety. After the root cause of the deviation is addressed, if the risk diffusion index falls below the warning threshold, the normal dispatch plan is gradually restored.

[0138] The closed-loop feedback and strategy adaptive tuning unit classifies and organizes the accumulated Level 1 mild deviation data, Level 2 moderate deviation data, and Level 3 severe deviation data in each scheduling cycle according to deviation type and execution object, generating a deviation feedback data package. The structure of the deviation feedback data package includes: a sequence of deviation occurrence timestamps, a sequence of scheduling instruction identifiers corresponding to the deviation, a sequence of execution deviation rates, a sequence of predicted deviation metrics, a sequence of deviation root cause classification results, and a sequence of dynamic compensation measures. The deviation feedback data package is uploaded to the energy efficiency-safety joint evolution prediction and management cost accounting module and the energy flow-risk flow bidirectional coupling network construction module through the data bus, providing real feedback data for the online fine-tuning and parameter updates of the upstream model.

[0139] In the energy efficiency-safety joint evolution prediction and management cost accounting module, the spatiotemporal graph attention network triggers an online fine-tuning mechanism after receiving a deviation feedback data packet. Using the actual supply-demand imbalance probability and actual risk diffusion index in the deviation feedback data packet as monitoring signals, the parameters of the spatiotemporal graph attention network are updated with small-step gradients. The loss function for online fine-tuning consists of two terms: a prediction deviation penalty term and a parameter change regularization term. ,in, For online fine-tuning of the loss function; This represents the current parameter vector of the spatiotemporal graph attention network; This is the parameter vector before online fine-tuning; The regularization strength coefficient is used to constrain the parameter update magnitude and prevent catastrophic forgetting. The fine-tuned prediction model performs inference with the updated parameters in the next prediction cycle, improving the prediction accuracy for future scheduling cycles.

[0140] The Actor-Critic network in the multi-objective resource scheduling decision module triggers a policy adaptive tuning mechanism after receiving a deviation feedback data packet; it converts the instruction-level execution deviation rate in the deviation feedback data packet into a negative reward signal, correcting the value estimate of the corresponding state-action pair in the experience replay buffer. ,in, This is the state-action value estimate after bias feedback correction; State-action value estimation of the original output of the Critic network; This is the deviation feedback correction coefficient; This is the weighted average of the instruction-level execution deviation rate for the state-action pair; additional rounds of parameter updates based on corrected value estimation are performed on the Critic network, enabling the decision network to gradually learn to avoid generating scheduling actions that are difficult to execute accurately or are prone to causing large deviations in the industrial field, thereby improving the field executability and robustness of the strategy.

[0141] When the deviation metric in the deviation feedback data packets of multiple consecutive scheduling cycles remains stable within the range of Level 1 mild deviation and no Level 3 severe deviation events occur, the closed-loop iteration is determined to have entered the convergence state. At this time, the module generates a closed-loop convergence confirmation report, which includes the number of consecutive stable operation cycles, the average deviation level after convergence, the convergence stability index of the policy network, and the latest prediction accuracy index of the prediction model. This report serves as the basis for system operation and maintenance personnel to judge the health status of the system.

[0142] Historical strategy repository accumulation and reuse unit: The scheduling schemes whose deviation metric values ​​remain within the first-level slight deviation range throughout the closed-loop feedback process are marked, and the corresponding system observation state vectors, scheduling action sequences and scheduling effect evaluation indicators are feature extracted and vectorized to form a record of successful scheduling strategies; the successful scheduling strategy records are indexed according to four dimensions: energy medium type, process type, risk level and scheduling time period characteristics and then stored in the historical strategy repository;

[0143] When a new scheduling cycle begins, before performing a full iterative solution on the multi-objective resource scheduling decision module, a similarity state retrieval mechanism is first used to search the historical strategy database for the historical successful scheduling strategy with the highest cosine similarity to the current system's observed state vector. The cosine similarity calculation formula is as follows: ,in, The current system observation state vector With historical state vector The cosine similarity between them; if the maximum cosine similarity exceeds the preset reuse threshold, the scheduling action sequence corresponding to the historical successful scheduling strategy is injected as the initial feasible solution into the Actor network of the multi-objective resource scheduling decision module, replacing the randomly initialized strategy output, thereby accelerating the convergence speed of the strategy network and realizing the effective reuse of historical experience.

[0144] Time-sensitivity weights are assigned to policy records in the historical policy repository. Each record's time-sensitivity weight is initialized to its maximum value upon entry into the repository and decreases exponentially with storage time. ,in, Record the policy duration in storage The subsequent timeliness weight; This is the initial value for the timeliness weight; This is the time-related attenuation coefficient. When the time weight is lower than the attenuation threshold, the strategy record is removed from the strategy library to ensure that the strategies in the strategy library always reflect the current operating conditions and equipment status, and to avoid the failure of historical strategies due to equipment aging or process changes.

[0145] In industrial control systems, the coexistence of multiple communication protocols is a common feature of digital smelting plants. The decision-making scheme protocol adaptation unit automatically converts abstract scheduling parameters into data frame formats that can be recognized by various control systems by establishing a protocol adaptation knowledge base and template mapping mechanism, thus achieving seamless connection between the intelligent decision layer and the heterogeneous control layer. The dual-channel redundant delivery mechanism is based on the redundancy and fault-tolerant design in communication reliability theory. It transmits the same instruction in parallel through two independent communication links with separate physical routes. Combined with the heartbeat confirmation mechanism and timeout retransmission strategy, it improves the communication success rate of instruction delivery from the probability level of a single link to a higher level after dual-channel parallel redundancy. The high-priority broadcast mechanism of safety emergency stop signals is based on the interrupt nesting priority design of real-time systems, ensuring that safety instructions can be transmitted to the execution equipment with the lowest delay under the worst operating conditions.

[0146] The deviation classification and diagnosis mechanism draws on the hierarchical approach to fault detection and diagnosis in industrial process control. Level 1, mild deviation, refers to random fluctuations and minor model mismatches within the normal range, which can be smoothly eliminated in the closed loop using proportional-integral fine-tuning compensation without interrupting the scheduling process. Level 2, moderate deviation, usually points to a locatable single-point fault or local model offset, requiring the suspension of the affected section and root cause orientation processing. After processing, scheduling can be quickly restored through local strategy evaluation. Level 3, severe deviation, indicates a major anomaly in the system or a serious model mismatch, necessitating the activation of global safety degradation and emergency scheduling plans, with safety as the top priority for a comprehensive restructuring of the scheduling scheme. This three-level progressive deviation response mechanism maintains scheduling continuity to the maximum extent while ensuring system safety, avoiding unnecessary production interruptions caused by over-response.

[0147] The core idea of ​​the closed-loop feedback mechanism is to use the deviation between the execution result and the prediction result as a feedback signal, which is then transmitted back to the upstream prediction model and decision model to drive the online adaptive update of model parameters. The online fine-tuning of the prediction model adopts a small-sample incremental learning strategy, which constrains the update magnitude through parameter change regularization terms. This improves the fitting accuracy of new data while maintaining the memory of historical knowledge, solving the problems of long time consumption and large sample requirements of traditional full retraining. The policy adaptive tuning of the decision model is based on the value estimation of deviation correction. By quantifying the execution deviation into a negative reward signal, the value perception of the Critic network is corrected, enabling the Actor network to automatically avoid scheduling actions that are difficult to execute accurately in subsequent policy generation. This achieves policy transfer and robustness improvement from simulation optimization to field application. The mechanism of accumulating and reusing historical policy libraries belongs to the engineering application of case reasoning. By extracting and indexing historical successful experiences, high-quality initial policies are provided for reinforcement learning under similar working conditions, which significantly accelerates the convergence process of policy search and reduces the computational load of online solution.

[0148] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An intelligent decision-making system for energy consumption optimization and safety accident prevention in the smelting industry, characterized by: include: The multi-source production and risk data perception module is used to acquire real-time energy medium consumption data, material flow data, equipment operating status data, and environmental safety monitoring data for the smelting process. The energy flow-risk flow bidirectional coupling network construction module constructs a directed energy flow graph containing energy nodes and material nodes based on real-time energy medium consumption data and material flow data. Based on equipment operation status data and environmental safety monitoring data, a risk propagation directed graph containing hazard source nodes and risk-bearing body nodes is constructed; by identifying the physical spatial mapping relationship and energy release association between energy nodes in the energy flow directed graph and hazard source nodes in the risk propagation directed graph, cross-graph coupling edges are established to generate a bidirectional coupling network of energy flow and risk flow. The energy efficiency-safety joint evolution prediction and management cost accounting module inputs the bidirectional coupling network of energy flow and risk flow in the current time window into a pre-trained spatiotemporal graph attention network to predict the probability of medium supply and demand imbalance of each energy node and the risk diffusion index of each hazard source node in the future scheduling cycle; calculates the energy release and shortage penalty cost based on the medium supply and demand imbalance probability, and calculates the expected loss cost of safety accidents based on the risk diffusion index, and constructs a comprehensive scheduling management objective function that includes the energy release and shortage penalty cost and the expected loss cost of safety accidents; The multi-objective resource scheduling decision module aims to minimize the overall scheduling management objective function. It uses the pipeline safety pressure threshold and the permissible concentration of hazardous sources for each energy medium as initial constraints, constructing a Markov decision process model. During the solution process, a dynamic shrinkage mechanism for the safety boundary is introduced: when the predicted risk diffusion index exceeds a preset warning threshold, the scheduling action space of the corresponding energy medium is dynamically shrunk according to the risk diffusion gradient, and the weight coefficient of the expected loss cost of safety accidents in the overall scheduling management objective function is increased. A multi-objective reinforcement learning algorithm is used to iteratively optimize within the dynamically shrunk action space, generating a safety-energy efficiency collaborative decision scheme that includes cross-process allocation instructions for energy mediums and capacity downgrading scheduling instructions for high-risk processes.

2. The intelligent decision-making system for optimization of energy consumption and prevention of safety accidents in smelting industry according to claim 1, characterized in that: The multi-source production and risk data perception module acquires real-time energy consumption data, material flow data, equipment operating status data, and environmental safety monitoring data for the smelting process; specifically including: By deploying energy metering sensor groups at key energy flow nodes in the smelting process, material tracking sensor groups at material flow nodes, operating condition monitoring sensor groups at equipment operating stations, and environmental monitoring sensor groups in environmentally sensitive areas, raw energy signals, raw material signals, raw operating condition signals, and raw environmental signals are collected respectively, resulting in a heterogeneous multi-source raw signal set with timestamps. The sampling rate is normalized and cross-channel crosstalk is suppressed on the heterogeneous multi-source raw signal set. After eliminating electromagnetic coupling interference between signals, the timing is aligned to obtain the energy raw timing sequence, material raw timing sequence, operating condition raw timing sequence and environmental raw timing sequence under a unified timing reference. The original time series of energy data is analyzed using a multi-parameter joint analysis of flow rate, pressure, and temperature, and dynamic metering calibration is performed to obtain a structured energy feature sequence. The original time series of materials data is traced by batch and mapped by composition to obtain a structured material feature sequence. The original time series of operating conditions data is analyzed by operating mode identification and degradation feature extraction to obtain a structured operating condition feature sequence. The original time series of environmental data is analyzed by multi-factor concentration gradient analysis to obtain a structured environmental feature sequence. Multi-source cross-validation and spatiotemporal consistency checks were performed on the energy structured feature sequence, material structured feature sequence, operating condition structured feature sequence, and environmental structured feature sequence. After removing abnormal feature values, the data were standardized and encapsulated to obtain real-time energy medium consumption data, material flow data, equipment operating status data, and environmental safety monitoring data.

3. The intelligent decision-making system for optimization of energy consumption and prevention of safety accidents in smelting industry according to claim 1, characterized in that: The energy flow-risk flow bidirectional coupling network construction module constructs a directed energy flow graph containing energy nodes and material nodes based on real-time energy medium consumption data and material flow data; specifically, it includes: Real-time energy medium consumption data is classified into energy medium types and analyzed into pipeline topology to obtain energy transmission topology and energy medium classification results. Material flow data is analyzed into material batch identification and process routing to obtain material transfer topology and material batch identification results. Spatiotemporal feature aggregation processing is performed on the energy transmission topology and energy medium classification results to generate energy node feature vectors containing spatial location, medium type and dynamic flow attributes. Spatiotemporal feature aggregation processing is also performed on the material transfer topology and material batch identification results to generate material node feature vectors containing spatial location, batch identifier and dynamic quality attributes. Based on energy transmission topology and energy flow direction, directed energy transmission edges with transmission loss and delay time as edge weights are constructed between energy node feature vectors. Based on material transfer topology and material flow direction, directed material transfer edges with mass loss and conversion time as edge weights are constructed between material node feature vectors. Based on the energy-material consumption correspondence between energy node feature vectors and material node feature vectors, energy-material driven directed edges with energy intensity as edge weights are constructed from energy node feature vectors to material node feature vectors. All energy node feature vectors, material node feature vectors, directed edges for energy transmission, directed edges for material transfer, and directed edges for energy-material drive are aggregated into a graph structure to generate an energy flow directed graph.

4. The intelligent decision-making system for optimization of energy consumption and prevention of safety accidents in smelting industry according to claim 3, characterized in that: Based on equipment operation status data and environmental safety monitoring data, a directed risk propagation graph is constructed, including hazard source nodes and risk-bearing body nodes; specifically including: Equipment deterioration pattern identification and fault feature extraction are performed on equipment operation status data to obtain equipment health status sequence. Multi-factor concentration gradient analysis and diffusion trend prediction are performed on environmental safety monitoring data to obtain environmental hazard concentration distribution sequence. The equipment health status sequence is used to determine the equipment risk type, identify equipment-related hazards and equipment-related risk carriers, and the environmental hazard concentration distribution sequence is used to determine the environmental risk type, identify environmental hazards and personnel exposure areas and risk carriers. Hazard source fusion processing is performed on equipment-type hazard sources and environmental hazard sources to generate hazard source feature vectors containing attributes such as equipment deterioration degree, hazardous substance concentration and potential energy release intensity. Hazard-bearing bodies and personnel exposure area hazard-bearing bodies are fusion processing to generate hazard-bearing body feature vectors containing attributes such as entity type, vulnerability degree and exposure duration. Based on the spatial distance relationship between the hazard source feature vector and the hazard-bearing body feature vector, the airflow diffusion direction and the energy release propagation path, a directed risk propagation edge from the hazard source feature vector to the hazard-bearing body feature vector is constructed with propagation delay time, attenuation coefficient and exposure intensity as edge weights. All hazard source feature vectors, risk-bearing body feature vectors, and directed edges of risk propagation are aggregated into a graph structure to generate a directed graph of risk propagation.

5. The intelligent decision-making system for optimization of energy consumption and prevention of safety accidents in smelting industry according to claim 4, characterized in that: The process involves identifying the physical spatial mapping relationship and energy release association between energy nodes in the directed energy flow graph and hazard source nodes in the directed risk propagation graph, establishing cross-graph coupling edges, and generating a bidirectional coupling network between energy flow and risk flow; specifically including: Spatial coordinates are extracted from all energy nodes in the directed energy flow graph to generate a set of spatial coordinates for energy nodes, and spatial coordinates are extracted from all hazard source nodes in the directed risk propagation graph to generate a set of spatial coordinates for hazard source nodes. Spatial distance calculation and proximity matching are performed on the spatial coordinate sets of energy nodes and hazardous source nodes to identify co-located node pairs. The physical spatial mapping relationship of the co-located node pairs is then verified to obtain the verified physical spatial mapping relationship. Based on the energy medium type and hazard source type of co-located node pairs, the energy release potential of the verified physical space mapping relationship is analyzed to quantify the energy release intensity, and the risk feedback impact analysis is conducted based on the impact of the risk propagation intensity of hazard source nodes on the energy transmission path to quantify the risk feedback intensity. Based on the verified physical space mapping relationship, a forward coupled directed edge from the energy node to the hazard source node is constructed with the energy release intensity as the edge weight, and a reverse coupled directed edge from the hazard source node to the energy node is constructed with the risk feedback intensity as the edge weight, thus obtaining the cross-graph coupled edge. The directed graph of energy flow, the directed graph of risk propagation, and all cross-graph coupling edges are integrated into a unified network structure to generate a bidirectional coupled network of energy flow and risk flow.

6. The intelligent decision-making system for optimization of energy consumption and prevention of accidents in the smelting industry as claimed in claim 1 wherein: The energy efficiency-safety joint evolution prediction and management cost accounting module inputs the bidirectional coupling network of energy flow and risk flow in the current time window into a pre-trained spatiotemporal graph attention network to predict the probability of medium supply-demand imbalance at each energy node and the risk diffusion index of each hazardous source node within the future scheduling cycle; specifically including: The energy flow-risk flow bidirectional coupling network of the current time window and the energy flow-risk flow bidirectional coupling network of multiple historical time windows are sampled by time-series slices according to a preset time step to obtain a spatiotemporal evolution graph sequence that includes energy node state characteristics, material node state characteristics, hazard source node state characteristics, risk-bearing body node state characteristics and cross-graph coupling characteristics. For each graph structure in the spatio-temporal evolution graph sequence, based on the edge type discrimination mechanism, independent attention parameter matrices are respectively assigned to the energy transmission directed edges, material transfer directed edges, energy-material drive directed edges, risk propagation directed edges, and cross-graph coupling directed edges. Then, based on the edge weights, source node features, and target node features of each edge type, multi-head graph attention calculation is performed to obtain the in-graph spatial attention weight matrices corresponding to each edge type. According to the in-graph spatial attention weight matrices, the node features are weighted and aggregated to obtain the node spatial dependence features at each time step. Time attention calculation is performed on the node spatial dependence features at each time step in the spatio-temporal evolution graph sequence. Through the self-attention mechanism in the time dimension, the time evolution patterns of the energy flow state and the risk flow state are captured to obtain the node time evolution features of each energy node and each hazard source node. Cross-graph feature interaction calculation is performed on the node time evolution features of each energy node and each hazard source node through the cross-graph coupling directed edges to obtain the bidirectional influence features reflecting the positive driving effect of the energy flow on the risk flow and the reverse constraint effect of the risk flow on the energy flow. Then, the node time evolution features and the bidirectional influence features are feature-fused to obtain the joint spatio-temporal evolution features of each energy node and each hazard source node. The joint spatio-temporal evolution features of each energy node are input into the medium supply and demand prediction branch for multi-layer perceptron mapping and sigmoid activation processing to obtain the medium supply and demand imbalance probabilities of each energy node. The joint spatio-temporal evolution features of each hazard source node are input into the risk diffusion prediction branch for multi-layer perceptron mapping and exponential activation processing to obtain the risk diffusion indices of each hazard source node.

7. The intelligent decision-making system for optimization of energy consumption and prevention of accidents in the smelting industry as claimed in claim 1 wherein: Based on the medium supply and demand imbalance probabilities, the energy dissipation and shortage penalty costs are calculated. Based on the risk diffusion indices, the expected loss costs of safety accidents are calculated. A comprehensive scheduling management objective function including the energy dissipation and shortage penalty costs and the expected loss costs of safety accidents is constructed. Specifically, it includes: The medium supply and demand imbalance probabilities of each energy node are combined with the rated supply capacity, real-time inventory level, and historical supply and demand fluctuation characteristics of each energy node for supply and demand deviation quantification processing to obtain the energy dissipation prediction amounts and energy shortage prediction amounts of each energy node in the future scheduling period. The energy dissipation prediction amounts are dynamically priced according to the time-sharing energy pricing coefficients, peak-valley regulation factors, and energy medium type coefficients corresponding to the future scheduling period to obtain the energy dissipation penalty costs. The energy shortage prediction amounts are quantified for production stoppage losses according to the process energy demand coefficients, unit energy output value, and shortage urgency coefficients to obtain the energy shortage penalty costs. Then, the energy dissipation penalty costs and the energy shortage penalty costs are summed up to obtain the energy dissipation and shortage penalty costs. The risk diffusion indices of each hazard source node are combined with the vulnerability attributes, personnel exposure density, and emergency response capabilities of the risk-bearing entities within the influence range of each hazard source node for accident consequence level classification processing to obtain the accident severity levels and accident occurrence probability distributions of each hazard source node. The severity level of the accident is mapped to the direct economic loss using the asset vulnerability coefficient, asset replacement value coefficient, and environmental pollution control cost coefficient to obtain the expected value of direct economic loss. The probability distribution of the accident is mapped to the indirect economic loss using the process recovery time coefficient, output value per unit time, and supply chain interruption loss coefficient to obtain the expected value of indirect economic loss. The expected values ​​of direct and indirect economic loss are then weighted and summed to obtain the expected cost of safety accident loss. A comprehensive scheduling and management objective function is constructed by weighted summation of the penalty costs for energy release and shortage and the expected loss costs of safety accidents. The pipeline safety pressure threshold and the allowable concentration of hazardous sources are used as boundary constraints to embed the comprehensive scheduling and management objective function.

8. The intelligent decision-making system for optimization of energy consumption and prevention of accidents in the smelting industry as claimed in claim 1 wherein: The multi-objective resource scheduling decision module takes minimizing the comprehensive scheduling management objective function as the optimization objective and uses the pipeline safety pressure threshold and the permissible concentration of hazardous sources for each energy medium as initial constraints to construct a Markov decision process model; specifically, it includes: Feature extraction is performed on the energy node state characteristics, material node state characteristics, hazard source node state characteristics and risk-bearing body node state characteristics in the energy flow-risk flow bidirectional coupling network in the current time window. The current operating pressure of the pipeline network of each energy medium and the current hazard concentration of the hazard source are combined for time-series splicing to obtain the current observed state vector of the system. All possible current observed state vectors of the system are constructed into the state space of the Markov decision process model. Based on the cross-process allocation instructions for energy media and the capacity downgrade scheduling instructions for high-risk processes, the valve opening adjustment amount, flow distribution ratio, and capacity load adjustment rate of each energy medium are discretized and mapped to generate an initial scheduling action set. All initial scheduling action sets are then constructed as the initial action space of the Markov decision process model. Obtain the state transition records within the historical scheduling cycle. Based on the current observed state vector of the system and the initial scheduling action set, predict the system observed state vector of the next time step through the pre-trained smelting process state transition network. Calculate the transition probability distribution between adjacent time step states to obtain the state transition probability matrix of the Markov decision process model. The energy release and shortage penalty cost and the expected loss cost of safety accidents in the comprehensive scheduling and management objective function are negatively mapped, and the violation degree of the pipeline safety pressure threshold and the allowable concentration of dangerous sources is calculated by combining the current observed state vector of the system. The negative cost mapping value and the violation penalty term are summed to obtain the instantaneous reward function of the Markov decision process model. By setting the time step and discount decay coefficient for future scheduling cycles, the state space, initial action space, state transition probability matrix, immediate reward function, and discount decay coefficient are encapsulated in the model to generate a Markov decision process model for multi-objective reinforcement learning.

9. The intelligent decision-making system for optimization of energy consumption and prevention of safety accidents in smelting industry according to claim 8, characterized in that: In solving the Markov decision process model, a dynamic shrinkage mechanism for the safety boundary is introduced: when the predicted risk diffusion index exceeds the preset warning threshold, the scheduling action space of the corresponding energy medium is dynamically shrunk according to the risk diffusion gradient, and the weight coefficient of the expected loss cost of safety accidents in the comprehensive scheduling management objective function is increased; a multi-objective reinforcement learning algorithm is used to iteratively optimize within the dynamically shrunk action space to generate a safety-energy efficiency collaborative decision-making scheme that includes cross-process allocation instructions for energy mediums and capacity downgrade scheduling instructions for high-risk processes. Specifically, it includes: The risk diffusion index of the hazard source node is monitored in real time during the operation of the Markov decision process model. When the risk diffusion index exceeds the preset warning threshold, the edge weight of the risk propagation directed edge in the risk propagation directed graph is extracted, the risk diffusion gradient of the hazard source node is calculated, and the risk evolution trend vector is obtained. Based on the risk evolution trend vector, the associated energy nodes affected by the risk are identified. According to the energy consumption intensity of the energy-driven directed edges of the associated energy nodes, the maximum allowable load reduction rate of the associated energy nodes is calculated. The maximum allowable load reduction rate is used to truncate the upper and lower limits of the values ​​of the high-risk process capacity downgrade scheduling instructions and the energy medium cross-process allocation instructions in the initial action space of the Markov decision process model, so as to obtain the dynamically shrunk action space. Based on the deviation between the risk diffusion index and the preset warning threshold, the safety risk penalty gain coefficient is calculated through a nonlinear mapping function. The initial weight coefficient of the expected loss cost of safety accidents in the comprehensive scheduling management objective function is amplified and updated using the safety risk penalty gain coefficient, resulting in the updated comprehensive scheduling management objective function. The dynamically shrunk action space and the updated comprehensive scheduling management objective function are input into the Actor-Critic network architecture of the multi-objective reinforcement learning algorithm. The Actor network samples and generates exploratory scheduling actions within the dynamically shrunk action space, and the Critic network calculates the expected value of the exploratory scheduling actions based on the updated comprehensive scheduling management objective function. The Actor network parameters are updated through policy gradients, and the Critic network parameters are updated through temporal difference errors. After multiple rounds of iteration and convergence, the optimal scheduling policy network is obtained. The current observation state vector of the system at the current time step is input into the optimal scheduling strategy network, and the output is a target scheduling action sequence containing the optimal pipeline valve opening adjustment amount, the optimal pipeline flow allocation ratio, and the optimal capacity load adjustment rate. The target scheduling action sequence is parsed and encapsulated into a safety-energy efficiency collaborative decision-making scheme containing cross-process allocation instructions for energy media and capacity downgrade scheduling instructions for high-risk processes.

10. The intelligent decision-making system for optimization of energy consumption and prevention of accidents in the smelting industry as claimed in claim 1 wherein: It also includes a decision-making scheme distribution and closed-loop feedback module, used to output the safety-energy efficiency collaborative decision-making scheme to the smelting manufacturing execution system or energy management system for scheduling and execution, and to obtain the actual energy efficiency indicators and safety status indicators after execution, used to update the model parameters of the spatiotemporal graph attention network and multi-objective reinforcement learning algorithm; specifically including: The safety-energy efficiency collaborative decision-making scheme is adapted to the instruction protocol and encrypted with industrial security to obtain a standardized executable instruction sequence. The standardized executable instruction sequence is then distributed to the smelting and manufacturing execution system and the energy management system through the industrial communication network for scheduling and execution, resulting in a real-time status data stream of the execution process. Multi-channel synchronous acquisition and noise filtering are performed on the real-time status data stream during the execution process to obtain the actual energy medium consumption data, actual material flow data, actual equipment operation status data and actual environmental safety monitoring data after execution. Based on the actual energy medium consumption data and actual material flow data after execution, energy efficiency indicators are calculated to obtain the actual energy efficiency indicators. Based on the actual equipment operation status data and actual environmental safety monitoring data, safety status assessment is performed to obtain the actual safety status indicators. The probability of medium supply and demand imbalance at each energy node and the risk diffusion index at each hazardous source node within the future scheduling cycle are compared with the actual energy efficiency index and the actual safety status index to perform prediction-actual deviation quantification processing, resulting in the energy flow prediction error sequence and the risk flow prediction error sequence. Error contribution tracking analysis is performed on the energy flow prediction error sequence and risk flow prediction error sequence along the back propagation path of the directed energy flow graph and the directed risk propagation graph. Key energy nodes, key hazard source nodes and key cross-graph coupling edges that cause prediction errors are identified. Parameter update gradient vectors are constructed based on the error contribution of key energy nodes, key hazard source nodes and key cross-graph coupling edges. The gradient descent of the intra-graph spatial attention weight matrix corresponding to each edge type of the spatiotemporal graph attention network is updated based on the parameter update gradient vector, and the policy gradient of the network parameters of the Actor-Critic network architecture of the multi-objective reinforcement learning algorithm is updated to obtain the updated spatiotemporal graph attention network and the updated multi-objective reinforcement learning algorithm.