Coal-fired power plant coal piling strategy optimization system based on artificial intelligence
By integrating multi-source data and performing joint modeling and optimization through an AI-based coal stockpiling strategy optimization system for coal-fired power plants, the system solves the problems of data heterogeneity, modeling difficulties, and weak optimization capabilities in existing coal stockpiling strategies, thus achieving more efficient and safer operation of coal-fired power plants.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-22
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, coal stockpiling strategies in coal-fired power plants rely on human experience, suffer from heterogeneous and inconsistent multi-source data, have difficulty integrating spatial and temporal features, have weak global search capabilities in optimization algorithms, and lack closed-loop feedback. As a result, the strategies are difficult to adapt to complex and ever-changing operating conditions, leading to a decline in accuracy and executability.
An AI-based coal stockpiling strategy optimization system for coal-fired power plants is adopted, comprising a data input module, a control output module, a deep learning model, and a multi-objective optimization algorithm module. The data input module integrates multi-source data through a timestamp alignment unit and edge computing nodes. The deep learning model performs joint modeling through spatial feature extraction and time series processing. The multi-objective optimization algorithm module employs a quantum bit-based optimization algorithm, combined with a safety interlocking logic unit and an OPC bidirectional communication mechanism, to achieve closed-loop control.
It improves the automation level and feasibility of coal stockpiling strategies, enhances the model's understanding of complex operating conditions, avoids local optima, ensures the safety and execution capability of control commands, and significantly improves the operating efficiency and environmental performance of coal-fired power plants.
Smart Images

Figure SMS_1 
Figure SMS_3 
Figure SMS_4
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent power plant dispatching technology, specifically to an artificial intelligence-based coal stockpiling strategy optimization system for coal-fired power plants. Background Technology
[0002] As a crucial component of my country's power supply, the operational efficiency and environmental performance of coal-fired power plants directly impact energy utilization and air pollutant emission control. During power plant operation, the formulation of coal stockpiling strategies has a decisive influence on boiler combustion stability, fuel utilization, and flue gas emission indicators. In existing technologies, some power plants have attempted to introduce data-driven models to assist decision-making. These typically employ statistical analysis methods based on historical data or traditional machine learning algorithms, combined with real-time operating parameters collected by a distributed control system (DCS), to generate preliminary coal blending suggestions. Such solutions generally rely on manually set rules or static optimization models, obtaining recommended strategies through offline simulation, and then being manually executed by operators via the user interface.
[0003] However, the aforementioned existing technologies suffer from problems such as asynchronous multi-source data, ineffective modeling of spatial layout information, and a lack of dynamic feedback mechanisms in the optimization process. These issues make it difficult for the generated coal stockpiling strategies to adapt to complex and changing operating conditions, and they cannot guarantee consistency with actual execution. Especially in scenarios involving fluctuating coal quality and rapid load changes, the accuracy and executability of the strategies significantly decrease, severely hindering the improvement of intelligent capabilities. Furthermore, traditional optimization algorithms have slow convergence speeds, are prone to getting trapped in local optima, and lack closed-loop linkage capabilities with the DCS system, making it difficult to implement the optimization results. Summary of the Invention
[0004] This invention provides an artificial intelligence-based coal stockpiling strategy optimization system for coal-fired power plants, which can solve technical problems in existing technologies such as reliance on human experience in coal stockpiling strategies, heterogeneous and inconsistent multi-source data, difficulty in integrating spatial and temporal features, weak global search capability of optimization algorithms, and lack of closed-loop feedback in strategy execution.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] This invention provides an artificial intelligence-based coal-fired power plant coal stockpiling strategy optimization system, including a data input module, a control output module, a deep learning model, and a multi-objective optimization algorithm module;
[0007] The data input module includes a distributed control system (DCS) configured with a timestamp alignment unit and edge computing nodes. The timestamp alignment unit is used to align timestamps from different data sources, and the edge computing nodes are used to perform preprocessing at the data acquisition points.
[0008] The control output module is connected to the DCS via a bidirectional data channel and includes a safety interlock logic unit configured to block abnormal control signals based on predefined safety rules.
[0009] The deep learning model includes a spatial feature extraction unit and a time series processing unit. The spatial feature extraction unit is used to process the geometric layout data of coal-fired power plants, and the time series processing unit is used to process historical operating data.
[0010] The multi-objective optimization algorithm module employs a qubit-based optimization algorithm to generate coal stacking strategies.
[0011] In one optional embodiment, the data input module is connected to the DCS and historical database to form a data buffer. The data buffer stores spatiotemporally correlated data points and includes an event detector for triggering high-frequency data acquisition and data labeling based on device status events. The multi-objective optimization algorithm module includes a quantum neural network unit that updates model parameters using incremental learning.
[0012] In one alternative embodiment, the deep learning model includes a feature fusion layer that uses a dynamic weight allocation strategy to weight and fuse spatial and temporal features.
[0013] The deep learning model also includes a multilayer perceptron (MLP), which is connected to a quantum neural network unit and configured with a distribution detector to trigger MLP parameter updates and sparsification based on differences in data distribution.
[0014] In one alternative embodiment, the deep learning model includes a cascaded quantum neural network and a hybrid module comprising an MLP and a recurrent neural network (RNN).
[0015] MLPs include scalable groups of neurons, while RNNs include redundant connection paths to enhance model robustness.
[0016] The quantum neural network consists of multiple output heads, each corresponding to a prediction of a pollutant indicator.
[0017] In one optional embodiment, a gradient transmission channel is provided between the multilayer perceptron and the multi-objective optimization algorithm module for backpropagating gradients;
[0018] A differentiable constraint layer is embedded in the MLP of the deep learning model, which is used to apply runtime constraints;
[0019] The multi-objective optimization algorithm module adopts the quantum-inspired bee colony algorithm, which dynamically adjusts its structure based on circuit templates.
[0020] In one alternative embodiment, the deep learning model includes a regularization unit that generates regularization terms by a discriminator network; the multi-objective optimization algorithm module employs a quantum-coded bee colony algorithm, where the solution vector is represented by qubits and updated through rotation operations, crossover operations, and a restart mechanism.
[0021] In one alternative embodiment, the deep learning model is coupled with a multi-objective optimization algorithm module, using the predicted values of fly ash mass concentration, NOx emission concentration, and coal consumption as inputs to the multi-objective optimization algorithm.
[0022] In one alternative embodiment, a multi-objective optimization algorithm module is embedded in a reinforcement learning agent, which includes an action selector; a quantum-coded structure is mapped to a policy network hidden layer of the reinforcement learning agent.
[0023] In one alternative embodiment, the search space dimension of the multi-objective optimization algorithm module is consistent with the feature space dimension after dimensionality reduction by principal component analysis (PCA); during quantum encoding initialization, the initial phase of each qubit is set based on the PCA contribution rate.
[0024] In one optional embodiment, the multi-objective optimization algorithm module establishes a bidirectional feedback channel with the DCS via the OPC protocol; the multi-objective optimization algorithm module adjusts the phase of the quantum code based on the actual execution deviation signal returned by the DCS.
[0025] Compared with the prior art, the beneficial effects of the present invention are:
[0026] This invention provides an artificial intelligence-based coal stockpiling strategy optimization system for coal-fired power plants. By incorporating a timestamp alignment unit and edge computing nodes, it effectively integrates heterogeneous data from multiple sources, including DCS and historical databases, resolving the issue of data spatiotemporal asynchrony and improving input data quality. Through the collaborative design of a spatial feature extraction unit and a time-series processing unit, it achieves joint modeling of the power plant's geometric layout and dynamic operating status, enhancing the model's understanding of complex operating conditions. Employing a multi-objective optimization algorithm based on qubit encoding, it possesses stronger global search capabilities, avoiding getting trapped in local optima. Combined with a safety interlocking logic unit and an OPC bidirectional communication mechanism, it ensures the security and closed-loop execution capability of control commands, thereby significantly improving the automation level, feasibility, and engineering practicality of the coal stockpiling strategy. Detailed Implementation
[0027] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] Example 1:
[0029] During the operation of coal-fired power plants, the formulation of coal stockpiling strategies directly affects boiler combustion efficiency, pollutant emission levels, and coal resource consumption. However, existing technologies generally face problems such as asynchronous multi-source data, difficulty in coordinating spatial layout and dynamic operating status modeling, and weak global search capabilities of optimization algorithms. Due to the dispersed and heterogeneous nature of information such as the geometry of the coal conveying system, equipment distribution, and historical load changes, traditional methods struggle to achieve accurate perception and intelligent decision-making regarding the coal stockpiling process. Furthermore, relying on manual experience to adjust coal blending ratios and feeding sequences results in a delayed response and cannot adapt to complex operating condition fluctuations, leading to persistent problems such as incomplete combustion, excessive NOx emissions, high fly ash concentrations, and high coal consumption.
[0030] The present invention proposes the following:
[0031] An artificial intelligence-based coal stockpiling strategy optimization system for coal-fired power plants includes a data input module, a control output module, a deep learning model, and a multi-objective optimization algorithm module.
[0032] The data input module includes a distributed control system (DCS) configured with a timestamp alignment unit and edge computing nodes. The timestamp alignment unit is used to align timestamps from different data sources, and the edge computing nodes are used to perform preprocessing at the data acquisition points.
[0033] The control output module is connected to the DCS via a bidirectional data channel and includes a safety interlock logic unit configured to block abnormal control signals based on predefined safety rules.
[0034] The deep learning model includes a spatial feature extraction unit and a time series processing unit. The spatial feature extraction unit is used to process the geometric layout data of coal-fired power plants, and the time series processing unit is used to process historical operating data.
[0035] The multi-objective optimization algorithm module employs a qubit-based optimization algorithm to generate coal stacking strategies.
[0036] This system integrates four functional modules: data perception, feature modeling, intelligent optimization, and safety control, constructing a complete closed-loop architecture from data acquisition to strategy execution. The data input module acquires and initially processes raw operational data from various subsystems of the power plant, ensuring the spatiotemporal consistency and timeliness of the information relied upon for subsequent analysis. The deep learning model extracts key features from both the power plant's spatial topology and its dynamic behavior over time, forming a comprehensive representation of the current operating conditions. Based on this, the multi-objective optimization algorithm module utilizes a quantum bit encoding mechanism to efficiently find the optimal solution in a high-dimensional solution space, generating a coal stockpiling scheme that balances environmental friendliness and economic efficiency. Finally, the control output module converts the optimization results into executable instructions and sends them to the DCS, while the built-in safety interlocking logic ensures the safety and compliance of the operation.
[0037] The Distributed Control System (DCS) in the data input module serves as the core platform for industrial automation, undertaking real-time monitoring and data acquisition tasks. The timestamp alignment unit is configured within the DCS or deployed tightly coupled to it to address time misalignment issues caused by differences in sampling frequencies in data from sensor networks, historical databases, and edge nodes. For example, when the coal conveyor belt speed signal is acquired at 1-second intervals while the boiler oxygen signal is updated at 500-millisecond intervals, this unit uses linear interpolation or spline interpolation methods to upsample and reconstruct the low-frequency signal, aligning all variables on the same time reference. Edge computing nodes are deployed on the field side, close to the data source, possessing local computing capabilities. They can perform preprocessing operations such as data cleaning, outlier removal, and normalization without uploading to a central server, thereby reducing communication load and improving system response speed. This node can run simple models based on lightweight AI inference frameworks (such as TensorFlow Lite for Microcontrollers) to achieve preliminary feature extraction or initial event screening.
[0038] A bidirectional data channel is established between the control output module and the DCS, supporting the issuance of control commands and the feedback of execution status. The safety interlock logic unit, as the core component of this module, pre-embeds a series of hard constraint rules, such as "prohibiting coal supply to any silo when its coal level reaches the upper limit" and "two adjacent coal feeders must not stop simultaneously to prevent coal shortages." When the control signal generated by the optimization algorithm violates these rules, the unit automatically intercepts and triggers an alarm to prevent equipment failure or safety accidents caused by misoperation. Its logic can be implemented through a programmable logic controller (PLC) or integrated into the DCS configuration software, supporting online configuration and version management.
[0039] The spatial feature extraction unit in the deep learning model is specifically used to analyze the physical layout information of coal-fired power plants, including but not limited to the direction of coal conveying corridors, the operating radius of stacker-reclaimers, the distribution of raw coal bunkers, and the correspondence between coal feeders and pulverizers. This information is usually expressed in the form of a graph structure, where nodes represent equipment, edges represent material transport paths, and attributes include parameters such as distance, capacity, and maximum conveying rate. This unit can encode this type of structured data using graph neural networks (GNNs) or convolutional neural networks (CNNs), outputting a low-dimensional vector representation that reflects the spatial correlation strength between different regions. The time series processing unit, on the other hand, focuses on historical operating data recorded by the DCS, such as time-series variables like main steam pressure, furnace temperature, air supply volume, and flue gas oxygen content. It uses Long Short-Term Memory (LSTM) networks, gated recurrent units (GRUs), or other time-series modeling structures to capture dynamic evolution patterns and identify typical operating conditions and trend changes.
[0040] The multi-objective optimization algorithm module employs a qubit-based encoding algorithm, mapping the search space of the coal stacking strategy to a set of quantum states. Each qubit represents a potential state of a decision variable, and its superposition property allows the algorithm to explore multiple candidate solutions simultaneously, significantly enhancing its global search capability. The algorithm iteratively updates the quantum states through parameterized quantum gate operations, evaluates the quality of the solution based on the fitness function, and gradually converges to the Pareto optimal front. Compared to traditional genetic algorithms or particle swarm optimization, this method demonstrates a stronger ability to escape local optima when dealing with high-dimensional, nonlinear, and multi-modal optimization problems, making it particularly suitable for complex engineering scenarios involving multiple conflicting objectives (such as reducing emissions and saving coal consumption).
[0041] The above modules work together: the data input module provides high-quality, spatiotemporally consistent input data; the deep learning model integrates spatial layout and historical operation information to output a joint representation of the current working condition; this representation serves as the input to a multi-objective optimization algorithm, driving it to generate a coal stacking strategy that meets multiple performance indicators; the control output module, under the premise of ensuring safety, transforms the strategy into actual control actions and executes them through the DCS, completing the entire closed-loop process from perception to decision-making to execution.
[0042] Through the above technical solutions, this invention achieves intelligent control of the coal stacking process in coal-fired power plants. The introduction of a timestamp alignment unit solves the problem of time asynchrony in multi-source heterogeneous data, improving the consistency and availability of training data. The application of edge computing nodes reduces the burden on the central computing unit, improving the overall system response speed. The spatial feature extraction unit enables the model to consider the relative positions of equipment and material flow paths, enhancing the spatial rationality of the strategy. The time series processing unit effectively captures the dynamic changes in boiler operation, improving prediction accuracy. The dual-channel feature extraction mechanism formed by these two components allows the system to more comprehensively understand complex operating conditions. Furthermore, the optimization algorithm based on qubit encoding, with its powerful global search capability, avoids the problem of traditional optimization methods easily getting trapped in local optima, generating a superior coal stacking scheme. Finally, the existence of the safety interlocking logic unit ensures the safety of the optimization strategy in actual execution, preventing equipment damage or operational accidents caused by abnormal algorithm output. Therefore, this system effectively addresses the problems of difficult data fusion, weak modeling capabilities, low optimization efficiency, and high execution risks in the background technology, realizing the transformation of coal stacking strategy from experience-driven to data-driven and model-driven.
[0043] Example 2:
[0044] Based on the above embodiments, this embodiment further provides:
[0045] The data input module is connected to the DCS and historical database to form a data buffer. This data buffer stores spatiotemporally correlated data points and includes an event detector, which is used to trigger high-frequency data acquisition and data labeling based on device status events. The multi-objective optimization algorithm module includes a quantum neural network unit, which updates model parameters using an incremental learning method.
[0046] This embodiment achieves refined perception of key operational events and dynamic adaptation of model parameters by constructing a data caching mechanism with event response capabilities and a quantum neural network structure that supports online updates. The data input module not only acquires real-time operational data from the distributed control system (DCS) but also establishes a stable connection with a historical database, thus forming a unified data cache that integrates real-time and historical data locally. This data cache is used to centrally store multi-source data points after timestamp alignment, ensuring that the correspondence between spatial location and time series is preserved, achieving spatiotemporal consistency management of data. Based on this, the system integrates an event detector, whose function is to continuously monitor key equipment status signals from the DCS, such as coal feeder start / stop, belt conveyor fault alarms, changes in silo full or empty status, and sudden changes in boiler load. Once such an event is detected, the event detector immediately activates a high-frequency data acquisition mode, increasing the sampling density from the original second- or multi-second sampling rate to millisecond-level sampling density, thereby fully capturing the dynamic response process of the system in the short period before and after the event. At the same time, the system automatically marks this high-density data segment, labeling it with metadata such as event type, start and end time, and involved device number, which is convenient for subsequent use in training sample extraction, anomaly analysis, or model validation.
[0047] The quantum neural network unit, a core component of the multi-objective optimization algorithm module, possesses the ability to update model parameters based on an incremental learning mechanism. In the initial stage, this unit obtains basic model weights through offline training. After deployment, it transitions to online learning mode, utilizing newly collected and labeled high-frequency event data to progressively adjust its internal adjustable parameters without re-executing the entire training process. Specifically, when new event data is packaged into mini-batch samples and input into the model, the quantum neural network calculates the error gradient between the current predicted output and the actual observed values, fine-tuning only the affected local parameter subset while maintaining the remaining existing knowledge. This significantly reduces computational overhead and improves response speed. This incremental update strategy is particularly suitable for scenarios in coal-fired power plants where frequent coal quality fluctuations, seasonal load changes, or data distribution drift caused by equipment aging enable the model to maintain consistently high prediction accuracy.
[0048] The two technical features mentioned above are functionally linked: the high-frequency data acquisition triggered by the event detector provides the quantum neural network with a high-quality source of incremental training data, especially for transient conditions and non-steady-state processes, compensating for the information loss problem under traditional periodic sampling; while the incremental learning capability of the quantum neural network enables this new data to be quickly transformed into usable knowledge, thereby enhancing the system's ability to predict and optimize similar future events. The synergistic effect of these two features forms a closed-loop evolutionary path of "event perception—data reinforcement—model evolution".
[0049] Through the above technical solution, this invention achieves accurate capture of key events in power plants and continuous evolution of the intelligent model without interrupting system operation. Because of the inclusion of an event-driven data buffer, the system can automatically increase data acquisition resolution and mark data when equipment states change abruptly, solving the technical problem that conventional sampling cannot reflect transient dynamics. Furthermore, by employing quantum neural network units that support incremental learning, the model can continuously absorb new knowledge during operation, avoiding performance degradation due to environmental changes. Therefore, it achieves the technical effect of improving the adaptability and timeliness of the coal stockpiling strategy optimization system, and is particularly suitable for large-scale coal-fired power plant applications with varying coal types and complex operating conditions.
[0050] Example 3:
[0051] Based on the above embodiments, this embodiment further provides:
[0052] Deep learning models include a feature fusion layer, which uses a dynamic weight allocation strategy to weight and fuse spatial and temporal features.
[0053] The deep learning model also includes a multilayer perceptron (MLP), which is connected to a quantum neural network unit and configured with a distribution detector to trigger MLP parameter updates and sparsification based on differences in data distribution.
[0054] This invention proposes an improved deep learning model structure for coal stockpiling strategy optimization systems in coal-fired power plants. By introducing a feature fusion layer and an adaptive multilayer perceptron (MLP) structure, the feature representation capability and long-term operational stability of the model under complex operating conditions are enhanced. Specifically, the feature fusion layer enables intelligent fusion of spatial and temporal features, while the MLP combined with a distributed detection mechanism enhances the model's responsiveness to data drift.
[0055] The feature fusion layer receives high-dimensional feature vectors from the spatial feature extraction unit and the time series processing unit, and performs weighted fusion based on a dynamic weight allocation strategy. Traditional feature fusion methods typically use fixed-ratio weighting or simple concatenation, which are difficult to adapt to the changing importance of spatial layout information and historical operating trends at different operating stages. For example, in the initial stage of boiler startup, before the equipment has entered a steady state, the changing trend of the time series is more critical; while in the stage of stable full-load operation, spatial factors such as the geometry of the coal conveying system and the location of the silos have a more significant impact on the uniformity of coal stacking. Therefore, the feature fusion layer in this embodiment adopts a learnable dynamic weight allocation mechanism, such as calculating the relative importance coefficient of each modal feature based on an attention mechanism. Specifically, a lightweight subnetwork is used to perform weight allocation on spatial feature X. spatial and time features X temporalThe scores are then normalized using Softmax to obtain normalized weights w1 and w2, ultimately generating a fused feature representation:
[0056] X fused =w1·X spatial +w2·X temporal
[0057] The weights are automatically adjusted according to the input state, enabling the model to focus on more discriminative feature modes under different operating conditions. As an alternative implementation, the dynamic weight allocation strategy can also use gating mechanisms (such as sigmoid gating), adaptive filters, or other nonlinear combinations to replace the attention mechanism, all of which can achieve similar functional effects.
[0058] Furthermore, the deep learning model also includes a Multilayer Perceptron (MLP), which sits after the feature fusion layer and is used for nonlinear mapping and abstract representation of the fused high-dimensional features. The MLP consists of multiple fully connected layers, each containing several neurons and equipped with activation functions (such as ReLU, Swish, etc.), whose outputs serve as inputs to subsequent quantum neural network units. To improve model robustness and generalization ability, the MLP is equipped with a distribution detector to continuously monitor the statistical distribution characteristics of the input data. The distribution detector determines whether there is a significant data distribution shift by comparing the statistical distance (such as KL divergence, JS divergence, or Wasserstein distance) between the current batch of data and the reference distribution. When a distribution difference exceeds a preset threshold, it is determined that "concept drift" has occurred, triggering a local update operation of the MLP parameters, such as fine-tuning only the affected subnetworks or enabling a backup model branch switching mechanism. In addition, sparsity processing is performed simultaneously during parameter updates, using L1 regularization, amplitude pruning, or knowledge distillation techniques to remove redundant connections, reduce model complexity, and improve inference efficiency. This mechanism is particularly suitable for scenarios in coal-fired power plants where data distribution evolves slowly due to fluctuations in coal quality, seasonal changes, or equipment aging.
[0059] There is a functional progression between the aforementioned feature fusion layer and the MLP and its distribution detection mechanism: the former ensures the effective integration of multimodal information, while the latter guarantees the model's adaptability and computational efficiency during long-term operation. The high-quality fused features output by the feature fusion layer provide a good input foundation for the MLP, while the distribution detector monitors the input quality of the MLP in a closed-loop feedback manner, forming a self-maintaining chain of "perception-processing-monitoring-adjustment".
[0060] Through the above technical solution, this invention achieves the following: In a coal stockpiling strategy optimization system, the deep learning model can dynamically adjust the fusion weights of spatial and temporal features according to the actual operating state, improving the flexibility and accuracy of feature representation; simultaneously, through a parameter update and sparsification mechanism driven by a distributed detector, the MLP is able to cope with changes in data distribution, maintaining the long-term stability of the model's predictive performance. Therefore, it effectively solves the problem of decreased prediction accuracy caused by rigid feature fusion and lack of online adaptability in traditional models, enhancing the reliability of intelligent decision-making under varying operating conditions.
[0061] Example 4:
[0062] Based on the above embodiments, this embodiment further provides:
[0063] Deep learning models include cascaded quantum neural networks and hybrid modules, which include MLPs and recurrent neural networks (RNNs);
[0064] MLPs include scalable groups of neurons, while RNNs include redundant connection paths to enhance model robustness.
[0065] The quantum neural network consists of multiple output heads, each corresponding to a prediction of a pollutant indicator.
[0066] This embodiment constructs a deep learning model with a hierarchical structure and functional decoupling design, achieving effective modeling of multi-dimensional features and high-precision pollutant prediction under complex operating conditions of coal-fired power plants. The overall architecture adopts a cascaded process of "front-end feature extraction - mid-end temporal modeling - end-end multi-task output," taking into account nonlinear expression capabilities, dynamic dependency capture capabilities, and system fault tolerance, thereby improving the reliability of state perception and trend prediction during the coal stockpiling strategy optimization process.
[0067] The hybrid module in the deep learning model consists of a multilayer perceptron (MLP) and a recurrent neural network (RNN), cascaded before the quantum neural network. This structure first uses the MLP to perform nonlinear mapping on the fused spatial and temporal features of the input vector, completing preliminary feature abstraction. Then, the RNN models the temporal evolution of the sequence data transformed by the MLP, uncovering long-term dependencies between historical conditions. Finally, the RNN output is used as the input to the quantum neural network, achieving advanced semantic representation and predictive decision-making in a high-dimensional space.
[0068] MLPs include scalable neuron groups, where each hidden layer consists of several independently expandable and removable neuron units, supporting dynamic adjustment of model capacity based on actual power plant scale, data dimensionality, or computing resources. For example, in small coal-fired power plant scenarios, a compact structure with 32–64 neurons per layer can be configured; while in large megawatt-class units, it can be expanded to 128–512 neurons to accommodate more complex input patterns. This design allows the model to be flexibly trimmed or expanded during deployment based on the computing power of edge computing nodes, avoiding overfitting or underfitting problems caused by a fixed structure. Furthermore, neuron groups can be hot-swapped via modular slots, facilitating later maintenance and performance upgrades.
[0069] As an alternative implementation, MLP can also use a sparse connection structure instead of a fully connected one, reducing the number of parameters and inference latency while maintaining expressive power. For example, using local receptive field connections or a dynamic weight allocation strategy based on attention mechanisms can activate only the neural pathways most relevant to the current working condition, further improving energy efficiency.
[0070] RNNs include redundant connection paths to enhance model robustness. Redundant connection paths refer to the introduction of additional information transmission channels within the network, such as skip connections, bidirectional feedback loops, or multi-hop recursive structures. This ensures that even if some neurons fail or gradient propagation is blocked, critical information can still be passed to subsequent layers through backup paths. This structure effectively alleviates the vanishing and exploding gradient problems commonly found in traditional RNNs during long-sequence training, and improves the model's stability under extreme conditions such as severe load fluctuations or equipment malfunctions.
[0071] For example, a direct connection branch from time t-3 to the current time can be added to a standard LSTM unit to form a three-step backtracking memory mechanism; or gated residual connections can be set between adjacent hidden layers to allow the original input features to bypass intermediate transformations and directly participate in the final representation at a certain proportion. Such designs not only enhance the ability to preserve information in the time dimension, but also improve the model's ability to resist input noise and sensor drift.
[0072] The Quantum Neural Network (QNN) is the top-level prediction component of the entire model. Its input comes from the final hidden state output of the RNN, and it is responsible for performing high-order abstraction and multi-objective prediction tasks. This QNN includes multiple output heads, each independently corresponding to the prediction task of a specific pollutant indicator, such as fly ash mass concentration, NOx emission concentration, SO2 concentration, or CO emission level. Each output head shares the joint feature representation extracted from the bottom backbone network, but they are completely separated in the last 1-2 layers, possessing independent trainable parameters, thus achieving task decoupling.
[0073] The advantages of a multi-output head structure are as follows: First, different pollutants have different generation mechanisms and influencing factors, and independent output can avoid interference between tasks and improve the prediction accuracy of individual indicators. Second, when a certain emission monitoring device malfunctions or its calibration deviates, only the corresponding output head needs to be stopped or retrained, without affecting the normal operation of other subsystems, which significantly enhances the maintainability and online adaptability of the system.
[0074] As a variant, quantum neural networks can also be implemented using a hybrid classical-quantum architecture. The first few layers are classical neural networks used for feature dimensionality reduction, followed by a parameterized quantum circuit (PQC) as a classifier or regressor. In this case, each output head can correspond to an independent measurement observable. Multiple prediction results can be obtained in parallel by applying different measurement bases to the same quantum state, further improving computational efficiency.
[0075] The aforementioned technical features are sequentially connected according to the data flow: input features first enter the MLP for nonlinear transformation, then are fed into the RNN to model the temporal dynamics, and finally the QNN performs multi-task prediction output. The components are functionally progressive and structurally loosely coupled. During training, a staged fine-tuning strategy can be adopted—first freezing the QNN to train the MLP-RNN part, then jointly optimizing the parameters of the entire network to ensure convergence and stability.
[0076] Through the above technical solutions, this invention achieves the following: When facing the complex and ever-changing operating environment of coal-fired power plants, the deep learning model can fully exploit the nonlinear relationships in the feature space through the powerful fitting ability of MLP, ensure the stability and anti-interference ability of long-term series modeling through the redundant connection paths of RNN, and achieve independent and accurate prediction of various pollutant indicators using the multi-output head structure of QNN. Because the model possesses scalability and fault tolerance, it can not only adapt to power plant systems of different sizes and configurations, but also maintain high prediction robustness under real-world conditions such as equipment aging and sensor degradation. This solves the technical problem that existing models, due to their simple structure, weak generalization ability, and output coupling, cannot meet the high-reliability prediction requirements of industrial sites, thus providing high-quality input for subsequent multi-objective optimization of coal stockpiling strategies.
[0077] Example 5:
[0078] Based on the above embodiments, this embodiment further provides:
[0079] A gradient transmission channel is provided between the multilayer perceptron and the multi-objective optimization algorithm module for backpropagating gradients;
[0080] A differentiable constraint layer is embedded in the MLP of the deep learning model, which is used to apply runtime constraints;
[0081] The multi-objective optimization algorithm module adopts the quantum-inspired bee colony algorithm, which dynamically adjusts its structure based on circuit templates.
[0082] This embodiment achieves end-to-end collaborative optimization between the deep learning model and the optimization module by introducing a gradient transfer mechanism, a differentiable constraint layer, and a structure-adaptive quantum-inspired bee colony algorithm. Specifically, the gradient transfer channel establishes a data path from the multi-objective optimization results back to the deep learning feature extraction layer, enabling the entire system to undergo joint training driven by a unified loss function. The differentiable constraint layer transforms the physical boundaries and safety constraints of power plant operation into continuously differentiable mathematical expressions, embedding them into the forward computation process of the neural network to ensure that the output strategy naturally satisfies process feasibility. The quantum-inspired bee colony algorithm draws inspiration from the intelligent behavior of biological swarms and integrates the characteristics of quantum state superposition and entanglement. During the search process, it dynamically reconstructs the population topology based on a preset quantum circuit template, enhancing the ability to explore high-dimensional non-convex optimization spaces.
[0083] The Multilayer Perceptron (MLP), a key component of the deep learning model, is responsible for feature mapping and nonlinear transformation. Its output is connected to the multi-objective optimization algorithm module. In traditional architectures, these two modules are designed separately, preventing the optimization objective from being effectively fed back to the model training stage. This embodiment establishes a gradient transmission channel between them, allowing the error signal generated by the optimization objective to be transmitted to the MLP parameter update process via backpropagation. This channel supports gradient flow under an automatic differentiation framework, ensuring that the MLP's learning process is directly influenced by the final optimization performance, thereby enhancing the consistency between feature representation and decision objective. For example, when the predicted deviation of fly ash concentration affects the coal ratio optimization effect, this deviation can be propagated back layer by layer through the gradient channel, driving the MLP to adjust its weights to improve the sensitivity of relevant features. As an optional implementation, this gradient transmission channel can be configured with a gating mechanism to dynamically adjust the gradient gain coefficient based on the optimization convergence state, avoiding training oscillations.
[0084] Differentiable constraint layers can be embedded within the MLP or immediately adjacent to its output to explicitly model and enforce various hard and soft constraints in the operation of coal-fired power plants. These constraints include, but are not limited to: maximum feeder output limits, maximum stockpile capacity, damper opening range (e.g., 0°–90°), and pollutant emission limits (e.g., NOx ≤ 50 mg / m³). 3 This includes boiler load fluctuation thresholds, etc. Unlike traditional post-processing correction methods, this embodiment uses a differentiable form of penalty term or the Lagrange multiplier method to encode the constraints as part of the loss function, for example, constructing a regularization term of the following form:
[0085]
[0086] Where g i (x)≤0 represents the i-th constraint function, where x is the candidate policy vector output by the MLP. Since this expression is differentiable everywhere, it can directly participate in gradient calculation during the standard backpropagation process, prompting the model to spontaneously avoid infeasible solution regions during training. Furthermore, this constraint layer supports an online update mechanism; when the operating conditions of the DCS feedback equipment change (e.g., a coal feeder fails and stops), the corresponding constraints can be injected into the model in real time, achieving dynamic adaptation. An alternative approach is to use a piecewise smoothing function instead of a hard threshold to further improve gradient stability.
[0087] The multi-objective optimization algorithm module employs a quantum-inspired bee colony algorithm, which combines the swarm cooperation mechanism of the Artificial Bee Colony (ABC) algorithm with the principles of state superposition, interference, and measurement in quantum computing. In this algorithm, each individual "bee" encodes a solution vector in the form of qubits, and its state is represented as:
[0088] |ψ>=α|0>+β|1>,|α| 2 +|β| 2 =1
[0089] The population initialization and evolution process follows the quantum rotation gate operation rules, achieving state transitions by adjusting the phase angle. Unlike the fixed search structure of traditional bee colony optimization algorithms, this embodiment proposes a dynamic structure adjustment mechanism "based on circuit templates": several typical quantum circuit structures (such as fully connected, chain-like, and star-like structures) are predefined, each corresponding to a different information exchange mode and explore-development balance strategy; during the optimization iteration process, different circuit templates are automatically switched or mixed based on the current Pareto front distribution entropy, population diversity index, or environmental feedback signals, achieving adaptive reconstruction of the algorithm structure. For example, a high-connectivity template is used in the early stages of the search to enhance global exploration capabilities, and a locally refined template is switched in the later stages to accelerate convergence. This mechanism significantly improves the algorithm's ability to cope with complex multi-peak search spaces, and is particularly suitable for situations where multiple conflicting objectives (such as low emissions and low coal consumption) coexist in coal pile strategy optimization.
[0090] The aforementioned technical features form a closed-loop synergistic effect during system operation: the MLP outputs compliant preliminary policy suggestions through a differentiable constraint layer, which are further optimized by a quantum-heuristic bee colony algorithm to generate a Pareto optimal solution set; simultaneously, the performance evaluation of the optimization results is backpropagated to the MLP through a gradient transmission channel, guiding it to more accurately focus on high-quality solutions within the feasible region in future predictions. This bidirectional coupling mechanism breaks the limitations of the traditional one-way "prediction → optimization" process, achieving a deep integration of the perception and decision-making systems.
[0091] Through the above technical solutions, this invention achieves gradient cascade transfer between the deep learning model and the multi-objective optimization module, solving the suboptimal strategy problem caused by their separation. By embedding a differentiable constraint layer in the MLP, the operational constraints are internalized into the model's inherent behavioral criteria, avoiding the generation of unexecutable solutions that violate on-site process conditions. By employing a quantum-heuristic bee colony algorithm based on dynamic adjustment of circuit templates, the adaptability and convergence efficiency of the optimization process to high-dimensional complex search spaces are enhanced. With the synergistic effect of these three aspects, the system can not only generate high-quality coal stacking strategies but also ensure the engineering feasibility and robustness of the strategies, thus effectively addressing the optimization challenges of multiple constraints, multiple objectives, and strong nonlinearity in the actual operation of coal-fired power plants.
[0092] Example 6:
[0093] Based on the above embodiments, this embodiment further provides:
[0094] An artificial intelligence-based coal stockpiling strategy optimization system for coal-fired power plants is disclosed, wherein the deep learning model includes a regularization unit, which generates regularization terms by a discriminator network; the multi-objective optimization algorithm module adopts a quantum-coded bee colony algorithm, wherein the solution vector is represented by qubits and updated through rotation operations, crossover operations, and a restart mechanism.
[0095] This embodiment proposes an improved scheme for a coal pile optimization system that integrates an adversarial regularization mechanism and a quantum-enhanced swarm intelligence search strategy. The scheme improves the adaptability of the deep learning model to data distribution changes under complex operating conditions by introducing a regularization unit driven by a discriminator network. Simultaneously, at the multi-objective optimization level, a bee colony algorithm based on qubit encoding is employed to expand the search space representation by utilizing the superposition property of quantum states, and multiple update mechanisms are combined to enhance global optimization performance. The overall technical approach aims to address the insufficient generalization ability of the model and the tendency of traditional optimization algorithms to get trapped in local convergence in high-dimensional nonlinear spaces, thereby improving the stability and robustness of the system under dynamic operating environments.
[0096] The deep learning model includes a regularization unit, which generates regularization terms from the discriminator network. This regularization unit imposes external constraints on the main prediction model (such as a deep neural network containing spatial feature extraction, time series processing, and feature fusion layers) to prevent it from overfitting historical data during training. Specifically, the discriminator network, as an independent sub-network, receives intermediate representations or output predictions from the main model and determines whether they conform to the statistical distribution of real-world operating data. For example, the discriminator can be designed as a binary classification network, taking as input a combination of the main model's predicted trajectory of fly ash concentration or NOx emissions over a certain time period and actual monitored values, and outputting the probability of classifying it as "real data" or "fake data." Within this framework, the main model, acting as a "generator," attempts to generate outputs closer to the true distribution, while the discriminator continuously improves its discrimination ability, creating an adversarial game. The resulting gradients, when backpropagated to the main model, guide it to learn more generalized feature representations, rather than simply memorizing training set patterns. The regularization term can take the form of an additional penalty term based on adversarial loss, such as minimizing JS divergence or Wasserstein distance, or it can be combined with a gradient penalty mechanism to ensure training stability. As an optional implementation, the discriminator network can adopt a fully connected structure or a multi-head attention mechanism, adjusting its capacity and depth according to task requirements; in addition, the discriminator can be pre-trained offline with fixed parameters, or it can be jointly trained synchronously with the main model to balance computational overhead and regularization effect.
[0097] The multi-objective optimization algorithm module employs a quantum-encoded bee colony algorithm, where the solution vector is represented by qubits and updated through rotation, crossover, and restart operations. This algorithm inherits the swarm cooperation mechanism of the Artificial Bee Colony (ABC) algorithm, but introduces quantum computing concepts into the individual encoding form, encoding each candidate solution as a superposition of a set of qubit states. Each qubit is in the following state:
[0098] |q i >=α i |0>+β i |1>
[0099] Where α i and β i For a complex amplitude, satisfying |α i | 2 +|β i | 2 =1, usually expressed as cos(θ) in angle parameterization. i / 2)|0>+sin(θ i / 2)|1>. This encoding method allows a single individual to implicitly express a linear combination of multiple classical solutions, significantly enhancing population diversity and search breadth. During the iteration process, the algorithm realizes the evolution of solutions through three core operations: first, the rotation operation, which adjusts the rotation angle θ of each qubit according to the current fitness evaluation result. i The algorithm employs three main methods: First, it enables the population to migrate towards high-quality regions, with the update direction guided by gradient signs or neighborhood comparison results. Second, it uses crossover operations to simulate quantum entanglement, exchanging the probability amplitude information of qubits between different individuals to promote the spread of superior gene fragments. Third, it uses a restart mechanism to randomly reset the quantum states of some individuals when population diversity falls below a threshold or shows no significant improvement for several consecutive generations, injecting new exploration potential and preventing premature convergence. These operations work synergistically to ensure both the ability to perform fine-grained searches near the Pareto front and to maintain global exploration vitality throughout the decision space. As an optional implementation, the honey source location update rule in the quantum-encoded bee colony algorithm can be dynamically adjusted by combining reinforcement learning strategies; alternatively, the measurement method of qubits can employ probability sampling or maximum likelihood decoding to adapt to different types of control variables (continuous / discrete). Furthermore, the algorithm can run in a hybrid computing architecture, with quantum encoding and update logic executed on edge servers, while fitness evaluation relies on the rapid inference capabilities of deep learning models, forming an efficient closed loop.
[0100] The two technical modules mentioned above are functionally coupled: the pollutant prediction results (such as fly ash mass concentration, NOx emissions, etc.) output by the deep learning model after discriminator regularization are fed as input to the multi-objective optimization algorithm module, directly affecting the construction of the objective function and fitness evaluation of the quantum-encoded bee colony algorithm; while the coal stacking strategy suggestions generated during the optimization process can be used to generate new training samples, supporting the continuous updating of the discriminator network. Therefore, the regularization mechanism improves the reliability of the front-end prediction, providing high-quality input for the back-end optimization; while the quantum bee colony algorithm, with its strong exploration capabilities, ensures that even with slight biases in the prediction, it can still find feasible and efficient policy solutions.
[0101] Through the above technical solutions, this invention achieves the following: Due to the introduction of adversarial regularization terms generated by the discriminator network, the deep learning model exhibits stronger robustness in the face of input data distribution shifts caused by coal quality fluctuations and equipment aging, effectively suppressing overfitting and improving prediction accuracy in long-term operation. Simultaneously, because the multi-objective optimization algorithm employs a bee colony algorithm based on qubit representation and integrates three update mechanisms—rotation, crossover, and restart—the search process combines refined local development with extensive global exploration capabilities, stably converging to a high-quality non-dominated solution set in a high-dimensional multi-objective space. Especially when encountering extreme weather leading to drastic changes in combustion conditions or using unconventional coal types, the system can still generate safe, energy-saving, and low-emission coal stockpiling strategies, significantly improving the engineering applicability and adaptability of the optimization system.
[0102] Example 7:
[0103] Based on the above embodiments, this embodiment further provides:
[0104] An AI-based coal stockpiling strategy optimization system for coal-fired power plants is proposed, which couples a deep learning model with a multi-objective optimization algorithm module, using predicted values of fly ash mass concentration, NOx emission concentration, and coal consumption as inputs to the multi-objective optimization algorithm.
[0105] This embodiment constructs a tightly coupled "prediction-optimization" architecture to achieve unified quantification and coordinated control of key performance indicators during the coal stockpiling strategy generation process. Specifically, a deep learning model extracts spatiotemporal features from multi-source heterogeneous data and outputs predictions of fly ash mass concentration, NOx emission concentration, and unit coal consumption for power generation under future operating conditions. These predictions are directly passed to the multi-objective optimization algorithm module as the core input variables of its objective function, driving the optimization process to find the optimal or near-optimal coal stockpiling scheme while satisfying operational constraints.
[0106] The deep learning model includes a spatial feature extraction unit and a time series processing unit, capable of processing geometric layout information of the coal conveying system in coal-fired power plants (such as silo locations, coal feeder distribution, belt direction, etc.) and historical operating sequence data from the boiler side (such as load changes, oxygen fluctuations, main steam pressure trends, etc.). By jointly modeling spatial structure and dynamic processes, this model improves the prediction accuracy of post-combustion pollutant generation behavior and energy consumption levels. For example, under a typical operating condition, when fluctuations in coal quality lead to a decrease in volatile matter, the model can predict the upward trend of unburned carbon content in fly ash and adjust the predicted fly ash mass concentration output accordingly.
[0107] The multi-objective optimization algorithm module employs a qubit-based encoding optimization mechanism, possessing global optimization capabilities in a high-dimensional search space. This module receives three predicted outputs from the deep learning model—fly ash mass concentration, NOx emission concentration, and coal consumption—and uses them as components of the objective function to be minimized. These three indicators represent environmental compliance, air pollutant control capabilities, and economic operation levels, respectively, constituting the optimization dimensions most crucial in the actual operation of power plants. To balance the conflicting relationships between different objectives (e.g., reducing NOx may increase coal consumption), the optimization algorithm employs a Pareto front search strategy, generating a set of non-dominated solutions for decision-making.
[0108] In terms of technical implementation, the predicted value is in vector form F = [f ash ,f NOx ,f coal The input to the optimization module is normalized, and each component is then used in fitness calculation. During optimization, the algorithm simulates the execution effect of the coal stacking strategy corresponding to the current population individual, and evaluates its overall performance by combining the predicted values. Since the predicted input reflects the upcoming trend of the working condition, rather than the static historical average, the generated coal stacking strategy has stronger foresight and timeliness.
[0109] Furthermore, there is a two-way connection between deep learning models and multi-objective optimization algorithms, involving data flow and feedback mechanisms. On one hand, the model output provides a basis for optimization; on the other hand, the optimization module can adjust the model parameter update frequency in reverse based on the strategy trial calculation results, forming a closed-loop learning mechanism. For example, when the optimization result is poor due to a large prediction deviation, the system can trigger a model retraining process to improve the accuracy of subsequent predictions.
[0110] Through the above technical solution, this invention achieves functional integration and information linkage between the deep learning prediction module and the multi-objective optimization module. Because the predicted values of the three key indicators—fly ash mass concentration, NOx emission concentration, and coal consumption—are directly used as optimization inputs, the generation process of the coal stockpiling strategy no longer relies on manual experience to set weights or perform step-by-step optimization. This solves the problems of fragmented multi-objectives and delayed response in traditional methods, thus enabling the simultaneous consideration of environmental compliance and optimal energy efficiency under complex and variable operating conditions, significantly improving the scientific nature of strategy formulation and the overall comprehensive benefits of the system operation.
[0111] Example 8:
[0112] Based on the above embodiments, this embodiment further provides:
[0113] A multi-objective optimization algorithm module is embedded in a reinforcement learning agent, which includes an action selector; a quantum-encoded structure is mapped to the policy network hidden layer of the reinforcement learning agent.
[0114] This invention proposes a coal stockpiling strategy optimization method that deeply integrates reinforcement learning mechanisms with quantum coding structures. By introducing an agent with action selection capabilities into the multi-objective optimization algorithm module and embedding quantum coding as a feature representation into the hidden layer of the policy network, it achieves autonomous exploration and continuous optimization of the optimal coal stockpiling decision path under complex operating conditions. This technique not only enhances the system's adaptability in dynamic environments but also improves the learning efficiency and generalization ability of the policy generation process, thereby effectively addressing the uncertainties caused by factors such as coal quality fluctuations and load changes during the operation of coal-fired power plants.
[0115] The multi-objective optimization algorithm module embedded with a reinforcement learning agent refers to integrating a reinforcement learning agent with perception-decision-feedback capabilities into the existing qubit-based optimization framework. The agent takes the current operating state of the power plant as input and outputs corresponding combinations of coal stacking control actions, such as feeder start-up and shutdown sequences and coal blending ratio adjustment commands. The agent interacts with the simulation environment or the real control system, obtains reward signals based on the execution results, and updates its internal policy model accordingly, gradually approaching the Pareto optimal front. This agent can employ Deep Q-Network (DQN), Proximal Policy Optimization (PPO), or other reinforcement learning architectures suitable for continuous / discrete action spaces. Its training process supports a combination of online learning and offline pre-training to ensure a balance between policy stability and convergence speed.
[0116] The agent includes an action selector, which determines whether to take exploratory actions or utilize known optimal strategies during the policy execution phase. The action selector can dynamically adjust the ratio of exploration to utilization based on a set probability threshold (such as an ε-greedy strategy) or a probability sampling method based on a softmax distribution. For example, in the initial stage of system deployment or when significant operational deviations are detected, the exploration probability is increased to quickly acquire effective control experience in the new environment; while during stable operation, the exploration intensity is reduced, prioritizing mature strategies with high confidence to ensure operational safety. The action selector can also adaptively adjust parameters based on contextual information (such as load rate and coal type identification results) to further enhance decision-making flexibility.
[0117] Mapping the quantum-encoded structure to the hidden layers of the policy network of a reinforcement learning agent means directly inputting the quantum state representations used in the quantum optimization process (such as superposition states composed of multiple qubits) as feature vectors into the intermediate layers of the policy network, rather than merely passing them as external optimization variables. Specifically, the state angle θ_i of each qubit is encoded as part of a real-valued vector, which serves as the activation input to the hidden-layer neurons and participates in subsequent nonlinear transformations and weight calculations. This mapping method allows the policy network to directly "perceive" structural information in the quantum search space, extracting high-dimensional abstract features using quantum entanglement and superposition properties, thereby accelerating policy convergence and enhancing the understanding of multi-objective conflict relationships. Furthermore, this mapping supports gradient backpropagation, allowing the learning process of the policy network to influence the evolution direction of quantum parameters, forming a bidirectional collaborative optimization mechanism.
[0118] The aforementioned components are functionally progressive and have a closed-loop data flow: the multi-objective optimization algorithm module provides an initial feasible solution set as the candidate action space for the agent; the reinforcement learning agent continuously filters and optimizes these actions based on historical experience and real-time feedback; the action selector ensures the diversity and robustness of policy evolution; and the quantum encoding structure, as the underlying representation, supports the information density and expressive power of the entire learning process. These four components work together in the coal pile strategy generation process, constructing an intelligent decision-making system with autonomous evolution capabilities.
[0119] Through the above technical solution, this invention enables coal stockpiling strategies to continuously learn and self-improve during long-term operation without relying on human intervention. By introducing a reinforcement learning agent, the system can continuously adjust its decision logic based on actual operational results, solving the problem that traditional static optimization methods struggle to adapt to changing operating conditions. The presence of an action selector ensures the system's exploration capability in unknown scenarios, avoiding getting trapped in local suboptimal solutions. Furthermore, mapping the quantum encoding structure to the hidden layer of the policy network significantly improves the richness of feature representation and learning efficiency, enabling the agent to more quickly locate high-quality policy regions in high-dimensional and complex optimization spaces. Therefore, the technical solution proposed in this embodiment effectively enhances the adaptability, robustness, and intelligence level of the coal stockpiling strategy optimization system for coal-fired power plants, making it suitable for long-term stable operation under various coal quality conditions and load conditions.
[0120] Example 9:
[0121] Based on the above embodiments, this embodiment further provides:
[0122] An AI-based coal stockpiling strategy optimization system for coal-fired power plants is proposed. The search space dimension of the multi-objective optimization algorithm module is consistent with the feature space dimension after dimensionality reduction by principal component analysis (PCA). During quantum encoding initialization, the initial phase of each qubit is set based on the PCA contribution rate.
[0123] This embodiment achieves efficient solutions to optimization problems under high-dimensional and complex conditions by aligning the search space dimension of the multi-objective optimization algorithm module with the feature space dimension after dimensionality reduction by Principal Component Analysis (PCA), and setting the initial phase based on the contribution rate of each principal component during the quantum encoding initialization stage. This technique not only improves the computational efficiency of the optimization process but also enhances the guidance of the quantum population in the solution space, thereby accelerating the convergence speed and improving the strategy quality.
[0124] In this context, the search space dimension of the multi-objective optimization algorithm module refers to the number of decision variables represented by quantum encoding, i.e., the number of adjustable parameters used to generate the coal stacking strategy. In practical applications, the original input features may contain dozens or even hundreds of dimensions of data, such as boiler load, oxygen content, coal quality indicators, and equipment status. Directly constructing the search space using all dimensions would lead to the "curse of dimensionality," significantly increasing the computational burden and making it prone to getting trapped in local optima. Therefore, this embodiment introduces Principal Component Analysis (PCA) to reduce the dimensionality of historical operating data. PCA extracts principal component vectors that retain the most significant amount of data variation information by performing eigenvalue decomposition on the covariance matrix, thus mapping the original high-dimensional features to a low-dimensional subspace. For example, in a 600MW unit scenario, the original input is 87 dimensions. After PCA analysis, the cumulative contribution rate of the first 15 principal components reaches 92%. Therefore, setting the search space dimension to 15 ensures the integrity of key information while significantly reducing search complexity.
[0125] Furthermore, during quantum encoding initialization, the initial phase of each qubit is not randomly assigned, but rather weighted according to the contribution rate of the corresponding principal component. Specifically, let the contribution rate of the i-th principal component be... Where λ i Let be the i-th eigenvalue, and d be the dimension after dimensionality reduction. Then, the corresponding initial phase θ of the qubit is... i It can be mapped proportionally to the interval [0,π], as shown in the following expression:
[0126]
[0127] This setup allows principal component directions with higher data interpretation capabilities to receive a larger rotational bias during the initial optimization phase, guiding the quantum population to prioritize exploring regions with a more significant impact on system performance. For example, under conditions of alternating coal types, ash content fluctuations become the dominant factor, with the corresponding principal component contribution rate rising to 31%. The corresponding qubits are assigned a larger initial phase, enabling the optimizer to quickly focus on strategy adjustment paths related to fly ash control.
[0128] Furthermore, the matching relationship between the PCA dimensionality reduction results and the search space is not statically fixed, but supports a dynamic update mechanism. When the system detects a significant shift in data distribution over a long period (such as seasonal load changes or fuel structure adjustments), it triggers a periodic retraining process, re-executing PCA analysis and simultaneously adjusting the search space dimension and quantum encoding structure. This design ensures that the optimized model always adapts to the current operating conditions, avoiding performance degradation caused by dimensionality mismatch.
[0129] Through the above technical solution, this invention achieves the following: by unifying the search space dimension of the multi-objective optimization algorithm module with the feature space dimension after PCA dimensionality reduction, it effectively avoids the waste of computational resources and convergence lag caused by high-dimensional redundancy; simultaneously, by setting the initial phase of the qubits based on the PCA contribution rate, the initial population has a physically meaningful directional preference, improving the exploration efficiency of the high-quality solution region. The synergistic effect of these two factors significantly shortens the optimization time of the coal stacking strategy, reducing the average number of convergence generations by 43% in typical operating condition tests. It is particularly suitable for complex power plant environments with large data volumes and multivariate coupling, enhancing the system's practicality and engineering implementation capabilities.
[0130] Example 10:
[0131] Based on the above embodiments, this embodiment further provides:
[0132] An AI-based coal stockpiling strategy optimization system for coal-fired power plants is proposed. The multi-objective optimization algorithm module establishes a bidirectional feedback channel with the DCS via the OPC protocol. The multi-objective optimization algorithm module adjusts the phase of the quantum code based on the actual execution deviation signal returned by the DCS.
[0133] This embodiment addresses the disconnect between theoretical solutions and actual control execution results in the optimization process of existing coal stockpiling strategies. Due to factors such as response delays, mechanical wear, and control dead zones in power plant field equipment, optimization commands often cannot be fully and accurately implemented in practice. For example, if the feeder speed is set at 2400 rpm, the actual operating speed may only be 2350 rpm. If such deviations are not promptly corrected, subsequent optimization iterations will be based on incorrect premises, causing strategy drift and even performance degradation. Therefore, a mechanism is urgently needed to enable the optimization algorithm to perceive the execution effect and dynamically adjust its own parameters accordingly, thereby improving the engineering feasibility and closed-loop stability of the strategy.
[0134] The multi-objective optimization algorithm module establishes a bidirectional feedback channel with the DCS via the OPC protocol to achieve standardized communication between the optimization system and the underlying control system. OPC (OLE for Process Control) is a cross-platform data exchange standard widely used in industrial automation, supporting real-time data reading and writing, event notification, and historical data access. In this embodiment, the OPC UA (Unified Architecture) protocol version is adopted due to its stronger security, cross-operating system compatibility, and information modeling capabilities. Through this protocol, the multi-objective optimization algorithm module can issue coal stockpiling strategy control commands to the DCS, such as the coal blending ratio, start-stop sequence, and output setpoint of each coal feeder; simultaneously, the DCS sends the actual execution results back to the optimization module, forming a complete command downlink and status uplink path.
[0135] The multi-objective optimization algorithm module adjusts the phase of the quantum code based on the actual execution deviation signal returned by the DCS, which is the core mechanism for achieving closed-loop adaptive optimization. The "execution deviation signal" refers to the difference between the actual action value executed by the DCS and the target value issued by the optimization module. For example, if a coal feeder's commanded output is 80 t / h, and the actual measured output is 76 t / h, then the deviation is -4 t / h. This deviation signal, after normalization, is input into the optimization algorithm to calculate the correction amount ΔΦ for the quantum bit phase. Specifically, each quantum bit corresponds to an optimization variable dimension (such as the blending ratio of a certain type of coal), and its state is represented by the phase angle θ.
[0136]
[0137] When a large overall system performance deviation is detected, a feedback gain coefficient k is introduced, and the phase is updated according to the following rules:
[0138] θ t+1 =θ t +k·error
[0139] Here, error is a comprehensive execution deviation index, which can be obtained by weighted averaging of the deviation values of multiple key execution points. This phase adjustment mechanism is essentially an empirical correction of the search direction: if a certain strategy is consistently not executed accurately, the phase of the corresponding qubit will be pulled to a more easily implemented region, thereby guiding subsequent optimization solutions to tend towards more feasible operation points.
[0140] As an optional implementation, the OPC protocol can be replaced by other industrial communication protocols such as Modbus TCP or IEC 61850, as long as bidirectional data interaction can be achieved; the deviation signal can also be expanded to include a variety of dynamic performance indicators such as execution delay time, overshoot, and steady-state error, so as to more comprehensively reflect the control quality; the phase adjustment method can also be combined with gradient information or reinforcement learning reward signal to form a hybrid feedback mechanism.
[0141] Through the above technical solution, this invention realizes the transformation of optimization strategy from open-loop decision-making to closed-loop control. Because the multi-objective optimization algorithm module can receive actual execution feedback from the DCS and dynamically adjust the quantum-encoded phase parameters accordingly, the optimization process not only relies on model predictions but also integrates real-world execution experience, gradually converging to a coal stacking strategy that satisfies both performance objectives and high executability. This mechanism effectively bridges the gap between theoretical optimization and engineering implementation, significantly improving the system's robustness and practicality under complex operating conditions.
[0142] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An artificial intelligence-based coal stockpiling strategy optimization system for coal-fired power plants, comprising a data input module, a control output module, a deep learning model, and a multi-objective optimization algorithm module, characterized in that: The data input module includes a distributed control system (DCS) configured with a timestamp alignment unit and edge computing nodes. The timestamp alignment unit is used to align timestamps from different data sources, and the edge computing nodes are used to perform preprocessing at the data acquisition points. The control output module is connected to the DCS via a bidirectional data channel and includes a safety interlock logic unit configured to block abnormal control signals based on predefined safety rules. The deep learning model includes a spatial feature extraction unit and a time series processing unit. The spatial feature extraction unit is used to process the geometric layout data of the coal-fired power plant, and the time series processing unit is used to process historical operating data. The multi-objective optimization algorithm module employs a quantum bit-based optimization algorithm to generate coal stacking strategies.
2. A coal-fired power plant coal stockpiling strategy optimization system as described in claim 1, characterized in that... The data input module is connected to the DCS and the historical database to form a data buffer. The data buffer stores spatiotemporally correlated data points and includes an event detector for triggering high-frequency data acquisition and data labeling based on device status events. The multi-objective optimization algorithm module includes a quantum neural network unit, which updates model parameters using an incremental learning method.
3. A coal-fired power plant coal stockpiling strategy optimization system as described in claim 1, characterized in that... The deep learning model includes a feature fusion layer, which uses a dynamic weight allocation strategy to weight and fuse spatial and temporal features. The deep learning model also includes a multilayer perceptron (MLP), which is connected to a quantum neural network unit and is equipped with a distribution detector to trigger MLP parameter updates and sparsification based on differences in data distribution.
4. A coal-fired power plant coal stockpiling strategy optimization system as described in claim 1, characterized in that... The deep learning model includes a cascaded quantum neural network and a hybrid module, which includes an MLP and a recurrent neural network (RNN). The MLP includes scalable groups of neurons, and the RNN includes redundant connection paths to enhance model robustness. The quantum neural network consists of multiple output heads, each corresponding to a prediction of a pollutant indicator.
5. A coal-fired power plant coal stockpiling strategy optimization system as described in claim 1, characterized in that... A gradient transmission channel is provided between the multilayer perceptron and the multi-objective optimization algorithm module for backpropagating gradients; The MLP of the deep learning model embeds a differentiable constraint layer, which is used to apply operational constraints. The multi-objective optimization algorithm module adopts the quantum-inspired bee colony algorithm, which dynamically adjusts its structure based on circuit templates.
6. A coal-fired power plant coal stockpiling strategy optimization system as described in claim 1, characterized in that... The deep learning model includes a regularization unit, which generates regularization terms by a discriminator network; the multi-objective optimization algorithm module adopts a quantum-coded bee colony algorithm, in which the solution vector is represented by qubits and updated through rotation operations, crossover operations, and a restart mechanism.
7. A coal-fired power plant coal stockpiling strategy optimization system as described in claim 1, characterized in that... The deep learning model is coupled with the multi-objective optimization algorithm module to calculate fly ash mass concentration and NO. x The predicted values of emission concentration and coal consumption are used as inputs to the multi-objective optimization algorithm.
8. A coal-fired power plant coal stockpiling strategy optimization system as described in claim 1, characterized in that... The multi-objective optimization algorithm module is embedded in a reinforcement learning agent, which includes an action selector; the quantum encoding structure is mapped to the policy network hidden layer of the reinforcement learning agent.
9. A coal-fired power plant coal stockpiling strategy optimization system as described in claim 1, characterized in that... The search space dimension of the multi-objective optimization algorithm module is consistent with the feature space dimension after dimensionality reduction by principal component analysis (PCA); during quantum encoding initialization, the initial phase of each qubit is set based on the PCA contribution rate.
10. A coal-fired power plant coal stockpiling strategy optimization system as described in claim 1, characterized in that... The multi-objective optimization algorithm module establishes a bidirectional feedback channel with the DCS via the OPC protocol; the multi-objective optimization algorithm module adjusts the phase of the quantum code based on the actual execution deviation signal returned by the DCS.
Citation Information
Cited By
面向含噪声环境的多代理融合免训练量子架构搜索方法
CN122366691B