Carbon reduction path combination scheduling method based on reinforcement learning and related equipment

By using deep reinforcement learning-based methods and real-world carbon factor production data for deep learning and reinforcement training, the problem of poor carbon reduction path combination and scheduling optimization in existing technologies is solved, achieving efficient optimization and intelligent scheduling of carbon emissions.

CN120579792BActive Publication Date: 2026-01-06SHENZHEN XINRUN FULIAN DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511074065.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2026-01-06
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

Existing carbon reduction technology solutions lack systematic overall planning. Each carbon reduction path is independent and no collaborative mechanism has been established. Traditional solutions, which are based on experience and simple data analysis, do not match the complex dynamic changes in actual production, resulting in poor performance in the combination and scheduling of carbon reduction paths.

Method used

By employing a deep reinforcement learning-based approach, measured carbon factor production data is obtained and input into a pre-trained carbon reduction path combination scheduling model for deep learning and reinforcement training, thereby achieving precise scheduling of the optimal carbon reduction path for carbon factors.

Benefits of technology

It enables efficient processing and optimization of complex and dynamically changing carbon emission data, determines the best carbon reduction path, improves the efficiency of intelligent combination and scheduling optimization of carbon reduction paths, and effectively reduces carbon emissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579792B_ABST
    Figure CN120579792B_ABST
Patent Text Reader

Abstract

The application discloses a carbon reduction path combination scheduling method based on reinforcement learning and related equipment, and relates to the technical field of low-carbon production. The carbon reduction path combination scheduling method based on reinforcement learning comprises the following steps: obtaining carbon factor production measured data; inputting the carbon factor production measured data into a pre-trained carbon reduction path combination scheduling model to obtain the optimal carbon reduction path of the current carbon factor, wherein the carbon reduction path combination scheduling model is a deep reinforcement learning model. The application adopts a deep reinforcement learning model. When the carbon factor production measured data is obtained and input into the pre-trained carbon reduction path combination scheduling model, the optimal carbon reduction path of the carbon factor is accurately scheduled through the deep learning and reinforcement training of the model, so that the complex and dynamically changing carbon emission data is efficiently processed and optimized, the optimal carbon reduction path is determined, and the effective reduction of carbon emission is finally completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of low-carbon production technology, and in particular to a carbon reduction path combination scheduling method and related equipment based on reinforcement learning. Background Technology

[0002] As global warming intensifies, environmental problems caused by emissions of greenhouse gases such as carbon dioxide are becoming increasingly prominent.

[0003] In industrial production, the exploration and implementation of carbon reduction pathways are crucial. Existing carbon reduction technologies have significant limitations. Most current carbon reduction practices lack systematic and coordinated planning. Different carbon reduction pathways operate independently, without an organic link or synergistic mechanism. For example, a company may simultaneously undertake energy-saving renovations and green electricity procurement, but the lack of coordination in implementation pace and resource allocation prevents the maximization of carbon reduction benefits. Moreover, traditional carbon reduction solutions are often based on experience and simple data analysis, assuming that sample data are independent and identically distributed. This does not reflect the strong correlation and complex dynamic changes in actual production carbon reduction data, resulting in poor effectiveness in combining and optimizing carbon reduction pathways. Summary of the Invention

[0004] The main purpose of this application is to provide a carbon reduction path combination scheduling method and related equipment based on reinforcement learning, which aims to solve the technical problem of poor performance in carbon reduction path combination and scheduling optimization.

[0005] To achieve the above objectives, this application proposes a carbon reduction path combination scheduling method based on reinforcement learning, the method comprising:

[0006] Obtain measured data on carbon factor production;

[0007] The measured carbon factor production data is input into a pre-trained carbon reduction path combination scheduling model to schedule the optimal carbon reduction path for the current carbon factor. The carbon reduction path combination scheduling model is a deep reinforcement learning model.

[0008] In one embodiment, the step of inputting the measured carbon factor production data into a pre-trained carbon reduction path combination scheduling model includes:

[0009] The training set is formed by randomly selecting quadruples of data from the playback memory units during the dynamic training process;

[0010] The reinforcement learning network is trained based on the training set, wherein gradient descent optimization is performed on the network weight parameters of the reinforcement learning network, and after a preset number of learning rounds, the weights of the target action value function of the reinforcement learning network are updated to the weights of the current action value function.

[0011] In one embodiment, the step of randomly retrieving quadruples of data from the playback memory unit during dynamic training to form the training set includes:

[0012] Initialize the playback memory unit, set its data length to the first preset value, and determine the number of carbon factors and the set of measured data corresponding to each carbon factor.

[0013] The action value function is initialized, its weight is set to the second random value, and the weight of the target action value function is initialized to the preset target initial weight.

[0014] First carbon factor production observation data is obtained, and the first carbon factor production observation data is preprocessed and shuffled to obtain a training data sequence, wherein each training data represents one of the collected observation data.

[0015] For each time step during the training process, a random scheduling action is selected with random probability; otherwise, the optimal scheduling action is selected.

[0016] Obtain the reward value after executing the scheduling action and the updated second carbon factor production observation data;

[0017] The first training data, which includes the training data sequence corresponding to the first carbon factor production observation data, the executed scheduling action, the reward value, and the second training data corresponding to the updated second carbon factor production observation data, are combined into a quadruple data and stored in the playback memory unit.

[0018] In one embodiment, the step of randomly retrieving quadruples of data from the playback memory unit during dynamic training to form the training set includes:

[0019] A reward value calculation rule is selected that is inversely correlated with carbon emissions, wherein the reward value decreases as carbon emissions increase;

[0020] The probability of selecting a random scheduling action decreases exponentially with the time step according to a preset rule.

[0021] In one embodiment, the step of obtaining measured data on carbon factor production includes:

[0022] Simultaneously acquire energy consumption measurement data, raw material carbon content detection data, and production process parameters;

[0023] The energy consumption measurement data, raw material carbon content detection data, and production process parameters at different sampling frequencies are aligned with timestamps and then stored in a distributed database according to carbon factor categories. Each carbon factor's data set is maintained and updated independently.

[0024] In one embodiment, the step of classifying and storing the data by carbon factor category in a distributed database includes:

[0025] Outlier detection is performed on the collected measured data, and outliers that meet the preset outlier conditions are identified and removed using the sliding window standard deviation method.

[0026] Data is automatically classified into timeliness levels based on its generation time, with near-real-time data being prioritized for the training set and historical data being downgraded to the validation set.

[0027] Furthermore, to achieve the above objectives, this application also proposes a reinforcement learning-based carbon reduction path combination scheduling device, which includes:

[0028] The acquisition module is used to acquire measured data on carbon factor production.

[0029] The scheduling module is used to input the measured carbon factor production data into a pre-trained carbon reduction path combination scheduling model to schedule the optimal carbon reduction path for the current carbon factor. The carbon reduction path combination scheduling model is a deep reinforcement learning model.

[0030] Furthermore, to achieve the above objectives, this application also proposes a carbon reduction path combination scheduling device based on reinforcement learning. The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the carbon reduction path combination scheduling method based on reinforcement learning as described above.

[0031] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the carbon reduction path combination scheduling method based on reinforcement learning as described above.

[0032] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the reinforcement learning-based carbon reduction path combination scheduling method described above.

[0033] One or more technical solutions proposed in this application have at least the following technical effects:

[0034] In contrast to related technologies, traditional carbon reduction schemes are mostly based on experience and simple data analysis, assuming that sample data are independent and identically distributed. This does not match the strong correlation and complex dynamic changes in carbon reduction data in actual production, resulting in poor performance in carbon reduction path combination and scheduling optimization. This application obtains actual carbon factor production data and inputs it into a pre-trained carbon reduction path combination scheduling model to obtain the optimal carbon reduction path for the current carbon factor. The carbon reduction path combination scheduling model is a deep reinforcement learning model. It is understood that this application uses a deep reinforcement learning model. When actual carbon factor production data is obtained and input into the pre-trained carbon reduction path combination scheduling model, the model's deep learning and reinforcement training achieve precise scheduling of the optimal carbon reduction path for the carbon factor. This enables efficient processing and optimization of complex and dynamically changing carbon emission data. Therefore, based on the deep reinforcement learning model, intelligent combination and optimized scheduling of carbon reduction paths can be achieved, thereby determining the optimal carbon reduction path and ultimately achieving effective reduction of carbon emissions. Attached Figure Description

[0035] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0036] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 A flowchart is provided for Embodiment 1 of the carbon reduction path combination scheduling method based on reinforcement learning in this application;

[0038] Figure 2 A diagram illustrating the overall data governance architecture for the reinforcement learning-based carbon reduction path combination scheduling method in this application;

[0039] Figure 3 A simplified flowchart illustrating the reinforcement learning-based carbon reduction path combination scheduling method provided in this application;

[0040] Figure 4 This is a schematic diagram of the module structure of the carbon reduction path combination scheduling device based on reinforcement learning, as described in an embodiment of this application.

[0041] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the carbon reduction path combination scheduling method based on reinforcement learning in the embodiments of this application.

[0042] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0043] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0044] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0045] The main solution in this application's embodiments is:

[0046] Obtain measured data on carbon factor production;

[0047] The measured carbon factor production data is input into a pre-trained carbon reduction path combination scheduling model to schedule the optimal carbon reduction path for the current carbon factor. The carbon reduction path combination scheduling model is a deep reinforcement learning model.

[0048] In this embodiment, the present application uses a carbon reduction path combination scheduling device based on reinforcement learning as the execution subject. For ease of description, it will be referred to as "device" in the following detailed description.

[0049] Because existing technologies are mostly based on experience and simple data analysis, assuming that sample data are independent and identically distributed, they do not match the characteristics of strong correlation and complex dynamic changes in carbon reduction data in actual production, resulting in poor performance in carbon reduction path combination and scheduling optimization.

[0050] This application provides a solution that employs a deep reinforcement learning model. When measured carbon factor production data is acquired and input into a pre-trained carbon reduction path combination and scheduling model, the model achieves precise scheduling of the optimal carbon reduction path through deep learning and reinforcement training. This enables efficient processing and optimization of complex and dynamically changing carbon emission data. Therefore, based on the deep reinforcement learning model, intelligent combination and optimized scheduling of carbon reduction paths can be achieved, thereby determining the optimal carbon reduction path and ultimately achieving effective reduction of carbon emissions.

[0051] Based on this, embodiments of this application provide a reinforcement learning-based method for coordinating and scheduling carbon reduction paths, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the carbon reduction path combination scheduling method based on reinforcement learning in this application.

[0052] In this embodiment, the reinforcement learning-based carbon reduction path combination scheduling method includes steps S10~S20:

[0053] Step S10: Obtain actual production data of carbon factor;

[0054] It should be noted that the measured carbon factor production data refers to various data related to carbon emissions actually measured during the production process, including but not limited to energy consumption measurement data, raw material carbon content detection data, and production process parameters. This data accurately reflects the carbon emissions in production activities and provides a basis for subsequent carbon reduction pathway planning.

[0055] Understandably, by acquiring actual carbon factor production data, we can provide accurate input to the carbon reduction pathway combination scheduling model, enabling the model to formulate the optimal carbon reduction pathway combination scheduling scheme based on actual production conditions, thereby achieving effective carbon emission reduction.

[0056] For example, refer to Figure 2 Obtaining actual production data for carbon factors requires the following steps:

[0057] 1.1 Confirm the data source. Confirm the data source to be collected based on the carbon emission survey results and quality control plan. The major categories include, but are not limited to: energy output and consumption, raw material usage, auxiliary material usage, output of main and by-products, solid and hazardous waste generation, total environmental emissions, transportation methods and mileage, auxiliary calculation process data, and measured emission factor data.

[0058] 1.2 Establishing Data Standards

[0059] Confirm data collection frequency: Determine the granularity of data collection time (real-time, hourly, daily, monthly, yearly, etc.) for each data source by considering the specific specifications of the organization's carbon and product carbon accounting business modules and the company's own needs.

[0060] Data collection standards: For example, enterprise energy output and consumption data should be collected uniformly and balanced with the data from the finance department.

[0061] Unified Interface Specifications: To facilitate subsequent data collection and data traceability, we should strive to unify the interface specifications during data collection.

[0062] 1.3 Data Storage and Governance

[0063] Data storage and governance mainly involves platform storage and preliminary governance of the data collected at the underlying level.

[0064] Data storage primarily involves interfacing with other systems and underlying acquisition devices through platform interfaces to store data on the platform, which is then provided to the dual-carbon platform's business modules. The acquired data is time-series data, and time-series databases are mainly used to process data with time tags (data that changes sequentially over time, i.e., time-series data). Storing time-series data requires consideration of large storage capacity and high-concurrency I / O. Time-series big data solutions utilize specialized storage methods to efficiently store and rapidly process massive amounts of time-series data, representing a crucial technology for solving massive data processing challenges. The time-series database employs a special data storage method, significantly improving the processing capability of time-related data. Compared to relational databases, it requires half the storage space and offers a substantial increase in query speed.

[0065] Data governance refers to the preprocessing, integration, and caching of numerous and diverse data items across three dimensions: production processes, emission sources, and their timelines. This results in a coherent data content set for the dual-carbon platform, providing convenient data services for various upper-level analytical applications within the collaborative center. Especially for applications involving large-scale, low-level data analysis, the data service model can improve analytical efficiency.

[0066] 1.4 Data from the same source, accessed on demand

[0067] Enterprises establish a unified data pool (standardized unified data source) based on carbon emission survey results and quality control plans. Various types of data are aggregated, calculated and cross-validated according to the accounting requirements of business modules (organizational carbon accounting, product carbon footprint accounting). The processed carbon factor production measurement data is then provided to each business module for automatic accounting and analysis.

[0068] In one feasible implementation, the step of obtaining measured data on carbon factor production includes:

[0069] Simultaneously acquire energy consumption measurement data, raw material carbon content detection data, and production process parameters;

[0070] The energy consumption measurement data, raw material carbon content detection data, and production process parameters at different sampling frequencies are aligned with timestamps and then stored in a distributed database according to carbon factor categories. Each carbon factor's data set is maintained and updated independently.

[0071] It should be noted that energy consumption metering data refers to the consumption data of various energy sources (such as electricity, natural gas, and fuel oil) during the production process. This data is typically collected in real-time by energy meters installed on production equipment, reflecting the energy usage in production activities. Raw material carbon content detection data refers to the results of testing the carbon content in the raw materials used in production. This is obtained through chemical analysis or other testing methods, reflecting the carbon content inherent in the raw materials themselves. Production process parameters refer to the parameters of various process conditions during production, such as temperature, pressure, flow rate, and time. These parameters directly affect carbon emissions during production. A distributed database is a database system that distributes data across multiple independent computers, enabling high availability and high-performance storage and access. The independent maintenance and updating of each carbon factor's dataset means that data for different carbon factors is managed independently, facilitating data maintenance and updates.

[0072] Understandably, by synchronously acquiring multiple carbon emission-related data and categorizing and storing these data from different sampling frequencies in a distributed database after aligning them with timestamps, the accuracy and consistency of the data can be ensured. This provides high-quality data support for subsequent carbon reduction path combination scheduling models, enabling the models to effectively schedule and optimize carbon reduction paths based on comprehensive and accurate production measurement data.

[0073] For example, in the steel production process, energy consumption data is collected in real time by energy metering instruments such as electricity meters and gas meters installed on equipment such as blast furnaces and converters. Simultaneously, carbon content data is obtained by detecting the carbon content of raw materials such as iron ore and coke input into the blast furnace. In addition, process parameters such as temperature and pressure inside the blast furnace, and smelting time and oxygen volumetric flow rate in the converter are also collected. These data from different sampling frequencies (e.g., energy consumption data collected hourly, raw material carbon content data detected daily, and process parameters collected every minute) are aligned with timestamps and stored in a distributed database, categorized by carbon factor. Each carbon factor dataset is maintained and updated independently. Thus, when carbon reduction path combination scheduling is required, the model can accurately obtain comprehensive and accurate measured carbon factor production data from the distributed database, providing a reliable data foundation for developing effective carbon reduction strategies.

[0074] In one feasible implementation, the step of classifying and storing the data by carbon factor category in a distributed database is followed by:

[0075] Outlier detection is performed on the collected measured data, and outliers that meet the preset outlier conditions are identified and removed using the sliding window standard deviation method.

[0076] Data is automatically classified into timeliness levels based on its generation time, with near-real-time data being prioritized for the training set and historical data being downgraded to the validation set.

[0077] It's important to note that outlier detection refers to the process of identifying data points in a dataset that significantly deviate from the normal range using statistical analysis or data mining techniques. The sliding window standard deviation method is a method that dynamically identifies outliers based on the standard deviation of a fixed-length window in a data sequence. It determines whether a data point is an outlier by calculating the standard deviation of the data within the window and setting a threshold. Data timeliness levels are classifications of data based on their generation time. Near real-time data refers to recently generated data with high timeliness, typically reflecting the current production status. Historical data, on the other hand, refers to data generated in the past, with a relatively long time span, used to verify and evaluate the stability and accuracy of models.

[0078] Understandably, outlier detection and data timeliness classification of collected measured data can effectively improve data quality and usability. Outlier detection helps remove errors or anomalies from the data, making it more accurate and reliable; data timeliness classification ensures a reasonable allocation of training and validation sets. Prioritizing near-real-time data for the training set can improve the model's adaptability to the current production state and its predictive accuracy, while using historical data as the validation set helps evaluate the model's long-term stability and generalization ability.

[0079] Step S20: Input the measured carbon factor production data into a pre-trained carbon reduction path combination scheduling model to schedule the optimal carbon reduction path for the current carbon factor. The carbon reduction path combination scheduling model is a deep reinforcement learning model.

[0080] It should be noted that the carbon reduction path combination scheduling model refers to a model built based on deep reinforcement learning algorithms, used to analyze carbon factor data and generate the optimal combination of carbon reduction paths. Deep reinforcement learning is an advanced algorithm combining deep learning and reinforcement learning, capable of learning optimal decision-making strategies through interaction with the environment. In this technical solution, the model is trained using a large amount of historical and simulated data, learning the optimal carbon reduction path scheduling strategies under different combinations of carbon factors. The optimal carbon reduction path refers to the combination of paths that minimizes carbon emissions or maximizes carbon reduction benefits while meeting production needs. Through intelligent scheduling by the model, customized carbon reduction solutions can be provided for different production scenarios.

[0081] Understandably, by inputting measured carbon factor production data into a pre-trained deep reinforcement learning model, the model can quickly and accurately determine the optimal carbon reduction path for each carbon factor under current production conditions based on its learned knowledge and strategies. This deep reinforcement learning-based approach can not only handle complex multivariate and nonlinear relationships but also dynamically adapt to changes in the production process, achieving real-time optimization and intelligent scheduling of carbon reduction paths. This effectively reduces carbon emissions during production and improves corporate carbon management efficiency and environmental performance.

[0082] This embodiment provides a carbon reduction path combination scheduling method based on reinforcement learning. It adopts a deep reinforcement learning model. When the measured carbon factor production data is obtained and input into the pre-trained carbon reduction path combination scheduling model, the model achieves accurate scheduling of the optimal carbon reduction path through deep learning and reinforcement training. This enables efficient processing and optimization of complex and dynamically changing carbon emission data. Therefore, based on the deep reinforcement learning model, intelligent combination and optimized scheduling of carbon reduction paths can be achieved, thereby determining the optimal carbon reduction path and ultimately achieving effective reduction of carbon emissions.

[0083] In one feasible implementation, the step of inputting the measured carbon factor production data into a pre-trained carbon reduction path combination scheduling model includes:

[0084] The training set is formed by randomly selecting quadruples of data from the playback memory units during the dynamic training process;

[0085] The reinforcement learning network is trained based on the training set, wherein gradient descent optimization is performed on the network weight parameters of the reinforcement learning network, and after a preset number of learning rounds, the weights of the target action value function of the reinforcement learning network are updated to the weights of the current action value function.

[0086] It's important to note that replay memory units are storage units used to store historical experience data during reinforcement learning training. This historical experience data is typically stored in the form of quadruples (state, action, reward, next state) for subsequent training. A quadruple is a data combination containing the current state, the action taken, the reward obtained, and the next state; it is the basic unit of reinforcement learning training. Gradient descent optimization is a commonly used optimization algorithm that updates the weight parameters by calculating the gradient of the loss function with respect to the network weights to minimize prediction error. The target action value function is the value function used in reinforcement learning to estimate the long-term reward obtained after taking a certain action; its weight updates are to enable the model to better learn the optimal policy.

[0087] Understandably, by randomly selecting quadruples of data from the replay memory unit to form the training set, and using this data to train the reinforcement learning network, historical experience data can be effectively utilized to improve the model's learning efficiency and stability. Gradient descent optimization helps to precisely adjust the network weight parameters, enabling the model to better fit the training data and learn the optimal carbon reduction path scheduling strategy. Regularly updating the weights of the objective action value function ensures that the model's objective function can be updated in accordance with changes in the current policy, thereby improving the model's convergence speed and decision-making performance.

[0088] In one feasible implementation, the step of randomly retrieving quadruplets of data from the playback memory unit during dynamic training to form the training set includes the following prior steps:

[0089] Initialize the playback memory unit, set its data length to the first preset value, and determine the number of carbon factors and the set of measured data corresponding to each carbon factor.

[0090] The action value function is initialized, its weight is set to the second random value, and the weight of the target action value function is initialized to the preset target initial weight.

[0091] First carbon factor production observation data is obtained, and the first carbon factor production observation data is preprocessed and shuffled to obtain a training data sequence, wherein each training data represents one of the collected observation data.

[0092] For each time step during the training process, a random scheduling action is selected with random probability; otherwise, the optimal scheduling action is selected.

[0093] Obtain the reward value after executing the scheduling action and the updated second carbon factor production observation data;

[0094] The first training data, which includes the training data sequence corresponding to the first carbon factor production observation data, the executed scheduling action, the reward value, and the second training data corresponding to the updated second carbon factor production observation data, are combined into a quadruple data and stored in the playback memory unit.

[0095] It should be noted that the first preset value refers to the upper limit of data storage capacity set when initializing the replay memory unit, ensuring that the replay memory unit can store a sufficient amount of historical experience data for training. The number of carbon factors refers to the number of types of carbon emission-related factors involved in the production process. The actual measurement data set corresponding to each carbon factor includes the actual measurement data of that carbon factor under different times and conditions. The action value function is a function used in reinforcement learning to evaluate the expected cumulative reward of taking an action in a specific state. Its weights are initialized to random values ​​to provide the model with an initial learning starting point. The weights of the target action value function are initialized to preset target initial weights to provide a relatively stable target in the early stages of training, helping the model converge faster. The first carbon factor production observation data refers to the initial carbon factor data acquired at the beginning of training. This data, after preprocessing and shuffling, forms a training data sequence for the model's initial learning. Each time step refers to the discrete unit of time in the reinforcement learning training process. In each time step, the model selects and executes an action based on the current state. Random probability refers to the probability of selecting a random scheduled action during training, used to explore different action spaces. The optimal scheduled action, on the other hand, is the best action selected based on the current model policy, used to utilize learned knowledge. The reward value is the immediate feedback obtained after executing the scheduled action, reflecting the effectiveness of that action in reducing carbon emissions. Second carbon factor production observation data refers to the updated carbon factor data after executing the scheduled action, reflecting the impact of the scheduled action on carbon emissions.

[0096] Understandably, initializing the replay memory unit and related functions ensures the rationality of the data foundation and initial state for model training. Acquiring and preprocessing carbon factor production observation data, and shuffling the order to form a training data sequence, helps eliminate the potential influence of data order on model training, enabling the model to generalize better. Selecting actions with a certain random probability at each time step ensures that the model fully explores the action space while utilizing learned knowledge by selecting the optimal scheduling action, achieving a balance between exploration and utilization. Acquiring reward values ​​and updated carbon factor data, and combining these data into quadruples stored in the replay memory unit, provides rich empirical data for subsequent model training, helping the model learn better carbon reduction path scheduling strategies and improving model performance and adaptability.

[0097] For example, this application discloses a carbon reduction path combination scheduling algorithm based on reinforcement learning, the algorithm steps are as follows:

[0098] a. Initialize the playback memory unit D, with a data length of N. .

[0099] Where M is the carbon factor number, This represents the number of measured data points collected for the k-th carbon factor.

[0100] b. Initialize the action function Q with random weights. .

[0101] c. Initialize the target action function For weight = .

[0102] d. For each data point in the collected measured data, perform the following steps:

[0103] e. Initialize the sequence ,in, One of the collected measured data points, after being shuffled and preprocessed, is denoted as [data point name]. = .

[0104] f. For times t = 1 to T, perform the following steps:

[0105] g. Select a random action with probability e. .

[0106] h. Set with probability 1 - e = .

[0107] i. Obtain the action to be executed The reward after and data .

[0108] j. place = , , and preprocessing = .

[0109] k. Storage ( , , , ) to D.

[0110] l. Randomly select a batch of data from D ( , , , ).

[0111] m. place =

[0112] n.

[0113] o. Train the network, adjust the network weights Parameters perform gradient descent optimization .

[0114] p. After each C round of learning, set = Q.

[0115] In one feasible implementation, the step of randomly retrieving quadruplets of data from the playback memory unit during dynamic training to form the training set includes the following prior steps:

[0116] A reward value calculation rule is selected that is inversely correlated with carbon emissions, wherein the reward value decreases as carbon emissions increase;

[0117] The probability of selecting a random scheduling action decreases exponentially with the time step according to a preset rule.

[0118] It should be noted that the reward value calculation rule refers to the specific method or formula used to determine the reward value obtained after executing a certain scheduling action. In this technical solution, the selected reward value calculation rule is inversely related to carbon emissions; that is, the lower the carbon emissions, the higher the reward value, and vice versa. This rule can effectively incentivize the model to select scheduling actions that are beneficial to reducing carbon emissions. The probability of selecting a random scheduling action refers to the probability of selecting a random action rather than the current optimal action during training. Decreasing according to a preset exponential law means that this probability will gradually decrease as the time step increases, following an exponential function. This design aims to allow for sufficient exploration in the early stages of training, while making greater use of learned knowledge in the later stages, thereby improving learning efficiency and model performance.

[0119] Understandably, selecting a reward calculation rule that is inversely correlated with carbon emissions can guide the model to prioritize scheduling strategies that reduce carbon emissions, ensuring that the model's learning objective aligns with the core need for carbon reduction. Controlling the probability of selecting random scheduling actions to decrease exponentially with time steps achieves a balance between exploration and utilization during training. In the early stages of training, higher random probabilities help the model explore different action spaces extensively, avoiding getting trapped in local optima. As training progresses, the random probabilities gradually decrease, allowing the model to make more effective decisions based on learned knowledge, thereby improving training efficiency and ultimately achieving better carbon reduction path scheduling results.

[0120] For example, to help understand the implementation process of the reinforcement learning-based carbon reduction path combination scheduling method obtained by combining this embodiment with the above embodiment one, please refer to... Figure 3 , Figure 3 A simplified flowchart of a reinforcement learning-based carbon reduction path combination scheduling method is provided, specifically:

[0121] I. Data Acquisition and Preprocessing

[0122] Simultaneously acquire energy consumption measurement data, raw material carbon content detection data, and production process parameters.

[0123] Data from different sampling frequencies are aligned with timestamps and then stored in a distributed database according to carbon factor categories. The dataset for each carbon factor is maintained and updated independently.

[0124] Outlier detection is performed on the collected data, and outliers are identified and removed using the sliding window standard deviation method. Data timeliness is automatically classified according to the data generation time, with near real-time data being given priority for use in the training set and historical data being downgraded to the validation set.

[0125] II. Model Initialization

[0126] Set the data length to the first preset value, and determine the number of carbon factors and the actual set of measured data.

[0127] Set the action value function weight to the second random value, and the target action value function weight to the preset target initial weight.

[0128] III. Training Data Preparation

[0129] The production observation data of the first carbon factor was obtained, and after preprocessing, the order was shuffled to obtain the training data sequence.

[0130] At each time step, a random scheduling action or the optimal scheduling action is selected with random probability. The reward value after the action is executed and the updated observation data are obtained and combined into a quadruple data storage and stored in the playback memory unit.

[0131] IV. Model Training and Optimization

[0132] The training set is formed by randomly selecting quadruples of data from the playback memory unit.

[0133] The reinforcement learning network is trained using the training set and gradient descent optimization is performed. After a preset number of learning rounds, the weights of the target action value function are updated to the weights of the current action value function.

[0134] V. Model Application and Scheduling

[0135] Actual carbon factor production data are input into a pre-trained carbon reduction path combination scheduling model.

[0136] The model scheduling obtains the optimal carbon reduction path for the current carbon factor, thereby achieving effective carbon reduction in the production process.

[0137] VI. Continuous Optimization and Implementation

[0138] In practical applications, the model continuously learns from new production data and dynamically adjusts its carbon reduction strategies to ensure it remains in an optimal state. The optimal carbon reduction path obtained through scheduling is applied to the production process, and carbon emissions are monitored in real time to ensure the achievement of carbon reduction goals.

[0139] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the carbon reduction path combination scheduling method based on reinforcement learning in this application. Any simple modifications based on this technical concept are within the scope of protection of this application.

[0140] This application also provides a carbon reduction path combination scheduling device based on reinforcement learning, please refer to... Figure 4 The reinforcement learning-based carbon reduction path combination scheduling device includes:

[0141] Module 10 is used to acquire measured data on carbon factor production.

[0142] The scheduling module 20 is used to input the measured carbon factor production data into a pre-trained carbon reduction path combination scheduling model to schedule the optimal carbon reduction path for the current carbon factor. The carbon reduction path combination scheduling model is a deep reinforcement learning model.

[0143] And / or, the reinforcement learning-based carbon reduction path combination scheduling device includes:

[0144] The first selection module is used to randomly select quadruples of data from the replay memory unit during the dynamic training process to form a training set.

[0145] The first training module is used to train the reinforcement learning network based on the training set, wherein gradient descent optimization is performed on the network weight parameters of the reinforcement learning network, and after a preset number of learning rounds, the weights of the target action value function of the reinforcement learning network are updated to the weights of the current action value function.

[0146] And / or, the reinforcement learning-based carbon reduction path combination scheduling device includes:

[0147] The first initialization module is used to initialize the playback memory unit, set its data length to a first preset value, and determine the number of carbon factors and the set of measured data corresponding to each carbon factor.

[0148] The second initialization module is used to initialize the action value function, set its weight to a second random value, and initialize the weight of the target action value function to a preset target initial weight.

[0149] The first processing module is used to acquire the first carbon factor production observation data, preprocess the first carbon factor production observation data, and obtain a training data sequence after shuffling the order, wherein each training data represents one of the collected observation data.

[0150] The first selection module is used to select a random scheduling action with random probability for each time step in the training process; otherwise, the optimal scheduling action is selected.

[0151] The first acquisition module is used to acquire the reward value after executing the scheduling action and the updated second carbon factor production observation data;

[0152] The first combination module is used to combine the first training data, which includes the training data sequence corresponding to the first carbon factor production observation data, the executed scheduling action, the reward value, and the second training data corresponding to the updated second carbon factor production observation data, into a quadruple data and store it in the playback memory unit.

[0153] And / or, the reinforcement learning-based carbon reduction path combination scheduling device includes:

[0154] The first selection module is used to select a reward value calculation rule that is inversely related to carbon emissions, wherein the reward value decreases as carbon emissions increase;

[0155] The first control module is used to control the probability of selecting random scheduling actions to decrease exponentially with the time step according to a preset rule.

[0156] And / or, the acquisition module 10 includes:

[0157] The second acquisition module is used to simultaneously acquire energy consumption metering data, raw material carbon content detection data, and production process parameters.

[0158] The first storage module is used to align the energy consumption measurement data, the raw material carbon content detection data, and the production process parameters at different sampling frequencies with timestamps, and then classify and store them in a distributed database according to carbon factor categories. The data set for each carbon factor is maintained and updated independently.

[0159] And / or, the first storage module includes:

[0160] The first removal module is used to perform outlier detection on the collected measured data, and uses the sliding window standard deviation method to identify and remove outliers that meet the preset outlier conditions.

[0161] The first classification module is used to automatically classify the timeliness level of data based on the data generation time. Near real-time data is given priority for use in the training set, while historical data is downgraded to the validation set.

[0162] The carbon reduction path combination scheduling device provided in this application, employing the carbon reduction path combination scheduling method based on reinforcement learning in the above embodiments, can solve the technical problem of poor performance in carbon reduction path combination and scheduling optimization. Compared with the prior art, the beneficial effects of the carbon reduction path combination scheduling device based on reinforcement learning provided in this application are the same as those of the carbon reduction path combination scheduling method based on reinforcement learning provided in the above embodiments, and other technical features in the carbon reduction path combination scheduling device based on reinforcement learning are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0163] This application provides a reinforcement learning-based carbon reduction path combination scheduling device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the reinforcement learning-based carbon reduction path combination scheduling method in Embodiment 1 above.

[0164] The following is for reference. Figure 5 The diagram illustrates a structural schematic of a reinforcement learning-based carbon reduction path combination scheduling device suitable for implementing embodiments of this application. The reinforcement learning-based carbon reduction path combination scheduling device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, tablets, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital televisions and desktop computers. Figure 5 The reinforcement learning-based carbon reduction path combination scheduling device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0165] like Figure 5As shown, the reinforcement learning-based carbon reduction path combination scheduling device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1002 or a program loaded from storage device 1003 into random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the reinforcement learning-based carbon reduction path combination scheduling device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the reinforcement learning-based carbon reduction path combination scheduling device to communicate wirelessly or wiredly with other devices to exchange data. Although a reinforcement learning-based carbon reduction path combination scheduling device with various systems is shown in the figure, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0166] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0167] The carbon reduction path combination scheduling device provided in this application, employing the carbon reduction path combination scheduling method based on reinforcement learning in the above embodiments, can solve the technical problem of poor performance in carbon reduction path combination and scheduling optimization. Compared with the prior art, the beneficial effects of the carbon reduction path combination scheduling device based on reinforcement learning provided in this application are the same as those of the carbon reduction path combination scheduling method based on reinforcement learning provided in the above embodiments, and other technical features in this carbon reduction path combination scheduling device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0168] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0169] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0170] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the reinforcement learning-based carbon reduction path combination scheduling method in the above embodiments.

[0171] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0172] The aforementioned computer-readable storage medium may be included in a reinforcement learning-based carbon reduction path combination scheduling device; or it may exist independently and not be assembled into a reinforcement learning-based carbon reduction path combination scheduling device.

[0173] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the reinforcement learning-based carbon reduction path combination scheduling device, cause the reinforcement learning-based carbon reduction path combination scheduling device to: acquire measured carbon factor production data;

[0174] The measured carbon factor production data is input into a pre-trained carbon reduction path combination scheduling model to schedule the optimal carbon reduction path for the current carbon factor. The carbon reduction path combination scheduling model is a deep reinforcement learning model.

[0175] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0176] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0177] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0178] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described reinforcement learning-based carbon reduction path combination and scheduling method, thereby solving the technical problem of poor performance in carbon reduction path combination and scheduling optimization. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the reinforcement learning-based carbon reduction path combination and scheduling method provided in the above embodiments, and will not be repeated here.

[0179] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the reinforcement learning-based carbon reduction path combination scheduling method described above.

[0180] The computer program product provided in this application can solve the technical problem of poor performance in carbon reduction path combination and scheduling optimization. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the reinforcement learning-based carbon reduction path combination and scheduling method provided in the above embodiments, and will not be repeated here.

[0181] All acquisition of signals, information, or actions in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with the authorization of the relevant device owner.

[0182] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for scheduling a carbon reduction path combination based on reinforcement learning, characterized in that, The method comprises: acquiring carbon factor production measured data; inputting the carbon factor production measured data into a pre-trained carbon reduction path combination scheduling model to schedule an optimal carbon reduction path for the current carbon factor, wherein the carbon reduction path combination scheduling model is a deep reinforcement learning model; Before the step of inputting the carbon factor production measured data into the pre-trained carbon reduction path combination scheduling model, the method comprises: randomly taking out quadruple data from a replay memory unit in a dynamic training process to form a training set; Before the step of randomly taking out quadruple data from the replay memory unit in the dynamic training process to form the training set, the method comprises: initializing the replay memory unit and setting its data length to a first preset value, while determining the number of carbon factors and the measured data set corresponding to each carbon factor; initializing the action value function and setting its weight to a second random value, and initializing the weight of the target action value function to a preset target initial weight; acquiring first carbon factor production observation data, preprocessing the first carbon factor production observation data, and obtaining a training data sequence after shuffling the order, wherein each training data represents one of the collected observation data; for each time step in the training process, a random scheduling action is selected with a random probability, otherwise an optimal scheduling action is selected; acquiring a reward value after executing the scheduling action and updated second carbon factor production observation data; combining the first training data containing the training data sequence corresponding to the first carbon factor production observation data, the executed scheduling action, the reward value, and the second training data corresponding to the updated second carbon factor production observation data into quadruple data and storing them into the replay memory unit; Before the step of randomly taking out quadruple data from the replay memory unit in the dynamic training process to form the training set, the method comprises: selecting a reward value calculation rule inversely related to carbon emissions, wherein the reward value decreases with the increase of carbon emissions; controlling the selection probability of the random scheduling action to decrease with the time step according to a preset exponential law; training the deep reinforcement learning model according to the training set, wherein the network weight parameters of the deep reinforcement learning model are optimized by gradient descent, and the weight of the target action value function of the deep reinforcement learning model is updated to the weight of the current action value function after a preset number of learning rounds.

2. The method of claim 1, wherein, The step of acquiring carbon factor production measured data comprises: synchronously acquiring energy consumption metering data, raw material carbon content detection data, and production process parameters; After aligning the energy consumption metering data, the raw material carbon content detection data, and the production process parameters of different sampling frequencies by time stamp, storing them into a distributed database by carbon factor category, wherein each carbon factor data set is independently maintained and updated.

3. The method of claim 2, wherein, After the step of storing into the distributed database by carbon factor category, the method comprises: performing outlier detection on the collected measured data, identifying and removing outliers that meet the preset abnormal conditions using the sliding window standard deviation method; According to the data generation time, the data time effectiveness level is automatically divided, and the near real-time data is preferentially used for the training set, and the historical data is degraded to the verification set.

4. A device for combined scheduling of decarbonization paths based on reinforcement learning, characterized in that, The device comprises: The acquisition module is configured to acquire carbon factor production measured data. The scheduling module is configured to input the carbon factor production measured data into a pre-trained carbon reduction path combination scheduling model to obtain an optimal carbon reduction path of the current carbon factor, wherein the carbon reduction path combination scheduling model is a deep reinforcement learning model. The step of inputting the carbon factor production measured data into the pre-trained carbon reduction path combination scheduling model comprises the following steps: Randomly taking out quadruple data from the replay memory unit in the dynamic training process to form a training set; The step of randomly taking out quadruple data from the replay memory unit in the dynamic training process to form a training set comprises the following steps: Initializing the replay memory unit and setting its data length to a first preset value, and determining the number of carbon factors and the measured data set corresponding to each carbon factor; Initializing the action value function and setting its weight to a second random value, and initializing the weight of the target action value function to a preset target initial weight; Acquiring first carbon factor production observation data, preprocessing the first carbon factor production observation data, and obtaining a training data sequence after shuffling the order, wherein each training data represents one of the collected observation data; For each time step in the training process, a random scheduling action is selected with a random probability, or an optimal scheduling action is selected; Obtaining a reward value after executing the scheduling action and updated second carbon factor production observation data; Combining the first training data containing the training data sequence corresponding to the first carbon factor production observation data, the executed scheduling action, the reward value, and the second training data corresponding to the updated second carbon factor production observation data into quadruple data and storing them into the replay memory unit; The step of randomly taking out quadruple data from the replay memory unit in the dynamic training process to form a training set comprises the following steps: Selecting a reward value calculation rule inversely related to carbon emissions, wherein the reward value decreases with the increase of carbon emissions; Controlling the selection probability of the random scheduling action to decrease with the time step according to a preset exponential law; Training the deep reinforcement learning model according to the training set, wherein the network weight parameters of the deep reinforcement learning model are optimized by gradient descent, and the weight of the target action value function of the deep reinforcement learning model is updated to the weight of the current action value function after a preset number of learning rounds.

5. A device for combined scheduling of carbon reduction paths based on reinforcement learning, characterized in that The device comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, which is configured to implement the steps of the reinforcement learning-based carbon reduction path combination scheduling method according to any one of claims 1 to 3.

6. A storage medium, characterized by The storage medium is a computer readable storage medium, and the storage medium stores a computer program, which is executed by a processor to implement the steps of the reinforcement learning-based carbon reduction path combination scheduling method according to any one of claims 1 to 3.

7. A computer program product, characterised in that, The computer program product comprises a computer program which, when executed by a processor, implements the steps of the method for scheduling a carbon reduction path combination based on reinforcement learning according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Industrial enterprise emission reduction operation scheduling model training method and scheduling method

    CN116011751A

  • Park carbon emission monitoring and early warning system

    CN117708538A