Carbon reduction path combination scheduling method based on reinforcement learning and related equipment
Through the deep reinforcement learning model combined with carbon factor measurement data, the carbon emission path combination is optimized, and the problems of independent carbon reduction paths and poor scheduling optimization effects in the existing technology are solved, and efficient carbon emission reduction and management are achieved.
Patent Information
- Application Number
- CN202511074065.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-08-01
AI Technical Summary
The existing carbon reduction technical solutions lack systematic overall planning, and each carbon reduction path is independent and no coordinated mechanism is established. The traditional solutions are based on experience and simple data analysis, and cannot effectively optimize the combination of carbon emission paths, resulting in the failure to maximize the carbon reduction benefits.
The carbon reduction path combination scheduling method based on deep reinforcement learning is adopted. By obtaining the actual carbon factor production measured data, using the deep reinforcement learning model for dynamic training and scheduling, the carbon emission path is optimized, and intelligent combination and scheduling is realized.
Efficient processing and optimization of complex and dynamically changing carbon emission data is achieved, optimal carbon reduction path is determined, and carbon management efficiency and environmental performance are improved.
Smart Images

Figure CN120579792A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of low-carbon production technology, and in particular to a carbon reduction path combination scheduling method and related equipment based on reinforcement learning. Background Art
[0002] As global warming intensifies, environmental problems caused by emissions of greenhouse gases such as carbon dioxide are becoming increasingly prominent.
[0003] In industrial production, the exploration and implementation of carbon reduction pathways are key. Existing carbon reduction technology solutions have obvious limitations. Most existing carbon reduction practices lack systematic overall planning. Different carbon reduction pathways are independent of each other, and no organic connection or coordination mechanism has been established. For example, a company may simultaneously carry out energy-saving transformation and green electricity procurement, but the two lack coordination in terms of implementation rhythm and resource allocation, resulting in the failure to maximize carbon reduction benefits. Moreover, traditional carbon reduction solutions are mostly based on experience and simple data analysis, assuming that sample data are independent and identically distributed. This is inconsistent with the strong correlation and complex dynamic changes of carbon reduction data in actual production, resulting in poor results in carbon reduction path combination and scheduling optimization. Summary of the Invention
[0004] The main purpose of this application is to provide a carbon reduction path combination scheduling method and related equipment based on reinforcement learning, aiming to solve the technical problem of poor effect of carbon reduction path combination and scheduling optimization.
[0005] To achieve the above objectives, this application proposes a carbon reduction path combination scheduling method based on reinforcement learning, which includes: Obtaining measured data on carbon factor production; The measured production data of the carbon factor is input into a pre-trained carbon reduction path combination scheduling model to obtain the optimal carbon reduction path for the current carbon factor, wherein the carbon reduction path combination scheduling model is a deep reinforcement learning model.
[0006] In one embodiment, before the step of inputting the measured carbon factor production data into the pre-trained carbon reduction pathway combination scheduling model, the step includes: Randomly extract quadruple data from the replay memory unit during dynamic training to form a training set; The reinforcement learning network is trained according to the training set, wherein gradient descent optimization is performed on the network weight parameters of the reinforcement learning network, and after each preset number of rounds of learning, the weight of the target action-value function of the reinforcement learning network is updated to the weight of the current action-value function.
[0007] In one embodiment, the step of randomly extracting quadruple data from the replay memory unit in the dynamic training process to form a training set includes: Initializing the playback memory unit, setting its data length to a first preset value, and determining the number of carbon factors and a set of measured data corresponding to each carbon factor; Initialize the action-value function, set its weight to the second random value, and initialize the weight of the target action-value function to the preset target initial weight; Obtaining first carbon factor production observation data, preprocessing the first carbon factor production observation data, and shuffling the order to obtain a training data sequence, wherein each training data represents one of the collected observation data; For each time step during training, a random scheduling action is selected with random probability, otherwise the optimal scheduling action is selected; Obtain the reward value after executing the scheduling action and the updated second carbon factor production observation data; The first training data of the training data sequence corresponding to the first carbon factor production observation data, the executed scheduling action, the reward value and the second training data corresponding to the updated second carbon factor production observation data are combined into four-tuple data and stored in the replay memory unit.
[0008] In one embodiment, the step of randomly extracting quadruple data from the replay memory unit in the dynamic training process to form a training set includes: Select a reward value calculation rule that is inversely correlated with carbon emissions, where the reward value decreases as carbon emissions increase; The selection probability of the random scheduling action is controlled to decrease according to the preset exponential law as the time step length.
[0009] In one embodiment, the step of obtaining carbon factor production measured data includes: Synchronously obtain energy consumption measurement data, raw material carbon content detection data and production process parameters; The energy consumption metering data of different sampling frequencies, the raw material carbon content detection data and the production process parameters are aligned by timestamp and stored in a distributed database according to the carbon factor category, wherein the data set of each carbon factor is independently maintained and updated.
[0010] In one embodiment, the step of storing the data in a distributed database according to the carbon factor classification includes: Perform outlier detection on the collected measured data, using the sliding window standard deviation method to identify and remove outliers that meet the preset abnormal conditions; Data timeliness levels are automatically divided according to the time of data generation. Near-real-time data is preferentially used for training sets, and historical data is downgraded to validation sets.
[0011] In addition, to achieve the above objectives, the present application also proposes a carbon reduction path combination scheduling device based on reinforcement learning, which includes: Acquisition module, used to obtain carbon factor production measured data; The scheduling module is used to input the measured data of carbon factor production into a pre-trained carbon reduction path combination scheduling model to schedule the optimal carbon reduction path for the current carbon factor, wherein the carbon reduction path combination scheduling model is a deep reinforcement learning model.
[0012] In addition, to achieve the above-mentioned purpose, the present application also proposes a carbon reduction path combination scheduling device based on reinforcement learning, which includes: a memory, a processor, and a computer program stored on the memory and runnable on the processor, and the computer program is configured to implement the steps of the carbon reduction path combination scheduling method based on reinforcement learning as described above.
[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the carbon reduction path combination scheduling method based on reinforcement learning are implemented as described above.
[0014] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the carbon reduction path combination scheduling method based on reinforcement learning as described above.
[0015] One or more technical solutions proposed in this application have at least the following technical effects: Compared with the related art, traditional carbon reduction schemes are mostly based on experience and simple data analysis, assuming that sample data are independent and identically distributed, which is inconsistent with the strong correlation and complex dynamic changes of carbon reduction data in actual production, resulting in poor results in carbon reduction path combination and scheduling optimization. In comparison, the present application obtains the measured data of carbon factor production; inputs the measured data of carbon factor production into a pre-trained carbon reduction path combination scheduling model, and schedules to obtain the optimal carbon reduction path for the current carbon factor, wherein the carbon reduction path combination scheduling model is a deep reinforcement learning model. It is understandable that the present application adopts a deep reinforcement learning model. When the measured data of carbon factor production is obtained and input into the pre-trained carbon reduction path combination scheduling model, the deep learning and reinforcement training of the model are used to achieve accurate scheduling of the optimal carbon reduction path for the carbon factor, thereby achieving efficient processing and optimization of complex and dynamically changing carbon emission data. Therefore, based on the deep reinforcement learning model, it is possible to achieve intelligent combination and optimized scheduling of carbon reduction paths, thereby determining the optimal carbon reduction path, and ultimately completing the effective reduction of carbon emissions. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 A flowchart illustrating the first embodiment of the carbon reduction path combination scheduling method based on reinforcement learning in this application is provided; Figure 2 The overall data governance architecture diagram provided for this application's reinforcement learning-based carbon reduction path combination scheduling method; Figure 3 A brief flowchart of the carbon reduction path combination scheduling method based on reinforcement learning provided in this application; Figure 4 This is a schematic diagram of the module structure of the carbon reduction path combination scheduling device based on reinforcement learning in an embodiment of the present application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the carbon reduction path combination scheduling method based on reinforcement learning in the embodiment of the present application.
[0019] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0020] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0021] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0022] The main solutions of the embodiments of this application are: Obtaining measured data on carbon factor production; The measured production data of the carbon factor is input into a pre-trained carbon reduction path combination scheduling model to obtain the optimal carbon reduction path for the current carbon factor, wherein the carbon reduction path combination scheduling model is a deep reinforcement learning model.
[0023] In this embodiment, the present application uses a carbon reduction path combination scheduling device based on reinforcement learning as the execution body. For the convenience of expression, it is specifically described below as "device".
[0024] Since existing technologies are mostly based on experience and simple data analysis, they assume that sample data are independent and identically distributed, which is inconsistent with the strong correlation and complex dynamic changes of carbon reduction data in actual production, resulting in poor results in carbon reduction path combination and scheduling optimization.
[0025] This application provides a solution that adopts a deep reinforcement learning model. When the actual measured data of carbon factor production is obtained and input into a pre-trained carbon reduction path combination scheduling model, the model's deep learning and reinforcement training are used to achieve accurate scheduling of the optimal carbon reduction path for the carbon factor, thereby achieving efficient processing and optimization of complex and dynamically changing carbon emission data. Therefore, based on the deep reinforcement learning model, the intelligent combination and optimized scheduling of carbon reduction paths can be achieved, and the optimal carbon reduction path can be determined, ultimately achieving effective reduction of carbon emissions.
[0026] Based on this, the embodiment of the present application provides a carbon reduction path combination scheduling method based on reinforcement learning, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the carbon reduction path combination scheduling method based on reinforcement learning in this application.
[0027] In this embodiment, the carbon reduction path combination scheduling method based on reinforcement learning includes steps S10 to S20: Step S10, obtaining carbon factor production measured data; It should be noted that measured carbon factor production data refers to various types of carbon emission-related data actually measured during the production process, including but not limited to energy consumption measurement data, raw material carbon content detection data, and production process parameters. This data can truly reflect the carbon emissions in production activities and provide a basis for subsequent carbon reduction path scheduling.
[0028] It is understandable that by obtaining the actual measured data on carbon factor production, accurate input can be provided for the carbon reduction path combination scheduling model, enabling the model to formulate the optimal carbon reduction path combination scheduling plan based on actual production conditions, thereby achieving effective carbon emission reduction.
[0029] For example, referring to Figure 2 To obtain the measured data of carbon factor production, the following steps are required: 1.1 Confirm the data source. The data source that needs to be collected should be confirmed based on the carbon emission investigation results and the content of the quality control plan. The major categories include but are not limited to: energy output and consumption, raw material usage, auxiliary material usage, main and by-product output, solid and hazardous waste generation, total environmental emissions, transportation methods and mileage, auxiliary calculation process data, and measured emission factor data.
[0030] 1.2 Establish data specifications Confirm data collection frequency: Determine the collection time granularity (real-time, hourly, daily, monthly, annual, etc.) for each data source based on the specific specifications of accounting business modules such as organizational carbon and product carbon, as well as the company's own needs.
[0031] Data collection specifications: For example, the energy output and consumption data of an enterprise should be collected uniformly and balanced with the data of the finance department.
[0032] Unified interface specifications: To facilitate later data collection and data traceability, try to unify interface specifications during data collection.
[0033] 1.3 Data Storage and Governance Data storage and management mainly carry out platform storage and preliminary management of data collected at the bottom layer.
[0034] Data storage is primarily connected to other systems and underlying collection devices through platform interfaces, with the data stored in the platform and subsequently made available to the dual-carbon platform business modules. The collected data is time series data, and time series databases are primarily used to process data with time tags (data that changes in chronological order, i.e., time serialization). The storage of time series data requires consideration of large storage capacity and high concurrent I / O. Time series big data solutions utilize special storage methods to efficiently store and quickly process massive amounts of time series big data, making them an important technology for addressing massive data processing. Time series databases utilize special data storage methods, significantly improving their processing capabilities for time-related data. Compared to relational databases, their storage space is halved, significantly increasing query speeds.
[0035] Data governance involves preprocessing, integrating, and caching the vast and diverse array of activity data from three dimensions: production processes, emission sources, and their time history. This process forms a comprehensive data content portfolio relevant to the dual-carbon platform, providing convenient data services for various upper-level analytical applications within the collaboration center. The data service model can improve analytical efficiency, particularly for applications involving large-scale, underlying data analysis.
[0036] 1.4 Data from the same source and on-demand call Enterprises establish a unified data pool (standardized unified data source) based on carbon emission results and quality control plans. Various types of data are aggregated and cross-validated according to the accounting requirements of business modules (organizational carbon accounting, product carbon footprint accounting), and the processed carbon factor production measured data are provided to each business module for automatic accounting and analysis.
[0037] In a feasible embodiment, the step of obtaining carbon factor production measured data includes: Synchronously obtain energy consumption measurement data, raw material carbon content detection data and production process parameters; The energy consumption metering data of different sampling frequencies, the raw material carbon content detection data and the production process parameters are aligned by timestamp and stored in a distributed database according to the carbon factor category, wherein the data set of each carbon factor is independently maintained and updated.
[0038] It should be noted that energy consumption metering data refers to the consumption data of various energy sources (such as electricity, gas, and fuel oil) during the production process. It is usually collected in real time through energy meters installed on production equipment and reflects the energy usage in production activities. Raw material carbon content detection data refers to the test results of the carbon content in the raw materials used in production. It is obtained through chemical analysis of the raw materials or other detection methods and reflects the amount of carbon contained in the raw materials themselves. Production process parameters refer to the parameters of various process conditions in the production process, such as temperature, pressure, flow, time, etc. These parameters directly affect the carbon emissions of the production process. A distributed database is a database system that stores data in a distributed manner across multiple independent computers, which can achieve high availability and high-performance storage and access of data. The independent maintenance and update of the data set of each carbon factor means that the data of different carbon factors are managed independently, which facilitates data maintenance and updating.
[0039] It is understandable that by synchronously acquiring a variety of data related to carbon emissions, and classifying and storing these data with different sampling frequencies in a distributed database after aligning them by timestamps, the accuracy and consistency of the data can be ensured, providing high-quality data support for the subsequent carbon reduction path combination scheduling model, so that the model can effectively schedule and optimize the carbon reduction path based on comprehensive and accurate production measured data.
[0040] For example, during steel production, energy metering instruments such as electricity and gas meters installed on equipment like blast furnaces and steelmaking converters collect real-time energy consumption data. Simultaneously, raw materials such as iron ore and coke entering the blast furnaces are tested for carbon content, generating raw material carbon content data. Furthermore, process parameters such as blast furnace temperature and pressure, converter smelting time, and oxygen volume flow rate are collected. These data, sampled at different frequencies (e.g., hourly energy consumption data, daily raw material carbon content data, and minutely process parameters), are aligned according to timestamps and stored in a distributed database, categorized by carbon factor. Data sets for each carbon factor are independently maintained and updated. This allows the model to accurately retrieve comprehensive and accurate carbon factor production data from the distributed database when combined scheduling of carbon reduction pathways is required, providing a reliable data foundation for developing effective carbon reduction strategies.
[0041] In a feasible embodiment, the step of storing the data in a distributed database according to the carbon factor classification includes: Perform outlier detection on the collected measured data, using the sliding window standard deviation method to identify and remove outliers that meet the preset abnormal conditions; Data timeliness levels are automatically divided according to the time of data generation. Near-real-time data is preferentially used for training sets, and historical data is downgraded to validation sets.
[0042] It should be noted that outlier detection refers to the process of identifying data points in a dataset that significantly deviate from the normal range through statistical analysis or data mining techniques. The sliding window standard deviation method is a method for dynamically identifying outliers based on the standard deviation of a fixed-length window in a data sequence. It determines whether a data point is an outlier by calculating the standard deviation of the data within the window and setting a threshold. Data timeliness is a classification of data timeliness based on the chronological order of data generation. Near-real-time data refers to recently generated data with high timeliness, typically reflecting the current production status, while historical data refers to data generated in the past, spanning a relatively long time period, and is used to verify and evaluate the stability and accuracy of the model.
[0043] It's understandable that outlier detection and data timeliness classification of collected measured data can effectively improve data quality and availability. Outlier detection helps remove errors or anomalies in the data, making it more accurate and reliable. Data timeliness classification ensures the proper allocation of training and validation sets. Prioritizing near-real-time data for training improves the model's adaptability to current production conditions and predictive accuracy, while using historical data as a validation set helps assess the model's long-term stability and generalization capabilities.
[0044] Step S20: input the measured carbon factor production data into a pre-trained carbon reduction path combination scheduling model to schedule the optimal carbon reduction path for the current carbon factor, wherein the carbon reduction path combination scheduling model is a deep reinforcement learning model.
[0045] It should be noted that the carbon reduction path combination scheduling model refers to a model built based on a deep reinforcement learning algorithm, which is used to analyze carbon factor data and generate the best carbon reduction path combination. The deep reinforcement learning model is an advanced algorithm that combines deep learning and reinforcement learning. It can learn the optimal decision-making strategy through interaction with the environment. In this technical solution, the model is trained with a large amount of historical data and simulation data to learn the optimal carbon reduction path scheduling strategy under different carbon factor combinations. The optimal carbon reduction path refers to a path combination that can minimize carbon emissions or maximize carbon reduction benefits while meeting production needs. Through the intelligent scheduling of the model, customized carbon reduction solutions can be provided for different production scenarios.
[0046] It's understandable that by inputting measured carbon factor production data into a pre-trained deep reinforcement learning model, the model can quickly and accurately schedule the optimal carbon reduction path for each carbon factor under current production conditions based on its learned knowledge and strategies. This deep reinforcement learning-based approach not only handles complex multivariable and nonlinear relationships, but also dynamically adapts to changes in the production process, achieving real-time optimization and intelligent scheduling of carbon reduction paths, thereby effectively reducing carbon emissions during production and improving the company's carbon management efficiency and environmental performance.
[0047] This embodiment provides a carbon reduction path combination scheduling method based on reinforcement learning, which adopts a deep reinforcement learning model. When the actual production data of carbon factors is obtained and input into the pre-trained carbon reduction path combination scheduling model, the optimal carbon reduction path of the carbon factor is accurately scheduled through the deep learning and reinforcement training of the model, thereby achieving efficient processing and optimization of complex and dynamically changing carbon emission data. Therefore, based on the deep reinforcement learning model, the intelligent combination and optimized scheduling of carbon reduction paths can be realized, and the optimal carbon reduction path can be determined, thereby ultimately achieving effective reduction of carbon emissions.
[0048] In a feasible embodiment, before the step of inputting the measured carbon factor production data into the pre-trained carbon reduction path combination scheduling model, the step includes: Randomly extract quadruple data from the replay memory unit during dynamic training to form a training set; The reinforcement learning network is trained according to the training set, wherein gradient descent optimization is performed on the network weight parameters of the reinforcement learning network, and after each preset number of rounds of learning, the weight of the target action-value function of the reinforcement learning network is updated to the weight of the current action-value function.
[0049] It should be noted that replay memory units are used to store historical experience data during reinforcement learning training. This historical experience data is typically stored as a four-tuple (state, action, reward, next state) for subsequent training. A four-tuple data set is a combination of data consisting of the current state, the action taken, the reward received, and the next state, and is the basic unit of reinforcement learning training. Gradient descent optimization is a commonly used optimization algorithm that updates weight parameters by calculating the gradient of the loss function with respect to the network weights to minimize prediction error. The target action-value function is a value function used in reinforcement learning to estimate the long-term reward obtained after taking a certain action. The weight updates are intended to enable the model to better learn the optimal strategy.
[0050] It's understandable that by randomly extracting quadruple data from the replay memory unit to form a training set and using this data to train the reinforcement learning network, we can effectively leverage historical experience data and improve the model's learning efficiency and stability. Gradient descent optimization helps precisely adjust the network weight parameters, enabling the model to better fit the training data and learn the optimal carbon reduction path scheduling strategy. Regularly updating the weights of the target action value function ensures that the model's objective function updates with changes in the current policy, thereby improving the model's convergence speed and decision-making performance.
[0051] In a feasible implementation manner, the step of randomly extracting quadruple data from the replay memory unit in the dynamic training process to form a training set includes: Initializing the playback memory unit, setting its data length to a first preset value, and determining the number of carbon factors and a set of measured data corresponding to each carbon factor; Initialize the action-value function, set its weight to the second random value, and initialize the weight of the target action-value function to the preset target initial weight; Obtaining first carbon factor production observation data, preprocessing the first carbon factor production observation data, and shuffling the order to obtain a training data sequence, wherein each training data represents one of the collected observation data; For each time step during training, a random scheduling action is selected with random probability, otherwise the optimal scheduling action is selected; Obtain the reward value after executing the scheduling action and the updated second carbon factor production observation data; The first training data of the training data sequence corresponding to the first carbon factor production observation data, the executed scheduling action, the reward value and the second training data corresponding to the updated second carbon factor production observation data are combined into four-tuple data and stored in the replay memory unit.
[0052] It should be noted that the first preset value refers to the upper limit of data storage capacity set when the replay memory unit is initialized, ensuring that the replay memory unit can store a sufficient amount of historical experience data for training. The number of carbon factors refers to the number of carbon emission-related factors involved in the production process. The measured data set corresponding to each carbon factor contains the actual measurement data of that carbon factor at different times and under different conditions. The action-value function is a function used in reinforcement learning to evaluate the expected cumulative reward of taking an action in a specific state. Its weights are initialized to random values to provide an initial learning starting point for the model. The weights of the target action-value function are initialized to the preset target initial weights to provide a relatively stable target in the early stages of training, helping the model converge faster. The first carbon factor production observation data refers to the initial carbon factor data obtained at the beginning of training. This data is preprocessed and shuffled to form the training data sequence used for initial model learning. Each time step is a discrete unit of time during reinforcement learning training. At each time step, the model selects and executes an action based on the current state. Random probability refers to the probability of selecting a random scheduling action during training, used to explore different action spaces. Optimal scheduling action, on the other hand, is the optimal action selected based on the current model strategy, leveraging learned knowledge. Reward value refers to the immediate feedback obtained after executing a scheduling action, reflecting the effectiveness of that action in reducing carbon emissions. Secondary carbon factor production observation data refers to the updated carbon factor data after executing a scheduling action, reflecting the impact of the scheduling action on carbon emissions.
[0053] It can be understood that by initializing the replay memory unit and related functions, the rationality of the data basis and initial state of the model training is ensured. Acquiring and preprocessing the carbon factor production observation data and disrupting the order to form a training data sequence can help eliminate the potential impact of the data order on model training and enable the model to generalize better. Selecting actions with a certain random probability in each time step not only ensures that the model fully explores the action space, but also utilizes the learned knowledge by selecting the optimal scheduling action, achieving a balance between exploration and utilization. Obtaining the reward value and the updated carbon factor data, and combining these data into a quadruple and storing them in the replay memory unit, provides rich experience data for subsequent model training, helps the model learn a better carbon reduction path scheduling strategy, and improves the performance and adaptability of the model.
[0054] For example, this application discloses a carbon reduction path combination scheduling algorithm based on reinforcement learning, and the algorithm steps are as follows: a. Initialize the playback memory unit D, whose data length is N, .
[0055] Where M is the carbon factor number, is the number of measured data collected for the kth carbon factor.
[0056] b. Initialize the action function Q to random weights .
[0057] c. Initialize the target action function Weight = .
[0058] d. Perform the following steps for each piece of collected measured data: e. Initialization sequence ,in, is one of the collected measured data, which is shuffled and preprocessed and recorded as = .
[0059] f. For time t = 1 to T, perform the following steps: g. Choose a random action with probability e .
[0060] h. Place with probability 1 - e = .
[0061] i. Get execution action Reward after and data .
[0062] j. Place = , , , and preprocess = .
[0063] k. Storage ( , , , ) to D.
[0064] l. Randomly take a batch of data from D ( , , , ).
[0065] m. Place =
[0066] n.
[0067] o. Train the network and adjust the network weights Perform gradient descent optimization on the parameters .
[0068] p. After each C round of learning, set =Q.
[0069] In a feasible implementation manner, the step of randomly extracting quadruple data from the replay memory unit in the dynamic training process to form a training set includes: Select a reward value calculation rule that is inversely correlated with carbon emissions, where the reward value decreases as carbon emissions increase; The selection probability of the random scheduling action is controlled to decrease according to the preset exponential law as the time step length.
[0070] It should be noted that the reward value calculation rule refers to the specific method or formula used to determine the reward value obtained after executing a certain scheduling action. In the present technical solution, the selected reward value calculation rule is inversely correlated with carbon emissions, that is, the lower the carbon emissions, the higher the reward value obtained; conversely, the more carbon emissions, the lower the reward value. This rule can effectively motivate the model to select scheduling actions that are conducive to reducing carbon emissions. The selection probability of a random scheduling action refers to the probability of selecting a random action instead of the current optimal action during the training process. Decreasing according to a preset exponential law means that this probability will gradually decrease in the form of an exponential function as the time step increases. This design is to conduct sufficient exploration in the early stages of training and to make greater use of learned knowledge in the later stages of training to improve learning efficiency and model performance.
[0071] It is understandable that selecting a reward value calculation rule that is inversely correlated with carbon emissions can guide the model to prioritize scheduling strategies that can reduce carbon emissions from a mechanistic perspective, so that the model's learning objectives are consistent with the core needs of carbon reduction. By controlling the probability of selecting random scheduling actions to decrease according to a preset exponential law over time steps, a balance between exploration and utilization is achieved during the training process. In the early stages of training, a higher random probability helps the model to explore different action spaces extensively and avoid falling into local optimality; as training deepens, the random probability gradually decreases, and the model gradually makes more effective decisions based on the knowledge it has learned, thereby improving training efficiency and ultimately achieving better carbon reduction path scheduling effects.
[0072] For example, to help understand the implementation process of the carbon reduction path combination scheduling method based on reinforcement learning obtained by combining this embodiment with the above embodiment 1, please refer to Figure 3 , Figure 3 This paper provides a brief flowchart of a combined scheduling method for carbon reduction paths based on reinforcement learning. Specifically: 1. Data Acquisition and Preprocessing Synchronously obtain energy consumption metering data, raw material carbon content detection data and production process parameters.
[0073] After aligning data with different sampling frequencies through timestamps, they are classified and stored in a distributed database according to carbon factor categories, and the data set of each carbon factor is independently maintained and updated.
[0074] Outlier detection is performed on the collected data, and the sliding window standard deviation method is used to identify and remove outliers; the data timeliness level is automatically divided according to the data generation time, and near-real-time data is preferentially used for the training set, and historical data is downgraded to the validation set.
[0075] 2. Model Initialization The data length is set to a first preset value, and the number of carbon factors and their measured data set are determined.
[0076] Set the action-value function weight to the second random value and the target action-value function weight to the preset target initial weight.
[0077] 3. Training Data Preparation The first carbon factor production observation data is obtained, and after preprocessing, the order is disrupted to obtain a training data sequence.
[0078] At each time step, a random scheduling action or the optimal scheduling action is selected with random probability, and the reward value after executing the action and the updated observation data are obtained, combined into four-tuple data and stored in the replay memory unit.
[0079] 4. Model Training and Optimization Randomly take out four-tuple data from the replay memory unit to form a training set.
[0080] The reinforcement learning network is trained according to the training set, and gradient descent optimization is performed. After each preset number of rounds of learning, the target action-value function weight is updated to the current action-value function weight.
[0081] 5. Model Application and Scheduling The measured carbon factor production data are input into the pre-trained carbon reduction path combination scheduling model.
[0082] The model scheduling obtains the optimal carbon reduction path for the current carbon factor, achieving effective carbon reduction in the production process.
[0083] 6. Continuous Optimization and Implementation In practice, the model continuously learns from new production data and dynamically adjusts carbon reduction strategies to ensure they are always in optimal condition. The optimal carbon reduction path derived from scheduling is applied to the production process, and carbon emissions are monitored in real time to ensure carbon reduction results are achieved.
[0084] It should be noted that the above examples are only used to understand this application and do not constitute a limitation on the carbon reduction path combination scheduling method based on reinforcement learning in this application. More forms of simple transformations based on this technical concept are all within the scope of protection of this application.
[0085] This application also provides a carbon reduction path combination scheduling device based on reinforcement learning, please refer to Figure 4 The carbon reduction path combination scheduling device based on reinforcement learning includes: An acquisition module 10 is used to obtain carbon factor production measured data; The scheduling module 20 is used to input the measured production data of the carbon factor into a pre-trained carbon reduction path combination scheduling model to schedule the optimal carbon reduction path for the current carbon factor, wherein the carbon reduction path combination scheduling model is a deep reinforcement learning model.
[0086] And / or, the carbon reduction path combination scheduling device based on reinforcement learning includes: The first selection module is used to randomly extract four-tuple data from the playback memory unit during the dynamic training process to form a training set; The first training module is used to train the reinforcement learning network according to the training set, wherein gradient descent optimization is performed on the network weight parameters of the reinforcement learning network, and the weight of the target action-value function of the reinforcement learning network is updated to the weight of the current action-value function after each preset number of rounds of learning.
[0087] And / or, the carbon reduction path combination scheduling device based on reinforcement learning includes: A first initialization module is used to initialize the playback memory unit, set its data length to a first preset value, and determine the number of carbon factors and the measured data set corresponding to each carbon factor; A second initialization module is used to initialize the action-value function, set its weight to a second random value, and initialize the weight of the target action-value function to a preset target initial weight; a first processing module for obtaining first carbon factor production observation data, preprocessing the first carbon factor production observation data, and scrambling the sequence to obtain a training data sequence, wherein each training data represents one of the collected observation data; The first selection module is used to select a random scheduling action with random probability for each time step in the training process, otherwise the optimal scheduling action is selected; The first acquisition module is used to obtain the reward value after executing the scheduling action and the updated second carbon factor production observation data; The first combination module is used to combine the first training data of the training data sequence corresponding to the first carbon factor production observation data, the executed scheduling action, the reward value, and the second training data corresponding to the updated second carbon factor production observation data into four-tuple data and store them in the playback memory unit.
[0088] And / or, the carbon reduction path combination scheduling device based on reinforcement learning includes: A first selection module is used to select a reward value calculation rule that is inversely correlated with carbon emissions, wherein the reward value decreases as carbon emissions increase; The first control module is used to control the selection probability of the random scheduling action to decrease according to a preset exponential law as the time step increases.
[0089] And / or, the acquisition module 10 includes: The second acquisition module is used to synchronously obtain energy consumption measurement data, raw material carbon content detection data and production process parameters; The first storage module is used to align the energy consumption metering data of different sampling frequencies, the raw material carbon content detection data and the production process parameters through timestamps, and store them in a distributed database according to carbon factor categories, wherein the data set of each carbon factor is independently maintained and updated.
[0090] And / or, the first storage module includes: The first removal module is used to perform outlier detection on the collected measured data, and use the sliding window standard deviation method to identify and remove outliers that meet the preset abnormal conditions; The first classification module is used to automatically classify data timeliness levels according to the data generation time. Near-real-time data is preferentially used for training sets, and historical data is downgraded to validation sets.
[0091] The reinforcement learning-based carbon reduction path combination scheduling device provided in this application utilizes the reinforcement learning-based carbon reduction path combination scheduling method described in the aforementioned embodiments to address the technical issue of poor performance in carbon reduction path combination and scheduling optimization. Compared to the prior art, the reinforcement learning-based carbon reduction path combination scheduling device provided in this application achieves the same beneficial effects as the reinforcement learning-based carbon reduction path combination scheduling method described in the aforementioned embodiments. Other technical features of the reinforcement learning-based carbon reduction path combination scheduling device are the same as those disclosed in the aforementioned embodiments and are not further elaborated upon here.
[0092] The present application provides a carbon reduction path combination scheduling device based on reinforcement learning. The carbon reduction path combination scheduling device based on reinforcement learning includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the carbon reduction path combination scheduling method based on reinforcement learning in the above-mentioned embodiment one.
[0093] Reference below Figure 5, which shows a schematic diagram of the structure of a reinforcement learning-based carbon reduction path combination scheduling device suitable for implementing the embodiments of the present application. The reinforcement learning-based carbon reduction path combination scheduling device in the embodiments of the present application can include, but is not limited to, mobile terminals such as mobile phones, tablet computers, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital televisions and desktop computers. Figure 5 The reinforcement learning-based carbon reduction path combination scheduling device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0094] like Figure 5 As shown, the reinforcement learning-based carbon reduction path combination scheduling device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the reinforcement learning-based carbon reduction path combination scheduling device. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, hard disk, etc.; and communication devices 1009. Communication devices 1009 can allow the reinforcement learning-based carbon reduction pathway combination scheduling device to communicate wirelessly or wired with other devices to exchange data. Although the figure shows a reinforcement learning-based carbon reduction pathway combination scheduling device with various systems, it should be understood that not all of the illustrated systems are required to be implemented or present. More or fewer systems may be implemented or present instead.
[0095] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0096] The reinforcement learning-based carbon reduction path combination scheduling device provided in this application utilizes the reinforcement learning-based carbon reduction path combination scheduling method described in the aforementioned embodiment to address the technical issue of poor performance in carbon reduction path combination and scheduling optimization. Compared to the prior art, the reinforcement learning-based carbon reduction path combination scheduling device provided in this application achieves the same beneficial effects as the reinforcement learning-based carbon reduction path combination scheduling method described in the aforementioned embodiment. Other technical features of this reinforcement learning-based carbon reduction path combination scheduling device are the same as those disclosed in the aforementioned embodiment and are not further elaborated upon here.
[0097] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0098] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0099] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the reinforcement learning-based carbon reduction path combination scheduling method in the above-mentioned embodiment.
[0100] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0101] The above-mentioned computer-readable storage medium may be included in the carbon reduction path combination scheduling device based on reinforcement learning; or it may exist independently without being assembled into the carbon reduction path combination scheduling device based on reinforcement learning.
[0102] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the carbon reduction path combination scheduling device based on reinforcement learning, the carbon reduction path combination scheduling device based on reinforcement learning: obtains measured carbon factor production data; The measured production data of the carbon factor is input into a pre-trained carbon reduction path combination scheduling model to obtain the optimal carbon reduction path for the current carbon factor, wherein the carbon reduction path combination scheduling model is a deep reinforcement learning model.
[0103] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0104] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0105] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0106] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned reinforcement learning-based carbon reduction path combination scheduling method. This computer-readable storage medium can address the technical issue of poor performance in carbon reduction path combination and scheduling optimization. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the reinforcement learning-based carbon reduction path combination scheduling method provided in the aforementioned embodiments, and are not further elaborated here.
[0107] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned reinforcement learning-based carbon reduction path combination scheduling method.
[0108] The computer program product provided in this application can address the technical issue of poor performance in carbon reduction path combination and scheduling optimization. Compared to existing technologies, the computer program product provided in this application offers the same beneficial effects as the reinforcement learning-based carbon reduction path combination scheduling method provided in the aforementioned embodiments, and will not be further elaborated here.
[0109] All acquisition of signals, information or actions in this application are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0110] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A carbon reduction path combination scheduling method based on reinforcement learning, characterized in that: The method includes: Obtaining measured data on carbon factor production; The measured production data of the carbon factor is input into a pre-trained carbon reduction path combination scheduling model to obtain the optimal carbon reduction path for the current carbon factor, wherein the carbon reduction path combination scheduling model is a deep reinforcement learning model.
2. The method according to claim 1, wherein Before the step of inputting the measured carbon factor production data into the pre-trained carbon reduction path combination scheduling model, the method includes: Randomly extract quadruple data from the replay memory unit during dynamic training to form a training set; The reinforcement learning network is trained according to the training set, wherein gradient descent optimization is performed on the network weight parameters of the reinforcement learning network, and after each preset number of rounds of learning, the weight of the target action-value function of the reinforcement learning network is updated to the weight of the current action-value function.
3. The method according to claim 2, wherein The step of randomly extracting quadruple data from the replay memory unit in the dynamic training process to form a training set includes: Initializing the playback memory unit, setting its data length to a first preset value, and determining the number of carbon factors and a set of measured data corresponding to each carbon factor; Initialize the action-value function, set its weight to the second random value, and initialize the weight of the target action-value function to the preset target initial weight; Obtaining first carbon factor production observation data, preprocessing the first carbon factor production observation data, and shuffling the order to obtain a training data sequence, wherein each training data represents one of the collected observation data; For each time step during training, a random scheduling action is selected with random probability, otherwise the optimal scheduling action is selected; Obtain the reward value after executing the scheduling action and the updated second carbon factor production observation data; The first training data of the training data sequence corresponding to the first carbon factor production observation data, the executed scheduling action, the reward value and the second training data corresponding to the updated second carbon factor production observation data are combined into four-tuple data and stored in the replay memory unit.
4. The method according to claim 2, wherein The step of randomly extracting quadruple data from the replay memory unit in the dynamic training process to form a training set includes: Select a reward value calculation rule that is inversely correlated with carbon emissions, where the reward value decreases as carbon emissions increase; The selection probability of the random scheduling action is controlled to decrease according to the preset exponential law as the time step length.
5. The method according to claim 1, wherein The step of obtaining carbon factor production measured data comprises: Synchronously obtain energy consumption measurement data, raw material carbon content detection data and production process parameters; The energy consumption metering data of different sampling frequencies, the raw material carbon content detection data and the production process parameters are aligned by timestamp and stored in a distributed database according to the carbon factor category, wherein the data set of each carbon factor is independently maintained and updated.
6. The method according to claim 5, wherein The step of storing the data in a distributed database according to the carbon factor classification includes: Perform outlier detection on the collected measured data, using the sliding window standard deviation method to identify and remove outliers that meet the preset abnormal conditions; Data timeliness levels are automatically divided according to the time of data generation. Near-real-time data is preferentially used for training sets, and historical data is downgraded to validation sets.
7. A carbon reduction path combination scheduling device based on reinforcement learning, characterized in that: The device comprises: Acquisition module, used to obtain carbon factor production measured data; The scheduling module is used to input the measured data of carbon factor production into a pre-trained carbon reduction path combination scheduling model to schedule the optimal carbon reduction path for the current carbon factor, wherein the carbon reduction path combination scheduling model is a deep reinforcement learning model.
8. A carbon reduction path combination scheduling device based on reinforcement learning, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the carbon reduction path combination scheduling method based on reinforcement learning as described in any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the carbon reduction path combination scheduling method based on reinforcement learning are implemented as described in any one of claims 1 to 6.
10. A computer program product, characterized in that The computer program product includes a computer program, which, when executed by a processor, implements the steps of the carbon reduction path combination scheduling method based on reinforcement learning as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Industrial enterprise emission reduction operation scheduling model training method and scheduling method
CN116011751A
Park carbon emission monitoring and early warning system
CN117708538A
Power distribution network source network load storage low-carbon optimization scheduling increment reinforcement learning method and system
CN119623567A
Multimodal transport path optimization method for deep reinforcement learning
CN120069723A
Industrial technology-based pollution reduction and carbon reduction collaborative path optimization method and system
CN120218369A