Multi-objective energy consumption analysis and dynamic optimization system based on deep reinforcement learning

Through a multi-objective energy consumption analysis and dynamic optimization system based on deep reinforcement learning, the problem of state representation lag of industrial equipment during instantaneous load changes is solved, accurate capture and optimization decision-making of equipment operating status are achieved, and the system's adaptability and energy utilization efficiency are improved.

CN120469244BActive Publication Date: 2025-09-19南京迅集科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510954287.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-19
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

Existing technologies lack dynamic characterization capabilities in the energy consumption analysis and optimization process of industrial equipment, especially when facing instantaneous load changes of equipment. This leads to delayed state characterization, affecting the system's responsiveness and optimization decision-making effects, especially in high-frequency oscillation scenarios.

Method used

A multi-objective energy consumption analysis and dynamic optimization system based on deep reinforcement learning is adopted. The instantaneous load and operating status data are obtained through the data acquisition module, a dynamic state transfer chain is constructed, the instantaneous load response priority is determined, and the operating parameters are adjusted based on the deep reinforcement learning model to achieve real-time regulation and energy consumption optimization.

Benefits of technology

It improves the ability to accurately capture the operating status of industrial equipment, enhances the system's adaptability and response efficiency to complex working conditions, optimizes the pertinence and adaptability of decision-making, and improves energy utilization efficiency and equipment operation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469244B_ABST
    Figure CN120469244B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of industrial automation technology. The present invention discloses a multi-objective energy consumption analysis and dynamic optimization system based on deep reinforcement learning; it aims to solve the problems of delayed state characterization and insufficient optimization decision-making of industrial equipment under instantaneous load changes. The system accurately characterizes the law of equipment state changes by constructing a dynamic state transfer chain; scientifically determines the response priority by analyzing the state transfer path and load change trend; optimizes equipment operating parameters by combining the dynamic decision-making logic of deep reinforcement learning; balances energy consumption, response speed and stability through multi-objective optimization strategies; and finally achieves efficient operation through real-time regulation. The present invention improves the system's adaptability to complex working conditions, the scientific nature of the response, and energy utilization efficiency, thereby reducing production costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial automation technology, and more specifically, to a multi-objective energy consumption analysis and dynamic optimization system based on deep reinforcement learning. Background Art

[0002] With the rapid development of industry and intelligent manufacturing, efficient operation and energy optimization of industrial equipment have become important research areas in modern manufacturing. In complex industrial production environments, equipment operating conditions are influenced by a variety of factors, which directly determine the equipment's energy consumption and operating efficiency. To improve energy efficiency and reduce production costs, data-driven energy consumption analysis and optimization technologies have garnered widespread attention in recent years. These technologies collect operational data from equipment and, in combination with mathematical modeling and machine learning methods, analyze the equipment's energy consumption characteristics and attempt to optimize energy consumption by adjusting operating parameters.

[0003] However, existing technologies suffer from significant lags in analyzing and optimizing industrial equipment energy consumption, particularly when characterizing equipment loads during transient load changes. This problem is particularly acute in real-world industrial scenarios. For example, in heavy machinery processing workshops or electrically powered chemical production lines, equipment loads often experience dramatic transient changes due to workpiece switching, process adjustments, or external power supply fluctuations. These changes can occur, for example, when the load suddenly jumps from low-power, stable operation to a high-power surge, or from a sustained high load to a no-load state. However, existing technologies typically rely on static or semi-static energy consumption analysis models and lack the ability to dynamically characterize these transient load changes. This results in the system being unable to promptly capture the transient characteristics of load changes, such as rate of change, acceleration, or trend inflection points. For example, when a large CNC machine tool is processing materials of varying hardness, the load can surge from a low value to a peak value within seconds. Existing systems are often limited to coarse-grained state assessments based on historical average data or fixed sampling periods, making it difficult to accurately identify the triggering moment of the load surge and its immediate impact on the equipment's operating status. This hysteresis means that when constructing a state transition model, the system can only rely on lagged or smoothed data, which cannot truly reflect the dynamic behavior of the equipment under instantaneous load changes. For example, it cannot distinguish whether the load change is a short-term fluctuation or a trend change, nor can it accurately characterize the direction and magnitude of the impact of load changes on equipment operating parameters (such as temperature, pressure, and vibration frequency). This hysteresis problem is particularly serious in high-frequency oscillation scenarios, such as frequent load switching in power systems or fluid pulsation in chemical equipment. The system may even completely ignore the existence of high-frequency oscillations, resulting in a serious disconnect between the state representation and the actual operating state, which in turn affects subsequent optimization decisions and real-time control effects. This hysteresis in state representation not only reduces the system's adaptability to complex working conditions, but can also lead to energy waste, increased equipment wear, and even safety hazards, especially in critical industrial scenarios that require rapid response.

[0004] In view of this, the present invention proposes a multi-objective energy consumption analysis and dynamic optimization system based on deep reinforcement learning to solve the above problems. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned objectives, the present invention provides the following technical solution: a multi-objective energy consumption analysis and dynamic optimization system based on deep reinforcement learning, comprising:

[0006] Data acquisition module, used to obtain instantaneous load data, energy consumption data and operating status data of industrial equipment at each sampling moment;

[0007] A transfer chain construction module, configured to construct a dynamic state transfer chain of the industrial equipment according to the instantaneous characteristics of the load change in the instantaneous load data and the state transfer logic of the parameters in the operating state data;

[0008] a priority determination module, configured to determine the instantaneous load response priority of the industrial equipment at the corresponding sampling moment according to the topological characteristics of the state transition path in the dynamic state transition chain and the change trend of the instantaneous load data;

[0009] A model optimization module, configured to adjust the operating parameters of the industrial equipment at the next sampling moment based on the instantaneous load response priority and the dynamic decision logic of the deep reinforcement learning model to obtain optimized operating parameters;

[0010] A multi-strategy fitting module is used to determine a multi-objective energy consumption optimization strategy for industrial equipment at each sampling moment based on the correlation characteristics between the optimized operating parameters and the energy consumption data;

[0011] The real-time control module is used to control the operating parameters of the industrial equipment in real time based on the multi-objective energy consumption optimization strategy to achieve dynamic characterization of instantaneous load changes and energy consumption optimization.

[0012] Preferably, the method for constructing the dynamic state transfer chain includes:

[0013] Identifying a trend inflection point of the load change based on instantaneous characteristics of the load change in the instantaneous load data as a trigger condition for state transition;

[0014] Determining a state node at each sampling moment according to the discrete state of the operating parameters in the operating state data;

[0015] According to the trigger condition, the state nodes are connected in chronological order to form an initial state transfer chain;

[0016] According to the consistency of parameter change directions between adjacent state nodes in the initial state transfer chain, transfer paths that meet the consistency condition are screened out to form the dynamic state transfer chain.

[0017] Preferably, the method for determining the transient load response priority includes:

[0018] Identifying a closed-loop structure of state nodes in a state transition path according to a topological characteristic of the state transition path in the dynamic state transition chain;

[0019] Determining the stability level of the state transition according to the number of cycles of the state node in the closed-loop structure;

[0020] Determining whether the load change enters a high-frequency oscillation mode according to a change trend of the instantaneous load data;

[0021] The transient load response priority is determined according to a matching degree between the stability level and the high-frequency oscillation mode.

[0022] Preferably, the method for constructing the dynamic decision logic includes:

[0023] determining a prioritized subset of action selections in a deep reinforcement learning model based on the instantaneous load response priorities;

[0024] Constructing action selection constraints based on the transition path from the current state node to the next state node in the dynamic state transition chain; and screening out a set of feasible actions based on the priority subset and the constraints;

[0025] According to the contribution of the actions in the set of feasible actions to the energy consumption optimization strategy, a final action is determined as the output of the dynamic decision logic.

[0026] Preferably, the method for adjusting the optimized operating parameters includes:

[0027] Predicting a candidate set of next state nodes based on the transition path of the current state node in the dynamic state transition chain;

[0028] According to the instantaneous load response priority, a state node with a matching priority is screened out from the candidate set as a target state node;

[0029] Determining the direction of parameter adjustment according to the operating parameter range corresponding to the target state node;

[0030] An adjustment step size is determined based on a deviation between the direction and the current operating parameters to obtain the optimized operating parameters.

[0031] Preferably, the method for determining the multi-objective energy consumption optimization strategy includes:

[0032] Predicting the energy consumption change trend of the industrial equipment at the next sampling moment based on the optimized operating parameters;

[0033] Determining a balance point between energy consumption optimization and load response based on the correlation characteristics between the energy consumption change trend and the instantaneous load data;

[0034] According to the balance point, the priority of the energy consumption optimization strategy is determined; according to the priority, the control order of the operating parameters is adjusted to obtain the multi-objective energy consumption optimization strategy.

[0035] Preferably, the method for identifying trend inflection points includes:

[0036] extracting continuous segments of load change according to instantaneous features of load change in the instantaneous load data;

[0037] identifying candidate inflection points based on a reversal of a load change direction in the continuous segment;

[0038] According to the duration of the load change before and after the candidate inflection point, the candidate inflection point with a duration greater than a preset time threshold is screened out as the trend inflection point.

[0039] Preferably, the method for determining the high-frequency oscillation mode includes:

[0040] extracting an oscillation period of the load change according to a change trend of the instantaneous load data;

[0041] determining a stable segment of the oscillation based on the continuity of the oscillation period;

[0042] Determining whether the average length of the oscillation period in the stable segment is less than a preset period threshold;

[0043] If it is less than the preset period threshold, it is determined that the load changes and enters the high-frequency oscillation mode.

[0044] Preferably, the method for screening the set of possible actions comprises:

[0045] determining an initial scope of action selection based on the priority subset;

[0046] According to the constraint conditions, determine whether the action within the initial range causes the state transition path to deviate from the target state node; based on the degree of deviation, eliminate actions whose deviation is greater than a preset deviation threshold;

[0047] The set of feasible actions is determined according to the contribution of the remaining actions to the energy consumption optimization strategy.

[0048] Preferably, the training method of the deep reinforcement learning model includes:

[0049] Constructing a training data set of the instantaneous load data, the energy consumption data, and the operating status data at historical sampling moments;

[0050] generating a state-action pair sequence according to a state transition path of the dynamic state transition chain in the training data set;

[0051] determining a training priority of a state-action pair according to the instantaneous load response priority;

[0052] Adjusting the sampling order of state-action pairs according to the training priority, wherein state-action pairs with higher priority are sampled first;

[0053] The output of the dynamic decision logic is used as the target, and a policy gradient-based deep reinforcement learning algorithm is adopted for training to obtain the deep reinforcement learning model.

[0054] The technical effects and advantages of the multi-objective energy consumption analysis and dynamic optimization system based on deep reinforcement learning of the present invention are as follows:

[0055] The present invention improves the ability to accurately capture the operating status of industrial equipment, and can more comprehensively and timely grasp the dynamic changes of equipment under complex working conditions, thereby providing a reliable basis for optimization decisions. At the same time, it effectively enhances the system's ability to characterize changes in equipment status, so that the dynamic behavior of equipment operation is more realistically reflected, thereby improving the system's adaptability to changing environments. In addition, the present invention optimizes the response priority determination process in a scientific way, enabling the system to quickly and reasonably allocate resources at critical moments, greatly improving the efficiency and accuracy of the response. The present invention also significantly improves the pertinence of optimization decisions, and can better meet the personalized needs under different working conditions, thereby achieving better operating results. Not only that, the present invention enhances the adaptability and sustainability of the system by balancing multiple optimization objectives, enabling it to maintain efficient operation in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 Schematic diagram of the multi-objective energy consumption analysis and dynamic optimization system based on deep reinforcement learning of the present invention. DETAILED DESCRIPTION

[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0058] This application example provides a multi-objective energy consumption analysis and dynamic optimization system based on deep reinforcement learning. The execution entities of the multi-objective energy consumption analysis and dynamic optimization system based on deep reinforcement learning include but are not limited to: industrial equipment monitoring platforms, energy consumption management systems, equipment control systems, data analysis platforms, intelligent decision support systems, etc. equipped with the system, which can be regarded as general computing nodes of this application, and data processing platforms include but are not limited to: equipment status monitoring systems, load analysis systems, and parameter optimization systems.

[0059] See also Figure 1 The present invention provides a multi-objective energy consumption analysis and dynamic optimization system based on deep reinforcement learning, including a data acquisition module, a transfer chain construction module, a priority determination module, a model optimization module, a multi-strategy fitting module and a real-time control module.

[0060] Data acquisition module, used to obtain instantaneous load data, energy consumption data and operating status data of industrial equipment at each sampling moment;

[0061] A transfer chain construction module is used to construct a dynamic state transfer chain of industrial equipment based on the instantaneous characteristics of load changes in the instantaneous load data and the state transfer logic of parameters in the operating state data;

[0062] A priority determination module is used to determine the instantaneous load response priority of the industrial equipment at the corresponding sampling moment based on the topological characteristics of the state transition path in the dynamic state transition chain and the change trend of the instantaneous load data;

[0063] The model optimization module is used to adjust the operating parameters of industrial equipment at the next sampling moment based on the dynamic decision logic of the instantaneous load response priority and deep reinforcement learning model to obtain the optimized operating parameters;

[0064] The multi-strategy fitting module is used to determine the multi-objective energy consumption optimization strategy of industrial equipment at each sampling moment based on the correlation characteristics of the optimized operating parameters and energy consumption data;

[0065] The real-time control module is used to control the operating parameters of industrial equipment in real time based on a multi-objective energy consumption optimization strategy to achieve dynamic characterization of instantaneous load changes and energy consumption optimization.

[0066] The present invention achieves accurate capture of the equipment operating status by acquiring the instantaneous load data, energy consumption data and operating status data of the industrial equipment, constructs a dynamic state transfer chain based on the instantaneous characteristics of the load change and the state transfer logic of the parameters in the operating status data, so that the system has the ability to characterize state changes, determines the instantaneous load response priority according to the topological characteristics of the state transfer path in the dynamic state transfer chain and the change trend of the instantaneous load data, thereby improving the scientific nature of the system response, adjusts the operating parameters based on the instantaneous load response priority and the dynamic decision logic of the deep reinforcement learning model, and makes the optimization decision more targeted, determines the multi-objective energy consumption optimization strategy according to the correlation characteristics of the optimized operating parameters and the energy consumption data, thereby enhancing the adaptability and sustainability of the system, and performs real-time regulation of the operating parameters of the industrial equipment through the multi-objective energy consumption optimization strategy, thereby improving the system operation efficiency and energy utilization efficiency.

[0067] During the operation of industrial equipment, data acquisition devices such as smart sensors, power monitors, and industrial IoT devices are used to collect various data and information from the equipment in real time. Instantaneous load data reflects the load intensity and utilization rate of industrial equipment at a specific moment; energy consumption data reflects the amount of energy, such as electricity and heat, consumed by industrial equipment during operation; and operating status data includes process parameters such as equipment temperature, vibration frequency, and pressure, as well as their changing trends.

[0068] It should be noted that the data acquisition device adopts a distributed deployment method to ensure that it can fully cover the key operating components of industrial equipment. The data acquisition device is used to continuously collect data at each sampling moment to ensure the continuity and integrity of the data.

[0069] In one implementation of the embodiment of the present invention, the sampling frequency is set to once every second.

[0070] In one implementation of an embodiment of the present invention, the collected data undergoes preprocessing steps such as denoising, outlier detection, and normalization to ensure data quality. Denoising is performed using a wavelet transform, outlier detection is performed using the 3σ principle, and data normalization is performed using the Z-score normalization method. The specific preprocessing methods are not described here and are well known to those skilled in the art. Other data preprocessing algorithms may also be used and are not limited here.

[0071] The following steps all use pre-processed instantaneous load data, energy consumption data, and operating status data for analysis.

[0072] In an embodiment of the present invention, a method for constructing a dynamic state transfer chain includes:

[0073] Based on the instantaneous characteristics of load changes in the instantaneous load data, the trend inflection point of load change is identified as the trigger condition for state transition;

[0074] Determine the state node at each sampling moment according to the discrete state of the operating parameters in the operating state data;

[0075] According to the triggering conditions, the state nodes are connected in chronological order to form the initial state transfer chain;

[0076] According to the consistency of parameter change directions between adjacent state nodes in the initial state transfer chain, the transfer paths that meet the consistency conditions are screened out to form a dynamic state transfer chain.

[0077] In this embodiment, the dynamic characteristics of load changes in instantaneous load data are first analyzed, and the first-order derivative (rate of change) and second-order derivative (acceleration of change) of the load are calculated to capture the speed and acceleration characteristics of the load change. A change point detection algorithm (such as the CUSUM algorithm, the PELT algorithm, the Bayesian change point detection algorithm, etc.) is applied to identify significant change points in the load curve. The amplitude, duration, and frequency characteristics of the load change are analyzed to determine whether the change is a temporary fluctuation or a trend shift. Based on the above characteristics, key trend inflection points in the load change are identified. These inflection points mark the transition of the load from one change mode to another (such as from rising to stable, from stable to falling, etc.). These inflection points are defined as trigger conditions for state transitions, and the timestamp and characteristic attributes of each trigger condition are recorded. Then, continuous parameters (such as temperature, pressure, flow, etc.) in the operating state data are processed and converted into discrete state values ​​through interval partitioning or clustering methods. Parameters that are themselves discrete (such as switch state, operating mode, etc.) are processed to ensure consistency in state representation. A unique state code is assigned to each discrete state. A state encoding mapping table is established to combine multiple parameter states at each sampling moment to form a state vector. Each state vector represents Each state node is assigned a unique identifier, and a state node library is established. Based on the previously identified trigger conditions (load trend inflection points), the time when the state transition occurs is determined. At the time corresponding to each trigger condition, the previous and next state nodes are recorded, and a transition relationship from the previous state node to the next state node is established. These state nodes are connected in chronological order to form an initial state transition chain. This initial chain contains all possible state transitions and requires further screening and analysis of parameter changes between adjacent state nodes in the initial state transition chain. For each pair of adjacent state nodes, the change direction of each parameter is calculated (increase, decrease, or remain unchanged). The consistency of the parameter change direction is determined, that is, whether it conforms to physical laws and process requirements. Consistency evaluation indicators are set, such as the coordination of parameter change direction and the rationality of the change amplitude. Transfer paths that meet the consistency conditions are screened based on the evaluation indicators, and transfer paths that do not conform to physical laws or have extremely low probability of occurrence are eliminated. The screened state nodes and transfer paths are reorganized to form a dynamic state transition chain. This transfer chain accurately reflects the state transition laws of industrial equipment under different load conditions and provides a topological model for subsequent priority determination and decision optimization.

[0078] In an embodiment of the present invention, a method for determining the transient load response priority includes:

[0079] According to the topological characteristics of the state transfer path in the dynamic state transfer chain, the closed-loop structure of the state nodes in the path is identified;

[0080] Determine the stability level of state transition according to the number of cycles of state nodes in the closed-loop structure;

[0081] According to the changing trend of instantaneous load data, determine whether the load change enters the high-frequency oscillation mode;

[0082] The priority of transient load response is determined according to the matching degree between the stability level and the high-frequency oscillation mode, wherein the priority is highest when the stability level is low and the high-frequency oscillation mode is in place.

[0083] In this embodiment, a topological analysis is first performed on the dynamic state transition chain. A graph theory algorithm (such as depth-first search, breadth-first search, etc.) is used to traverse the entire transition chain to identify special structures in the transition chain, such as linear paths, branching paths, and convergence points. The focus is on identifying closed-loop structures, that is, loops formed by connecting the end to the end of a state node sequence, such as "state A→state B→state C→state A". A loop detection algorithm (such as Tarjan algorithm, Johnson algorithm, etc.) is used to find all closed-loop structures. The closed-loop length, the type of state nodes included, and other features are analyzed. The identified closed-loop structures are classified and marked. , forming a closed-loop structure set, and then conducting in-depth analysis on each closed-loop structure, statistically analyzing the frequency and number of cycles of the closed-loop structure in historical data. The number of cycles refers to the number of cycles in which the system continuously runs in the same closed loop, and analyzing the stability of state parameters during the cycle, such as parameter fluctuation range, average level, etc. According to the number of cycles and parameter stability, a stability evaluation system is established, and the stability is divided into multiple levels, such as high stability (many cycles, stable parameters), medium stability (moderate cycles, controllable parameter fluctuations), and low stability (few cycles, large parameter fluctuations). A corresponding stability level is assigned. Simultaneously, the time series characteristics of the instantaneous load data are analyzed, and the frequency characteristics of the load changes are calculated, such as through spectrum analysis and wavelet transform. The main frequency components and energy distribution of the load changes are determined, and criteria for high-frequency oscillation are set, such as frequency thresholds and amplitude thresholds. When the dominant frequency of the load change exceeds the set threshold and the amplitude meets certain conditions, the system is judged to have entered a high-frequency oscillation mode. The start time, duration, and intensity characteristics of the high-frequency oscillation are recorded. Finally, the stability level is matched with the high-frequency oscillation mode, and a matching evaluation matrix is ​​established. Response priorities are determined based on different combinations. In principle, the lower the stability level and the more obvious the high-frequency oscillation, the higher the response priority. This is because low stability means the system is unstable, and high-frequency oscillation indicates drastic load changes, requiring priority response. Priority grading criteria are designed, such as highest priority (red warning), high priority (orange warning), medium priority (yellow warning), and low priority (green normal). The corresponding instantaneous load response priority is determined based on the current system state's position in the evaluation matrix. This priority guides the decision-making process of the subsequent model optimization module, ensuring that the system responds promptly to critical state changes.

[0084] In an embodiment of the present invention, a method for constructing dynamic decision logic includes:

[0085] Determine the priority subset of actions to be selected in the deep reinforcement learning model based on the instantaneous load response priority;

[0086] Construct action selection constraints based on the transition path from the current state node to the next state node in the dynamic state transition chain;

[0087] Filter out the feasible action set based on the priority subset and constraints;

[0088] According to the contribution of the actions in the set of feasible actions to the energy consumption optimization strategy, the final action is determined as the output of the dynamic decision logic.

[0089] In this embodiment, first, based on the previously determined instantaneous load response priority, a subset of actions with matching priorities is screened out from the action space of the deep reinforcement learning model. Different action screening strategies are set according to different priorities. For example, in a high-priority state, actions with fast response speed tend to be selected. In a medium-priority state, response speed and energy efficiency are balanced. In a low-priority state, more emphasis is placed on energy efficiency optimization. A subset of actions that meet the current priority requirements is extracted from the complete set of action space to form a priority subset for action selection. Then, the transfer rules in the dynamic state transfer chain are analyzed to determine the set of next state nodes that may be transferred from the current state node. The probability distribution of each transfer path is statistically analyzed based on historical data, and the control parameter adjustment range required to achieve a specific transfer is calculated. These transfer rules are converted into mathematical constraints, such as upper and lower limits of parameter changes, change rate constraints, coupling relationships between parameters, etc., to construct a set of constraints as boundary conditions for action selection. Then, the priority subset is combined with the constraints to screen the feasibility of the actions, check whether each action in the priority subset meets the constraints, calculate the predicted value of the system state after the action is executed, verify whether the predicted state is within the target range, eliminate actions that do not meet the constraints or cause the system to deviate from the expected state, retain actions that meet all constraints, and form a set of feasible actions. Finally, the contribution of each action in the feasible action set to energy consumption optimization is evaluated, and an evaluation function is constructed. Taking into account multiple indicators such as energy consumption reduction, response speed, and state stability, a deep reinforcement learning model (such as DQN, DDPG, SAC, etc.) is used to predict the long-term benefits of each action, considering the immediate effect and long-term impact of the action, and selecting the action with the highest evaluation function value as the final decision. This decision result is used as the output of the dynamic decision logic to guide the subsequent parameter adjustment process. The model will continue to learn from the actual operation results, optimize its decision-making strategy, and improve the energy consumption optimization effect.

[0090] In an embodiment of the present invention, a method for optimizing an operating parameter adjustment includes:

[0091] Predict the candidate set of the next state node based on the transfer path of the current state node in the dynamic state transfer chain;

[0092] According to the instantaneous load response priority, the state node with matching priority is selected from the candidate set as the target state node;

[0093] Determine the direction of parameter adjustment based on the operating parameter range corresponding to the target state node;

[0094] According to the deviation between the direction and the current operating parameters, the adjustment step size is determined to obtain the optimized operating parameters.

[0095] In this embodiment, all possible transfer paths of the current state node are first extracted from the dynamic state transfer chain. Based on the topological structure and transfer probability distribution of the transfer chain, all next state nodes that the current state may transfer to are identified. A Markov prediction model or a sequence prediction neural network (such as RNN, LSTM, etc.) is used to predict the probability distribution of system transfer under different control conditions. According to the predicted probability sorting, several state nodes with higher probabilities are selected to form a candidate set of next state nodes. This set contains the states that the system is most likely to transfer to. Then, the instantaneous load response priority is matched with the candidate state node set for analysis. In the high priority state, the state node that can quickly respond to load changes is given priority. In the medium priority state, the response speed and energy efficiency are balanced. In the low priority state, the state node with the highest energy efficiency is given priority. According to the matching analysis results, the state node that best suits the current priority is selected from the candidate set as the target state node of the system. This target node represents the ideal state that the system should transfer to. Further analysis is performed. The operating parameter characteristics corresponding to the target state node are analyzed by querying historical data for typical parameter ranges for that state node, such as temperature range, pressure range, and speed range. The target value or target range for each parameter is determined. The current operating parameters are compared with the target parameters to determine the direction in which each parameter should be adjusted, such as increase, decrease, or remain unchanged. This form a parameter adjustment direction set, which specifies the adjustment direction for each parameter. Finally, based on the determined adjustment direction, the deviation between the current parameter value and the target parameter value (or the midpoint of the target range) is calculated. The parameter adjustment step size is determined based on the deviation size and system response characteristics. The step size design considers various factors, such as response speed requirements (derived from priority), system stability requirements, and physical limitations of parameter adjustment. A specific adjustment amount is calculated for each parameter to be adjusted, such as "temperature +2°C," "pressure -0.5MPa," or "speed +200rpm." These specific parameter adjustment values ​​are combined to form an optimized operating parameter set, which guides the system in executing specific parameter adjustments to achieve transition to the target state.

[0096] In an embodiment of the present invention, a method for determining a multi-objective energy consumption optimization strategy includes:

[0097] Based on the optimized operating parameters, predict the energy consumption trend of industrial equipment at the next sampling moment;

[0098] Determine the balance point between energy consumption optimization and load response based on the correlation characteristics between energy consumption change trends and instantaneous load data;

[0099] Determine the priority of energy consumption optimization strategy based on the balance point;

[0100] According to the priority, the control order of the operating parameters is adjusted to obtain a multi-objective energy consumption optimization strategy.

[0101] In this embodiment, first, based on the optimized operating parameters, an energy consumption prediction model (such as a regression model, a neural network model, a physical model, etc.) is used to predict the energy consumption status of the industrial equipment at the next sampling moment, calculate the expected consumption of different energy types (electricity, heat, fuel, etc.), predict the change trend of energy efficiency indicators, such as unit output energy consumption, energy utilization efficiency, etc., analyze the changes in energy consumption composition, such as the production energy consumption ratio, the auxiliary energy consumption ratio, etc., and form a predicted energy consumption change trend report, which describes the expected change in energy consumption after parameter adjustment. Then, the correlation between the energy consumption change trend and the instantaneous load data is analyzed, and the correlation analysis, gray correlation analysis and other methods are used to quantify the correlation between energy consumption changes and load response. An energy consumption-load sensitivity matrix is ​​established to represent the sensitivity of energy consumption to load changes. A multi-objective evaluation function is constructed, and the energy consumption optimization goal and the load response goal are considered at the same time. Through mathematical optimization methods (such as Pareto optimization, multi-objective evolutionary algorithm, etc.), the optimal balance point between energy consumption optimization and load response is found. A balance point represents the operating state with optimal energy consumption under the premise of meeting the load response requirements. Then, based on the found balance point and the current system status, the priority order of the energy consumption optimization strategy is determined. Under high load response requirements, the load response priority is higher than the energy consumption optimization. Under low load response requirements, the energy consumption optimization priority is increased. The priority is dynamically adjusted according to the real-time system status to ensure that the system can make reasonable decisions under different conditions. Finally, based on the determined priority, the execution order of parameter control is designed. For high-priority targets, the relevant parameters are adjusted first and a larger adjustment range is allocated. For low-priority targets, the relevant parameters are adjusted later and the adjustment range is limited. The timing arrangement of parameter adjustment is designed to clarify which parameters are adjusted first and which parameters are adjusted later. The coordination mechanism between parameters is designed to handle the coupling relationship and conflict between parameters. These control strategies are integrated into a complete set of multi-objective energy consumption optimization strategies. This strategy not only considers energy consumption optimization, but also takes into account load response, and can find the best balance between multiple objectives.

[0102] In an embodiment of the present invention, a method for identifying a trend inflection point includes:

[0103] Extracting continuous segments of load change according to the instantaneous characteristics of load change in the instantaneous load data;

[0104] Identify candidate inflection points based on the reversal of load change direction in the continuity segment;

[0105] According to the duration of the load change before and after the candidate inflection point, the candidate inflection point with a duration greater than a preset time threshold is screened out as the trend inflection point.

[0106] In this embodiment, first, a time series analysis is performed on the instantaneous load data, and the first-order difference of the load data (the amount of change at adjacent moments) is calculated to obtain the direction and amplitude information of the load change. The load data is divided into multiple continuous segments according to the change characteristics (such as directional consistency and similarity of the change rate). The load change in each segment has similar trend characteristics, such as a continuous rising segment, a continuous falling segment, a stable fluctuation segment, etc. The boundary marking and attribute annotation of the continuous segments obtained by segmentation are performed, including the start time, end time, average change rate, etc., and then the connection between the continuous segments is analyzed, focusing on the position where the load change direction is reversed. The direction reversal refers to the change pattern from rising to falling, or from falling to rising. Numerical analysis methods (such as extreme value detection, derivative sign change detection, etc.) are used to identify these direction reversal points, and each point is recorded. The time position of each reversal point, the rate of change before and after the reversal, the amplitude of the reversal and other characteristics are used to mark these reversal points as candidate inflection points. These candidate inflection points are the locations where the load trend may change. Finally, the persistence of the load change before and after each candidate inflection point is analyzed, and the trend duration before and after the inflection point is calculated. A preset time threshold is set. This threshold is determined according to the system characteristics and application requirements, usually ranging from a few sampling cycles to dozens of sampling cycles. Candidate inflection points whose trend duration before and after are greater than the preset time threshold are screened out. This can filter out short-term direction changes caused by random fluctuations or measurement noise, retain the inflection points that truly represent the trend change, and determine the candidate inflection points after screening as trend inflection points. These trend inflection points mark the real change in the load change pattern and are important trigger conditions for building a dynamic state transfer chain.

[0107] In an embodiment of the present invention, a method for determining a high-frequency oscillation mode includes:

[0108] According to the changing trend of instantaneous load data, the oscillation period of load change is extracted;

[0109] According to the continuity of the oscillation period, the stable segment of the oscillation is determined;

[0110] Determine whether the average length of the oscillation period in the stable segment is less than a preset period threshold;

[0111] If it is less than the preset period threshold, it is determined that the load changes and enters the high-frequency oscillation mode.

[0112] In this embodiment, the instantaneous load data is first subjected to frequency domain analysis, and methods such as fast Fourier transform (FFT) or wavelet transform are applied to convert the time domain signal into frequency domain representation, analyze the spectrum characteristics, identify the main frequency components and energy distribution, and use methods such as zero crossing detection and peak detection to directly measure the oscillation period in the time domain. The time intervals between adjacent peaks or valleys are recorded to form an oscillation period sequence, which contains the periodic variation information of the load oscillation. Then, the continuity and stability of the oscillation period sequence are analyzed, and the difference between adjacent periods, such as the standard deviation or coefficient of variation of the period length, is calculated. A stability judgment standard is set, such as "the coefficient of variation of N consecutive periods is less than M%". According to the judgment standard, time segments where the oscillation is relatively stable are identified. The oscillations within these segments show regularity and continuity, and the stability of each stable segment is recorded. The start time, end time, number of cycles included, and other features are then statistically analyzed for the oscillation cycles within each stable segment. Statistics such as the average cycle length and median cycle length are calculated. Based on system characteristics and application requirements, a preset cycle threshold is set. This threshold is the critical value for judging high-frequency oscillations. For example, for some industrial equipment, an oscillation cycle of less than 5 seconds may be defined as high-frequency oscillation. The average cycle length is compared with the preset cycle threshold to determine whether it is less than the threshold. Finally, if the average oscillation cycle of the stable segment is less than the preset cycle threshold and the oscillation amplitude meets certain conditions (such as exceeding twice the normal fluctuation range), it is determined that the load change enters the high-frequency oscillation mode. The characteristic parameters such as the start time, duration, average cycle and average amplitude of the high-frequency oscillation are recorded. These parameters will be used for subsequent priority determination and decision optimization.

[0113] In an embodiment of the present invention, a method for screening a set of actionable items includes:

[0114] Determine the initial scope of action selection based on the priority subset;

[0115] Based on the constraints, determine whether the actions within the initial range cause the state transition path to deviate from the target state node; based on the degree of deviation, eliminate actions whose deviation is greater than the preset deviation threshold;

[0116] According to the contribution of the remaining actions to the energy consumption optimization strategy, the feasible action set is determined.

[0117] In this embodiment, first, the initial range of action selection is set according to the previously determined priority subset. The priority subset is a set of actions screened according to the priority of the instantaneous load response, which represents the action type that is prioritized in the current system state. This initial range limits the search area of ​​the action space, avoiding exhaustive search in the complete action space and improving decision-making efficiency. Then, the constraints are applied to evaluate each action within the initial range. The constraints come from the transfer rules and physical limitations in the dynamic state transfer chain. A system model (such as a physical model, a data-driven model, etc.) is used to predict the state change of the system after executing each action, simulate the state transfer path, compare the predicted transfer path with the target state node, and calculate the degree of deviation. The degree of deviation can be measured using the Euclidean distance, Manhattan distance or Mahalanobis distance of the state vector. The larger the metric value, the more serious the deviation. Then set A preset deviation threshold is set, which is determined according to the system control accuracy requirements and application scenarios. The deviation degree of each action is compared with the preset threshold, and actions with a deviation degree greater than the threshold are eliminated. These actions may cause the system state to deviate from the expected target. Actions with a deviation degree less than or equal to the threshold are retained. These actions can guide the system to the target state. Finally, the energy consumption optimization contribution of the remaining actions is evaluated, and an evaluation function is constructed. Taking into account multiple factors such as energy consumption reduction effect, execution cost, response speed, etc., a deep reinforcement learning model is used to predict the long-term benefits of each action, such as Q value or value function. The remaining actions are sorted according to the evaluation results, and several actions with high contribution rankings are selected to form the final set of feasible actions. The actions in this set not only meet the constraints but also have a high energy consumption optimization effect, providing the system with multiple optional decision options.

[0118] In an embodiment of the present invention, a training method for a deep reinforcement learning model includes:

[0119] Construct a training dataset of instantaneous load data, energy consumption data, and operating status data at historical sampling moments;

[0120] Generate a sequence of state-action pairs based on the state transition path of the dynamic state transition chain in the training dataset; determine the training priority of the state-action pairs based on the instantaneous load response priority;

[0121] Adjust the sampling order of state-action pairs according to the training priority, where state-action pairs with high priority are sampled first;

[0122] Taking the output of dynamic decision logic as the goal, a deep reinforcement learning algorithm based on policy gradient is used for training to obtain a deep reinforcement learning model.

[0123] In this embodiment, the historical data of industrial equipment operation is first collected, including instantaneous load data, energy consumption data and operating status data at multiple sampling moments. The collected historical data are preprocessed, including missing value filling, outlier processing, data standardization, etc. The processed data are organized in chronological order to form a time series training data set to ensure data quality and representativeness, covering the operating status of the equipment under different working conditions. Then, the state transition path information of the dynamic state transition chain is extracted from the training data set, the state nodes and state transitions in the historical operation are identified, and the state nodes are mapped to the "state" representation in reinforcement learning, which is usually a multi-dimensional feature vector. The execution process of the state transition is identified. The control actions are mapped into the "action" representation in reinforcement learning, and the state and the corresponding action are combined to generate a sequence of state-action pairs. These sequences record the actions taken by the system in different states and their effects, forming the basic experience data of reinforcement learning. Then, according to the previously defined transient load response priority evaluation system, a training priority is assigned to each state-action pair in the training data set. State-action pairs with high response priority (such as states in low stability and high-frequency oscillation mode) receive higher training priority, state-action pairs with medium response priority receive medium training priority, and state-action pairs with low response priority receive lower training priority. The priority allocation reflects the importance and learning urgency of different states. Then, according to the assigned training priority, the experience replay strategy in the reinforcement learning process is adjusted, and a priority sampling mechanism is designed. State-action pairs with high priority have a higher probability of being sampled, which increases the learning frequency of important experience. A sampling probability calculation formula is designed, such as using a softmax function to convert priority into sampling probability, and implementing a "priority experience replay" mechanism to ensure that the model learns more from key experience. Finally, a deep reinforcement learning algorithm based on policy gradient (such as DDPG, TD3, SAC, PPO, etc.) is used to build a model architecture and design a network structure, which usually includes a policy network (Act or) and value network (Critic), design the reward function, comprehensively consider factors such as energy consumption optimization effect, load response speed, system stability, etc., use the output of dynamic decision logic as the training target, guide the model to learn the optimal decision strategy, use priority sampling experience data for model training, optimize model parameters through multiple rounds of iteration, use experience replay, target network and other technologies to improve training stability, evaluate model performance regularly, and test the generalization ability of the model through verification data set. After sufficient training, a deep reinforcement learning model that can adapt to different working conditions and make optimization decisions is obtained. This model will serve as the core decision-making engine of the system to guide the energy consumption optimization control of industrial equipment.

[0124] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art will be able to modify the technical solutions described in the foregoing embodiments or to substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

[0125] The thresholds in this specification are selected and set by those skilled in the art according to actual conditions.

[0126] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

Claims

1. A multi-objective energy consumption analysis and dynamic optimization system based on deep reinforcement learning, characterized by: include: Data acquisition module, used to obtain instantaneous load data, energy consumption data and operating status data of industrial equipment at each sampling moment; A transfer chain construction module, configured to construct a dynamic state transfer chain of the industrial equipment according to the instantaneous characteristics of the load change in the instantaneous load data and the state transfer logic of the parameters in the operating state data; a priority determination module, configured to determine the instantaneous load response priority of the industrial equipment at the corresponding sampling moment according to the topological characteristics of the state transition path in the dynamic state transition chain and the change trend of the instantaneous load data; A model optimization module, configured to adjust the operating parameters of the industrial equipment at the next sampling moment based on the instantaneous load response priority and the dynamic decision logic of the deep reinforcement learning model to obtain optimized operating parameters; The method for constructing the dynamic decision logic includes: determining a prioritized subset of action selections in a deep reinforcement learning model based on the instantaneous load response priorities; Constructing action selection constraints based on the transition path from the current state node to the next state node in the dynamic state transition chain; and screening out a set of feasible actions based on the priority subset and the constraints; Determining a final action as an output of the dynamic decision logic based on the contribution of the actions in the set of possible actions to the energy consumption optimization strategy; The method for adjusting the optimized operating parameters includes: Predicting a candidate set of next state nodes based on the transition path of the current state node in the dynamic state transition chain; According to the instantaneous load response priority, a state node with a matching priority is screened out from the candidate set as a target state node; Determining the direction of parameter adjustment according to the operating parameter range corresponding to the target state node; Determining an adjustment step size based on a deviation between the direction and the current operating parameters to obtain the optimized operating parameters; A multi-strategy fitting module is used to determine a multi-objective energy consumption optimization strategy for industrial equipment at each sampling moment based on the correlation characteristics between the optimized operating parameters and the energy consumption data; The real-time control module is used to control the operating parameters of the industrial equipment in real time based on the multi-objective energy consumption optimization strategy to achieve dynamic characterization of instantaneous load changes and energy consumption optimization.

2. The system according to claim 1, wherein: The method for constructing the dynamic state transfer chain includes: Identifying a trend inflection point of the load change based on instantaneous characteristics of the load change in the instantaneous load data as a trigger condition for state transition; Determining a state node at each sampling moment according to the discrete state of the operating parameters in the operating state data; According to the trigger condition, the state nodes are connected in chronological order to form an initial state transfer chain; According to the consistency of parameter change directions between adjacent state nodes in the initial state transfer chain, transfer paths that meet the consistency condition are screened out to form the dynamic state transfer chain.

3. The system according to claim 1, wherein: The method for determining the transient load response priority includes: Identifying a closed-loop structure of state nodes in a state transition path according to a topological characteristic of the state transition path in the dynamic state transition chain; Determining the stability level of the state transition according to the number of cycles of the state node in the closed-loop structure; Determining whether the load change enters a high-frequency oscillation mode according to a change trend of the instantaneous load data; The transient load response priority is determined according to a matching degree between the stability level and the high-frequency oscillation mode.

4. The system according to claim 1, wherein: The method for determining the multi-objective energy consumption optimization strategy includes: Predicting the energy consumption change trend of the industrial equipment at the next sampling moment based on the optimized operating parameters; Determining a balance point between energy consumption optimization and load response based on the correlation characteristics between the energy consumption change trend and the instantaneous load data; According to the balance point, the priority of the energy consumption optimization strategy is determined; according to the priority, the control order of the operating parameters is adjusted to obtain the multi-objective energy consumption optimization strategy.

5. The system according to claim 2, wherein: The method for identifying a trend inflection point includes: extracting continuous segments of load change according to instantaneous features of load change in the instantaneous load data; identifying candidate inflection points based on a reversal of a load change direction in the continuous segment; According to the duration of the load change before and after the candidate inflection point, the candidate inflection point with a duration greater than a preset time threshold is screened out as the trend inflection point.

6. The system according to claim 3, wherein: The method for determining the high-frequency oscillation mode includes: extracting an oscillation period of the load change according to a change trend of the instantaneous load data; determining a stable segment of the oscillation based on the continuity of the oscillation period; Determining whether the average length of the oscillation period in the stable segment is less than a preset period threshold; If it is less than the preset period threshold, it is determined that the load changes and enters the high-frequency oscillation mode.

7. The system according to claim 1, wherein: The method for screening the set of available actions comprises: determining an initial scope of action selection based on the priority subset; According to the constraint conditions, determine whether the action within the initial range causes the state transition path to deviate from the target state node; based on the degree of deviation, eliminate actions whose deviation is greater than a preset deviation threshold; The set of feasible actions is determined according to the contribution of the remaining actions to the energy consumption optimization strategy.

8. The system according to claim 1, wherein: The training method of the deep reinforcement learning model includes: Constructing a training data set of the instantaneous load data, the energy consumption data, and the operating status data at historical sampling moments; generating a state-action pair sequence according to a state transition path of the dynamic state transition chain in the training data set; determining a training priority of a state-action pair according to the instantaneous load response priority; Adjusting the sampling order of state-action pairs according to the training priority, wherein state-action pairs with higher priority are sampled first; The output of the dynamic decision logic is used as the target, and a policy gradient-based deep reinforcement learning algorithm is adopted for training to obtain the deep reinforcement learning model.

Citation Information

Patent Citations

  • GPU dynamic energy efficiency optimization operation method and system based on deep reinforcement learning

    CN116909378A

  • Systems, apparatus, and methods for dynamic cell state management for energy saving

    US20240291629A1