Hydraulic machine production line energy consumption optimization scheduling method and system based on digital twinning and reinforcement learning

CN122779554APending Publication Date: 2026-09-18CHENGDU ZHENGXI INTELLIGENT EQUIPMENT GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611226330.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-13
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

(1)依赖人工经验或固定工艺:操作人员手动分配任务、调用预设的压制压力与速度,无法根据实时电价、负载波动或紧急任务动态调整,造成设备空转、功率不匹配等浪费,

Benefits of technology

1、通过数字孪生模型中的效率系数η在线修正机制,利用实测功率与估算功率的差异,采用递归最小二乘法实时调整效率系数η,使功率估算能够自动跟随设备老化、油温变化等实际工况漂移,避免效率系数固定不变导致长期运行后精度下降的问题,确保能耗计算始终逼近真实值;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122779554A_ABST
    Figure CN122779554A_ABST
Patent Text Reader

Abstract

The application discloses a hydraulic machine production line energy consumption optimization scheduling method and system based on digital twinning and reinforcement learning, and relates to the technical field of hydraulic machine production scheduling.The method comprises the following steps: a digital twinning model is established for each hydraulic machine, real-time operation data is input, the efficiency coefficient is corrected online, and estimated power and cumulative energy consumption are output; the production line scheduling problem is modeled as a Markov decision process, including a state space, an action space and a reward function; the current state of each hydraulic machine is obtained at a fixed period, the trained scheduling strategy network is input, scheduling instructions are output and are sent to the corresponding programmable logic controller.The application can dynamically respond to time-of-use electricity prices and task changes while meeting process constraints, optimize production line energy consumption in real time, and significantly reduce comprehensive costs through the fusion of digital twinning and reinforcement learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hydraulic press production scheduling technology, and in particular to a method and system for optimizing energy consumption scheduling of hydraulic press production lines based on digital twins and reinforcement learning. Background Technology

[0002] Hydraulic presses, as core processing equipment in metal forming, composite material pressing, and powder metallurgy, are characterized by high power and high energy consumption. In industries such as automotive parts, aerospace, and home appliances, multiple hydraulic press production lines or multiple hydraulic presses often operate simultaneously, with their electrical energy consumption accounting for a significant proportion of the factory's total energy consumption. Therefore, how to reduce the energy consumption and electricity costs of hydraulic press production lines while ensuring timely completion of production tasks has become a pressing technical problem for the industry.

[0003] Currently, the scheduling and energy consumption management of hydraulic press production lines have the following main shortcomings: (1) Reliance on manual experience or fixed processes: Operators manually assign tasks and call preset pressing pressure and speed, which cannot be dynamically adjusted according to real-time electricity prices, load fluctuations or emergency tasks, resulting in waste such as equipment idling and power mismatch. (2) Energy-saving technologies for a single hydraulic press lack global coordination: measures such as frequency conversion drive and energy recovery only optimize a single piece of equipment, which cannot achieve multi-machine coordinated scheduling at the production line level, and it is also difficult to further reduce electricity costs by taking advantage of time-of-use pricing policies; (3) Most existing digital twin models are offline static models, which do not take into account the time-varying characteristics of the hydraulic system efficiency coefficient with changes in working conditions such as oil temperature and wear. This leads to a gradual decrease in the accuracy of energy consumption estimation as the running time increases, making it difficult to provide an accurate energy consumption benchmark for real-time scheduling decisions. (4) Lack of dynamic constraint capability on process parameters (such as pressure and speed) makes it impossible to ensure that the hydraulic press is always within the safe working window allowed by the process during dynamic adjustment, which poses a risk of substandard product quality; In summary, existing technologies lack methods and systems that can use the current time-of-use electricity price as a decision variable, have rolling optimization capabilities, output continuous target pressure and target speed values ​​to achieve global energy consumption optimization for multiple hydraulic presses, and simultaneously correct efficiency coefficients online while ensuring product quality. Therefore, there is an urgent need for an energy consumption optimization scheme that can integrate digital twin and reinforcement learning technologies, respond in real time to changes in electricity prices and task urgency, and dynamically schedule hydraulic press production lines. Summary of the Invention

[0004] This invention provides a method and system for optimizing energy consumption scheduling in hydraulic press production lines based on digital twins and reinforcement learning. It constructs a digital twin model with online correction capabilities to output high-precision power estimates in real time. Furthermore, it combines Markov decision processes with adjustable pressure and speed window constraints to achieve time-of-use pricing-oriented optimization. Production line-level rolling optimization scheduling to reduce overall manufacturing costs.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: A method for optimizing the scheduling of energy consumption in a hydraulic press production line based on digital twins and reinforcement learning is applied to a production line consisting of multiple hydraulic presses, each with its own programmable logic controller (PLC). The method includes the following steps: S1. Real-time estimation using digital twins: A digital twin model is established for each hydraulic press. Each digital twin model receives real-time operating data from the corresponding hydraulic press and outputs the estimated power and cumulative energy consumption of the hydraulic press. S2. Construction of reinforcement learning environment: The production line scheduling problem is modeled as a Markov decision process, which includes a state space, action space and reward function; The state space includes at least the main cylinder pressure, slider speed, hydraulic oil temperature, start / stop status, cumulative energy consumption, list of tasks to be processed, and current time-of-use electricity price for each hydraulic press. The motion space includes the target pressure value, target speed value, and start / stop commands for each hydraulic press; The reward function is a negative overall cost, which includes at least the energy consumption cost calculated based on the estimated power. S3. Training the scheduling policy network: Training the scheduling policy network within the Markov decision process framework; S4. Rolling Optimization and Command Issuance: Perform a scheduling operation at a fixed period to obtain the current status of each hydraulic press, input the trained scheduling strategy network, output the target pressure value, target speed value, and start / stop command of each hydraulic press, and issue them to the corresponding programmable logic controller of the hydraulic press.

[0006] Preferably, the digital twin model calculates and estimates the power based on the hydraulic system power formula using real-time operating data; the hydraulic system power formula includes an efficiency coefficient. The efficiency coefficient is updated online using a data-driven correction method, which specifically includes: collecting the measured power of each hydraulic press, and correcting the efficiency coefficient online based on the difference between the estimated power and the measured power.

[0007] Preferably, the data-driven correction employs recursive least squares or Kalman filtering.

[0008] Preferably, each task in the list of tasks to be processed includes a target pressure range, a target speed range, a deadline, an adjustable pressure window, and an adjustable speed window, all determined by the process. The adjustable pressure window is the upper and lower fluctuation range of the target pressure range, and the adjustable speed window is the upper and lower fluctuation range of the target speed range. The total cost also includes at least one of equipment changeover penalties and order delay penalties.

[0009] Preferably, in step S3, a near-end policy optimization algorithm or a deep Q-network algorithm is used to train the scheduling policy network.

[0010] Preferably, the target pressure value and target speed value output by the scheduling strategy network are constrained within the adjustable pressure window and adjustable speed window of the corresponding task, respectively, so as to adjust the process parameters while ensuring the product processing quality.

[0011] Preferably, the fixed period in step S4 is 1-10 minutes; Furthermore, when an urgent task is detected, a dispatch is triggered immediately without waiting for a fixed period. Urgent tasks refer to tasks marked by users in the system, or tasks whose deadline is less than or equal to the current time and have a fixed period.

[0012] Preferably, the state space is constituted by the current state of each hydraulic press; The action space is limited to the same scheduling cycle, and the target pressure value, target speed value and start / stop command are simultaneously input to the programmable logic controller of the corresponding hydraulic press. The energy cost in the reward function is the sum of the energy costs of the multiple hydraulic presses.

[0013] Preferably, the scheduling system for implementing the above-described energy consumption optimization scheduling method for hydraulic press production lines based on digital twins and reinforcement learning includes: Data acquisition module: It communicates with the pressure sensors, smart meters and programmable logic controllers of each hydraulic press to collect the main cylinder pressure, slider speed, hydraulic oil temperature, start and stop status and measured power of each hydraulic press; Digital twin module: Deployed on industrial edge computing devices and connected to the data acquisition module, it is used to calculate the estimated power and cumulative energy consumption of each hydraulic press based on the real-time operating data of each hydraulic press, and update the efficiency coefficient in the digital twin model online based on the difference between the measured power and the estimated power. Reinforcement learning scheduling module: Deployed on industrial edge computing devices and connected to the digital twin module, it is pre-deployed with a scheduling strategy network trained in step S3, which is used to output normalized action values ​​based on the current state through forward inference. The normalized action values ​​are de-normalized and constrained by adjustable pressure windows and adjustable speed windows to generate scheduling instructions. The current status includes the main cylinder pressure, slide speed, hydraulic oil temperature, start / stop status, cumulative energy consumption, list of tasks to be processed, and current time-of-use electricity price for each hydraulic press; the scheduling instructions include the target pressure value, target speed value, and start / stop instructions. Command issuance module: It communicates with the reinforcement learning scheduling module and the programmable logic controller and is used to write scheduling commands into the set value register of the programmable logic controller of the corresponding hydraulic press. Human-computer interaction module: Connected to the reinforcement learning scheduling module, it is used to display real-time status and receive human intervention instructions.

[0014] Preferably, the reinforcement learning scheduling module further includes an emergency task detection subunit, which continuously monitors the list of tasks to be processed. When an emergency task is detected, the reinforcement learning scheduling module is immediately triggered to perform a scheduling operation without waiting for a fixed period. Urgent tasks refer to tasks marked by users through the human-computer interaction module, or tasks whose deadline is less than or equal to the current time and have a fixed period.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By using the online correction mechanism of efficiency coefficient η in the digital twin model, the difference between measured power and estimated power is used to adjust efficiency coefficient η in real time using the recursive least squares method. This enables power estimation to automatically follow the drift of actual working conditions such as equipment aging and oil temperature changes, avoiding the problem of decreased accuracy after long-term operation due to a fixed efficiency coefficient, and ensuring that energy consumption calculation always approximates the true value. 2. Deploying the digital twin module and reinforcement learning scheduling module on industrial edge computing devices enables closed-loop decision-making to be completed on-site in the workshop, without relying on cloud networks, thus reducing communication latency and network outage risks. Simultaneously, the emergency task detection subunit immediately triggers scheduling when the deadline approaches or when manually marked, balancing the economic efficiency of long-term rolling optimization with the timeliness of handling unexpected tasks. 3. The energy cost in the reward function is calculated by multiplying the estimated power by the current time-of-use electricity price, enabling the scheduling strategy network to sense electricity price fluctuations. During peak electricity price periods, it can proactively reduce pressure and speed setpoints to suppress instantaneous power, and during flat or valley electricity price periods, it can appropriately increase operating parameters to strive for task progress, thereby minimizing the overall electricity cost of the production line, rather than simply reducing the total energy consumption. Attached Figure Description

[0016] Figure 1This is a structural framework diagram of the scheduling system of the present invention; Figure 2 This is the data stream of the scheduling system of the present invention; Figure 3 This is a flowchart illustrating the scheduling method of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Example 1: This example provides a hydraulic press production line energy consumption optimization scheduling system based on digital twins and reinforcement learning, used to execute the scheduling method described in this invention, such as... Figure 1 and Figure 2 As shown, it includes: Data acquisition module: It communicates with the pressure sensors, smart meters and programmable logic controllers of each hydraulic press to collect the main cylinder pressure, slider speed, hydraulic oil temperature, start and stop status and measured power of each hydraulic press; Digital twin module: used to run the digital twin model, deployed on industrial edge computing equipment, connected to the data acquisition module, used to calculate the estimated power and cumulative energy consumption of each hydraulic press based on the acquired real-time data, and update the efficiency coefficient in the digital twin model online based on the difference between the measured power and the estimated power; The digital twin module forms a one-to-one virtual-real mapping relationship with the hydraulic mechanism. Its output estimated power and cumulative energy consumption are updated in real time according to the operating status of the hydraulic press, providing an accurate energy consumption benchmark for the reinforcement learning scheduling module.

[0019] Reinforcement learning scheduling module: Deployed on industrial edge computing devices and connected to the digital twin module, it is pre-deployed with a scheduling strategy network trained in step S3, which is used to output normalized action values ​​based on the current state through forward inference. The normalized action values ​​are inversely normalized and constrained by adjustable pressure windows and adjustable speed windows to generate scheduling instructions. The current status includes the main cylinder pressure, slide speed, hydraulic oil temperature, start / stop status, cumulative energy consumption, list of tasks to be processed, and current time-of-use electricity price for each hydraulic press; the scheduling instructions include target pressure value, target speed value, and start / stop instructions; Instruction issuing module: Communicatively connected to the reinforcement learning scheduling module and the programmable logic controller, used to write the scheduling instruction into the set value register of the programmable logic controller of the corresponding hydraulic press; Human-computer interaction module: connected to the reinforcement learning scheduling module, used to display real-time status and receive manual intervention instructions.

[0020] The scheduling system in this embodiment is deployed on the hydraulic press production line. An industrial edge computing device is installed in the electrical control cabinet on the production line. Specifically, the industrial edge computing device can be an industrial edge gateway or an industrial control computer, and the scheduling system runs on this industrial edge computing device.

[0021] In this embodiment, the production line includes: N hydraulic presses (N≥2), where N=3 in this embodiment, namely #1, #2, and #3. Each hydraulic press is equipped with an independent hydraulic system and a programmable logic controller (PLC), and a pressure sensor (range 0-40MPa) is installed on the main cylinder pipeline, a magnetostrictive displacement sensor is installed on the slider, a PT100 temperature sensor is installed on the hydraulic oil tank, and a smart meter is installed on the main power supply circuit.

[0022] The PLCs, smart meters, and industrial edge computing devices of each hydraulic press are interconnected via industrial Ethernet, using Profinet or Modbus TCP as the communication protocol.

[0023] Furthermore, the industrial edge computing devices communicate with a host computer that has a production management system deployed on the workshop LAN. The production management system maintains and manages the list of tasks awaiting processing on the production line and provides the current time-of-use electricity price information. The industrial edge computing devices retrieve this information from the production management system as needed, which is then used by the reinforcement learning scheduling module to construct the current state.

[0024] In a preferred embodiment, the scheduling system adopts a three-layer collaborative architecture of "end-edge-cloud". End-side architecture: This refers to the pressure sensors, displacement sensors, temperature sensors, smart meters, and PLCs configured on each of the aforementioned hydraulic presses.

[0025] Edge architecture: This refers to the aforementioned industrial edge computing device, which is equipped with a digital twin module and a reinforcement learning scheduling module.

[0026] Cloud-side architecture: Includes a cloud monitoring server for data backup and visualization. It receives and stores historical operating data and cumulative energy consumption reports uploaded from the edge architecture for remote viewing by management personnel. During production line shutdowns or maintenance, historical operating data is used to perform offline incremental training on the digital twin module, and the updated efficiency coefficients are distributed to edge computing devices.

[0027] The human-machine interaction module is deployed on the operation terminal on the production line and communicates with the industrial edge computing device to display the real-time status of each hydraulic press and receive manual intervention commands.

[0028] The data acquisition module obtains the following five types of real-time data from the PLCs and smart meters of each hydraulic press in a polling manner: The system collects real-time data including main cylinder pressure (sampling frequency 10Hz), slider speed (sampling frequency 10Hz), hydraulic oil temperature (sampling frequency 1Hz), and start / stop status (sampling frequency 1Hz). Before entering the digital twin module for processing, all collected real-time data is appended with device numbers and unified timestamps by the data acquisition module. This data is then aggregated to form a real-time status set for the entire production line, which can be accessed by the digital twin module and the reinforcement learning scheduling module using device numbers as an index.

[0029] In this embodiment, the digital twin module is deployed in an edge computing device. Internally, this module constructs and independently runs a digital twin model for each hydraulic press. It executes cyclically with a calculation step size of 100 milliseconds (corresponding to a 10Hz frequency). Within each execution cycle, the digital twin module extracts the latest master cylinder pressure, slider speed, hydraulic oil temperature, and start / stop status data for each hydraulic press from the real-time status set, categorized by device number. This data is input into the corresponding digital twin model, and the module outputs the estimated power and updated cumulative energy consumption of the hydraulic press for the current cycle.

[0030] For each hydraulic press, the estimated power calculation is as follows: P est =(p×Q) / η, Among them, P est To estimate the power (unit: kW), p is the main cylinder pressure (unit: MPa), which is collected by a pressure sensor; Q is the hydraulic oil flow rate (unit: L / min), which is calculated from the slider speed and the effective area of ​​the main cylinder; η is the efficiency coefficient, the initial value of which is set by the equipment factory parameters (e.g., 0.85), and is continuously updated through online correction.

[0031] Cumulative energy consumption is obtained by integrating the estimated power over time: E est (t)=E est (t-△t)+P est (t)×△t, Among them, E est (t) represents the cumulative energy consumption at the current time t (unit: kWh). E est (t-△t) represents the cumulative energy consumption at the previous moment (unit: kWh), with an initial value of E. est (t)=0; P est (t) represents the estimated power at the current time t (unit: kW); △t is the sampling interval (0.1 seconds in this embodiment); The digital twin module uses the recursive least squares (RLS) method to update the efficiency coefficient η online. The specific steps are as follows: (1) Read the measured power P collected by the smart meter meas (k); (2) Calculate the prediction error: e(k) = P meas (k)-P est (k); (3) Calculate the gain coefficient: K(k)=[C(k-1)×φ(k)] / [λ+φ(k)²×C(k-1)]; (4) Update efficiency coefficient: η(k) = η(k-1) + K(k) × e(k); (5) Update the covariance matrix: C(k)=(1 / λ)×[1-K(k)×φ(k)]×C(k-1).

[0032] Where k is the sampling time number; P meas (k) represents the measured power at the k-th sampling time; P est (k) represents the estimated power at the k-th sampling time; e(k) is the prediction error at the k-th sampling time; φ(k) is the regression value, which is the theoretical hydraulic power, i.e., P(k)×Q(k); P(k) is the pressure at the k-th sampling time, and Q(k) is the flow rate at the k-th sampling time; C(k) is the covariance matrix value at the k-th sampling time, with an initial value of P(0)=1; K(k) is the gain coefficient at the k-th sampling time; η(k) is the efficiency coefficient updated at the kth sampling time. The initial value η(0) is set by the factory parameters of the equipment, for example, 0.85, and is continuously updated through online correction.

[0033] Through the above recursion, the efficiency coefficient η can be adjusted in real time according to factors such as equipment aging and oil temperature changes, so that the estimated power can continuously approach the measured power.

[0034] The digital twin module performs the above calculations at a frequency of 10Hz. To ensure time consistency, each round of calculation uses the input data with the most recent timestamp, ensuring that the estimated power and the measured power are aligned at the same time.

[0035] The estimated power and cumulative energy consumption of each hydraulic press output in real time by the digital twin module are sent to the reinforcement learning scheduling module as energy consumption status information. Together with the main cylinder pressure, slider speed, hydraulic oil temperature, start-stop status of each hydraulic press provided by the data acquisition module, as well as the list of tasks to be processed and time-of-use electricity price information from the production management system, they constitute the state observation basis required for the reinforcement learning scheduling module to make scheduling decisions.

[0036] In this embodiment, the reinforcement learning scheduling module is deployed in an edge computing device and connected to the digital twin module. This reinforcement learning scheduling module follows the Markov Decision Process (MDP) framework and is defined as follows: State space: includes the main cylinder pressure, slide speed, hydraulic oil temperature, start / stop status, cumulative energy consumption, list of tasks to be processed, and current time-of-use electricity price for each hydraulic press; Action space: includes the target pressure value, target speed value, and start / stop commands output by each hydraulic press; Reward function: negative total cost, including energy cost, equipment switching penalty and order delay penalty, where: energy cost is calculated by multiplying the estimated power output by the digital twin module by the current time-of-use electricity price.

[0037] The reward function has minimizing energy consumption cost as its core optimization objective. Device switching penalties and order delay penalties are offered as optional supplementary options. The goal is to reduce electricity costs while ensuring equipment stability and timely delivery.

[0038] The reinforcement learning scheduling module deploys a pre-trained scheduling policy network. This network is a multilayer perceptron (MLP), consisting of an input layer, at least one hidden layer, and an output layer. The number of nodes in the input layer equals the dimension of the state space, and the number of nodes in the output layer equals the dimension of the action space. The connection weights between layers are the weight parameters of the scheduling policy network, determining the mapping relationship from input states to output actions. For a production line with N hydraulic presses, the output layer of the scheduling policy network has N×3 nodes, corresponding to the normalized target pressure value, normalized target speed value, and start / stop activation value of each hydraulic press. These three sets of outputs together constitute the original action vector. This original action vector, after inverse normalization and window constraints, generates the final scheduling instruction.

[0039] During the training phase, the scheduling policy network calculates the gradient based on the reward function and uses the gradient to update the network weights. After training, the network weight parameters are exported and deployed to industrial edge computing devices. During the runtime phase, the scheduling policy network only performs forward inference calculations and does not calculate the reward function or update the weight parameters.

[0040] During the operation phase, the reinforcement learning scheduling module obtains the current state of each hydraulic press in each scheduling cycle. The specific data sources for the current state are as follows: The main cylinder pressure, slider speed, hydraulic oil temperature, and start / stop status of each hydraulic press are obtained from the data acquisition module. The cumulative energy consumption of each hydraulic press is obtained from the digital twin module; Obtain the list of tasks to be processed and the current time-of-use electricity price from the production management system. Each task in the list of tasks to be processed includes a target pressure range, a target speed range, a deadline, an adjustable pressure window, and an adjustable speed window. Receive manual intervention instructions (such as emergency task markers) from the human-computer interaction module; All the above parameters, after being timestamped at the same sampling time, constitute the original state vector.

[0041] Before inputting the scheduling strategy network, to prevent the numerical range of the state parameters of each hydraulic press from differing too much and affecting the convergence performance of the scheduling strategy network, the reinforcement learning scheduling module normalizes the original state vector. The specific rules are as follows: Master cylinder pressure: divided by the maximum pressure range (40MPa), mapped to [0,1]; Slider speed: Divide by the maximum speed range (50mm / s), map to [0,1]; Hydraulic oil temperature: Subtract 30℃ and divide by 50℃, then map to [0,1]. Cumulative energy consumption: The cumulative energy consumption of each hydraulic press is divided by 1000kWh and mapped to [0,1]. Current time-of-use electricity price: divided by the maximum electricity price range of 2 yuan / kWh, mapped to [0,1].

[0042] For the list of tasks to be processed: each task includes a target pressure range, a target speed range, a deadline, an adjustable pressure window, and an adjustable speed window. The normalization method follows the same normalization benchmark as the state parameters of each hydraulic press (pressure is divided by 40MPa, speed is divided by 50mm / s, and the deadline is converted to the remaining time and then divided by 24h) and mapped to the [0,1] interval. The normalized task attributes are truncated and zero-padded according to the preset maximum number of tasks (e.g., 5), and concatenated into a fixed-dimensional feature vector.

[0043] After the above normalization process, the original state vector forms a complete state vector.

[0044] The complete state vector is input into the scheduling strategy network. The scheduling strategy network outputs the normalized target pressure value, normalized target speed value, and start / stop activation value for each hydraulic press. After inverse normalization and window constraints, the final scheduling instruction is generated. The specific rules are as follows: Target pressure value: Output range [-1, 1], after inverse normalization mapped to [0, 40 MPa], and then clipped to the adjustable pressure window of the corresponding task (for example, the target pressure range is 12~15 MPa, the adjustable pressure window is ±5%, that is, the final output is constrained to 11.4~15.75 MPa).

[0045] Target velocity value: Output range [-1, 1], after inverse normalization mapped to [0, 50 mm / s], then clipped to the adjustable velocity window of the corresponding task (for example, the target velocity range is 15~20 mm / s, the adjustable velocity window is ±5%, that is, the final output is constrained to 14.25~21 mm / s). Start / Stop Commands: The start / stop activation value is mapped to [0,1] by the Sigmoid function. After threshold judgment (e.g., threshold is 0.5), if it is greater than 0.5, the output is 1 to indicate running; otherwise, the output is 0 to indicate stopping.

[0046] In normal production mode, the reinforcement learning scheduling module performs rolling optimization at fixed intervals: it waits for the fixed interval to arrive, obtains the current state, normalizes it, inputs it into the scheduling strategy network, outputs scheduling instructions and sends them to the instruction issuing module, and so on.

[0047] Example: For a single hydraulic press, its state parameters can be represented as a five-dimensional vector: [Main cylinder pressure: 12.5MPa, slider speed: 18mm / s, hydraulic oil temperature: 48℃, start / stop status: 1 (running), cumulative energy consumption: 125.6kWh].

[0048] For a production line with three hydraulic presses, the five-dimensional vectors of the three hydraulic presses are first concatenated in order of their numbers to form a fifteen-dimensional vector (3 × five-dimensional vector = fifteen-dimensional vector). Then, the list of tasks to be processed is converted into a fixed-dimensional feature vector using the aforementioned normalization method, and this vector is concatenated with the normalized current time-of-use electricity price to the fifteen-dimensional vector. These three elements together constitute the complete state vector of the input layer of the scheduling strategy network.

[0049] Example 2, as Figure 3 As shown, this embodiment provides an energy consumption optimization scheduling method for a hydraulic press production line based on digital twins and reinforcement learning. This method is implemented based on the scheduling system described above. The production line includes multiple hydraulic presses, each equipped with an independent programmable logic controller (PLC).

[0050] The following describes in detail the specific implementation methods of each step, using the three hydraulic presses (numbered #1, #2, and #3) in the production line as an example.

[0051] S1. Real-time estimation using digital twins: A digital twin model is established for each hydraulic press. The real-time operating data of the hydraulic press corresponding to the digital twin model is used as input, and the estimated power and cumulative energy consumption of the hydraulic press are output. The real-time operating data includes master cylinder pressure, slider speed, hydraulic oil temperature, and start / stop status. The digital twin model calculates estimated power based on the hydraulic system power formula; this formula includes an efficiency coefficient; the efficiency coefficient η is updated online using a data-driven correction method, employing recursive least squares or Kalman filtering. Specifically, this includes: collecting the measured power of each hydraulic press, and using the difference between the estimated power and the measured power to correct the efficiency coefficient online.

[0052] In this embodiment, the efficiency coefficient η is updated online using the recursive least squares method. The difference between the measured power and the estimated power is used for recursive correction. The specific recursive steps are the same as those described in the digital twin module of the system embodiment, and will not be repeated here.

[0053] S2. Construction of reinforcement learning environment: The hydraulic press production line scheduling problem is modeled as a Markov decision process, including the definition of state space, action space and reward function.

[0054] The state space includes at least the main cylinder pressure, slider speed, hydraulic oil temperature, start / stop status, cumulative energy consumption, list of tasks to be processed, and current time-of-use electricity price for each hydraulic press. The motion space includes the target pressure value, target speed value, and start / stop commands for each hydraulic press; The overall cost with a negative reward function includes at least the energy cost calculated based on the estimated power.

[0055] Each task in the list of tasks to be processed includes a target pressure range, a target speed range, a deadline determined by the process, as well as an adjustable pressure window and an adjustable speed window; the adjustable pressure window is the upper and lower fluctuation range of the target pressure range, and the adjustable speed window is the upper and lower fluctuation range of the target speed range.

[0056] The state space is composed of the current states of the multiple hydraulic presses. The target pressure value, target speed value, and start / stop command in the action space are sent to the programmable logic controller of the corresponding hydraulic press within the same scheduling cycle. The energy cost in the reward function is the sum of the energy costs of the multiple hydraulic presses.

[0057] For a production line containing three hydraulic presses (numbered #1, #2, and #3), the state space is composed of the current states of each of the three hydraulic presses. At each scheduling moment, the main cylinder pressure, slide speed, hydraulic oil temperature, start / stop status, and accumulated energy consumption of hydraulic press #1 constitute its current state; the corresponding parameters of hydraulic presses #2 and #3 constitute their respective current states. These three sets of current states, along with the global list of tasks to be processed and the current time-of-use electricity price, are concatenated to form a complete original state vector, which serves as the input to the scheduling strategy network.

[0058] The action space is limited so that within the same scheduling cycle, the scheduling strategy network simultaneously outputs three sets of action values, corresponding to the target pressure value, target speed value, and start / stop command for hydraulic presses #1, #2, and #3, respectively, which are then input to the programmable logic controller (PLC) of the corresponding hydraulic press. The target pressure value and target speed value of each hydraulic press are constrained by the adjustable pressure window and adjustable speed window of its current task, respectively, and the start / stop command is 0 or 1.

[0059] The reward function is defined as the negative total cost R, i.e., R = -(C energy +C switch +C delay ).in, C energy The energy cost is calculated by multiplying the estimated power output from the digital twin model in step S1 by the current time-of-use electricity price; the energy cost in the reward function is the sum of the energy costs of the three hydraulic presses. Energy cost per unit of hydraulic press: Energy cost = P est ×price×△t, where P est The estimated power of the hydraulic press is output by the digital twin model, where price is the current time-of-use electricity price and Δt is the scheduling cycle length.

[0060] C switch This is a penalty for equipment switching; used to suppress frequent starts and stops of the hydraulic press and drastic adjustments to operating parameters. In one embodiment, a fixed penalty value is applied when the start / stop state of the hydraulic press changes; when the change in the target pressure value or target speed value relative to the previous scheduling cycle exceeds a preset threshold, a penalty proportional to the change is applied.

[0061] C delay Penalty for order delays. Used to ensure on-time delivery of tasks. In one embodiment, when the actual completion time of a task exceeds its deadline, a delay penalty is calculated by multiplying the excess time by a unit time penalty coefficient; if the task is completed before the deadline, then C... delay It is zero.

[0062] S3. Training the scheduling policy network: Under the Markov decision process framework, the scheduling policy network is trained using a deep reinforcement learning algorithm; the deep reinforcement learning algorithm is either the proximal policy optimization algorithm or the deep Q-network algorithm.

[0063] Based on the Markov decision process framework constructed in step S2, this step constructs and trains the scheduling policy network. Specifically: First, a scheduling policy network to be trained is constructed. The structure of this scheduling policy network adopts a multilayer perceptron, and its specific network architecture has been described in detail in the aforementioned system embodiments, and will not be repeated here.

[0064] In this embodiment, the deep reinforcement learning algorithm employs a proximal policy optimization algorithm. During the training phase, the scheduling policy network learns through interaction with the external environment. The external environment physically corresponds to the digital twin module, data acquisition module, and production management system of this system, and can also be replaced by a simulated environment constructed from historical data or a digital twin simulation model. After the scheduling policy network outputs a scheduling action based on the current state, the external environment calculates and returns the immediate reward value based on the reward function defined in step S2. Training aims to maximize the cumulative reward value, which is the sum of the immediate reward values ​​at each time point after being weighted by a discount factor. After training is completed, the network weight parameters are exported and deployed to the industrial edge computing device. During the runtime phase, the network only performs forward inference, without calculating the reward function or updating the weight parameters.

[0065] The specific training steps are as follows: (1) Initialize network parameters The Xavier method (uniform distribution initialization) is used to set the initial values ​​of the weight parameters of the scheduling policy network, ensuring that the variance of the output of each layer remains consistent. To balance the impact of near-term and long-term decisions on network weight updates, a discount factor γ (dimensionless, ranging from 0 to 1, set to 0.99 in this embodiment) is introduced; the closer to 1, the more emphasis is placed on long-term rewards. The learning rate is set to 3 × 10⁻⁶. -4 .

[0066] (2) Generate training data Training samples are generated using historical operating data of the hydraulic press production line or simulation data generated based on the digital twin model (simulating the production line operation process under different electricity price curves, task arrival rates and equipment states). Each training sample is a quadruple containing the current state, the action performed, the immediate reward value obtained, and the next state to which the action is transitioned after the action is performed.

[0067] (3) Execute the training of the near-end policy optimization algorithm In each training round, the scheduling policy network outputs a set of scheduling actions based on the current state. The external environment (i.e., the combination of the digital twin module, data acquisition module, and production management system) returns the corresponding immediate reward value according to the reward function defined in step S2. The generated quadruple data is stored in the experience replay pool. The experience replay pool is a fixed-capacity data buffer used to store historical interaction data for random sampling during training to break the temporal correlation between data and improve training stability. When the amount of data in the experience replay pool reaches the preset batch size (256 in this embodiment), a batch of quadruple data is randomly sampled for training. During training, the immediate reward values ​​at each time step are weighted by a discount factor γ and accumulated to obtain the cumulative reward value, calculated as follows:

[0068] Where t is the sequence number of the current decision step, and i is the index of the subsequent steps starting from the current step (the value ranges from 0 to Tt). G t Let be the cumulative reward value at step t. r t+i Let be the instantaneous reward value at step t+i. γ is the discount factor, and T is the length of the decision sequence. The gradient is calculated based on the cumulative reward value, and the weight parameters of the scheduling policy network are iteratively updated using the pruning objective function of the near-end policy optimization algorithm.

[0069] (4) Convergence judgment After every 100 rounds of training, the average cumulative reward value of the current scheduling policy network in the verification environment is calculated. When the fluctuation range of the average cumulative reward value is less than the preset threshold (±5% in this embodiment) for 10 consecutive rounds, the scheduling policy network training is considered to have converged and training is stopped; if it has not converged, steps (2) to (3) are repeated until the preset maximum number of training rounds (5000 rounds in this embodiment) is reached and then training stops.

[0070] (5) Network deployment After training, the weight parameters of the scheduling policy network are exported and deployed to the reinforcement learning scheduling module in the industrial edge computing device. During the actual operation phase (i.e., during the execution of step S4), the scheduling policy network only performs forward inference and no longer calculates the reward function or updates the weights.

[0071] S4. Rolling optimization and command issuance: Perform a scheduling operation at a fixed period, including obtaining the current status of each hydraulic press, inputting the trained scheduling strategy network, outputting the target pressure value, target speed value and start / stop command of each hydraulic press and issuing them to the corresponding programmable logic controller of the hydraulic press. The fixed cycle is 1-10 minutes; Furthermore, when an urgent task is detected, the rolling optimization and instruction issuance process is executed immediately without waiting for the fixed period; The urgent task refers to a task marked by the user through the human-computer interaction module, or a task whose deadline is less than or equal to the current time and has a fixed period.

[0072] During actual operation, the trained scheduling strategy network is deployed in industrial edge computing devices. At this point, the scheduling strategy network stops updating parameters and only performs forward inference. Its interaction object shifts to the real production line: in each scheduling cycle, the network receives the current global state from the data acquisition module and the digital twin module, and outputs specific scheduling instructions after normalization, inference, and denormalization. These instructions are then issued to the PLC via the instruction delivery module, changing the actual operating state of the hydraulic press and completing one interaction with the real environment.

[0073] This process is divided into two working modes: normal production mode and emergency task triggering mode. In one embodiment, the fixed scheduling cycle in normal production mode is set to 5 minutes, and the cycle length can be adjusted from 1 minute to 10 minutes according to the actual production line conditions. The specific operation sequences for the two working modes are described below.

[0074] In normal production mode, the scheduling system repeatedly performs the following operations at fixed intervals: (1) Obtaining State and Network Reasoning At the beginning of each scheduling cycle, the reinforcement learning scheduling module obtains the current state of each hydraulic press in the manner described in the system implementation, and forms an original state vector after timestamp alignment; the original state vector is normalized to form a complete state vector; the complete state vector is input into the scheduling strategy network trained and deployed in step S3, performs forward inference, and outputs the normalized action value corresponding to each hydraulic press, namely the normalized target pressure value, the normalized target speed value, and the start / stop activation value; (2) Action denormalization and instruction generation The reinforcement learning scheduling module processes each component of the normalized action value separately, performs inverse normalization and window constraints on the normalized target pressure value and normalized target velocity value, and performs threshold discrimination on the start and stop activation values ​​to generate the final scheduling instruction. The specific generation rules are as follows: Target pressure value: After being mapped to the pressure range, it is clipped to the adjustable pressure window of the current task; Target speed value: After being mapped to the speed range, it is clipped to the adjustable speed window of the current task; Start / Stop Commands: Start / stop activation values ​​are determined by a threshold (for example, if the threshold is 0.5, output 1 if it is greater than 0.5 to indicate running, otherwise output 0 to indicate stopping).

[0075] (3) Issuance of instructions The generated scheduling command is sent to the command issuing module, which then writes the target pressure value, target speed value, and start / stop command into the setpoint register of the corresponding hydraulic press programmable logic controller (PLC) via industrial Ethernet. Upon receiving the new command, the PLC adjusts the proportional valve opening and master cylinder motion parameters, and the actual operating state of the hydraulic press smoothly transitions to the new setpoint within seconds.

[0076] After completing the above operations, the scheduling system waits for the next fixed period to arrive and then executes the above process again, forming a rolling optimization closed loop.

[0077] In emergency task trigger mode: The reinforcement learning scheduling module has a built-in emergency task detection subunit that continuously monitors the list of tasks to be processed within each fixed period. A task is determined to be an emergency task when any of the following conditions are met: Condition 1: The user manually marks a task as "urgent" through the human-computer interaction module; Condition 2: The difference between the deadline of a task and the current time is less than or equal to the fixed period (i.e., 5 minutes).

[0078] Once an urgent task is detected, the system immediately triggers a scheduling operation without waiting for the current fixed cycle to end. The operations performed after triggering are as follows: the original state vector at the current moment is immediately acquired, normalized, and input into the scheduling strategy network to generate scheduling instructions, which are then sent to the programmable logic controllers (PLCs) of each hydraulic press. This mechanism ensures that urgent orders can be processed promptly, avoiding delivery delays caused by waiting for the fixed scheduling cycle.

[0079] In a preferred embodiment, the instruction issuing module employs a "write-read-compare" confirmation mechanism to ensure that the scheduling instructions are reliably executed. Specifically, after writing the instruction to the programmable logic controller's setpoint register, the instruction issuing module immediately reads the value of the register and compares it with the written value. If the comparison matches, the instruction is confirmed to have been successfully issued. If the comparison does not match or there is no response after a timeout, a retransmission mechanism is triggered, for example, retransmitting three times. If the retransmission still fails, a communication anomaly warning is pushed to the human-machine interaction module.

[0080] By combining the aforementioned fixed-cycle rolling optimization with emergency task triggering, this step can dynamically respond to changes in time-of-use electricity prices, operating conditions, and the list of tasks to be processed, while ensuring stable production line operation and timely task delivery. It continuously outputs scheduling instructions that minimize long-term comprehensive costs, thereby achieving real-time optimized scheduling of production line-level energy consumption.

[0081] In terms of the overall implementation timeline, this method is divided into an offline training phase and an online execution phase, such as... Figure 3 As shown.

[0082] The offline training phase, corresponding to steps S2 and S3, trains the scheduling strategy network in one go, based on historical production line operation data or simulation data generated from a digital twin model. After training, the network weight parameters are deployed to industrial edge computing devices. This phase only needs to be executed once in actual production, or only re-executed when the model is upgraded.

[0083] The online operation phase corresponds to steps S1 and S4. S1 runs continuously as a background process at a frequency of 10Hz, updating the estimated power and cumulative energy consumption of each hydraulic press in real time. S4 is triggered at a fixed period (e.g., 5 minutes), using the estimated power and cumulative energy consumption output in real time from S1 and the trained scheduling strategy network to generate scheduling instructions and send them to the PLC of each hydraulic press, forming a closed-loop optimization.

[0084] After the industrial edge computing device is powered on, the scheduling system automatically enters the online operation phase without manual intervention. Step S1 is then started and continuously executed at a frequency of 10Hz, and step S4 is triggered to execute when the first fixed cycle is reached.

[0085] Example 3 illustrates how a complete scheduling example demonstrates how the entire process, from data acquisition to instruction issuance, collaboratively optimizes production line energy consumption in real time. This example is based on a production line with two hydraulic presses (numbered #1 and #2), with a fixed scheduling cycle of 5 minutes. The scheduling process is as follows: (a) Initial state (10:00:00 AM) The scheduling system triggers a new scheduling cycle. The initial state parameters of each hydraulic press at the start of the scheduling cycle are shown in Table 1. These include main cylinder pressure, slide speed, hydraulic oil temperature, start / stop status, cumulative energy consumption, and current task.

[0086] Table 1: Initial State Parameters of Hydraulic Press #1 and Hydraulic Press #2 #1 14.2MPa 20mm / s 50℃ 1 (Run) 156.3kWh Task A #2 0MPa 0mm / s 46℃ 0 (Stop) 89.7kWh Pending allocation In the list of tasks to be processed, Task A is being executed on Hydraulic Press #1. Its target pressure range is 13~16MPa, and its target speed range is 18~22mm / s. The adjustable pressure window and adjustable speed window are both ±5%. The adjustable pressure window is [12.35, 16.8]MPa, and the adjustable speed window is [17.1, 23.1]mm / s. The deadline is 11:00:00, and there is ample time remaining.

[0087] Task B is pending assignment. The target pressure range is 9~11MPa, the target speed range is 15~18mm / s, and the adjustable pressure window and adjustable speed window are both ±4%. That is, the adjustable pressure window is [8.64,11.44]MPa, and the adjustable speed window is [14.4,18.72]mm / s. The deadline is 10:20:00.

[0088] Current time-of-use electricity price: 0.6 yuan / kWh (flat rate).

[0089] (ii) Execution: Real-time estimation using digital twins Following step S1, a digital twin model is established for each hydraulic press. Taking hydraulic press #1 as an example, the digital twin model inputs the main cylinder pressure p = 14.2 MPa, the slider speed v = 20 mm / s, and the effective area of ​​the main cylinder A = 0.0433 m². The hydraulic oil flow rate Q = 52 L / min is obtained by converting the slider speed and the effective area of ​​the main cylinder. The current efficiency coefficient η = 0.81. The estimated power is then calculated. P est =(p×Q) / η=(14.2×52) / 0.81≈911.6KW, Smart meter reads measured power P meas =930.0kW, calculate the prediction error: e(k)=P meas (k)-P est (k) = 930.0kW - 911.6kW = 18.4kW; The recursive least squares method is used to fine-tune the current efficiency coefficient η to 0.81335, based on the cumulative energy consumption formula in step S1: E est (t)=E est (t-△t)+P est (t)×△t, sampling interval △t=0.1 seconds. The estimated power P is calculated using this method. est Substituting 911.6kW into the calculation, the energy consumption increment for this cycle is: E est (t)=156.3+911.6×0.1 / 3600≈156.325kWh.

[0090] #2 hydraulic press is in a stopped state, with estimated power of zero and cumulative energy consumption unchanged. The current efficiency coefficient η2 is 0.85, which is set by the equipment factory parameters.

[0091] (III) Execution: State Construction and Normalization According to the state space defined in step S2, obtain the main cylinder pressure, slider speed, hydraulic oil temperature and start / stop status of each hydraulic press; the cumulative energy consumption, list of tasks to be processed and the current time-of-use electricity price of each hydraulic press; The above parameters are aligned with timestamps to form the original state vector, and then normalized.

[0092] In this example, the normalization benchmark is consistent with the system implementation: main cylinder pressure divided by 40MPa, slide block speed divided by 50mm / s, hydraulic oil temperature first subtracted by 30℃ and then divided by 50℃, cumulative energy consumption divided by 1000kWh, and current time-of-use electricity price divided by the maximum electricity price of 2 yuan / kWh. Based on this benchmark, the components of the five-dimensional state vectors of hydraulic press #1 and hydraulic press #2 after normalization are shown in Tables 2 and 3 below: Table 2: Original values ​​→ Normalized values ​​of the five-dimensional state vector of hydraulic press #1 First dimension Master cylinder pressure 14.2MPa 0.355 Second dimension Slider speed 20mm / s 0.400 The third dimension Hydraulic oil temperature 50℃ 0.400 4th dimension Start-stop status 1 (Run) 1 5th dimension Cumulative energy consumption 156.3kWh 0.156 Table 3: Original values ​​→ Normalized values ​​of the five-dimensional state vector of hydraulic press #2 First dimension Master cylinder pressure 0MPa 0 Second dimension Slider speed 0mm / s 0 The third dimension Hydraulic oil temperature 46℃ 0.320 4th dimension Start-stop status 0 (Stop) 0 5th dimension Cumulative energy consumption 89.7kWh 0.090 Feature vector of the task list to be processed: Task A has ample remaining time, 60 minutes remaining (until 11:00), Task B has 20 minutes remaining (until 10:20), and the remaining time is tight. After normalization, they are concatenated into a fixed-dimensional feature vector.

[0093] Current time-of-use electricity price: 0.6 / 2 = 0.3 The above components together form a complete state vector, which is then input into the scheduling policy network to generate scheduling instructions.

[0094] (iv) Execution: Scheduling strategy network reasoning and scheduling instruction generation Following step S4, the complete state vector is input into the pre-trained scheduling policy network. Based on the converged scheduling policy network (which learned scheduling rules during training to appropriately accelerate during flat electricity price periods, appropriately decelerate during peak electricity price periods, and take into account task urgency), two sets of normalized action values ​​are output.

[0095] After processing each component of the normalized action value, the scheduling instructions are generated as shown in Table 4: Table 4. Scheduling instructions for hydraulic presses #1 and #2

[0096] #1 134.5MPa 21mm / s 1 (Run) With lower electricity prices in the same period, it is advisable to accelerate the process to meet project deadlines. #2 9.5MPa 15mm / s 1 (Run) Initiate task #2 to execute task B and ensure on-time delivery. After verification, the target pressure of hydraulic press #1, 14.5 MPa, is within the adjustable pressure window [12.35, 16.8] MPa, and the target speed of 21 mm / s is within the adjustable speed window [17.1, 23.1] mm / s. The target pressure of the #2 hydraulic press, 9.5 MPa, falls within the adjustable pressure window [8.64, 11.44] MPa, and the target speed of 15 mm / s falls within the adjustable speed window [14.4, 18.72] mm / s. Therefore, this adjustment achieves energy savings while still ensuring normal product pressing.

[0097] During the generation of the aforementioned scheduling instructions, the reward function of the reinforcement learning scheduling module already includes equipment switching penalties and order delay penalties. Therefore, when the scheduling policy network attempts to change the start / stop state, the equipment switching penalty applies a negative reward to suppress frequent start / stop operations; when there is a risk of task delay, the order delay penalty applies a negative reward, prompting the scheduling policy network to prioritize ensuring tasks are completed on time. In this example, both hydraulic presses #1 and #2 remained operational without any start / stop switching, and both task deadlines were met—a result of the combined effect of these two penalties.

[0098] (v) Execution: Issuance and execution of instructions The instruction issuing module writes the scheduling instructions into the setpoint register of the corresponding hydraulic press's programmable logic controller. Specifically: Write the target pressure value of 14.5MPa, the target speed value of 21mm / s, and the start / stop command 1 (run) to the programmable logic controller of hydraulic press #1. Write the target pressure value of 9.5MPa, the target speed value of 15mm / s, and the start / stop command 1 (run) to the programmable logic controller of hydraulic press #2.

[0099] Upon receiving the instruction, the corresponding programmable logic controllers adjusted the opening of their respective proportional valves. Hydraulic press #1's pressure was adjusted to 14.5 MPa and its speed to 21 mm / s; hydraulic press #2's pressure was adjusted to 9.5 MPa and its speed to 15 mm / s. This adjustment put the production line into energy-saving operation, reducing instantaneous power consumption while meeting the adjustable window constraints.

[0100] (vi) Electricity price peak period triggers dispatch adjustment (10:05:00 AM) Five minutes later, the dispatch system triggered the dispatch cycle again. At this time, the time-of-use electricity price changed to 1.2 yuan / kWh (peak electricity price). The current actual operating status of each hydraulic press, fed back by the data acquisition module and the digital twin module, is shown in Table 5. Table 5: Actual operating status of hydraulic press #1 and hydraulic press #2 during peak periods #1 14.5MPa 21mm / s 158.1kWh Task A is approximately 35% complete. #2 10.0MPa 16mm / s 91.2kWh Task B is approximately 20% complete. The digital twin model updates the efficiency coefficients of each hydraulic press (the efficiency coefficient η1 of #1 is finely adjusted to 0.815 based on the recursive least squares method, and the efficiency coefficient η2 of #2 is corrected from the initial value of 0.85 to 0.84 after several rounds of adjustments), calculates the current estimated power, and updates the cumulative energy consumption.

[0101] After receiving the new state, the scheduling strategy network identifies that the electricity price has entered the peak period. After forward inference, it outputs a normalized action value. The reinforcement learning scheduling module performs inverse normalization and window constraints on the normalized action value and outputs a new scheduling instruction, as shown in Table 6 below: Table 6: New Dispatch Instructions for Hydraulic Presses #1 and #2 #1 13.5MPa 18mm / s 1 (Run) High peak electricity prices necessitate reducing main cylinder pressure and speed to decrease instantaneous power. #2 9.5MPa 15mm / s 1 (Run) Synchronous speed reduction, but still maintaining operation to ensure task B is completed on time. Verification showed that the target pressure of hydraulic press #1 (13.5 MPa) was within the adjustable pressure window [12.35, 16.8] MPa, and the target speed of hydraulic press #2 (18 mm / s) was within the adjustable speed window [17.1, 23.1] mm / s. Similarly, the target pressure of hydraulic press #2 (9.5 MPa) was within the adjustable pressure window [8.64, 11.44] MPa, and the target speed of hydraulic press #2 (15 mm / s) was within the adjustable speed window [14.4, 18.72] mm / s. Therefore, this peak-period adjustment, while meeting process constraints, reduced the instantaneous power consumption during peak electricity pricing periods, thus lowering electricity costs.

[0102] (vii) Comparison of energy consumption optimization effects To verify the energy-saving effect of the method of the present invention, the results of the scheduling method of the present invention are compared with the results of the fixed parameter scheduling. The fixed parameter scheduling means that each hydraulic press always operates at the median of the target pressure range and the median of the target speed range for the current task, without dynamic adjustment. The median refers to the arithmetic mean of the upper and lower limits of the corresponding range. For example, if the target pressure range of task A is 13~16 MPa, then under fixed parameter scheduling, hydraulic press #1 always operates at 14.5 MPa; if the target speed range of task B is 15~18 mm / s, then hydraulic press #2 always operates at 16.5 mm / s.

[0103] The total power of the two scheduling methods in each time period is compared in Table 7 below: Table 7: Comparison of Total Power between Fixed Parameter Scheduling and the Invention 10:00-10:05 (flat section) 0.6 yuan / kWh 1230kW 1285kW +55kW (slightly increased to allow for speed reduction during peak hours). 10:05-10:10 (Peak period) 1.2 yuan / kWh 1230kW (no adjustment) 1090kW -140kW (a decrease of approximately 11.4%) Total power for fixed parameter scheduling: refers to the sum of estimated power calculated by the digital twin model when both hydraulic presses #1 and #2 are running at the median of their respective target pressure and target speed ranges. Total power of this invention: refers to the sum of the estimated power calculated by the digital twin model when the two hydraulic presses #1 and #2 are running according to the target pressure and target speed values ​​output by the scheduling strategy network.

[0104] As shown in Table 7, during the flat electricity price period (0.6 yuan / kWh), this invention, by appropriately increasing the operating speed of hydraulic press #1 and starting hydraulic press #2, made the total power slightly higher than the fixed parameter scheduling scheme by 55 kW, reserving task progress space for subsequent peak-segment speed reduction operation; during the peak electricity price period (1.2 yuan / kWh), this invention, by reducing the pressure and speed of hydraulic press #1 (from 14.5 MPa to 13.5 MPa, 21 mm / s to 18 mm / s), and keeping hydraulic press #2 running at 9.5 MPa and 15 mm / s to ensure that task B is completed on time, made the total power lower than the fixed parameter scheduling scheme by about 140 kW, a reduction of about 11.4%.

[0105] As can be seen from the above comparison, the method of the present invention, through a dynamic scheduling strategy of appropriately increasing speed during flat electricity price periods and appropriately decreasing speed during peak electricity price periods, effectively reduces the instantaneous power during peak electricity price periods while meeting the task deadline requirements.

[0106] Quantifying the above effects from an energy consumption cost perspective: Taking the 10:05-10:10 peak period as an example, the energy consumption cost of fixed parameter scheduling = total power 1230kW × electricity price 1.2 yuan / kWh × period duration 5 / 60 hours = 123 yuan; the energy consumption cost of the method of this invention = total power 1090kW × electricity price 1.2 yuan / kWh × period duration 5 / 60 hours = 109 yuan, saving 14 yuan per period. This shows that the present invention, through peak period speed reduction, not only reduces instantaneous power but also directly reduces the energy consumption cost of the production line.

[0107] (viii) Example of detecting an emergency task trigger Suppose that at 10:08:00, the user marks task B as "urgent" through the human-computer interaction module. After detecting this mark, the emergency task detection subunit does not wait for the next 5-minute cycle, but immediately collects the current state (current speed of hydraulic press #2 is 15mm / s, current cumulative energy consumption is 91.2kWh, and task B has 12 minutes remaining). After normalization, it forms a complete state vector, which is then input into the scheduling strategy network to perform forward inference.

[0108] The scheduling strategy network determines the progress based on the current remaining time and the remaining workload of task B. If it determines that the task may not be completed on time at the current speed of 15 mm / s, it outputs a speed increase command. After inverse normalization and window constraint pruning, a temporary speed command of 17 mm / s is generated. This value is within the adjustable speed window [14.4, 18.72] mm / s of task B, which meets the process constraints.

[0109] After the command was issued, the speed of hydraulic press #2 increased from 15 mm / s to 17 mm / s. The scheduling system continuously monitors the progress of task B, and determines the completion of task B based on the cumulative energy consumption increment (i.e., the increase in energy consumption from the start time of task B to the current time) reaching the energy consumption required for task B. In this embodiment, the energy consumption required for task B is approximately 2.5 kWh. This value is determined by the process parameters of task B (pressure range 9~11 MPa, speed range 15~18 mm / s) and the estimated processing time. The specific calculation method is: Energy consumption = P est ×T B , where P est To estimate the power from the corresponding digital twin model based on the process parameters of Task B, T B This is the estimated processing time for Task B. Simultaneously, the corresponding digital twin module continuously updates the accumulated energy consumption; when the cumulative energy consumption increment reaches this value, Task B is considered complete.

[0110] Once task B is completed, the scheduling system automatically resumes normal fixed-cycle rolling optimization and waits for the next 5-minute cycle to arrive.

[0111] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for optimizing energy consumption scheduling in a hydraulic press production line based on digital twins and reinforcement learning, applied to a production line consisting of multiple hydraulic presses, each hydraulic press corresponding to a programmable logic controller (PLC), characterized in that... The method includes the following steps: S1. Real-time estimation using digital twins: A digital twin model is established for each hydraulic press. Each digital twin model receives real-time operating data of the corresponding hydraulic press and outputs the estimated power and cumulative energy consumption of the hydraulic press. S2. Construction of reinforcement learning environment: The production line scheduling problem is modeled as a Markov decision process, which includes a state space, action space and reward function; The state space includes at least the main cylinder pressure, slider speed, hydraulic oil temperature, start / stop status, cumulative energy consumption, list of tasks to be processed, and current time-of-use electricity price for each hydraulic press. The motion space includes the target pressure value, target speed value, and start / stop commands of each hydraulic press; The reward function is a negative overall cost, which includes at least the energy consumption cost calculated based on the estimated power. S3. Training the scheduling policy network: Under the framework of the Markov decision process, the scheduling policy network is trained using a deep reinforcement learning algorithm; S4. Rolling Optimization and Command Issuance: Perform a scheduling operation at a fixed period, including obtaining the current status of each hydraulic press, inputting the trained scheduling strategy network, outputting the target pressure value, target speed value, and start / stop command of each hydraulic press, and issuing them to the corresponding programmable logic controller of the hydraulic press.

2. The method according to claim 1, characterized in that, In step S1, the digital twin model calculates and estimates the power based on the real-time operating data and the hydraulic system power formula; the hydraulic system power formula includes an efficiency coefficient. The efficiency coefficient is updated online using a data-driven correction method, which specifically includes: collecting the measured power of each hydraulic press, and correcting the efficiency coefficient online based on the difference between the estimated power and the measured power.

3. The method according to claim 2, characterized in that, The data-driven correction employs recursive least squares or Kalman filtering.

4. The method according to claim 1, characterized in that, In step S2, each task in the list of tasks to be processed includes a target pressure range, a target speed range, a deadline, an adjustable pressure window, and an adjustable speed window, all determined by the process. The adjustable pressure window is the upper and lower fluctuation range of the target pressure range, and the adjustable speed window is the upper and lower fluctuation range of the target speed range. The overall cost also includes at least one of equipment switchover penalties and order delay penalties.

5. The method according to claim 1, characterized in that, In step S3, the deep reinforcement learning algorithm is either a proximal policy optimization algorithm or a deep Q-network algorithm.

6. The method according to claim 4, characterized in that, The target pressure value and target speed value output by the scheduling strategy network are constrained within the adjustable pressure window and adjustable speed window of the corresponding task, respectively.

7. The method according to claim 1, characterized in that, The fixed period in step S4 is 1-10 minutes; Furthermore, when an urgent task is detected, a scheduling event is triggered immediately without waiting for the fixed period. The urgent task refers to a task marked by the user in the system, or a task whose deadline is less than or equal to the current time and has a fixed period.

8. The method according to claim 1, characterized in that, The state space is composed of the current state of each hydraulic press. The action space is used to limit the target pressure value, target speed value and start / stop command to the corresponding hydraulic press programmable logic controller within the same scheduling cycle. The energy cost in the reward function is the sum of the energy costs of the multiple hydraulic presses.

9. A scheduling system for implementing the energy consumption optimization scheduling method for hydraulic press production lines based on digital twins and reinforcement learning as described in any one of claims 1-8, characterized in that, include: Data acquisition module: It communicates with the pressure sensors, smart meters and programmable logic controllers of each hydraulic press to collect the main cylinder pressure, slider speed, hydraulic oil temperature, start and stop status and measured power of each hydraulic press; Digital twin module: Deployed on industrial edge computing equipment and connected to the data acquisition module, it is used to calculate the estimated power and cumulative energy consumption of each hydraulic press based on the collected real-time operating data of each hydraulic press, and to update the efficiency coefficient in the digital twin model online based on the difference between the measured power and the estimated power. Reinforcement learning scheduling module: Deployed on industrial edge computing devices and connected to the digital twin module, it is pre-deployed with a scheduling strategy network trained in step S3, which is used to output normalized action values ​​based on the current state through forward inference. The normalized action values ​​are inversely normalized and constrained by adjustable pressure window and adjustable speed window to generate scheduling instructions. The current status includes the main cylinder pressure, slide speed, hydraulic oil temperature, start / stop status, cumulative energy consumption, list of tasks to be processed, and current time-of-use electricity price for each hydraulic press; the scheduling instructions include target pressure value, target speed value, and start / stop instructions; Instruction issuing module: Communicatively connected to the reinforcement learning scheduling module and the programmable logic controller, used to write the scheduling instruction into the set value register of the programmable logic controller of the corresponding hydraulic press; Human-computer interaction module: connected to the reinforcement learning scheduling module, used to display real-time status and receive manual intervention instructions.

10. The scheduling system according to claim 9, characterized in that, The reinforcement learning scheduling module also includes an emergency task detection subunit, which is used to continuously monitor the list of tasks to be processed. When an emergency task is detected, the reinforcement learning scheduling module is immediately triggered to perform a scheduling operation without waiting for the fixed period. The urgent task refers to a task marked by the user through the human-computer interaction module, or a task whose deadline is less than or equal to the current time and has a fixed period.