Elevator energy recovery application method and system based on supercapacitor
By adopting a deep reinforcement learning model based on supercapacitors in the elevator system and combining the optimal storage allocation solution of virtual energy storage units, the problem of low energy recovery efficiency in the existing technology is solved, and more efficient energy utilization and supercapacitor life extension are achieved.
Patent Information
- Application Number
- CN202510167782.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing elevator energy recovery technology is inefficient, the storage and distribution strategies are incomplete, and the lack of intelligent control strategies leads to low energy recovery efficiency.
The elevator energy recovery application method based on supercapacitors is adopted, and the state space and action space are constructed by collecting elevator operation data and supercapacitor status data, and the energy recovery strategy is optimized using the deep reinforcement learning model, combining the virtual energy storage unit to calculate the optimal storage allocation plan, and a layered control strategy is implemented.
It improves the efficiency and utilization rate of elevator energy recovery, reduces the total electricity bill expenditure for elevator operation, and extends the service life of supercapacitors.
Smart Images

Figure CN120185218A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to energy recovery technology, and in particular to an elevator energy recovery application method and system based on supercapacitors. Background Art
[0002] Elevator energy recovery technology aims to collect and reuse the regenerative braking energy generated during the operation of elevators to improve energy efficiency and reduce operating costs. The existing technologies have the following defects and deficiencies:
[0003] Traditional energy recovery methods have low efficiency. The method of reverse power feeding to the power grid is greatly affected by power grid fluctuations and has high grid connection difficulty, while resistive braking directly converts the regenerative braking energy into heat dissipation, resulting in energy waste.
[0004] The energy storage and distribution strategies are not perfect. Most of the existing energy storage schemes adopt a single energy storage medium, lacking optimization strategies for energy storage and distribution among multiple elevators, and unable to maximize the energy utilization efficiency.
[0005] There is a lack of intelligent control strategies. Traditional control methods are difficult to adapt to the complexity and dynamics of elevator operating conditions and cannot be dynamically adjusted according to the real-time operating state of the elevator and the state of the energy storage system, resulting in low energy recovery efficiency. Summary of the Invention
[0006] The embodiments of the present invention provide an elevator energy recovery application method and system based on supercapacitors, which can solve the problems in the existing technologies.
[0007] In the first aspect of the embodiments of the present invention,
[0008] An elevator energy recovery application method based on supercapacitors is provided, including:
[0009] Collecting the operation data of an elevator group and the state data of supercapacitors, where the operation data includes the speed, acceleration and deceleration states, real-time load weight, floor position and regenerative braking energy data during the operation of the elevator, and the state data includes the state of charge, temperature distribution, internal resistance change and terminal voltage data of the supercapacitors;
[0010] Constructing a state space based on the operation data and state data, constructing an action space based on the charging power, discharging power of the supercapacitors and the energy distribution ratio among multiple elevators, constructing a reward function based on the economic benefits under peak-valley electricity prices, the loss of the cycle life of the supercapacitors and the utilization efficiency of the regenerative braking energy, and constructing a deep reinforcement learning model based on a double Q-network structure; training the deep reinforcement learning model through a prioritized experience replay mechanism to obtain an energy recovery optimization strategy;
[0011] When the elevator is in the braking power generation state, the regenerative braking energy is stored in the supercapacitor. When the elevator is in the driving power consumption state, the stored energy is released for elevator traction. Calculate the optimal storage allocation scheme among multiple elevators through a virtual energy storage unit. Based on the energy recovery optimization strategy and the optimal storage allocation scheme, implement a hierarchical control strategy, where the upper layer determines the energy storage priority based on fuzzy adaptive weights, and the lower layer uses a distributed algorithm to calculate the energy distribution coefficient and target power of each elevator, and updates the charge and discharge rate according to the health state of the supercapacitor.
[0012] In an alternative embodiment,
[0013] The steps of constructing a state space based on operation data and state data, constructing an action space based on the charging power, discharging power of the supercapacitor, and the energy distribution ratio among multiple elevators, and constructing a reward function based on the economic benefit under peak-valley electricity prices, the loss of the supercapacitor cycle life, and the utilization efficiency of regenerative braking energy include:
[0014] Based on operation data and state data, use a sliding time window to construct a motion domain state vector of energy prediction and elevator motion state, use an internal resistance-temperature coupling model to construct an energy storage domain state vector of the health factor and capacity attenuation rate of the supercapacitor, use wavelet transform to construct a power grid domain state vector of electricity price fluctuation characteristics and load fluctuation characteristics, and fuse the motion domain state vector, energy storage domain state vector, and power grid domain state vector through an attention fusion function to form a state space;
[0015] Construct a hierarchically encoded action space. Set discrete actions of charging, discharging, and standby modes at the macro level, construct power distribution actions using the Dirichlet distribution at the middle level, construct power adjustment actions according to the capacitor impedance characteristics at the micro level, convert the discrete actions at the macro level into an action vector through one-hot encoding, map the power distribution actions at the middle level into a distribution coefficient vector through a probability density function, map the power adjustment actions at the micro level into an adjustment factor through an impedance characteristic curve, and use tensor outer product operation to combine the action vector, distribution coefficient vector, and adjustment factor to form a unified action tensor;
[0016] Collect power data and electricity price coefficients during the operation of the elevator group, and calculate the economic benefit reward; according to the discharge depth and cycle times of the supercapacitor, calculate the life loss penalty in combination with the reference cycle times; calculate the regenerative braking energy utilization efficiency reward based on the energy conversion efficiency and the ratio of recovered energy to the maximum braking power; weight the economic benefit reward, life loss penalty, and regenerative braking energy utilization efficiency reward through a dynamic balance weight coefficient to obtain a reward function, and the dynamic balance weight coefficient determines the initial value based on the fuzzy analytic hierarchy process and is dynamically adjusted according to peak-valley periods, the health state of the supercapacitor, and the elevator load level.
[0017] In an alternative embodiment,
[0018] The steps of constructing a deep reinforcement learning model based on a double Q-network structure include:
[0019] The deep reinforcement learning model includes a main Q-network for action selection and an evaluation Q-network for action evaluation. Both the main Q-network and the evaluation Q-network adopt a branch-fusion architecture. In the state encoding branch, three layers of residual blocks are set to extract state features. Each layer of residual block contains two convolutional layers and a skip connection; in the action encoding branch, two layers of graph attention layers are set to extract elevator group topology features. Each attention layer contains a multi-head self-attention mechanism and edge feature embedding; a bidirectional gated recurrent unit is used to perform temporal fusion on the state features and the elevator group topology features, and the fused features are output as value estimates through two fully connected layers;
[0020] Construct a target network with the same branch-fusion architecture as the main Q-network and the evaluation Q-network, and guide the parameter update of the target network through a knowledge transfer loss function. The knowledge transfer loss function includes a feature layer, a decision layer, and a value layer; in the feature layer, a feature layer loss is constructed by calculating the L2 distance between the intermediate layer feature representations of the target network and the source network; in the decision layer, a decision layer loss is constructed using the KL divergence between the policy distributions output by the target network and the source network; in the value layer, a smooth L1 loss is used to calculate the Q-value estimation difference between the target network and the source network to construct a value layer loss; the feature layer loss, the decision layer loss, and the value layer loss are weighted and combined to obtain the knowledge transfer loss function.
[0021] In an alternative embodiment,
[0022] The steps of training the deep reinforcement learning model through a prioritized experience replay mechanism to obtain an energy recovery optimization strategy include:
[0023] Based on the state space and action space of the deep reinforcement learning model, transfer samples are obtained. The economic benefits under peak-valley electricity prices, the cycle life loss of supercapacitors, and the regenerative braking energy utilization efficiency in the reward function are constructed into a target weight vector. The target weight vector is calculated according to the state of charge of the supercapacitor, the electricity price coefficient, and the life loss rate. The temporal difference error of the transfer samples is calculated using the target weight vector, and the sample priority is obtained by adaptively adjusting the threshold of the temporal difference error in combination with the target weight expected value;
[0024] An elevator group topology graph is constructed based on the energy distribution ratio among multiple elevators, the local importance of each elevator node is calculated according to the energy transmission relationship, and the sampling probability is obtained by probabilistically fusing the local importance and the sample priority;
[0025] Select training samples from the experience replay pool according to the sampling probability, construct a target network and a training network based on the double Q-network structure, calculate the target Q value using the target network, weight the temporal difference error in combination with the sample priority, and update the training network parameters;
[0026] Calculate the cumulative value of the reward function within a set number of steps. When the cumulative value exceeds a preset threshold, copy the training network parameters to the target network to obtain an energy recovery optimization strategy.
[0027] In an alternative embodiment,
[0028] The steps of calculating the optimal storage allocation scheme among multiple elevators through a virtual energy storage unit include:
[0029] Take the state of charge, terminal voltage, and internal resistance of the supercapacitors corresponding to multiple elevators as state variables, construct a dynamic model of the virtual energy storage unit. The virtual energy storage unit equivalent the distributed energy storage system to a centralized system through an energy carrying weight coefficient, and the energy carrying weight coefficient is determined according to the rated capacity and real-time energy storage efficiency of the supercapacitor;
[0030] Construct an elevator group energy flow diagram based on the energy transfer relationship among multiple elevators, calculate the energy inflow power and energy outflow power of each elevator node, and determine the importance of the node energy flow according to the energy inflow power and energy outflow power;
[0031] Establish a mapping relationship between the virtual energy storage unit and the physical supercapacitor using the importance of the node energy flow. Construct an optimization objective function based on the deviation between the target power and the actual allocated power of the virtual energy storage unit and the state of charge deviation, and establish constraint conditions in combination with the state of charge constraint and power constraint of the supercapacitor;
[0032] Convert the optimization objective function and constraint conditions into a quadratic programming problem, solve the quadratic programming problem using the interior point method to obtain the power distribution coefficient, and obtain the optimal storage allocation scheme among multiple elevators based on the power distribution coefficient.
[0033] In an alternative embodiment,
[0034] Based on the energy recovery optimization strategy and the optimal storage allocation scheme, implement a hierarchical control strategy. The steps in which the upper layer determines the energy storage priority based on fuzzy adaptive weights, and the lower layer uses a distributed algorithm to calculate the energy distribution coefficient and target power of each elevator, and updates the charge and discharge rate according to the health state of the supercapacitor include:
[0035] Take the elevator operation status, peak-valley electricity price, and energy storage status as the input variables of the fuzzy controller. Construct an adaptive weight adjustment mechanism based on the regenerative braking energy efficiency, electricity price cost change, and life loss of the supercapacitor. Obtain the energy storage priority of each elevator through fuzzy rule reasoning;
[0036] Construct a distributed optimization problem based on the communication weight matrix and local control input. Calculate the target power adjustment coefficient in combination with the energy storage priority, and use the consensus theory to solve the energy distribution coefficient of each elevator;
[0037] Obtain the evaluation result of the health status according to the internal resistance and capacity of the supercapacitor. Dynamically adjust the charge and discharge rate constraint based on the evaluation result, and feedback the charge and discharge rate constraint to the distributed optimization problem;
[0038] Construct a prediction window based on the elevator scheduling information. Design a predictive energy buffer based on the predicted regenerative braking energy and historical actual recovered energy, and use the predictive energy buffer to dynamically adjust the energy storage power distribution;
[0039] Construct a group coupling compensation term based on the speed difference, position distribution, and load distribution of adjacent elevators. Combine the group coupling compensation term with the energy storage priority and the evaluation result to establish a multi-objective collaborative optimization function;
[0040] Update the predictive energy buffer and the group coupling compensation term within the upper-layer control period, optimize the weight coefficient of the multi-objective collaborative optimization function, and calculate the power correction amount of each elevator according to the multi-objective collaborative optimization function; Execute power distribution correction and coupling dynamic compensation according to the power correction amount within the lower-layer control period, and output the final control instruction.
[0041] In an alternative embodiment,
[0042] The steps of constructing a group coupling compensation term based on the speed difference, position distribution, and load distribution of adjacent elevators, and combining the group coupling compensation term with the energy storage priority and the evaluation result to establish a multi-objective collaborative optimization function include:
[0043] Construct a speed coupling function according to the speed difference and physical distance between adjacent elevators. The speed coupling function uses a distance attenuation term to correct the influence of the speed difference; Construct a position coupling function based on the spatial distribution of the elevator group; Construct a load coupling function according to the load distribution of each elevator. The load coupling function combines load prediction information to design a feedforward compensation term and dynamically adjusts the load correlation coefficient; Combine the speed coupling function, position coupling function, and load coupling function through an adaptive weight coefficient to form a group coupling compensation term, and the adaptive weight coefficient is obtained by online identification using the recursive least squares method;
[0044] Construct a hierarchical coupling integrator, the hierarchical coupling integrator comprising: a health status assessment unit for determining upper and lower threshold values of the energy storage capacity according to the assessment results; a priority interval division unit for dividing the energy storage working range into three energy storage intervals, namely high, medium, and low, according to the upper and lower threshold values of the energy storage capacity, and allocating an energy storage weight coefficient to each energy storage interval, the energy storage weight coefficient being dynamically adjusted according to the system operation state; a coupling mapping processing unit for dividing the population coupling compensation term into a strong coupling region, a weak coupling region, and a transition region according to the coupling strength, wherein the strong coupling region is processed by a linear mapping function, the transition region is processed by a piecewise linear mapping function, and the weak coupling region is processed by a non-linear mapping function to obtain a mapping compensation value; an adaptive fusion unit for determining the energy storage interval where the current working point is located according to the current working point, calculating an influence coefficient in combination with the assessment results, and performing weighted averaging on the influence coefficient and the mapping compensation value to obtain a fusion result; a feedback optimization unit for collecting system operation data, constructing an evaluation function according to the energy storage efficiency and the population synergy index of the supercapacitor, and online correcting the fusion parameters through the evaluation function, the fusion parameters including the energy storage weight coefficient and the influence coefficient;
[0045] Construct a multi-objective collaborative optimization function according to the fusion result of the hierarchical coupling integrator, and the multi-objective collaborative optimization function processes the non-linear coupling term by a piecewise linearization method.
[0046] In the second aspect of the embodiment of the present invention,
[0047] Provide an elevator energy recovery application system based on a supercapacitor, comprising:
[0048] A first unit for collecting the operation data of an elevator group and the state data of the supercapacitor, the operation data including the speed, acceleration and deceleration state, real-time load, floor position, and regenerative braking energy data during the operation of the elevator, and the state data including the state of charge, temperature distribution, internal resistance change, and terminal voltage data of the supercapacitor;
[0049] A second unit for constructing a state space based on the operation data and the state data, constructing an action space based on the charging power, discharging power of the supercapacitor, and the energy distribution ratio among multiple elevators, constructing a reward function based on the economic benefit under peak-valley electricity prices, the loss of the supercapacitor cycle life, and the regenerative braking energy utilization efficiency, and constructing a deep reinforcement learning model based on a double Q-network structure; training the deep reinforcement learning model through a prioritized experience replay mechanism to obtain an energy recovery optimization strategy;
[0050] The third unit is used to store the regenerative braking energy in the supercapacitor when the elevator is in the braking power generation state, and release the stored energy for elevator traction when the elevator is in the driving power consumption state; calculate the optimal storage allocation scheme among multiple elevators through the virtual energy storage unit; based on the energy recovery optimization strategy and the optimal storage allocation scheme, implement a hierarchical control strategy, where the upper layer determines the energy storage priority based on fuzzy adaptive weights, and the lower layer uses a distributed algorithm to calculate the energy distribution coefficient and target power of each elevator, and updates the charge and discharge rate according to the health state of the supercapacitor.
[0051] In the third aspect of the embodiments of the present invention,
[0052] A kind of electronic device is provided, including:
[0053] A processor;
[0054] A memory for storing instructions executable by the processor;
[0055] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0056] In the fourth aspect of the embodiments of the present invention,
[0057] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0058] The present invention optimizes the energy recovery strategy through a deep reinforcement learning model based on peak-valley electricity prices, and combines the charge and discharge characteristics of the supercapacitor to maximize the utilization rate of regenerative braking energy, thereby reducing the total electricity cost of elevator operation.
[0059] The present invention takes into account the health state of the supercapacitor, and dynamically adjusts the charge and discharge rate through fuzzy adaptive weights and a distributed algorithm to avoid damage to the supercapacitor caused by overcharging and over-discharging, and effectively extends its service life.
[0060] The present invention uses a deep reinforcement learning model and a virtual energy storage unit to realize the intelligent distribution and optimized storage of regenerative braking energy among multiple elevators, and improves the overall energy recovery efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is a schematic flowchart of the method for applying elevator energy recovery based on a supercapacitor in the embodiments of the present invention;
[0062] Figure 2 It is a schematic structural diagram of the system for applying elevator energy recovery based on a supercapacitor in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0064] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0065] Figure 1 The flowchart of the application method of elevator energy recovery based on a supercapacitor according to the embodiments of the present invention is shown as Figure 1 shown, and the method includes:
[0066] S1. Collect the operation data of the elevator group and the state data of the supercapacitor. The operation data includes the speed, acceleration and deceleration states, real-time load weight, floor position, and regenerative braking energy data during the elevator operation. The state data includes the state of charge, temperature distribution, internal resistance change, and terminal voltage data of the supercapacitor;
[0067] S2. Construct a state space based on the operation data and state data, construct an action space based on the charging power, discharging power of the supercapacitor, and the energy distribution ratio among multiple elevators, construct a reward function based on the economic benefits under peak-valley electricity prices, the loss of the supercapacitor cycle life, and the utilization efficiency of regenerative braking energy, and construct a deep reinforcement learning model based on the double Q-network structure; train the deep reinforcement learning model through a prioritized experience replay mechanism to obtain an energy recovery optimization strategy;
[0068] S3. When the elevator is in the braking power generation state, store the regenerative braking energy in the supercapacitor. When the elevator is in the driving power consumption state, release the stored energy for elevator traction; calculate the optimal storage allocation scheme among multiple elevators through a virtual energy storage unit; based on the energy recovery optimization strategy and the optimal storage allocation scheme, implement a hierarchical control strategy, where the upper layer determines the energy storage priority based on fuzzy adaptive weights, and the lower layer uses a distributed algorithm to calculate the energy distribution coefficient and target power of each elevator, and updates the charge and discharge rate according to the health state of the supercapacitor.
[0069] In an alternative embodiment,
[0070] Steps for constructing a state space based on operation data and status data, constructing an action space based on the charging power, discharging power of a supercapacitor, and the energy distribution ratio among multiple elevators, and constructing a reward function based on the economic benefits under peak-valley electricity prices, the loss of the supercapacitor cycle life, and the energy regeneration braking utilization efficiency include:
[0071] Based on operation data and status data, use a sliding time window to construct a motion domain state vector of energy prediction and elevator motion state, use an internal resistance-temperature coupling model to construct a storage domain state vector of the health factor and capacity attenuation rate of the supercapacitor, use wavelet transform to construct a power grid domain state vector of electricity price fluctuation characteristics and load fluctuation characteristics, and fuse the motion domain state vector, storage domain state vector, and power grid domain state vector through an attention fusion function to form a state space;
[0072] Construct a hierarchically encoded action space. Set discrete actions of charging, discharging, and standby modes at the macro level, construct power distribution actions using the Dirichlet distribution at the middle level, construct power adjustment actions according to the capacitor impedance characteristics at the micro level, convert the discrete actions at the macro level into an action vector through one-hot encoding, map the power distribution actions at the middle level into a distribution coefficient vector through a probability density function, map the power adjustment actions at the micro level into adjustment factors through an impedance characteristic curve, and use tensor outer product operation to combine the action vector, distribution coefficient vector, and adjustment factors to form a unified action tensor;
[0073] Collect power data and electricity price coefficients during the operation of the elevator group, and calculate the economic benefit reward; according to the discharge depth and cycle times of the supercapacitor, calculate the life loss penalty in combination with the reference cycle times; calculate the energy regeneration braking utilization efficiency reward based on the energy conversion efficiency and the ratio of the recovered energy to the maximum braking power; weight the economic benefit reward, life loss penalty, and energy regeneration braking utilization efficiency reward through a dynamic balance weight coefficient to obtain a reward function, and the dynamic balance weight coefficient determines the initial value based on the fuzzy analytic hierarchy process and is dynamically adjusted according to peak-valley periods, the health state of the supercapacitor, and the elevator load level.
[0074] Exemplarily, first, construct a state space. The state space is a complete description of the current state of the system, and it includes three dimensions: the motion domain, the storage domain, and the power grid domain.
[0075] In the motion domain, use a sliding time window to collect elevator operation data, such as speed, acceleration, load, and floor position, etc., and the operation state of the elevator, such as upward, downward, stop, and standby, etc., to construct a motion domain state vector of energy prediction and elevator motion state. For example, select a 10-second sliding time window and collect data once per second, and a sequence containing data such as speed, acceleration, load, and floor position at 10 time points, as well as the corresponding operation state sequence, can be obtained.
[0076] In the energy storage domain, using the internal resistance-temperature coupling model of the supercapacitor, construct the state vector of the health factor and capacity attenuation rate of the supercapacitor in the energy storage domain. The internal resistance-temperature coupling model can calculate its health factor and capacity attenuation rate based on the real-time internal resistance and temperature data of the supercapacitor. For example, a mapping relationship between the internal resistance-temperature and the health factor and capacity attenuation rate can be established according to historical data.
[0077] In the power grid domain, use wavelet transform to extract the electricity price fluctuation characteristics and load fluctuation characteristics, and construct the state vector of the power grid domain. Wavelet transform can decompose the electricity price and load signals into components of different frequencies and extract their fluctuation characteristics. For example, the high-frequency components can be extracted to reflect short-term fluctuations, and the low-frequency components can be extracted to reflect long-term trends.
[0078] Finally, through the attention fusion function, fuse the state vectors of the motion domain, energy storage domain, and power grid domain to form a state space. The attention fusion function can dynamically allocate weights according to the importance of different domains, fuse the state vectors of the three domains into a unified state vector, and completely describe the current state of the system. For example, during the peak electricity consumption period, a higher weight can be assigned to the power grid domain.
[0079] Secondly, construct the action space. The action space is the set of all actions that the system can take, and it also contains three levels: the macro level, the meso level, and the micro level.
[0080] At the macro level, set three discrete action modes: charging, discharging, and standby. These three modes correspond to the charging, discharging, and non-participation in energy exchange states of the supercapacitor respectively.
[0081] At the meso level, use the Dirichlet distribution to construct the power distribution action and determine the energy distribution ratio among multiple elevators. The Dirichlet distribution can generate a set of positive numbers whose sum is 1, representing the power distribution coefficients of each elevator. For example, for three elevators, the Dirichlet distribution can generate the distribution coefficients of (0.4, 0.3, 0.3), indicating that 40%, 30%, and 30% of the power is allocated respectively.
[0082] At the micro level, construct the power regulation action according to the impedance characteristics of the supercapacitor to finely regulate the power allocated to each elevator. For example, according to the impedance characteristic curve of the supercapacitor, multiply the power allocated to each elevator by a regulation factor to achieve fine power regulation.
[0083] Finally, use the tensor outer product operation to combine the macro-level actions, meso-level power distribution coefficients, and micro-level power regulation factors to form a unified action tensor, representing the specific actions taken by the system.
[0084] Secondly, construct the reward function. The reward function is used to evaluate the benefits after the system takes actions, and it consists of three parts: economic benefit reward, life loss penalty, and regenerative braking energy utilization efficiency reward.
[0085] The economic benefit reward is calculated based on the power data of the elevator group and the electricity price coefficient. For example, the electricity consumption cost of the elevator group can be calculated at the current electricity price.
[0086] The life loss penalty is calculated based on the discharge depth and cycle times of the supercapacitor, combined with the reference cycle times. The deeper the discharge depth and the more the cycle times, the greater the life loss.
[0087] The regenerative braking energy utilization efficiency reward is calculated based on the energy conversion efficiency and the ratio of the recovered energy to the maximum braking power. The more the recovered energy and the higher the energy conversion efficiency, the greater the regenerative braking energy utilization efficiency reward.
[0088] Finally, the economic benefit reward, life loss penalty, and regenerative braking energy utilization efficiency reward are weighted and summed through the dynamic balance weight coefficient to obtain the final reward function. The initial value of the dynamic balance weight coefficient is determined based on the fuzzy analytic hierarchy process and is dynamically adjusted according to the peak-valley period, the health state of the supercapacitor, and the elevator load level. For example, during the peak period, the weight of the economic benefit reward can be increased; when the health state of the supercapacitor is poor, the weight of the life loss penalty can be increased.
[0089] Through intelligent energy management, the present invention can optimize the electricity consumption strategy of the elevator group under peak-valley electricity prices, reduce the electricity consumption cost, and improve the economic benefits; by controlling the discharge depth and cycle times of the supercapacitor, it can reduce its life loss, extend the service life, and reduce the maintenance cost; by effectively utilizing the regenerative braking energy, it can improve the energy utilization rate, reduce energy waste, and promote energy conservation and emission reduction.
[0090] In an optional implementation manner,
[0091] The steps of constructing a deep reinforcement learning model based on a double Q-network structure include:
[0092] The deep reinforcement learning model includes a main Q-network for action selection and an evaluation Q-network for action evaluation. Both the main Q-network and the evaluation Q-network adopt a branch-fusion architecture. In the state encoding branch, three layers of residual blocks are set to extract state features, and each layer of residual block contains two convolutional layers and a skip connection; in the action encoding branch, two layers of graph attention layers are set to extract the elevator group topology features, and each attention layer contains a multi-head self-attention mechanism and edge feature embedding; a bidirectional gated recurrent unit is used to perform temporal fusion on the state features and the elevator group topology features, and the fused features are output through two fully connected layers for value estimation.
[0093] Construct a target network with the same branch-fusion architecture as the main Q-network and the evaluation Q-network, and guide the parameter update of the target network through a knowledge transfer loss function, where the knowledge transfer loss function includes a feature layer, a decision layer, and a value layer; in the feature layer, construct a feature layer loss by calculating the L2 distance between the intermediate layer feature representations of the target network and the source network; in the decision layer, construct a decision layer loss using the KL divergence between the policy distributions output by the target network and the source network; in the value layer, use a smooth L1 loss to calculate the Q-value estimation difference between the target network and the source network to construct a value layer loss; weight and combine the feature layer loss, the decision layer loss, and the value layer loss to obtain the knowledge transfer loss function.
[0094] Exemplarily, first, construct the main Q-network and the evaluation Q-network. These two networks share the same structure, i.e., the branch-fusion architecture. In the state encoding branch, use three-layer residual blocks to extract state features. Each layer of the residual block contains two convolutional layers and a skip connection. For example, if the input state is a vector with a dimension of 20, after the first convolutional layer (with a kernel size of 3, a stride of 1, and a padding of 1), the output dimension remains 20. After the second convolutional layer (with a kernel size of 3, a stride of 1, and a padding of 1), the dimension is still 20. Finally, add the input of the first convolutional layer to the output of the second convolutional layer to complete the skip connection. And so on, complete the construction of the three-layer residual block. In the action encoding branch, set two-layer graph attention layers to extract the elevator group topology features. Each layer of the attention layer contains a multi-head self-attention mechanism and edge feature embedding. Suppose there are 5 elevators in the elevator group, then a 5x5 adjacency matrix can be constructed to represent the connection relationship between the elevators. The edge features can represent information such as the distance or communication delay between the elevators. Through the graph attention layer, the importance of the relationships between different elevators can be learned. Next, use a bidirectional gated recurrent unit to perform temporal fusion on the state features and the elevator group topology features. For example, input a sequence of vectors with a dimension of 20 for the state features and a sequence of vectors with a dimension of 5 for the topology features into the bidirectional gated recurrent unit, and output a sequence of vectors with a dimension of 30 for the fused features. Finally, the fused features pass through two fully connected layers to output a value estimate. For example, input a vector with a dimension of 30 into the first fully connected layer (with an output dimension of 16), and then into the second fully connected layer (with an output dimension of 1), and finally output a scalar value as the value estimate.
[0095] Secondly, construct a target network with a branch-fusion architecture that has the same structure as the main Q-network and the evaluation Q-network. The parameter update of the target network is guided by the knowledge transfer loss function. The knowledge transfer loss function includes a feature layer loss, a decision layer loss, and a value layer loss. The feature layer loss is constructed by calculating the distance between the intermediate layer feature representations of the target network and the source network (main Q-network or evaluation Q-network), such as calculating the mean squared error of the output features of the third residual block of the two networks. The decision layer loss is constructed using the difference between the policy distributions output by the target network and the source network, such as calculating the cross-entropy of the output action probability distributions of the two networks. The value layer loss is constructed using the difference in Q-value estimates between the target network and the source network, such as calculating the absolute error of the output Q-values of the two networks. The feature layer loss, decision layer loss, and value layer loss are weighted and combined to obtain the knowledge transfer loss function. For example, weights of 0.5, 0.25, and 0.25 are assigned to the three loss functions respectively, and they are added to obtain the final knowledge transfer loss function.
[0096] Finally, update the parameters of the target network by minimizing the knowledge transfer loss function, and periodically copy the parameters of the target network to the main Q-network and the evaluation Q-network, thereby achieving knowledge transfer and model training.
[0097] Through knowledge transfer, the target network can effectively guide the learning of the main Q-network and the evaluation Q-network, accelerate the convergence speed of the model, and reduce the training time; knowledge transfer can help the model learn more robust feature representations, thereby improving the adaptability and generalization ability of the model in different scenarios; through the branch-fusion architecture and the graph attention mechanism, the model can better capture the state information and the elevator group topology features, thereby making better decisions and improving the overall performance of the elevator group control system.
[0098] In an alternative embodiment,
[0099] The steps of training the deep reinforcement learning model through the prioritized experience replay mechanism to obtain the energy recovery optimization strategy include:
[0100] Obtain transfer samples based on the state space and action space of the deep reinforcement learning model, construct the economic benefits under peak-valley electricity prices, the cycle life loss of the supercapacitor, and the regenerative braking energy utilization efficiency in the reward function as the target weight vector, calculate the target weight vector according to the state of charge of the supercapacitor, the electricity price coefficient, and the life loss rate, calculate the temporal difference error of the transfer samples using the target weight vector, and perform adaptive threshold adjustment on the temporal difference error in combination with the target weight expectation value to obtain the sample priority;
[0101] Construct an elevator group topology map based on the energy distribution ratio among multiple elevators, calculate the local importance of each elevator node according to the energy transmission relationship, and perform probability fusion on the local importance and the sample priority to obtain the sampling probability;
[0102] Select training samples from the experience replay pool according to the sampling probability, construct a target network and a training network based on the double Q-network structure, calculate the target Q value using the target network, weight the temporal difference error in combination with the sample priority, and update the training network parameters;
[0103] Calculate the cumulative value of the reward function within a set number of steps. When the cumulative value exceeds the preset threshold, copy the training network parameters to the target network to obtain the energy recovery optimization strategy.
[0104] Exemplarily, first, prepare the training environment for the deep reinforcement learning model. This environment includes multiple elevators that can transfer energy among each other, forming an elevator group. It is necessary to clarify the state space of each elevator, such as the running speed, acceleration, load, and position of the elevator; and the action space, such as the control instructions for the elevator, including acceleration, deceleration, and stop. At the same time, it is necessary to define the reward function for evaluating the quality of the strategy. The reward function includes three aspects: economic benefits under peak-valley electricity prices, loss of the cycle life of the supercapacitor, and the utilization efficiency of regenerative braking energy.
[0105] Next, obtain the transition samples. Transition samples refer to the process records of the environment transitioning to the next state and obtaining the corresponding reward after performing a certain action in the current state. Specifically, control the elevator to run according to the set strategy, and record the state of the elevator, the actions performed, and the rewards obtained at different time points. For example, record that the speed of the elevator is 1 m / s at 10:00. After performing the acceleration action, the speed becomes 2 m / s at 10:01, and the obtained reward is 0.5. Store these records as transition samples in the experience replay pool.
[0106] Then, calculate the priority of the transition samples. First, calculate the target weight vector according to the state of charge of the supercapacitor, the electricity price coefficient, and the life loss rate. For example, the current time is 14:00, which belongs to the peak electricity price period, and the electricity price coefficient is 1.2; the state of charge of the supercapacitor is 80%, and the life loss rate is 0.01. According to the pre-set rules, calculate the target weight vector as [1.0, 0.2, 0.8], corresponding to the weights of economic benefits, life loss, and energy utilization efficiency respectively. Then, calculate the temporal difference error of the transition samples using the target weight vector. The temporal difference error reflects the gap between the current strategy and the optimal strategy. Finally, perform adaptive threshold adjustment on the temporal difference error in combination with the target weight expectation value to obtain the sample priority.
[0107] Subsequently, the sampling probability is calculated. An elevator group topology graph is constructed based on the energy distribution ratio in the elevator hall. For example, if the energy distribution ratio among three elevators is 1:2:1, a star topology graph can be constructed, with the middle node connecting the other two nodes, and the weights of the edges being 1, 2, and 1 respectively. The local importance of each elevator node is calculated according to the energy transfer relationship. For example, the local importance of the middle node is 1 + 2 + 1 = 4, and the local importance of the other two nodes is 1 and 1 respectively. The local importance is probabilistically fused with the sample priority to obtain the sampling probability.
[0108] After that, training samples are selected from the experience replay pool according to the sampling probability. For example, a random number between 0 and 1 is generated. If the random number is less than 0.24, the sample is selected. The target Q-value is calculated using the target network. The target Q-value represents the maximum cumulative reward that may be obtained in the future after performing a certain action in the current state. The temporal difference error is weighted in combination with the sample priority, and the training network parameters are updated. For example, the higher the sample priority, the greater the weight of the temporal difference error, and the greater the update amplitude of the training network parameters.
[0109] Finally, the cumulative value of the reward function is calculated within the set number of steps. For example, the set number of steps is 1000, and the cumulative value of the reward function within 1000 steps is calculated. When the cumulative value exceeds the preset threshold, for example, the preset threshold is 100 and the current cumulative value is 120, the training network parameters are copied to the target network to obtain the energy recovery optimization strategy.
[0110] Through the training of the deep reinforcement learning model, the present invention can obtain a better elevator energy recovery strategy, thereby improving the utilization efficiency of regenerative braking energy and reducing energy waste; by using the economic benefit under the peak-valley electricity price as part of the reward function, the model can be guided to learn to charge during low electricity price periods and discharge during high electricity price periods, thereby reducing the operating cost of the elevator.
[0111] In an alternative embodiment,
[0112] The steps of calculating the optimal storage allocation scheme among multiple elevators through the virtual energy storage unit include:
[0113] Taking the state of charge, terminal voltage, and internal resistance of the supercapacitors corresponding to multiple elevators as state variables, a dynamic model of the virtual energy storage unit is constructed. The virtual energy storage unit equivalent the distributed energy storage system to a centralized system through the energy-bearing weight coefficient, and the energy-bearing weight coefficient is determined according to the rated capacity and real-time energy storage efficiency of the supercapacitor;
[0114] Based on the energy transfer relationship among multiple elevators, an elevator group energy flow graph is constructed, the energy inflow power and energy outflow power of each elevator node are calculated, and the node energy flow importance is determined according to the energy inflow power and energy outflow power;
[0115] A mapping relationship between the virtual energy storage unit and the physical supercapacitor is established by using the importance of the node energy flow. An optimization objective function is constructed based on the deviation between the target power and the actual allocated power of the virtual energy storage unit and the state of charge deviation, and constraint conditions are established by combining the state of charge constraint and the power constraint of the supercapacitor.
[0116] The optimization objective function and the constraint conditions are transformed into a quadratic programming problem, and the interior point method is used to solve the quadratic programming problem to obtain the power distribution coefficient. Based on the power distribution coefficient, an optimal storage allocation scheme among multiple elevators is obtained.
[0117] Exemplarily, first, a dynamic model of the virtual energy storage unit is established. The core of this model is to equivalent the distributed energy storage system to a centralized system. Specifically, each elevator is equipped with a supercapacitor, and the state of charge, terminal voltage, and internal resistance of these supercapacitors are used as state variables. To integrate these dispersed energy storage units, a concept is introduced: the energy carrying weight coefficient. This coefficient is jointly determined by the rated capacity of the supercapacitor and the real-time energy storage efficiency. For example, if a certain supercapacitor has a large rated capacity and a high current energy storage efficiency, then its energy carrying weight coefficient is larger, which means it occupies a larger proportion in the virtual energy storage unit. Through these state variables and weight coefficients, we can construct a dynamic model that describes the overall operating state of the virtual energy storage unit.
[0118] Next, an energy flow diagram of the elevator group is constructed. We need to analyze the energy transfer relationship among multiple elevators. For example, which elevators are rising and consuming energy, and which elevators are descending and regenerating energy. Then, the energy inflow power and the energy outflow power of each elevator node are calculated. For example, if an elevator is descending, the energy generated by its regenerative braking is used as the energy inflow, and if the elevator is rising, the energy consumed by its motor is used as the energy outflow. According to the magnitudes of the energy inflow power and the energy outflow power, the importance of the energy flow of each node can be determined. The node with a large inflow power and a small outflow power has a higher importance, and vice versa.
[0119] Then, establish the mapping relationship between the virtual energy storage unit and the physical supercapacitor. Using the node energy flow importance calculated previously, it is possible to determine how the energy in the virtual energy storage unit is distributed to each physical supercapacitor. The higher the importance of the node, the more energy is allocated. The goal of the distribution is to minimize the deviation between the target power of the virtual energy storage unit and the actual allocated power, as well as the state of charge deviation. These deviations are constructed into an optimization objective function. At the same time, considering the state of charge and power limits of each supercapacitor, we need to set corresponding constraints. For example, the state of charge of the supercapacitor cannot exceed its maximum value or be lower than its minimum value; the charge and discharge power of the supercapacitor also cannot exceed its rated power.
[0120] Finally, solve the optimal storage allocation scheme. Transform the optimization objective function and constraints into a quadratic programming problem. Here, instead of using mathematical formulas, we understand it as a problem of finding the best solution under specific constraints. Solve this quadratic programming problem using the interior point method to obtain a set of power distribution coefficients. For example, a coefficient of 0.8 means that this supercapacitor will bear 80% of the power demand of the virtual energy storage unit. Based on this set of power distribution coefficients, we can obtain the optimal storage allocation scheme for multiple elevator shafts.
[0121] Through optimizing the energy storage allocation scheme, the present invention can maximize the utilization of the energy generated by elevator regenerative braking, reduce energy waste, and thus improve the overall energy utilization efficiency; through reasonable energy distribution, it can ensure the power demand of the elevator and improve the operation efficiency and stability of the elevator.
[0122] In an alternative embodiment,
[0123] Based on the energy recovery optimization strategy and the optimal storage allocation scheme, implement a hierarchical control strategy, where the upper layer determines the energy storage priority based on fuzzy adaptive weights, and the lower layer uses a distributed algorithm to calculate the energy distribution coefficients and target powers of each elevator, and updates the charge and discharge rate according to the health state of the supercapacitor. The steps include:
[0124] Take the elevator operation state, peak-valley electricity price, and energy storage state as the input variables of the fuzzy controller, construct an adaptive weight adjustment mechanism based on the regenerative braking energy efficiency, electricity price cost change, and life loss of the supercapacitor, and obtain the energy storage priority of each elevator through fuzzy rule reasoning;
[0125] Construct a distributed optimization problem based on the communication weight matrix and local control inputs, calculate the target power adjustment coefficient in combination with the energy storage priority, and use the consensus theory to solve the energy distribution coefficients of each elevator;
[0126] Obtain the health state evaluation result according to the internal resistance and capacity of the supercapacitor, dynamically adjust the charge and discharge rate constraint based on the evaluation result, and feedback the charge and discharge rate constraint to the distributed optimization problem;
[0127] Construct a prediction window according to the elevator dispatching information, design a predicted energy buffer based on the predicted regenerative braking energy and the historical actual recovered energy, and dynamically adjust the energy storage power distribution by using the predicted energy buffer;
[0128] Construct a population coupling compensation term based on the speed difference, position distribution and load distribution of adjacent elevators, combine the population coupling compensation term with the energy storage priority and the evaluation result, and establish a multi-objective collaborative optimization function;
[0129] Update the predicted energy buffer and the population coupling compensation term within the upper layer control period, optimize the weight coefficients of the multi-objective collaborative optimization function, and calculate the power correction amount of each elevator according to the multi-objective collaborative optimization function; execute power distribution correction and coupling dynamic compensation according to the power correction amount within the lower layer control period, and output the final control instruction.
[0130] Exemplarily, first, obtain the operation state information of each elevator, including speed, acceleration, load, and position, etc. At the same time, obtain the real-time peak-valley electricity price information and the state information of the energy storage system, such as the current voltage and remaining capacity of the supercapacitor, etc. These information will be used as the input of the fuzzy controller.
[0131] Then, construct a fuzzy controller. The input variables of this controller include elevator operation state, peak-valley electricity price, and energy storage state. The output variable is the energy storage priority of each elevator. The rule base of the fuzzy controller is preset according to factors such as regenerative braking energy efficiency, electricity price cost change, and life loss of the supercapacitor. For example, when the elevator is in the high-speed descending braking state, the electricity price is at the peak period and the remaining capacity of the supercapacitor is low, then the energy storage priority of this elevator is set to high. Through fuzzy inference, the energy storage priority of each elevator can be obtained.
[0132] Next, construct a distributed optimization problem. The goal of this optimization problem is to maximize the energy recovery efficiency of the entire elevator group. The constraint conditions include the power limit of each elevator, the charge and discharge rate limit of the supercapacitor, and energy conservation, etc. To solve this optimization problem, the consensus theory is adopted. Specifically, each elevator iteratively updates its own energy distribution coefficient and target power according to the state information of its neighboring elevators and its own state information.
[0133] During the optimization process, the health state of the supercapacitor needs to be considered. By measuring the internal resistance and capacitance of the supercapacitor, its health state can be evaluated. According to the evaluation result of the health state, the charge-discharge rate constraint is dynamically adjusted. For example, if the health state of the supercapacitor is poor, its charge-discharge rate constraint is reduced to extend its service life. The updated charge-discharge rate constraint will be fed back into the distributed optimization problem.
[0134] To improve the prediction accuracy of energy recovery, a prediction window is constructed. Based on the predicted regenerative braking energy and the historical actual recovered energy, a predictive performance energy buffer is designed. This buffer can dynamically adjust the energy storage power distribution to cope with prediction errors and real-time fluctuations.
[0135] To further optimize the control performance, a group coupling compensation term is constructed. This compensation term takes into account factors such as the speed difference, position distribution, and load distribution of adjacent elevators. The group coupling compensation term is combined with the energy storage priority and the evaluation result of the health state of the supercapacitor to establish a multi-objective collaborative optimization function.
[0136] During the upper-layer control period, the predictive performance energy buffer and the group coupling compensation term are updated, and the weight coefficients of the multi-objective collaborative optimization function are optimized. According to the optimized multi-objective collaborative optimization function, the power correction amount of each elevator is calculated. During the lower-layer control period, power distribution correction and coupling dynamic compensation are performed according to the power correction amount, and finally control instructions are output.
[0137] For example, assume an elevator group contains three elevators. Elevator 1 is in the high-speed descending braking state, elevator 2 is in the low-speed ascending state, and elevator 3 is in the stationary state. The current electricity price is at a peak period, and the remaining capacity of the supercapacitor is low. According to the fuzzy control rules, the energy storage priority of elevator 1 is set to high, and the energy storage priorities of elevator 2 and elevator 3 are set to low. Through the distributed optimization algorithm, the energy distribution coefficient and target power of each elevator are calculated. Assume that the health state of the supercapacitor of elevator 1 is poor, then its charge-discharge rate constraint is reduced. Finally, according to the calculated power correction amount, the power distribution of each elevator is corrected, and the final control instructions are output.
[0138] The present invention preferentially stores regenerative braking energy through fuzzy adaptive weights and distributed optimization algorithms, dynamically adjusts the charge-discharge strategy according to the energy storage state and electricity price, maximizes the energy recovery efficiency, dynamically adjusts the charge-discharge rate according to the health state of the supercapacitor, avoids overcharging and over-discharging, effectively extends the service life of the energy storage device, and the predictive performance energy buffer and group coupling compensation mechanism can effectively cope with prediction errors and real-time fluctuations, improving the stability and robustness of the elevator group control system.
[0139] In an alternative embodiment,
[0140] Steps for constructing a population coupling compensation term based on the speed difference, position distribution, and load distribution of adjacent elevators, and combining the population coupling compensation term with the energy storage priority and the evaluation result to establish a multi-objective collaborative optimization function include:
[0141] Construct a speed coupling function according to the speed difference and physical distance between adjacent elevators, and the speed coupling function corrects the influence of the speed difference using a distance attenuation term; construct a position coupling function based on the spatial distribution of the elevator group; construct a load coupling function according to the load distribution of each elevator, and the load coupling function designs a feedforward compensation term in combination with load prediction information to dynamically adjust the load correlation coefficient; form a population coupling compensation term by weighted combination of the speed coupling function, position coupling function, and load coupling function through an adaptive weight coefficient, and the adaptive weight coefficient is obtained by online identification using the recursive least squares method;
[0142] Construct a hierarchical coupling fusion device, and the hierarchical coupling fusion device includes: a health status evaluation unit that determines the upper and lower limit thresholds of the energy storage capacity according to the evaluation result; a priority interval division unit that divides the energy storage working range into three energy storage intervals, namely high, medium, and low, according to the upper and lower limit thresholds of the energy storage capacity, and assigns an energy storage weight coefficient to each energy storage interval, and the energy storage weight coefficient is dynamically adjusted according to the system operation state; a coupling mapping processing unit that divides the population coupling compensation term into a strong coupling area, a weak coupling area, and a transition area according to the coupling strength, where the strong coupling area is processed using a linear mapping function, the transition area is processed using a piecewise linear mapping function, and the weak coupling area is processed using a non-linear mapping function to obtain a mapping compensation value; an adaptive fusion unit that determines the energy storage interval where the current working point is located according to the current working point, calculates an influence coefficient in combination with the evaluation result, and performs a weighted average of the influence coefficient and the mapping compensation value to obtain a fusion result; a feedback optimization unit that collects system operation data, constructs an evaluation function according to the energy storage efficiency of the supercapacitor and the population synergy index, and online corrects the fusion parameters through the evaluation function, and the fusion parameters include the energy storage weight coefficient and the influence coefficient;
[0143] Construct a multi-objective collaborative optimization function according to the fusion result of the hierarchical coupling fusion device, and the multi-objective collaborative optimization function processes non-linear coupling terms using a piecewise linearization method.
[0144] Exemplarily, first, construct a population coupling compensation term. This step is further divided into the construction of a speed coupling function, the construction of a position coupling function, the construction of a load coupling function, and the formation of a population coupling compensation term.
[0145] When constructing the speed coupling function, first calculate the speed difference between adjacent elevators. Then, considering the attenuation effect of distance on the speed difference, introduce a distance attenuation term. The farther the distance, the smaller the impact of the speed difference. This attenuation term can select an appropriate functional form according to the actual situation, such as exponential attenuation or inverse proportional attenuation.
[0146] The construction of the position coupling function is based on the spatial distribution of the elevator group. The floors where the elevators are located can be used as position information, and the floor difference between the elevators is calculated. Suppose two elevators are located on the 10th floor and the 20th floor respectively, then the position coupling value is |10 - 20| = 10.
[0147] The construction of the load coupling function needs to consider the load distribution of each elevator. To more accurately reflect the load situation, a feed-forward compensation term can be designed by combining load prediction information. At the same time, introduce a load correlation coefficient and dynamically adjust this coefficient according to the system operation state. Suppose the loads of two elevators are 500 kg and 300 kg respectively, the correlation coefficient is 0.8, and the feed-forward compensation term is 100, then the load coupling value is 0.8×|500 - 300| + 100 = 260.
[0148] Finally, the speed coupling function, position coupling function, and load coupling function are weighted and combined through an adaptive weight coefficient to form a group coupling compensation term. The adaptive weight coefficient is obtained by online identification using the recursive least squares method.
[0149] Next, construct a hierarchical coupling fusion device. The fusion device consists of a health status assessment unit, a priority interval division unit, a coupling mapping processing unit, an adaptive fusion unit, and a feedback optimization unit.
[0150] The health status assessment unit determines the upper and lower limit thresholds of the energy storage capacity according to the assessment results. Suppose the assessment result is good, then the upper limit of the energy storage capacity is 1000, and the lower limit is 100.
[0151] The priority interval division unit divides the energy storage working range into three energy storage intervals: high, medium, and low according to the upper and lower limit thresholds of the energy storage capacity, and assigns an energy storage weight coefficient to each interval. The energy storage weight coefficient is dynamically adjusted according to the system operation state. Suppose the current energy storage capacity is 800, then it is in the high energy storage interval, and the weight coefficient is 0.8.
[0152] The coupling mapping processing unit divides the group coupling compensation term into a strong coupling area, a weak coupling area, and a transition area according to the coupling strength, and processes them respectively using a linear mapping function, a non-linear mapping function, and a piecewise linear mapping function to obtain a mapping compensation value. Suppose the group coupling compensation term is 133.48 and it is in the transition area, then the mapping compensation value obtained by processing with the piecewise linear mapping function is 100.
[0153] The adaptive fusion unit determines its energy storage interval according to the current working point, calculates the influence coefficient in combination with the evaluation result, and performs weighted averaging on the influence coefficient and the mapping compensation value to obtain the fusion result.
[0154] The feedback optimization unit collects the system operation data, constructs an evaluation function based on the energy storage efficiency and group cooperation index of the supercapacitor, and online corrects the fusion parameters (energy storage weight coefficient and influence coefficient) through the evaluation function.
[0155] Finally, a multi-objective collaborative optimization function is constructed according to the fusion result of the hierarchical coupling fusion device. This function uses the piecewise linearization method to process the non-linear coupling term.
[0156] By combining the energy storage priority and the evaluation result, the present invention establishes a multi-objective collaborative optimization function, realizes the optimized utilization of energy, and reduces the overall energy consumption of the system; through mechanisms such as the hierarchical coupling fusion device and the adaptive weight coefficient, the robustness and adaptability of the system are improved, enabling it to better cope with various complex working conditions and ensuring the stable operation of the system.
[0157] Figure 2 FIG. is a schematic structural diagram of an elevator energy recovery application system based on a supercapacitor according to an embodiment of the present invention, as Figure 2 shown, the system includes:
[0158] The first unit is used to collect the operation data of the elevator group and the state data of the supercapacitor. The operation data includes the speed, acceleration and deceleration state, real-time load, floor position and regenerative braking energy data during the elevator operation. The state data includes the state of charge, temperature distribution, internal resistance change and terminal voltage data of the supercapacitor.
[0159] The second unit is used to construct a state space based on the operation data and the state data, construct an action space based on the charging power, discharging power of the supercapacitor and the energy distribution ratio among multiple elevators, construct a reward function based on the economic benefit under the peak-valley electricity price, the loss of the supercapacitor cycle life and the utilization efficiency of the regenerative braking energy, and construct a deep reinforcement learning model based on the double Q network structure; train the deep reinforcement learning model through the prioritized experience replay mechanism to obtain an energy recovery optimization strategy.
[0160] A third unit is configured to store the regenerative braking energy in a supercapacitor when the elevator is in a braking power generation state, and release the stored energy for elevator traction when the elevator is in a driving power consumption state; calculate an optimal storage allocation scheme among multiple elevators through a virtual energy storage unit; implement a hierarchical control strategy based on the energy recovery optimization strategy and the optimal storage allocation scheme, where the upper layer determines the energy storage priority based on fuzzy adaptive weights, and the lower layer uses a distributed algorithm to calculate the energy distribution coefficient and target power of each elevator, and updates the charge and discharge rate according to the health state of the supercapacitor.
[0161] In a third aspect of the embodiments of the present invention,
[0162] a kind of electronic device is provided, including:
[0163] a processor;
[0164] a memory for storing instructions executable by the processor;
[0165] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0166] In a fourth aspect of the embodiments of the present invention,
[0167] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0168] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for performing various aspects of the present invention.
[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An elevator energy recovery application method based on supercapacitors, characterized in that: include: Collecting the operation data of the elevator group and the status data of the supercapacitor, wherein the operation data includes the speed, acceleration and deceleration status, real-time load, floor position and regenerative braking energy data of the elevator during operation, and the status data includes the charge state, temperature distribution, internal resistance change and terminal voltage data of the supercapacitor; A state space is constructed based on operation data and state data, an action space is constructed based on the charging power, discharging power and energy distribution ratio of the supercapacitor and multiple elevators, a reward function is constructed based on the economic benefits under peak and valley electricity prices, the cycle life loss of the supercapacitor and the utilization efficiency of regenerative braking energy, and a deep reinforcement learning model based on a dual-Q network structure is constructed; the deep reinforcement learning model is trained through a priority experience replay mechanism to obtain an energy recovery optimization strategy; When the elevator is in the braking power generation state, the regenerative braking energy is stored in the supercapacitor. When the elevator is in the driving power consumption state, the stored energy is released for elevator traction. Calculate the optimal storage allocation plan for multiple elevators through virtual energy storage units; Based on the energy recovery optimization strategy and the optimal storage allocation scheme, a hierarchical control strategy is implemented, in which the upper layer determines the energy storage priority based on the fuzzy adaptive weight, and the lower layer uses a distributed algorithm to calculate the energy allocation coefficient and target power of each elevator, and updates the charge and discharge rate according to the health status of the supercapacitor.
2. The method according to claim 1, characterized in that The steps of constructing a state space based on operation data and state data, constructing an action space based on the charging power, discharging power and energy distribution ratio of the supercapacitor and multiple elevators, and constructing a reward function based on the economic benefits under peak and valley electricity prices, the cycle life loss of the supercapacitor and the utilization efficiency of regenerative braking energy include: Based on the operation data and status data, the motion domain state vector of energy prediction and elevator motion state is constructed using a sliding time window. The energy storage domain state vector of the health factor and capacity attenuation rate of the supercapacitor is constructed using the internal resistance-temperature coupling model. The power grid domain state vector of electricity price fluctuation characteristics and load fluctuation characteristics is constructed using wavelet transform. The motion domain state vector, energy storage domain state vector and power grid domain state vector are fused to form a state space through the attention fusion function. Construct a hierarchically coded action space, set discrete actions of charging, discharging and standby modes at the macro level, construct power allocation actions using Dirichlet distribution at the meso level, and construct power regulation actions according to capacitor impedance characteristics at the micro level. Convert the macro-level discrete actions into action vectors through one-hot encoding, map the meso-level power allocation actions into allocation coefficient vectors through probability density functions, and map the micro-level power regulation actions into regulation factors through impedance characteristic curves. Use tensor outer product operations to combine the action vector, allocation coefficient vector and regulation factor to form a unified action tensor. The power data and electricity price coefficient during the operation of the elevator group are collected to calculate the economic benefit reward; the life loss penalty is calculated based on the discharge depth and cycle number of the supercapacitor in combination with the reference cycle number; the regenerative braking energy utilization efficiency reward is calculated based on the energy conversion efficiency and the ratio of recovered energy to maximum braking power; the economic benefit reward, life loss penalty and regenerative braking energy utilization efficiency reward are weighted by the dynamic balance weight coefficient to obtain a reward function, the dynamic balance weight coefficient is determined based on the fuzzy hierarchical analysis method to determine the initial value, and is dynamically adjusted according to the peak and valley time periods, the health status of the supercapacitor and the elevator load level.
3. The method according to claim 1, characterized in that The steps to build a deep reinforcement learning model based on the double Q network structure include: The deep reinforcement learning model includes a main Q network for action selection and an evaluation Q network for action evaluation. Both the main Q network and the evaluation Q network adopt a branch-fusion architecture. Three layers of residual blocks are set in the state encoding branch to extract state features, and each residual block includes two convolutional layers and a jump connection; two layers of graph attention layers are set in the action encoding branch to extract elevator group topology features, and each attention layer includes a multi-head self-attention mechanism and edge feature embedding; a bidirectional gated recurrent unit is used to perform temporal fusion of the state features and the elevator group topology features, and the fused features are output through two fully connected layers to estimate the value; A target network with the same branch-fusion architecture as the main Q network and the evaluation Q network is constructed, and the parameter update of the target network is guided by the knowledge transfer loss function, wherein the knowledge transfer loss function includes a feature layer, a decision layer and a value layer; at the feature layer, the feature layer loss is constructed by calculating the L2 distance between the feature representation of the target network and the intermediate layer of the source network; at the decision layer, the decision layer loss is constructed by using the KL divergence between the strategy distribution output by the target network and the source network; at the value layer, the value layer loss is constructed by using the smooth L1 loss to calculate the difference in Q value estimation between the target network and the source network; the knowledge transfer loss function is obtained by weighted combination of the feature layer loss, the decision layer loss and the value layer loss.
4. The method according to claim 1, characterized in that: The steps of training the deep reinforcement learning model through the priority experience replay mechanism to obtain the energy recovery optimization strategy include: Based on the state space and action space of the deep reinforcement learning model, transfer samples are obtained, the economic benefits under the peak and valley electricity prices in the reward function, the cycle life loss of the supercapacitor and the regenerative braking energy utilization efficiency are constructed as a target weight vector, the target weight vector is calculated according to the state of charge, electricity price coefficient and life loss rate of the supercapacitor, the time difference error of the transfer sample is calculated using the target weight vector, and the time difference error is adaptively adjusted by threshold value in combination with the target weight expected value to obtain the sample priority; An elevator group topology is constructed according to the energy distribution ratio between multiple elevators, the local importance of each elevator node is calculated according to the energy transmission relationship, and the local importance is probabilistically fused with the sample priority to obtain a sampling probability; Selecting training samples from the experience replay pool according to the sampling probability, constructing a target network and a training network based on the dual Q network structure, calculating a target Q value using the target network, weighting the temporal difference error in combination with the sample priority, and updating the training network parameters; The cumulative value of the reward function is calculated within a set number of steps. When the cumulative value exceeds a preset threshold, the training network parameters are copied to the target network to obtain an energy recovery optimization strategy.
5. The method according to claim 1, characterized in that The steps of calculating the optimal storage allocation scheme for multiple elevators through the virtual energy storage unit include: The state of charge, terminal voltage and internal resistance of the supercapacitors corresponding to multiple elevators are used as state variables to construct a dynamic model of a virtual energy storage unit. The virtual energy storage unit converts the distributed energy storage system into a centralized system through an energy carrying weight coefficient, and the energy carrying weight coefficient is determined according to the rated capacity and real-time energy storage efficiency of the supercapacitor; Building an energy flow diagram of an elevator group based on the energy transmission relationship between multiple elevators, calculating the energy inflow power and energy outflow power of each elevator node, and determining the energy flow importance of the node according to the energy inflow power and energy outflow power; The node energy flow importance is used to establish a mapping relationship between the virtual energy storage unit and the physical supercapacitor, an optimization objective function is constructed based on the deviation between the target power of the virtual energy storage unit and the actual allocated power and the state of charge deviation, and a constraint condition is established in combination with the state of charge constraint and power constraint of the supercapacitor; The optimization objective function and constraint conditions are converted into a quadratic programming problem, the quadratic programming problem is solved by an interior point method to obtain a power allocation coefficient, and an optimal storage allocation scheme for multiple elevators is obtained based on the power allocation coefficient.
6. The method according to claim 1, characterized in that Based on the energy recovery optimization strategy and the optimal storage allocation scheme, a hierarchical control strategy is implemented, wherein the upper layer determines the energy storage priority based on the fuzzy adaptive weight, and the lower layer uses a distributed algorithm to calculate the energy allocation coefficient and target power of each elevator, and the steps of updating the charge and discharge rate according to the health status of the supercapacitor include: The elevator operation status, peak and valley electricity prices and energy storage status are used as input variables of the fuzzy controller. An adaptive weight adjustment mechanism is constructed according to the regenerative braking energy efficiency, electricity price cost changes and supercapacitor life loss. The energy storage priority of each elevator is obtained through fuzzy rule reasoning. A distributed optimization problem is constructed based on the communication weight matrix and the local control input, the target power regulation coefficient is calculated in combination with the energy storage priority, and the energy allocation coefficient of each elevator is solved by using the consistency theory; Obtaining a health status assessment result according to the internal resistance and capacity of the supercapacitor, dynamically adjusting the charge and discharge rate constraints based on the assessment result, and feeding the charge and discharge rate constraints back to the distributed optimization problem; A prediction window is constructed according to the elevator dispatch information, a predictive energy buffer is designed based on the predicted regenerative braking energy and the historical actual recovered energy, and the predictive energy buffer is used to dynamically adjust the energy storage power distribution; constructing a group coupling compensation term based on the speed difference, position distribution and load distribution of adjacent elevators, combining the group coupling compensation term with the energy storage priority and the evaluation result, and establishing a multi-objective collaborative optimization function; The predictive energy buffer and group coupling compensation items are updated in the upper control cycle, the weight coefficient of the multi-objective collaborative optimization function is optimized, and the power correction amount of each elevator is calculated according to the multi-objective collaborative optimization function; in the lower control cycle, power distribution correction and coupling dynamic compensation are performed according to the power correction amount, and the final control instruction is output.
7. The method according to claim 6, characterized in that The steps of constructing a group coupling compensation term based on the speed difference, position distribution and load distribution of adjacent elevators, combining the group coupling compensation term with the energy storage priority and the evaluation result, and establishing a multi-objective collaborative optimization function include: A speed coupling function is constructed according to the speed difference and physical distance between adjacent elevators, and the speed coupling function uses a distance attenuation term to correct the speed difference effect; a position coupling function is constructed based on the spatial distribution of the elevator group; a load coupling function is constructed according to the load distribution of each elevator, and the load coupling function is combined with load prediction information to design a feedforward compensation term and dynamically adjust the load correlation coefficient; the speed coupling function, the position coupling function and the load coupling function are weighted and combined by an adaptive weight coefficient to form a group coupling compensation term, and the adaptive weight coefficient is obtained by online identification using a recursive least squares method; Construct a hierarchical coupling fusion device, which includes: a health status assessment unit, which determines the upper and lower thresholds of the energy storage capacity according to the assessment results; a priority interval division unit, which is used to divide the energy storage working range into three energy storage intervals of high, medium and low according to the upper and lower thresholds of the energy storage capacity, and allocate an energy storage weight coefficient to each energy storage interval, and the energy storage weight coefficient is dynamically adjusted with the system operation state; a coupling mapping processing unit, which is used to divide the group coupling compensation item into a strong coupling area, a weak coupling area and a transition area according to the coupling strength, wherein the strong coupling area is processed by a linear mapping function, the transition area is processed by a piecewise linear mapping function, and the weak coupling area is processed by a nonlinear mapping function to obtain a mapping compensation value; an adaptive fusion unit, which is used to determine the energy storage interval in which the current working point is located, calculate the influence coefficient in combination with the assessment result, and obtain a fusion result by weighted averaging the influence coefficient and the mapping compensation value; a feedback optimization unit, which is used to collect system operation data, construct an evaluation function according to the energy storage efficiency and group synergy index of the supercapacitor, and perform online correction on the fusion parameters through the evaluation function, and the fusion parameters include the energy storage weight coefficient and the influence coefficient; A multi-objective collaborative optimization function is constructed according to the fusion result of the hierarchical coupling fuser, and the multi-objective collaborative optimization function uses a piecewise linearization method to process nonlinear coupling terms.
8. An elevator energy recovery application system based on supercapacitors, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to collect the operation data of the elevator group and the status data of the supercapacitor, wherein the operation data includes the speed, acceleration and deceleration status, real-time load, floor position and regenerative braking energy data of the elevator during operation, and the status data includes the charge state, temperature distribution, internal resistance change and terminal voltage data of the supercapacitor; The second unit is used to construct a state space based on operation data and state data, to construct an action space based on the charging power, discharging power and energy distribution ratio of the supercapacitor and multiple elevators, to construct a reward function based on the economic benefits under peak and valley electricity prices, the cycle life loss of the supercapacitor and the utilization efficiency of regenerative braking energy, and to construct a deep reinforcement learning model based on a double Q network structure; to train the deep reinforcement learning model through a priority experience replay mechanism to obtain an energy recovery optimization strategy; The third unit is used to store the regenerative braking energy in the supercapacitor when the elevator is in a braking power generation state, and release the stored energy for elevator traction when the elevator is in a driving power consumption state; Calculate the optimal storage allocation plan for multiple elevators through virtual energy storage units; Based on the energy recovery optimization strategy and the optimal storage allocation scheme, a hierarchical control strategy is implemented, in which the upper layer determines the energy storage priority based on the fuzzy adaptive weight, and the lower layer uses a distributed algorithm to calculate the energy allocation coefficient and target power of each elevator, and updates the charge and discharge rate according to the health status of the supercapacitor.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Super capacitor management system and method of adaptive optimization algorithm
CN120601592A
Kinetic energy recovery management method and system for one-dragging-N energy storage type elevator
CN121097763A
1-n energy storage type elevator kinetic energy recovery management method and system
CN121097763B
Elevator control system and method based on dynamic passenger flow prediction and life cooperative control
CN121158617A
Off-grid super-capacitor elevator energy-saving one-driving-two regulation and control method and system
CN121689162A