Power storage dynamic storage allocation method and system integrating multi-dimensional factors

By combining edge computing and digital twin technology with reinforcement learning, a dynamic storage space allocation system for power warehousing is constructed, which solves the problem of low storage space allocation efficiency in traditional power warehousing management and realizes intelligent and efficient warehousing management.

CN120450375BActive Publication Date: 2025-09-19ANHUI XICHENG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510927476.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-19
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

In traditional power storage management, storage space allocation methods rely on static strategies and manual experience, which cannot adapt to complex scenarios with diverse power equipment types, sudden maintenance needs, and changing operating environments, resulting in low storage space utilization and poor equipment scheduling efficiency.

Method used

Edge computing is used to build a digital twin environment for power storage. Multi-source heterogeneous data is collected through the OPC-UA protocol, a three-dimensional data association model is established, a 28-dimensional state vector is extracted, a hierarchical reward function is designed, and a reinforcement learning model is used for dynamic storage allocation. The strategy is optimized through an hourly, daily, and monthly closed-loop evolution mechanism.

Benefits of technology

It has achieved intelligent upgrades in power storage management, improved storage space utilization and equipment scheduling efficiency, quickly responded to emergency orders, reduced operating costs and safety risks, and possessed self-learning and self-adaptation capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450375B_ABST
    Figure CN120450375B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of warehouse automation technology, and specifically discloses a method and system for dynamic storage space allocation of electric power warehouses that integrates multi-dimensional factors, including the following steps: constructing an electric power warehouse digital twin environment based on edge computing, collecting multi-source heterogeneous data in real time through the OPC-UA protocol and establishing a three-dimensional data association model; extracting a 28-dimensional state vector based on the three-dimensional data association model, and constructing a reinforcement learning state space. The method and system for dynamic storage space allocation of electric power warehouses that integrates multi-dimensional factors in an embodiment of the present invention construct a real-time mapped warehouse virtual environment through edge computing and digital twin technology, and combine reinforcement learning with cross-system linkage rules to achieve an intelligent upgrade of electric power warehouse management; this solution can dynamically optimize storage space allocation based on multi-dimensional parameters such as the remaining service life of equipment and maintenance cycle, significantly improve storage space utilization and equipment scheduling efficiency, and comprehensively improve the intelligence and automation level of electric power warehouses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of warehouse automation technology, and in particular to a method and system for dynamic storage space allocation of electric power warehouses that integrates multi-dimensional factors. Background Art

[0002] In the field of power storage automation management, traditional storage space allocation methods mostly rely on static strategies and manual experience, which makes it difficult to adapt to complex scenarios with diverse power equipment types, sudden maintenance needs, and changing operating environments.

[0003] Existing technologies typically process warehouse data independently, lacking cross-system integration and analysis of equipment lifecycle information, production plans, and grid operating status. This leads to a disconnect between storage planning and actual demand. Furthermore, traditional solutions lack effective dynamic assessment and optimization mechanisms, making them unable to respond in real time to changes in the frequency of power equipment in and out of the warehouse, load-bearing safety requirements, and the adjustment of urgent order priorities. This, in turn, leads to low warehouse space utilization and poor equipment scheduling efficiency. Summary of the Invention

[0004] The present invention aims to solve, at least to some extent, one of the technical problems in the related art. To this end, the present invention aims to propose a method and system for dynamic storage allocation of power storage that integrates multi-dimensional factors to improve storage utilization and equipment scheduling efficiency.

[0005] To achieve the above objectives, a first embodiment of the present invention proposes a method for dynamic storage allocation of power storage that integrates multi-dimensional factors, comprising the following steps:

[0006] S1. Build a digital twin environment for power storage based on edge computing, collect multi-source heterogeneous data in real time through the OPC-UA protocol, and establish a three-dimensional data association model;

[0007] S2. Extracting a 28-dimensional state vector based on the three-dimensional data association model and constructing a reinforcement learning state space containing equipment-storage location-order coupling constraints;

[0008] S3. Design a hierarchical reward function that integrates the base layer, strategy layer, and emergency layer, and adjust the weights of each layer in real time through a dynamic weight adaptation module;

[0009] S4. Use the improved proximal policy optimization algorithm to train the reinforcement learning model and use the digital twin environment to conduct strategy preview verification;

[0010] S5. Generate storage allocation decisions in real time based on the trained reinforcement learning model, and continuously optimize the strategy through an hourly, daily, and monthly closed-loop evolutionary mechanism.

[0011] To achieve the above objectives, a second embodiment of the present invention proposes a dynamic storage allocation system for power storage that integrates multi-dimensional factors, including:

[0012] The data acquisition and modeling module is used to build a digital twin environment for power storage based on edge computing. It collects multi-source heterogeneous data in real time through the OPC-UA protocol and establishes a three-dimensional data association model that includes equipment ID, production batch, maintenance cycle, and grid location.

[0013] a state space construction module for extracting a 28-dimensional state vector based on the three-dimensional data association model to construct a reinforcement learning state space containing equipment-storage-order coupling constraints, wherein the 28-dimensional state vector includes an equipment state vector, a storage location state vector, an order state vector, and a system state vector;

[0014] The reward function calculation module is used to design a hierarchical reward function that integrates the basic layer, policy layer, and emergency layer, and adjusts the weights of each layer in real time through the dynamic weight adaptation module;

[0015] A reinforcement learning training module is used to train the reinforcement learning model using an improved proximal policy optimization algorithm and conduct policy preview verification using a digital twin environment. The policy network uses three layers of deep separable convolutional layers combined with a bidirectional LSTM network.

[0016] A cross-system collaboration module integrates data from the equipment management system, production execution system, and smart grid monitoring system to enable collaborative decision-making on grid maintenance plans, production work order changes, and equipment failure events through a three-tiered linkage rules engine consisting of strategic, tactical, and execution layers.

[0017] The decision execution module is used to generate storage allocation decisions in real time based on the trained reinforcement learning model, execute storage adjustments after digital twin preview, and trigger the linkage response of the logistics system and production system;

[0018] A closed-loop evolution module is used to continuously optimize strategies through an hourly, daily, and monthly closed-loop evolution mechanism. The closed-loop evolution mechanism includes real-time rewards to adjust action parameters, cumulative rewards to update weights, and historical strategy performance to reconstruct the network structure.

[0019] To achieve the above-mentioned purpose, the third aspect of the present invention proposes an electronic device, including a memory, a processor and a computer program stored on the memory. When the computer program is executed by the processor, it implements the above-mentioned dynamic storage space allocation method for power storage that integrates multi-dimensional factors.

[0020] The dynamic storage space allocation method and system for power storage that integrates multi-dimensional factors in the embodiment of the present invention builds a real-time mapped storage virtual environment through edge computing and digital twin technology, and combines reinforcement learning with cross-system linkage rules to achieve intelligent upgrades in power storage management; this solution can dynamically optimize storage space allocation based on multi-dimensional parameters such as the remaining service life of the equipment and the maintenance cycle, significantly improving storage space utilization and equipment scheduling efficiency; and through a three-layer linkage rule engine, it can quickly adjust storage space strategies for events such as power grid maintenance plans and production work order changes, effectively shortening the response time for emergency orders; with the help of a closed-loop evolution mechanism, the reward function and strategy model are continuously optimized, so that the storage system has self-learning and self-adaptive capabilities, providing precise support for the full life cycle management of power equipment, and comprehensively improving the intelligence and automation level of power storage. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The disclosure of the present invention is described with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of the present invention. In the drawings, the same reference numerals are used to refer to the same components. Among them:

[0022] Figure 1 1 is a flow chart of a method for dynamic storage allocation of electric power storage integrating multi-dimensional factors in one embodiment of the present invention;

[0023] Figure 2 28-dimensional state vector feature importance heat map in one embodiment of the present invention;

[0024] Figure 3 is a histogram of influence weights of different feature categories in a 28-dimensional state vector in one embodiment of the present invention;

[0025] Figure 4 is a graph showing a dynamic response curve of a layered reward function according to an embodiment of the present invention;

[0026] Figure 5 is a graph of a weight adaptive update process according to an embodiment of the present invention;

[0027] Figure 6 1 is a comparison diagram of convergence curves during the training process of the improved Proximal Policy Optimization (PPO) algorithm in one embodiment of the present invention;

[0028] Figure 7 2 is a schematic structural diagram of a power storage dynamic storage allocation system integrating multi-dimensional factors in another embodiment of the present invention;

[0029] Figure 8 It is a structural diagram of an electronic device according to another embodiment of the present invention. DETAILED DESCRIPTION

[0030] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0031] The following describes, with reference to the accompanying drawings, a method and system for dynamic storage allocation of power storage that integrates multi-dimensional factors, and an electronic device according to an embodiment of the present invention.

[0032] like Figure 1 A flowchart of a dynamic storage allocation method for power storage that integrates multiple factors is presented. This method implements dynamic storage allocation for power storage through a closed-loop technology chain: digital twin environment construction - state space modeling - reward function design - reinforcement learning training - and decision optimization. The method mainly includes the following steps:

[0033] S1. Use edge computing nodes to collect multi-source heterogeneous data (such as the remaining service life of equipment, work order delivery time, etc.) from equipment management systems and production execution systems nearby, achieve cross-system data interoperability through the OPC-UA protocol, and build a three-dimensional digital twin model that is mapped in real time to the physical warehouse.

[0034] Among them, digital twin refers to a virtual warehousing model built through technologies such as the Internet of Things and edge computing, which can mirror the equipment status, logistics flow and other scenarios of physical warehousing in real time; and the OPC-UA protocol is a cross-platform industrial communication protocol that supports the standardized collection and transmission of multi-source heterogeneous data, ensuring data consistency in systems such as equipment management and production execution.

[0035] S2. Based on the three-dimensional digital twin model (three-dimensional data association model) mentioned above, a 28-dimensional state vector is extracted, including four types of state vectors: equipment, storage location, order, and system. These vectors are combined to construct a reinforcement learning state space containing equipment-storage location-order coupling constraints, thus converting the warehousing scenario into a computable mathematical problem.

[0036] In dynamic storage allocation for power storage, "coupling constraints" refer to the interrelated and mutually influential restrictive relationships between the three key elements of equipment, storage, and orders. Equipment attributes (such as size, weight, frequency of use, and remaining useful life) determine the type of storage location suitable for it. For example, large, heavy equipment requires a storage location with strong load-bearing capacity and ample space. Frequently used equipment should be placed in a location that allows for quick access. Order requirements (such as delivery time, priority, required equipment type, and quantity) influence the order in which equipment is shipped out and the storage allocation strategy. Equipment for urgent orders needs to be prioritized for storage locations near exits to ensure rapid shipment. The storage location's own characteristics (such as location, load-bearing capacity, and space) also constrain equipment storage options.

[0037] Therefore, the complex interdependence and constraints between equipment, storage locations, and orders are known as "coupling constraints." When extracting the 28-dimensional state vector, it is precisely by analyzing these coupling constraints that the relevant attributes of equipment, storage locations, orders, and the system are converted into specific state vectors. This constructs a mathematical representation that comprehensively reflects the warehouse's operational status, providing the foundational data for designing reward functions and conducting reinforcement learning training in subsequent steps.

[0038] The reinforcement learning state space refers to the set of all possible states that an agent can perceive during its interaction with the environment. In this method, the extracted 28-dimensional state vectors constitute the reinforcement learning state space for the dynamic storage slot allocation problem in power storage. Each state vector represents a specific state of the storage system at a given moment, such as the remaining useful life of the equipment, the load factor of the storage slot, and the delivery time of the order. Based on the current state, the agent (i.e., the reinforcement learning algorithm) takes appropriate actions (such as assigning a piece of equipment to a specific storage slot). It then receives feedback rewards based on the reward function designed for subsequent steps and transitions to the next state. By continuously looping through this state space (state perception - action selection - reward acquisition - state transition), the agent gradually learns the optimal storage slot allocation strategy.

[0039] It should be noted that the reinforcement learning state space is closely connected to the digital twin environment constructed in step S1. The digital twin environment provides the state space with real-time data of real warehousing scenarios. At the same time, the accurate construction of the state space provides effective input for the improved proximal policy optimization (PPO) algorithm in subsequent steps, enabling the algorithm to be trained and optimized within a reasonable state space, ultimately achieving efficient storage allocation decisions.

[0040] S3. Design a layered reward function that integrates the foundational layer, the strategic layer, and the emergency layer. A dynamic weight adaptation module adjusts the weights of each layer in real time. The foundational layer improves inbound and outbound efficiency; the strategic layer extends the lifespan of high-value equipment; and the emergency layer shortens fault recovery time (a core requirement of the power system).

[0041] S4: The reinforcement learning model is trained using an improved proximal policy optimization (PPO) algorithm, and then the strategy is rehearsed and verified using a digital twin environment. As can be seen in the previous steps, the digital twin environment constructed in step S1 provides a virtual simulation operating scenario for reinforcement learning; the 28-dimensional state vector extracted in step S2 constitutes the input state space of the reinforcement learning model; and the hierarchical reward function designed in step S3 provides goal guidance for the reinforcement learning model. Therefore, the introduction of the reinforcement learning model in step S4 builds on the foundation established in the previous steps. Finding the optimal storage allocation strategy through model training is a key transition in the entire method, from data preparation to policy generation.

[0042] In this approach, the reinforcement learning model represents an intelligent decision-making system. By continuously interacting with the digital twin environment, it learns how to make optimal storage allocation decisions based on the current storage state (i.e., the 28-dimensional state vector). During training, the model gradually optimizes its decision-making strategy with the goal of maximizing cumulative rewards (calculated using a hierarchical reward function), thereby achieving efficient dynamic storage allocation for power storage.

[0043] S5. Generate storage allocation decisions in real time based on the trained reinforcement learning model, and continuously optimize the strategy through an hourly-daily-monthly closed-loop evolution mechanism. The hourly, daily, and monthly levels constitute a progressive closed-loop evolutionary system from real-time response to long-term optimization. The hourly level focuses on rapid response to real-time operational indicators, adjusting storage allocation parameters through real-time rewards to solve current efficiency and cost issues. This is a short-term fine-tuning adjustment; the daily level is based on 24-hour accumulated rewards and adjusts the weights of the three-layer reward function in step S3 ( ) is updated to balance the priorities of different objectives (such as efficiency and long-term planning), representing mid-term strategic direction calibration. From a more macro perspective, the monthly level restructures the reinforcement learning model's network structure in step S4 based on historical and long-term strategic performance, addressing the model's adaptability in complex scenarios and representing long-term architectural optimization. These three steps progress in a progressive manner: hourly adjustments provide data accumulation for daily optimization, while daily weight updates provide direction for monthly structural reconstruction, jointly ensuring that the reinforcement learning model continuously adapts to changes in the warehouse environment.

[0044] As an example, at the hourly level: if transportation costs surge during a certain period, the allocation weight of storage locations farther from the exit will be dynamically reduced. The storage allocation weights here refer to the weights and thresholds of various factors that influence storage allocation decisions, which will be clearly displayed in the subsequent 28-dimensional state vector.

[0045] Daily level: If urgent orders are frequently processed, the weight of the emergency level will be automatically increased ( );

[0046] Monthly level: For example, in the improved proximal policy optimization algorithm in step S4, the policy network uses three depthwise separable convolutional layers to process storage space features. However, after one month of operation, if the reinforcement learning model is found to be inaccurately extracting spatial features when handling complex storage layouts, resulting in poor storage allocation strategy performance (as determined by the long-term cumulative reward value of the hierarchical reward function in step S3), the number of convolutional layers can be increased based on historical policy performance. More convolutional layers can extract more complex and abstract storage space features, enabling the policy network to better understand spatial relationships and thus optimize the storage allocation strategy. This process deeply optimizes the reinforcement learning model network structure in step S4 to adapt to the long-term changing needs of the warehouse environment.

[0047] The dynamic storage allocation scheme for power storage formed through the above steps S1-S5 builds a complete closed-loop system from real-time data collection to continuous strategy optimization. This method uses digital twin technology to achieve accurate mirroring of the storage environment. Combined with the reinforcement learning model and the hierarchical reward mechanism, it can comprehensively consider the coupling relationship between equipment attributes, storage characteristics and order requirements, and dynamically generate the optimal storage allocation strategy. At the same time, hourly real-time response, daily strategy calibration and monthly architecture optimization make the system adaptive, which can effectively respond to dynamic changes in the storage environment, significantly improve storage space utilization and equipment scheduling efficiency, ensure rapid response to emergency orders, reduce operating costs and safety risks, and realize intelligent, efficient and sustainable optimization of power storage management.

[0048] In some embodiments of the present invention, four types of system data are collected in real time by edge computing nodes to form the core input of the digital twin environment. These multi-source heterogeneous data include:

[0049] Equipment management system: collects equipment ID, remaining service life (e.g., a transformer has a remaining service life of 3.5 years), and maintenance cycle (e.g., the standard maintenance cycle is 1 year) to assess equipment health status and maintenance priority;

[0050] Production Execution System: This system obtains order urgency flags, the size of associated equipment groups (e.g., a work order involves five circuit breakers), and work order delivery times (e.g., requiring shipment within 48 hours), directly impacting the timeliness strategy for storage allocation.

[0051] Smart grid monitoring system: collects grid fault correlation (e.g., a fault affects eight substations) and substation locations (used to calculate equipment dispatch paths), supporting storage optimization in grid maintenance and fault emergency scenarios;

[0052] Warehouse management system: Collecting storage location load rate, space utilization, and distance to the exit are key parameters for evaluating storage location physical properties;

[0053] The above four types of system data are transmitted to the digital twin environment through the OPC-UA protocol, forming the basis of the three-dimensional data association model. They directly correspond to the device status, storage location status, etc. in the 28-dimensional state vector in step S2, providing real-time state input for reinforcement learning.

[0054] It should be noted that multi-source heterogeneous data needs to be pre-processed by edge computing nodes (such as data cleaning and format unification) before it can be integrated.

[0055] In addition, unlike existing technologies that process warehouse data independently, this solution achieves cross-system data linkage through the following mechanisms:

[0056] Timestamp alignment technology: Nanosecond-level timestamps are added to data from different systems. A sliding window algorithm (window size 500ms) is used to eliminate data inconsistencies caused by transmission delays, ensuring real-time mapping of device status and order requirements.

[0057] Semantic mapping model: Establishes a dictionary that associates device IDs with production batches and grid locations. For example, device ID "E-001" corresponds to production batch "P202506" and substation "SB-12," enabling full-chain tracking from order to equipment to storage location.

[0058] The construction of a three-dimensional data association model needs to meet the following requirements:

[0059] Equipment storage priority = α × maintenance urgency + β × production demand urgency + γ × grid fault correlation;

[0060] Among them, α, β, and γ are the weight coefficients of storage priority, and α + β + γ = 1. Through the dual-objective optimization of historical order completion rate and equipment maintenance timeliness rate, the weights of α, β, and γ are dynamically adjusted. For example, when the quarterly maintenance timeliness rate is less than 80%, α is automatically increased to 0.5 (original weight 0.3), strengthening the influence of maintenance priority on storage allocation.

[0061] Maintenance urgency: The difference between the remaining service life of the equipment and the standard maintenance cycle is standardized. The formula is: Maintenance urgency = 1- ;

[0062] For example, if the remaining service life of a certain equipment is 0.5 years and the standard cycle is 1 year, then the maintenance urgency = 0.5. The larger the value, the more urgent the maintenance need.

[0063] Production demand urgency: Calculated based on the difference between the work order delivery time and the current time, using an exponential decay function, production demand urgency =

[0064] If a work order has a remaining delivery time of 8 hours, the calculated urgency is ≈ 0.716, ensuring that urgent orders receive higher priority.

[0065] Grid fault correlation: linearly mapped to the [0,1] interval according to the number of substations affected by the fault. The formula is: Grid fault correlation =

[0066] If a fault affects 5 substations and the total number of regions is 20, then the correlation degree = 0.25, which is used to trigger the emergency storage adjustment.

[0067] This three-dimensional data association model transforms heterogeneous indicators of equipment health, production demand, and grid security into unified priority values, resolving the traditional reliance on human experience for multi-objective decision-making. The priority model output serves as a state preprocessing feature for the reinforcement learning model in step S4, narrowing the strategy search space. For example, high-priority equipment is directly assigned to a candidate storage location within 5 meters of an exit, improving decision-making efficiency.

[0068] In some embodiments of the present invention, the 28-dimensional state vector includes:

[0069] Equipment state vector: remaining service life of the equipment , maintenance cycle weight , Historical delivery frequency , Real-time power of the device and equipment failure warning level ;

[0070] Storage state vector: load-bearing rate , space utilization , Standardized value of distance from exit , storage height coefficient , Storage Continuity Index 、The partition to which the storage location belongs , storage location historical accident rate , temperature sensitivity , humidity sensitivity And moisture-proof status ;

[0071] Order status vector: urgent order flag , associated device group size , Estimated delivery time window , Special qualification requirements for orders , Order multiple batch split mark And the environmental protection requirement level of the order ;

[0072] System state vector: Shared storage ratio across storage areas , handling equipment busyness index , system real-time energy consumption , Number of security alarms in the warehouse area , storage management system load rate , Historical strategy reuse rate and channel congestion coefficient ;

[0073] Of the 28-dimensional state vector, 13 dimensions represent broad categories, requiring ambiguity to be resolved through detailed subdivisions (such as equipment type weight and storage location environmental status). For example, "storage location status" should not be solely based on "height coefficient" but also include "moisture resistance" (power transformers are susceptible to moisture). This 28-dimensional state vector can be further subdivided depending on the situation, such as channel congestion coefficient and storage location moisture resistance / temperature control status. This solution, through deep adaptation to power scenarios and dynamic decision support, further demonstrates its necessity.

[0074] For example, in the 28-dimensional state vector, the real-time power consumption of the equipment needs to be synchronized with the digital twin in real time, which reflects the "power scene restoration degree" of the digital twin and solves the problem of functional ambiguity of the digital twin.

[0075] like Figure 2 This figure shows a heat map of the importance of 28-dimensional state vector features. Red (9.0+) indicates the most important features; yellow (7.0-8.9) indicates moderately important features; and blue (<7.0) indicates minimally important features. In the figure, moisture resistance (8.6) and fault warning (8.9) are highly important, which aligns with the solution's goal of comprehensively improving the intelligent level of power storage.

[0076] And as Figure 3 As shown here, a histogram of the influence weights of different feature categories in the 28-dimensional state vector is shown. Among them, the equipment status accounts for 34.5%, which reflects the management of the equipment throughout its life cycle; the storage status accounts for 29.8%, which corresponds to the calculation of the storage load balancing index, and thus can support the technical effect of improving space utilization in this solution; the order status accounts for 23.7%; and the system status accounts for 12.0%. Figure 3 It can be seen that the advantages of this solution over the traditional method are: not only does it consider more dimensional features (28 dimensions vs. the traditional 10-15 dimensions), but more importantly, it makes storage allocation decisions more accurate and reasonable by scientifically quantifying the influence weight of each feature.

[0077] In some embodiments of the present invention, the hierarchical reward function serves as the core optimization objective of the reinforcement learning model and forms a closed loop with the entire process of the present method. The calculation formula of the hierarchical reward function is:

[0078]

[0079] in, 、 、 They are the reward weight coefficients of the basic layer, strategy layer, and emergency layer, which directly control the impact priority of the three types of rewards and meet the requirements. + + =1;

[0080] Base layer rewards The calculation formula is:

[0081]

[0082] Where, is the deviation rate of the delivery time, that is, the ratio of the actual delivery time to the predicted delivery time; is the transportation cost, which is equal to the product of the transportation distance and the weight of the equipment; is space utilization; The storage load balance index is defined as the standard deviation of the storage load rate of each storage area, which is used to balance the load-bearing safety requirements in the storage of power equipment;

[0083] It should be noted that the storage load balance index Calculated as the standard deviation of the storage load rate of each storage location in the reservoir area. For example, if the storage load rates of three storage locations in a reservoir area are 60%, 75%, and 80% respectively, the mean , . The smaller it is, the higher the reward is, avoiding safety accidents caused by local overload, and reducing the load-bearing accident rate by 85% compared with traditional solutions.

[0084] The space utilization Using the three-dimensional volume calculation method, if the total volume of a storage location is 10m³ and the volume of the stored equipment is 7m³, then , it can be increased to 85% through device mixed storage strategy (such as combining small-volume high-frequency devices with large-volume low-frequency devices), and the basic layer reward is +0.3.

[0085] Strategy layer rewards The calculation formula is:

[0086]

[0087] Where, The matching degree between the remaining service life of the equipment and the storage planning period; The continuity index of the associated group is used to count the continuous storage ratio of the same order equipment group in the storage location. For example, if an order requires 5 equipment and 4 of them are stored in the same area, then ,High continuity can reduce the crossing of transport paths and reduce the ,penalty value of the strategy layer; is the risk value of high-load storage location, which is calculated by combining the load rate and the importance of the equipment. For example, if the load rate of a storage location is 90% and a core transformer is stored, then (Importance coefficient 1.5), strategy layer reward -0.135, the system automatically triggers storage reallocation.

[0088] Emergency Level Rewards The calculation formula is:

[0089]

[0090] Where, is the emergency order response coefficient, which is 1 when the emergency order is successfully allocated to the storage position near the exit. For example, when a substation fails and a backup circuit breaker needs to be urgently called:

[0091] If it is allocated to a storage location within 5 meters of the exit within 10 minutes, , reward +0.5, synchronously triggering hourly adjustment, the storage allocation weight of this area is temporarily increased by 20%;

[0092] like Figure 2 As shown, the emergency order mark (9.8) is out of 10 points. When the emergency order is successfully allocated to the storage location near the export, =1, reward +0.5.

[0093] If the task is not completed within 30 minutes, the reward will be -0.3. When the daily weight is updated Increase from 0.2 to 0.3 to strengthen subsequent emergency response.

[0094] The input dependency of the above layered reward function: The ΔT of the base layer reward is directly related to the “estimated delivery time window” in the order status vector “Distance from the exit” to the storage state vector ; "Equipment Remaining Life Matching" at the strategy level Depends on the "remaining useful life" of the device state vector .

[0095] The output of the layered reward function affects: the reward value drives the closed-loop evolution mechanism. For example, when updating the weight at the daily level, if the emergency layer reward continues to be low, the system will automatically increase the reward value. To 0.4 (original weight 0.2), strengthen the response to emergency orders.

[0096] The hierarchical reward function in this solution is different from the single efficiency-oriented one in existing technologies. This solution achieves three-dimensional optimization through a three-layer structure:

[0097] Basic layer: Efficiency is prioritized. For example, if the transportation cost of a certain storage area surges, the basic layer rewards automatically reduce the weight of the remote storage location allocation, and the reinforcement learning model is linked to adjust the priority of convolutional layer feature extraction;

[0098] Strategy layer: long-term planning, such as the remaining service life of the transformer is 5 years, matching the storage time of the 5-year planning cycle =1, strategic layer reward +0.3, to avoid premature scrapping caused by mismatch between equipment and storage space life;

[0099] Emergency layer: safety guarantee. When the grid fault correlation is ≥0.5, the emergency layer is triggered compulsorily. If an emergency order is allocated to a storage location ≤3 meters away from the exit within 10 minutes, =1, reward +0.5, response time is shortened by 60% compared with the traditional solution.

[0100] In addition, you can also add , Energy consumption (kWh) of handling equipment. When AGV energy consumption is high during a certain period, the system prioritizes allocation to nearby storage locations, thus reducing energy consumption.

[0101] Or add it at the policy level , Based on sensor data such as equipment vibration and temperature, equipment with low health status is preferentially allocated to storage locations that are convenient for maintenance.

[0102] like Figure 4 The dynamic response curve of the layered reward function is shown. In the figure, the blue line (base layer) reflects the daily operation efficiency. Its peak value of 0.65 (9:00) corresponds to the peak period of power storage outbound, and its valley value of -0.15 (15:00) reflects the surge in transportation costs during power storage operation. The green line (strategy layer) reflects the effect of long-term planning and is stable in the range of 0.2-0.4, which can verify the sustained benefits of equipment life matching in this solution. The red line (emergency layer) represents the response to emergencies. For example, the peak value of 0.5 from 8:00 to 12:00 corresponds to the "emergency response to substation failure". At this time, the emergency layer reward jumps by 0.5 ( Figure 4 The moment corresponding to the yellow dot in the figure verifies the above formula: The total reward increase of 0.35 reflects the effect of the dynamic weight mechanism; and the -0.3 valley value from 18:00 to 20:00 reflects the "urgent order processing timeout".

[0103] In addition, from Figure 4 It can also be seen that when the work order fluctuation occurs at 14:20 ( Figure 4At the moment corresponding to the purple dot in the middle), the base layer reward drops by 0.2, which corresponds to the transportation cost in this solution. The increase in strategy layer rewards by 0.15 reflects the dynamic adjustment of the reserve ratio.

[0104] It is also worth noting that the total reward (black line) remains stable at >0.4 under sudden events, which proves the effectiveness of the three-layer reward complementary design in this scheme.

[0105] The hierarchical reward function in this scheme breaks down the "efficiency-planning-emergency" goals of power storage into independent reward layers, achieves dynamic balance of multiple objectives through dynamic weights, addresses the "one-size-fits-all" decision-making defects of existing technologies, and introduces power-specific indicators such as equipment remaining life matching and grid fault correlation. For example, when the deviation between the transformer storage space planning cycle and the remaining life exceeds 2 years, the strategy layer automatically triggers reallocation, reducing the equipment's full life cycle cost. This not only provides a precise optimization target for the reinforcement learning model, but also realizes a technological leap from "passive execution" to "active optimization" for power storage through the fusion of multi-dimensional indicators and a dynamic adjustment mechanism.

[0106] In some embodiments of the present invention, the dynamic weight adaptation module updates the weights by gradient descent:

[0107]

[0108] in, is the weight vector at the tth iteration, which is essentially the reward weight coefficient of the base layer, strategy layer, and emergency layer at the tth iteration 、 、 A dynamic set of, each iteration (t increases by 1), Based on the weight vector of the previous iteration Make updates; is the current iteration number; T is the transpose; is the adaptive learning rate; The historical average reward is stored in the daily memory module of the closed-loop evolution mechanism and is used to measure the quality of the current strategy. is the hierarchical reward function Weight vector The gradient vector of .

[0109] As an example, suppose that the initial iteration =0, (The emergency layer has a higher weight); at the 50th iteration, the system detects a decrease in emergency orders, and after calculating by the gradient descent method, ,Right now Improved from 0.3 to 0.45 (base layer priority increased); From 0.4 to 0.2 (emergency layer priority is reduced). This process reflects Dynamically adjusted with the number of iterations t The logic ensures that weight updates are consistent with changes in storage scenarios.

[0110] As an example, let's compute the gradient vector:

[0111] Assume that at some point , , the gradient vector , initial weight , adaptive learning rate ,but: .

[0112] As a result, the weight of the basic layer increases slightly, and the system automatically strengthens the deviation rate of the delivery time. With the attention of the company, the on-time delivery rate within the verification cycle has also been improved.

[0113] In the above formula, the adaptive learning rate The dynamic adjustment formula is:

[0114]

[0115] in, is the initial learning rate, ranging from 0.01 to 0.05, to ensure stability in the initial stage of training; is the maximum learning rate, accelerating convergence in the later stages of iteration; is the maximum number of iterations, is the current iteration progress. hour ,when hour , forming an adaptive rhythm of "slow at first and then fast".

[0116] When reconstructing the network structure at the monthly level (such as adding convolutional layers), reset , the learning rate changes from Re-grow linearly to avoid training shocks caused by changes in network structure.

[0117] As an example, when emergency orders surge: real-time rewards The reward contribution of the emergency layer is improved. And the gradient vector of The component is positive; the weight update formula is From 0.2 to 0.35, the emergency order response time was shortened from an average of 12 minutes to 7 minutes; after the daily cumulative reward verification, Maintain high weight to form a closed-loop optimization of "detection-response-solidification".

[0118] As an example, the dynamic weight mechanism also provides more accurate gradient guidance for the PPO algorithm mentioned above:

[0119] When the base layer weight When improving, the depth of the separable convolutional layer of the policy network enhances the distance between the storage and the exit Feature extraction; strategy layer weight When it is dominant, the bidirectional LSTM strengthens the remaining service life of the equipment Temporal dependency learning.

[0120] like Figure 5 A graph showing the weight adaptive update process is shown in the figure. (Base layer) from the initial 0.25 to the final 0.5, reflecting the shift from efficiency priority to multi-objective balance; (Strategy layer) From the initial 0.3 to the final 0.45, it reflects the increasing importance of long-term planning; (Emergency layer) jumps when an emergency occurs (+0.2 at 65 iterations), which verifies the sensitivity of the gradient update formula in this scheme.

[0121] exist Figure 5 In the case of a work order fluctuation event (35 iterations), it is represented by a blue dot in the figure. Down 0.1, It rises by 0.08, which reflects the dynamic adjustment of storage ratio strategy. When a power grid failure event occurs (65 iterations), it is represented by a red dot in the figure. The jump from 0.25 to 0.55 further verifies the weight formula updated by the gradient descent method above.

[0122] Figure 5 The final weight distribution in is as follows: =0.35, =0.45, =0.2, which reflects the reasonable balance between efficiency, planning and emergency response mentioned above.

[0123] It is also important to note that Figure 5 The response time for medium weight adjustment is <5 iterations, which shows the technical effect of hourly real-time adjustment in this solution.

[0124] In this scheme, the gradient of the reward function with respect to the weight is calculated and the weight is updated along the gradient direction, ensuring that each update is made in the direction of increasing the total reward. This is different from the blindness of traditional heuristic weight adjustment.

[0125] In some embodiments of the present invention, in the improved proximal policy optimization (PPO) algorithm:

[0126] 1. The policy network uses three layers of depthwise separable convolutional layers to process storage space features and a bidirectional LSTM to process order time series data. This is because the storage state vector extracted in step S2 contains spatial information such as the storage location and layout. The three layers of depthwise separable convolutional layers can efficiently extract and process these storage space features. Compared to traditional convolutional layers, depthwise separable convolutional layers have significantly fewer parameters and higher computational efficiency. They can quickly process the complex characteristics of the storage space, providing higher-quality input to the policy network, thereby helping it to more accurately generate storage allocation strategies. Once these strategies are generated, they are verified in the digital twin environment, and the policy network is optimized based on the verification results.

[0127] As an example, when processing a heat map of storage location layout, deep convolution is first used to extract spatial features (such as the distribution of high-load areas) channel by channel. These features are then fused through point convolution to improve the representation efficiency of spatial constraints such as "distance from the storage location to the exit" and "load balance." This process allows the seven-dimensional spatial features (load capacity, space utilization, etc.) in the storage location state vector to be standardized before being input into the convolutional layer, which outputs a storage location adaptability feature map that directly influences the policy network's storage location allocation decisions.

[0128] When bidirectional LSTM processes order time series, the “estimated delivery time window = "Historical delivery frequency " and other time series data, using bidirectional LSTM to simultaneously capture time dependencies in the past (such as historical order delivery patterns) and the future (such as upcoming order peaks). For example, when urgent orders appear continuously during a certain period of time, the bidirectional LSTM will learn the "urgent orders appear in large numbers" pattern and adjust the storage allocation strategy in advance to pre-allocate high-frequency equipment to areas near the exit.

[0129] As an example, order timing characteristics and emergency layer rewards When a peak in urgent orders is predicted, the strategy network prioritizes allocating storage locations near exports, improving the urgent order fulfillment rate and obtaining higher rewards.

[0130] 2. The value network outputs multi-objective value estimates and uses the Huber loss function to reduce the impact of outliers. The loss function expression is:

[0131]

[0132] in, is the target value, is the predicted value, is the threshold.

[0133] In the above scheme, the hierarchical reward function includes multiple objectives, such as the base layer, the policy layer, and the emergency layer. The multi-objective value output by the value network is precisely the value assessment of the possible value obtained by taking different actions in the current state across these multiple objective dimensions. For example, the value network will evaluate the value of a storage allocation strategy based on multiple objectives such as the deviation rate of outbound delivery time, handling costs, the matching degree of remaining equipment life with the storage planning cycle, and the processing of urgent orders. These value assessment results are used to guide the optimization of the policy network, helping it learn storage allocation strategies that better meet multi-objective requirements. They are also mutually verified with the reward feedback after the strategy preview verification in the digital twin environment, jointly promoting the training and optimization of the reinforcement learning model.

[0134] As an example, when a sudden abnormality in the storage load rate occurs (e.g., a storage location temporarily stores overweight equipment), the Huber loss function automatically switches to linear loss mode to avoid excessive penalty on outliers by the mean square error loss. For example, if the target value y = 0.8 (normal load rate), the predicted value ŷ = 1.2 (overload rate), and the threshold ε = 0.2, the loss is: L = 0.2 × (|1.2 - 0.8| - 0.2 / 2) = 0.2 × 0.3 = 0.06;

[0135] Compared with the mean square error loss of (1.2-0.8)² / 2=0.08, the Huber loss reduces the loss value by 25%, ensuring that the value network can still stably output value estimates in abnormal scenarios such as device failure.

[0136] like Figure 6 Shown is a comparison of the convergence curves of the improved proximal policy optimization (PPO) algorithm training process. In the figure, the blue solid line (improved PPO) represents the training process of the improved algorithm proposed in this patent, and its initial reward value is about 25 (after 10 iterations), which is consistent with the characteristics of the initial exploration stage of reinforcement learning; the reward value reaches 74.4 (marked with blue dots) at 40 iterations, verifying the effect of "deep separable convolution + bidirectional LSTM accelerated feature extraction" mentioned above; after 70 iterations, it enters the stable convergence zone (88±2) and finally stabilizes at around 92.5, corresponding to the dynamic weight adaptation module and closed-loop evolution mechanism mentioned above. In comparison, the red dotted line (traditional PPO) has a reward value of only 33.6 (marked with red dots) at the same 40 iterations, a gap of 45%, and its final convergence value is about 53, which is 26% lower than the improved algorithm. Through quantitative comparison, Figure 6 The performance improvement brought by depthwise separable convolution, bidirectional LSTM, dynamic weight mechanism and Huber loss function is intuitively verified.

[0137] In some embodiments of the present invention, the closed-loop evolution mechanism is implemented as follows:

[0138] ①Hourly real-time parameter adjustment:

[0139] Dynamic adjustment logic: Based on the real-time reward feedback from step S3, adjust the storage allocation action parameters. For example, the transportation cost in a certain period of time Surge (e.g., handling equipment failure causing the path to become longer, e.g. Figure 2 As shown, when the normalized value of the distance from the exit When the value reaches 9.3, the base layer rewards decrease, and the system automatically reduces the weight of distant storage locations by 15%, prioritizing nearby storage locations. This reduces the cost of transporting goods by 20% within 30 minutes. The adjusted parameters (such as the weight of storage distance) are input into the policy network as prior knowledge, accelerating the model's adaptation to real-time scenarios.

[0140] ② Daily update of the reward function weights of the base layer, strategy layer, and emergency layer based on the 24-hour cumulative rewards , the daily weight optimization example is as follows:

[0141] Weight update process: If the proportion of emergency orders on a certain working day reaches 30%, the contribution of emergency layer rewards will be increased. When updating the daily level, the gradient descent method will be used to The rate of emergency order response increased from 0.2 to 0.35. The next day, in similar scenarios, the emergency order response rate increased from 70% to 92%.

[0142] ③ At the monthly level, the strategy network and value network structures in the above scheme are reconstructed based on historical strategy performance. The monthly adaptive optimization example is as follows:

[0143] In a certain month, the average frequency of outbound shipments fluctuated by more than 40%, and the accuracy of the strategy network in allocating storage locations for high-frequency equipment dropped to 65%, triggering a monthly reconstruction. By adding a convolutional layer (from 3 layers), the accuracy of the strategy network in allocating storage locations for high-frequency equipment dropped to 65%. 4 layers), strengthen the "storage height coefficient Storage Continuity Index " feature extraction, the accuracy rate rose to 91% the next month.

[0144] As an example, monthly reconstruction is based on the sensitivity analysis of the state vector based on the historical strategy, such as finding that the “historical accident rate of storage position " is not effectively learned, a new convolution kernel is added to specifically process this feature dimension.

[0145] In the three-level closed-loop evolution, the hourly level quickly responds to short-term fluctuations, the daily level balances the priorities of multiple objectives, and the monthly level adapts to long-term trends, forming a complete self-optimization system, which is different from the static strategy of traditional solutions.

[0146] Because power warehousing involves multiple constraints, including the entire equipment lifecycle (e.g., the remaining service life of a transformer), grid operating conditions (e.g., substation failures), and production order requirements (e.g., emergency work orders), a single system cannot fully address these factors. For example, if a smart grid monitoring system detects a substation failure, it must connect with the warehouse management system to quickly retrieve backup equipment and simultaneously connect with the production execution system to adjust work order priorities. This relies on cross-system data interoperability.

[0147] Therefore, this method also includes a cross-system collaborative decision-making step, which integrates data from the equipment management system, production execution system, smart grid monitoring system and warehouse management system to achieve collaborative processing of power grid maintenance plans, production work order changes and equipment failure events.

[0148] During the digital twin construction process in step S1, multi-source data (such as device IDs and work order delivery times) from equipment management and production execution is collected via the OPC-UA protocol. This essentially lays the data foundation for cross-system collaboration. For example, the digital twin must simultaneously mirror both equipment health status and grid maintenance plans, inevitably involving inter-system data fusion. Furthermore, during the state vector extraction process in step S2, the 28-dimensional vector includes cross-system features such as "grid fault correlation" and "shared storage space ratio across storage areas" (for example, grid fault correlation requires the linkage calculation of smart grid and storage data), directly demonstrating the necessity of collaborative decision-making.

[0149] In the entire process logic chain from step S1 to step S5: data collection (S1) → multi-dimensional feature extraction (S2) → multi-objective reward design (S3) → reinforcement learning decision (S4) → closed-loop optimization (S5), each link implicitly involves cross-system requirements. For example, the emergency layer reward calculation in step S3 requires the "urgent order response coefficient." The order urgency comes from the production execution system, and the storage space allocation comes from the warehousing system, which must be coordinated.

[0150] The cross-system collaborative decision-making is achieved through a three-layer linkage rule engine:

[0151] 1. Strategic level: Power grid maintenance pre-dispatch mechanism

[0152] Trigger condition: When the smart grid monitoring system releases a maintenance plan (e.g., a substation is scheduled for maintenance in 72 hours), the corresponding spare equipment information (e.g., transformer ID, remaining service life) is extracted from the equipment management system and combined with the storage location distance data from the warehouse management system (e.g., a list of storage location IDs ≤ 5 meters from the exit).

[0153] Execution logic: 72 hours in advance, move the spare equipment from the regular storage location to the area near the exit, and link the storage location state vector of step S2 (update the standardized value of the distance from the exit). ) and S3 emergency level rewards (increase emergency order response coefficient ).

[0154] As an example: A power grid maintenance plan is notified three days in advance. The system automatically allocates three spare transformers to a storage location 3 meters away from the exit. The delivery time is shortened from the normal 15 minutes to 5 minutes, and the emergency layer reward is +0.5. The response efficiency of similar incidents within the verification period is improved by 60%.

[0155] 2. Tactical level: Dynamic position adjustment based on work order fluctuations

[0156] Trigger condition: The fluctuation of work order demand in the production execution system exceeds 20% (for example, the number of finished product orders surges by 25% in a certain period), and the size of the associated equipment group in the order status vector is analyzed. and estimated delivery time window .

[0157] Execution logic: Dynamically adjust the ratio of raw material and finished product storage space, and the adjustment range shall not exceed 30% of the total storage space. For example, 30% of the raw material storage space shall be converted to finished product storage space, and the space utilization rate of step S2 shall be linked. and the base layer reward of step S3.

[0158] As an example: When the demand for work orders in a certain warehouse center surges by 28%, the finished product storage utilization rate increases from 65% to 88% after tactical layer adjustments, the basic layer reward is +0.35, and the transportation cost is reduced by 18%.

[0159] 3. Execution layer: equipment failure emergency response

[0160]

[0161] The event urgency is an integer between 1 and 10, the impact range is the number of substations affected by the fault, and the response time window is the remaining time for fault handling (in hours).

[0162] Linkage mechanism: Mark the spare equipment as "priority outbound", update the equipment state vector (maintenance cycle weight) in step S2 Improvement) and the emergency layer reward of step S3 (+0.5 when reaching the target).

[0163] The three-layer linkage logic in this solution is that the strategic layer solves long-term planning (weekly level), the tactical layer handles medium-term fluctuations (daily level), and the execution layer responds to real-time events (seconds level), forming a complete time-dimensional decision-making system.

[0164] In some embodiments of the present invention, the feedback loop of the three-layer linkage rule engine includes:

[0165] 1. Execution layer - warehouse management system

[0166] Feedback parameter: the deviation rate of delivery time ΔT. For example, if the actual delivery time of a device is 8 minutes and the predicted delivery time is 10 minutes, ΔT = 0.8, which is fed back to the warehouse management system to optimize the transportation path (such as avoiding congested channels).

[0167] Optimization effect: After 3 consecutive days of feedback of ΔT data, the transportation path is shortened by an average of 12%, and the transportation cost item in the base layer reward is −0.3× Increased by 0.25, the overall on-time delivery rate within the verification cycle improved.

[0168] 2. Tactical layer - production execution system

[0169] Influence coefficient calculation:

[0170] in, is the rate of change of transport distance; is the production time change rate; 、 is the tactical layer efficiency impact weight coefficient and .

[0171] As an example, =0.6, =0.4, the change rate of transportation distance after storage adjustment ΔD=−15% (shortened by 15%), the change rate of production time =+8% (increase by 8%), then:

[0172]

[0173] A negative value indicates that production efficiency has improved. After feedback to the production execution system, subsequent storage adjustments will give priority to similar strategies.

[0174] 3. Strategic layer - Enterprise resource planning system

[0175] Safety stock factor:

[0176] in, It is the benchmark coefficient, with a value range of 1.0-1.5. The number of successful collaborations is the number of times the efficiency is improved by more than 10% after linkage execution.

[0177] As an example, =1.2, the total number of collaborations in a quarter is 20, and the number of successes is 15 (efficiency improvement > 10%), then:

[0178]

[0179] As a result, safety stock was reduced by 30%, which was fed back into the enterprise resource planning system to optimize inventory costs.

[0180] The feedback loop in this solution, through cross-system data exchange (e.g., feedback from the execution layer to the warehouse management system), enables continuous iteration of storage strategies, unlike the one-time decision-making of traditional solutions. The three-tiered, interconnected rules engine and feedback loop system achieve a transition from "passive response" to "active rehearsal," not only resolving the technical difficulties of cross-system coordination in power storage but also improving overall operational efficiency through quantitative decision-making models and self-optimization mechanisms.

[0181] Corresponding to the above method, such as Figure 7 As shown, this solution also proposes a dynamic storage allocation system for power storage that integrates multi-dimensional factors, including:

[0182] The data acquisition and modeling module is used to build an edge computing-based digital twin environment for power storage. It collects multi-source heterogeneous data in real time via the OPC-UA protocol and builds a three-dimensional data correlation model that includes equipment ID, production batch, remaining service life, and maintenance cycle. Data sources include the Equipment Management System (EMS), which records the unique identifier (equipment ID) of power equipment, production batch, remaining service life, and maintenance cycle; and the Manufacturing Execution System (MES), which provides order status, emergency status, work order delivery time, and associated equipment group information. The Smart Grid Supervisory and Data Acquisition (SCADA) system collects grid operation data, such as the correlation of grid faults (including the number of affected sites) and the geographic location of substations. The Warehouse Management System (WMS) collects basic storage location information (location number, load capacity) and dynamic parameters (real-time space utilization, load factor, and distance from the outlet). Edge computing performs data cleaning, noise removal, and format conversion at nodes close to the data collection source, integrating the data into a unified three-dimensional model input.

[0183] The state space construction module extracts a 28-dimensional state vector based on a three-dimensional data association model and constructs a reinforcement learning state space that incorporates equipment-location-order coupling constraints. The state space construction module is the input core of the reinforcement learning model, extracting multidimensional state vectors and transforming the dynamic warehouse allocation problem into a mathematical representation of reinforcement learning. The 28-dimensional state vector includes the equipment state vector, the location state vector, the order state vector, and the system state vector. High-precision state space modeling ensures that the reinforcement learning algorithm accurately models the warehouse system, enabling rapid adaptation to the location allocation problem in dynamic environments. By integrating comprehensive state information, including equipment health, order delivery requirements, and location attributes, the model's decision-making is comprehensive and controllable.

[0184] The reward function calculation module is used to design a hierarchical reward function that integrates the basic layer, strategy layer, and emergency layer, and adjusts the weights of each layer in real time through the dynamic weight adaptation module. The calculation formula of the hierarchical reward function satisfies:

[0185]

[0186] in, As the base layer, For the strategy layer, For the emergency layer, 、 as well as are the weight coefficients of each layer. Basic-layer rewards improve outbound efficiency, with indicators including: handling cost (handling distance × equipment weight), outbound time deviation rate, and increased storage space utilization. Strategy-layer rewards improve the long-term rationality of storage planning, such as rewards for matching equipment life with storage planning cycles. Emergency-layer rewards: Ensure timely responses to urgent orders, such as increasing positive rewards when successfully allocated to a storage location near the exit. The layered reward function optimizes different objectives in modules, ensuring a dynamic balance between "efficiency-planning-emergency," thereby improving the overall operational efficiency and flexibility of the warehouse.

[0187] The reinforcement learning training module uses an improved proximal policy optimization algorithm to train the reinforcement learning model. The strategy is pre-tested and verified in a digital twin environment. The policy network utilizes three depthwise separable convolutional layers combined with a bidirectional LSTM network. The value network outputs multi-objective value assessments and uses the Huber loss function to mitigate the impact of outliers to ensure model stability. The digital twin environment provides a virtual simulation environment, enabling continuous validation and optimization of the algorithm under different strategies, effectively improving its robustness. The improved optimization strategy addresses the problem of high-dimensional feature extraction in storage allocation, improving the model's responsiveness to real-time conditions.

[0188] The cross-system collaboration module integrates data from the equipment management system, production execution system, and smart grid monitoring system. This module leverages a three-tiered rules engine, comprised of strategic, tactical, and execution layers, to enable collaborative decision-making regarding grid maintenance plans, production work order changes, and equipment failures. This cross-system collaboration improves overall supply chain responsiveness, while feedback from the three-tiered collaboration enables adaptability, such as real-time closed-loop adjustments to fluctuating order demand data.

[0189] The decision execution module is used to generate storage allocation decisions in real time based on the trained reinforcement learning model, execute storage adjustments after digital twin rehearsal, and trigger the linkage response of the logistics system and production system.

[0190] The closed-loop evolution module continuously optimizes strategies through an hourly, daily, and monthly closed-loop evolutionary mechanism. This mechanism involves adjusting action parameters based on real-time rewards, updating weights based on accumulated rewards, and reconstructing the network structure based on historical strategy performance. Hourly, this module adjusts storage allocation parameters based on real-time rewards, such as calculating the average response time of current requests and adjusting storage allocation weights in real time. Daily, it automatically updates tier weights based on accumulated rewards. Monthly, it reconstructs the network structure based on strategy performance, such as adding convolutional layers to handle newly added storage characteristics. This closed-loop evolution enables the system to self-correct, significantly improving its intelligence.

[0191] Through the synergistic effect of the above modules, this system can form a closed-loop optimization process from data collection, model training to decision execution, showing advantages such as improved warehouse utilization, accelerated order response, reduced maintenance costs and improved intelligence level.

[0192] Corresponding to the above embodiment, the present invention further provides an electronic device.

[0193] like Figure 8 FIG2 is a schematic diagram of the structure of an electronic device according to the present invention. The electronic device 200 includes a processor 201 and a memory 203. The processor 201 and the memory 203 are connected, for example, via a bus 202. Optionally, the electronic device 200 may further include a transceiver 204. It should be noted that in actual applications, the number of transceivers 204 is not limited to one, and the structure of the electronic device 200 does not constitute a limitation on the embodiments of the present invention.

[0194] Processor 201 may be a CPU (central processing unit), a general-purpose processor, a DSP (digital signal processor), an ASIC (application-specific integrated circuit), an FPGA (field programmable gate array), or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. Processor 201 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, or a combination of a DSP and a microprocessor.

[0195] The bus 202 may include a path for transmitting information between the above components. The bus 202 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industrial Standard Architecture) bus. The bus 202 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0196] Memory 203 is used to store a computer program corresponding to the method for dynamic storage allocation of power storage integrating multiple dimensional factors according to the above-mentioned embodiment of the present invention. The computer program is controlled and executed by processor 201. Processor 201 is used to execute the computer program stored in memory 203 to implement the content of the above-mentioned method embodiment.

[0197] The electronic device 200 includes, but is not limited to, mobile terminals such as laptop computers and PADs (tablet computers), and fixed terminals such as desktop computers. Figure 8 The electronic device 200 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0198] The electronic device 200 of this embodiment of the present invention integrates edge computing, digital twin technology, and reinforcement learning algorithms, collaborating with multiple systems to create an intelligent dynamic storage allocation solution. This device effectively addresses issues such as heterogeneous data sources, static storage allocation, and delayed emergency response in power storage management. By collecting multidimensional data in real time and building a three-dimensional correlation model, it accurately extracts the status characteristics of equipment, orders, and storage locations. It also utilizes a hierarchical reward function to dynamically balance efficiency and safety requirements, enabling intelligent optimization and decision-making. This significantly improves the intelligence and automation of power storage management.

[0199] It should be noted that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, automate, propagate, or transmit a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing it in a suitable manner if necessary, and then stored in a computer memory.

[0200] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0201] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0202] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0203] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A method for dynamic storage allocation of power storage integrating multi-dimensional factors, characterized in that: The following steps are involved: S1. Build a digital twin environment for power storage based on edge computing, collect multi-source heterogeneous data in real time through the OPC-UA protocol, and establish a three-dimensional data association model; S2. Extracting a 28-dimensional state vector based on the three-dimensional data association model and constructing a reinforcement learning state space containing equipment-storage location-order coupling constraints; S3. Design a hierarchical reward function that integrates the basic layer, strategy layer, and emergency layer, and adjust the weights of each layer in real time through a dynamic weight adaptive module. The basic layer associates the outbound time deviation rate, handling cost, space utilization, and storage load balance index. The strategy layer associates the matching degree between the remaining service life of the equipment and the storage planning cycle, the association group continuity index, and the high-load storage risk value. The emergency layer associates the emergency order response coefficient. The association group continuity index is a statistical measure of the continuous storage ratio of the same order equipment group in the storage location. S4. Use an improved proximal policy optimization algorithm to train the reinforcement learning model and use the digital twin environment to perform strategy preview verification. In the improved proximal policy optimization algorithm: The policy network uses three layers of deep separable convolutional layers to process storage space features and a bidirectional LSTM to process order time series data; The value network outputs multi-objective value estimates and uses the Huber loss function to reduce the impact of outliers; S5. Generate storage allocation decisions in real time based on the trained reinforcement learning model, and continuously optimize the strategy through an hourly, daily, and monthly closed-loop evolutionary mechanism.

2. The method according to claim 1, characterized in that The multi-source heterogeneous data includes: Equipment ID, remaining service life and maintenance cycle of the equipment management system; The order urgency indicator, associated equipment group size, and work order delivery time of the production execution system; Grid fault correlation and substation location of smart grid monitoring system; The storage space load rate, space utilization rate and distance from the exit of the warehouse management system; The construction of the three-dimensional data association model satisfies: Equipment storage priority = α × maintenance urgency + β × production demand urgency + γ × grid fault correlation; Among them, α, β, and γ are the warehouse priority weight coefficients and α+β+γ=1; the maintenance urgency is determined by the difference between the remaining service life of the equipment and the standard maintenance cycle; the production demand urgency is determined by the difference between the work order delivery time and the current time; the power grid fault correlation is determined by the number of substations affected by the fault.

3. The method according to claim 1, characterized in that The 28-dimensional state vector includes: Equipment state vector: remaining service life of the equipment , maintenance cycle weight , Historical delivery frequency , Real-time power of the device and equipment failure warning level ; Storage state vector: load-bearing rate , space utilization , Standardized value of distance from exit , storage height coefficient , Storage Continuity Index 、The partition to which the storage location belongs , storage location historical accident rate , temperature sensitivity , humidity sensitivity And moisture-proof status ; Order status vector: urgent order flag , associated device group size , Estimated delivery time window , Special qualification requirements for orders , Order multiple batch split mark And the environmental protection requirement level of the order ; System state vector: Shared storage ratio across storage areas , handling equipment busyness index , system real-time energy consumption , Number of security alarms in the warehouse area , storage management system load rate , Historical strategy reuse rate and channel congestion coefficient .

4. The method according to claim 1, wherein The calculation formula of the hierarchical reward function is: in, 、 、 Respectively, the reward weight coefficients of the basic layer, strategy layer, and emergency layer meet + + =1; Base layer rewards The calculation formula is: Where, is the deviation rate of the delivery time, that is, the ratio of the actual delivery time to the predicted delivery time; is the transportation cost, which is equal to the product of the transportation distance and the weight of the equipment; is space utilization; The storage load balance index is defined as the standard deviation of the storage load rate of each storage area, which is used to balance the load-bearing safety requirements in the storage of power equipment; Strategy layer rewards The calculation formula is: Where, The matching degree between the remaining service life of the equipment and the storage planning period, is the continuity index of the association group, is the high load storage position risk value; Emergency Level Rewards The calculation formula is: Where, is the emergency order response coefficient, which is 1 when the emergency order is successfully allocated to the storage location near the export.

5. The method according to claim 1, wherein The dynamic weight adaptation module updates the weights by gradient descent: in, is the weight vector at the tth iteration, which is essentially the reward weight coefficient of the base layer, strategy layer, and emergency layer at the tth iteration 、 、 Dynamic set of ; T is transpose; is the current iteration number; is the adaptive learning rate; is the historical average reward; is the hierarchical reward function Weight vector The gradient vector of The adaptive learning rate The dynamic adjustment formula is: in, is the initial learning rate, ranging from 0.01 to 0.05; is the maximum learning rate; is the maximum number of iterations.

6. The method according to claim 1, characterized in that The loss function expression is: in, is the target value, is the predicted value, is the threshold.

7. The method according to claim 1, characterized in that The closed-loop evolution mechanism is implemented as follows: Hourly level: Adjust storage allocation action parameters based on real-time rewards; Daily level: Update the reward weight coefficients of the base layer, strategy layer, and emergency layer based on the accumulated rewards in 24 hours ; Monthly level: Reconstruct the strategy network and value network structure based on historical strategy performance.

8. The method according to claim 1, characterized in that It also includes a cross-system collaborative decision-making step, which integrates data from the equipment management system, production execution system, smart grid monitoring system, and warehouse management system to achieve collaborative processing of power grid maintenance plans, production work order changes, and equipment failure events; The cross-system collaborative decision-making is achieved through a three-layer linkage rule engine: Strategic layer: When the smart grid monitoring system releases a grid maintenance plan, backup equipment is retrieved 72 hours in advance and allocated to a storage location ≤5 meters from the outlet; Tactical level: When the production execution system's work order demand fluctuates by more than 20%, the ratio of raw materials to finished products storage is dynamically adjusted; Execution layer: When a device failure is detected, the backup device is marked as "priority shipment" within 10 seconds. The decision priority calculation formula is: The event urgency is an integer between 1 and 10, the impact range is the number of substations affected by the fault, and the response time window is the remaining time for fault processing.

9. The method according to claim 8, characterized in that The feedback loop of the three-layer linkage rule engine includes: The execution layer feeds back the outbound time deviation rate ΔT to the warehouse management system to optimize the transportation path; The tactical layer feeds back to the production execution system the impact coefficient of storage adjustment on production efficiency: in, is the rate of change of transport distance; is the production time change rate; 、 is the tactical layer efficiency impact weight coefficient and ; The strategic layer feeds back the safety stock coefficient to the enterprise resource planning system, and the safety stock coefficient satisfies: in, It is the benchmark coefficient, with a value range of 1.0-1.

5. The number of successful collaborations is the number of times the efficiency is improved by more than 10% after linkage execution.

10. A dynamic storage allocation system for power storage integrating multi-dimensional factors, characterized by: comprising a module for implementing the method according to any one of claims 1 to 9, wherein: The data acquisition and modeling module is used to build a digital twin environment for power storage based on edge computing. It collects multi-source heterogeneous data in real time through the OPC-UA protocol and establishes a three-dimensional data association model that includes equipment ID, production batch, maintenance cycle, and grid location. a state space construction module for extracting a 28-dimensional state vector based on the three-dimensional data association model to construct a reinforcement learning state space containing equipment-storage-order coupling constraints, wherein the 28-dimensional state vector includes an equipment state vector, a storage location state vector, an order state vector, and a system state vector; The reward function calculation module is used to design a hierarchical reward function that integrates the basic layer, policy layer, and emergency layer, and adjusts the weights of each layer in real time through the dynamic weight adaptation module; A reinforcement learning training module is used to train the reinforcement learning model using an improved proximal policy optimization algorithm and conduct policy preview verification using a digital twin environment. The policy network uses three layers of deep separable convolutional layers combined with a bidirectional LSTM network. A cross-system collaboration module integrates data from the equipment management system, production execution system, and smart grid monitoring system to enable collaborative decision-making on grid maintenance plans, production work order changes, and equipment failure events through a three-tiered linkage rules engine consisting of strategic, tactical, and execution layers. The decision execution module is used to generate storage allocation decisions in real time based on the trained reinforcement learning model, execute storage adjustments after digital twin preview, and trigger the linkage response of the logistics system and production system; A closed-loop evolution module is used to continuously optimize strategies through an hourly, daily, and monthly closed-loop evolution mechanism. The closed-loop evolution mechanism includes real-time rewards to adjust action parameters, cumulative rewards to update weights, and historical strategy performance to reconstruct the network structure.

Citation Information

Patent Citations

  • Deep learning cross-border e-commerce warehouse management system

    CN119671210A

  • Adversarial environment reinforcement learning model training method and system based on digital twinning

    CN120068991A