Storage resource adaptive collaborative distribution system based on multi-time-scale deep reinforcement learning

The adaptive collaborative allocation system of warehouse resources based on multi-time-scale deep reinforcement learning solves the problems of low efficiency and poor adaptability in traditional warehouse resource allocation methods, realizes adaptive collaborative allocation of resources, and improves the efficiency and flexibility of warehouse operations.

CN120672261AActive Publication Date: 2025-09-19LONGYAN UNIV +2

Patent Information

Application Number
CN202511171009.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-09-19
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

Traditional warehouse resource allocation methods rely on manual experience or simple rule-based algorithms, which are unable to cope with dynamic changes in large-scale, high-frequency operation scenarios, resulting in the coexistence of equipment idleness and overuse. In addition, existing intelligent means lack multi-time scale data processing and feature extraction, and cannot achieve adaptive and coordinated resource allocation.

Method used

The adaptive collaborative allocation system of warehouse resources adopts multi-time-scale deep reinforcement learning. Through data collection, state feature extraction, reinforcement learning state space construction and resource allocation strategy generation, combined with strategy adaptive optimization, it realizes the adaptive collaborative allocation of warehouse resources.

Benefits of technology

It improves warehousing operation efficiency, reduces resource waste, enhances the warehousing system's ability to cope with complex business scenarios, and achieves rational allocation and efficient utilization of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672261A_ABST
    Figure CN120672261A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of warehouse resource collaborative allocation, and discloses a warehouse resource adaptive collaborative allocation system based on multi-time scale deep reinforcement learning, and the system comprises a warehouse data collection unit which collects and standardizes equipment operation data and environmental parameters; the storage resource state feature extraction unit is used for extracting resource state features to form abnormal features; the reinforcement learning state space construction unit is used for analyzing features to form state space construction parameters; the resource allocation strategy generation unit is used for constructing a model generation strategy and dividing influence factors; and the strategy self-adaptive optimization unit is used for updating the optimization strategy by using an algorithm. In addition, a multi-source information fusion unit and a task queue dynamic adjustment module are arranged. The system realizes self-adaptive collaborative distribution of storage resources, improves storage operation efficiency, adapts to dynamic changes of a storage system, and meets intelligent operation requirements of modern storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of collaborative allocation of warehouse resources, and specifically to a warehouse resource adaptive collaborative allocation system based on multi-time-scale deep reinforcement learning. Background Art

[0002] With the booming e-commerce industry and the rapid advancement of the logistics sector, the scale and complexity of warehousing operations are increasing. Warehousing operations involve the coordinated operation of multiple types of equipment, such as automated guided vehicles (AGVs), stacker cranes, and forklifts. These operations also require consideration of multiple aspects, including cargo storage, handling, and sorting. This places extremely high demands on the rational allocation of warehousing resources. Traditional warehouse resource allocation methods mostly rely on manual experience or simple rule-based algorithms. When allocating resources manually, warehouse managers make decisions about equipment assignments and storage locations based on past experience. However, in large-scale, high-frequency operations, manual judgment is inefficient and prone to errors. For example, during shopping festivals, when order volumes surge, manual allocation struggles to quickly and accurately schedule equipment and storage locations. This can easily lead to a mix of idle equipment and overutilization, with some storage locations becoming overcrowded and others underutilized. While simple rule-based algorithms can improve allocation efficiency to a certain extent, they lack adaptability to the dynamic changes in warehouse systems. These algorithms are typically based on fixed parameters and logic, unable to perceive real-time factors such as equipment operating status, environmental changes, and task priority adjustments. When equipment malfunctions, storage conditions change, or urgent tasks are added to the warehouse, resource allocation plans based on simple rule-based algorithms often fail to adjust in a timely manner, resulting in reduced warehouse efficiency and significant resource waste. Some existing warehouse resource allocation technologies have also attempted to introduce intelligent means, such as using a single machine learning algorithm for resource scheduling. However, most of these methods only consider resource allocation problems under a single time scale and cannot fully handle the dynamic changes in different time dimensions in the warehouse system. Task cycles, equipment maintenance cycles, cargo storage cycles, etc. in warehouse operations all have different time scales. Algorithms with a single time scale have difficulty coordinating resource requirements under these different time scales and achieving optimal resource allocation. In addition, existing technologies often lack deep mining and effective integration of multi-dimensional data in terms of data processing and feature extraction, and are unable to fully extract key features that reflect the status of warehouse resources. As a result, the formulation of resource allocation strategies lacks comprehensive and accurate data support, making it difficult to achieve adaptive and coordinated allocation of warehouse resources and unable to meet the needs of efficient and intelligent operations in modern warehouses. Summary of the Invention

[0003] The purpose of the present invention is to provide a warehouse resource adaptive collaborative allocation system based on multi-time-scale deep reinforcement learning to solve the problems raised in the above background technology.

[0004] To achieve the above objectives, the present invention provides the following technical solution: a warehouse resource adaptive collaborative allocation system based on multi-time-scale deep reinforcement learning, the system comprising: The warehouse data collection unit is used to collect real-time operating data and environmental parameters of various types of equipment in the warehouse scene, and standardize the data to form structured data in a unified format; A warehouse resource status feature extraction unit is used to use the structured data in a unified format and the multi-dimensional resource status feature extraction method in the warehouse management center to complete the resource collaborative allocation task within the warehouse space and extract the resource status features of different warehouse locations and form abnormal features. The resource status features include: inventory turnover rate, storage space utilization rate, equipment idle rate, transportation path length, and task waiting time; The reinforcement learning state space construction unit is used to analyze and cluster the resource collaborative allocation tasks within the storage space and the resource state characteristics of different storage locations based on the abnormal characteristics of the unit time step. At the same time, it analyzes the device interaction frequency under the dynamic load of different storage locations to form the state space construction parameters of different storage location dynamic loads; A resource allocation strategy generation unit is used to model different storage locations in the storage space to form a resource allocation model, which includes a global strategy model based on deep reinforcement learning, a local optimization model, and a multi-objective collaborative model. Based on the resource allocation model and the structured data in a unified format, a resource allocation strategy is generated for the corresponding inventory scheduling, equipment collaboration, and path planning behaviors, and the factors affecting resource allocation are classified to achieve adaptive collaborative resource allocation in the storage space. The strategy adaptive optimization unit is used to model the resource allocation strategy of the storage space and identify the allocation strategy through the adaptive collaborative allocation parameters of the storage space resources using the deep reinforcement learning algorithm, and to update the parameters of the deep reinforcement learning algorithm using the experience replay mechanism and policy gradient optimization.

[0005] Preferably, the warehouse data collection unit includes: Multi-source data acquisition module, used to collect real-time operating data and environmental parameters of multiple types of equipment generated by different storage locations in the storage space per unit time step; A data standardization module is used to standardize the real-time operation data and environmental parameters of the multiple types of equipment into corresponding resource parameters according to a preset Z-score transformation; The outlier processing module is used to process outliers on the resource parameters obtained by standardization to obtain interference-free observation data; outlier processing includes deletion, interpolation and statistical learning methods; A dynamic load source tracking module is used to track the interactive impact degree data of different devices from the non-interference observation data, and standardize the interactive impact degree data of different devices according to the dynamic load source of the storage location to obtain the dynamic load source information of the device that matches the dynamic load source of the storage location; The environment impact dynamic load information parsing module is used to perform environment impact analysis on the equipment dynamic load source information to form different environment dynamic load source parameter variables as the structured data with a unified format.

[0006] Preferably, the storage resource status feature extraction unit includes: The edge computing node adaptation module is used to preset edge computing node parameters, adapt to storage throughput and operation complexity, and establish adaptation statistics tables; The warehouse management center module is used to adjust the extraction steps and preset the extraction parameters of the multi-dimensional resource status feature extraction method in the warehouse management center; A multi-storage location resource status feature extraction module is used to extract periodic resource status features of different storage locations based on the multi-dimensional resource status feature extraction method, the structured data in a unified format, and the adaptation statistical table, and generate abnormal resource status features of different storage locations; The resource status feature extraction module is used for edge computing nodes to analyze the execution status of resource collaborative allocation tasks, to judge and analyze resource collaborative allocation tasks that have discovered resource status features, and to generate collaborative allocation resource status abnormal features. The collaborative allocation resource status abnormal features and the abnormal features of resource status at different storage locations form the abnormal features.

[0007] Preferably, the reinforcement learning state space construction unit includes: The storage resource status parameter parsing module is used to locate the storage location based on abnormal characteristics of the collaborative allocation resource status and abnormal characteristics of the resource status of different storage locations, including storage resource status parameter load parsing, storage resource status parameter environment parsing and storage resource status parameter abnormality location; The resource status parameter classification and processing flow module is used to classify and form the warehouse resource status parameter information in the abnormal characteristics of the collaborative allocation resource status and the abnormal characteristics of the resource status of different storage locations, and obtain the data processing flow corresponding to the warehouse resource status parameter information, forming the warehouse resource status parameter reference benchmark, the coefficient to be optimized of the warehouse resource status parameter calculation model, and the noise type; The dynamic load module is used to detect the equipment interaction of the full-time-step dynamic load of different storage locations, and statistically analyze the sources of the dynamic load of equipment in different storage locations to form the state space construction parameters of the dynamic load of different storage locations.

[0008] Preferably, the resource allocation strategy generating unit includes: Resource allocation model management module, used to build and manage global strategy models, local optimization models, and multi-objective collaborative models based on deep reinforcement learning; The allocation strategy generation module is used to determine the resource allocation security of different storage locations, systems, and management terminals based on the real-time operation data and environmental parameters of multiple types of equipment, the global strategy model, the local optimization model, and the multi-objective collaborative model, using a preset state coding mechanism and strategy generation algorithm to obtain allocation strategy parameters; An operation difficulty identification module is used to quantitatively evaluate the operation difficulty of different storage locations based on the resource allocation security identification parameters and the dynamic load source of the equipment operation complexity and obtain the operation difficulty identification parameters; A resource allocation influencing factor classification module is used to automatically complete the resource allocation security identification of different storage locations, systems and management terminals based on the allocation strategy parameters and operation difficulty identification parameters, and to form a record of resource allocation influencing factors; The allocation process visualization module is used to visualize the system resource allocation process.

[0009] Preferably, the strategy adaptive optimization unit includes: The operation status recording module is used to continuously monitor the working status of different storage locations and the dynamic load source of the system equipment operation complexity, forming a record of the status of different storage locations and the operation complexity of the storage system, providing a reference basis for subsequent resource allocation and coordinated scheduling; The strategy effectiveness evaluation module is used to determine the resource allocation security of different storage locations to obtain resource allocation security test parameters, evaluate the duration of different storage locations or allocation strategy levels, and evaluate the allocation accuracy of different storage locations for future resource allocation tasks; The reinforcement learning network optimization module is used to allocate resources based on the security of different storage locations, while combining the accuracy of different storage locations for future resource allocation tasks, and using the experience replay mechanism and policy gradient optimization to update the parameters of the deep reinforcement learning algorithm.

[0010] Preferably, the system further comprises a multi-source information fusion unit for integrating equipment perception data, management system instruction data and external environment data in the storage space to form a fusion information set; The multi-source information fusion unit includes a time synchronization module for aligning the timestamps of data from different sources; a space calibration module for calibrating the spatial position of device perception data based on storage location coordinate information; and a conflict resolution module for identifying conflicting information from different data sources and fusing them through a voting mechanism or credibility weights. The fused information set serves as a supplementary input for the structured data in a unified format.

[0011] Preferably, the multi-objective collaborative model includes an inventory balancing sub-model, an equipment load balancing sub-model and a path overlap avoidance sub-model; the inventory balancing sub-model is used to dynamically adjust the cargo location allocation ratio so that the inventory level of each storage location approaches a preset mean; the equipment load balancing sub-model is used to make the difference in operating loads of different equipment less than a threshold through a task allocation strategy; the path overlap avoidance sub-model is used to reduce the probability of cross-overlap of equipment transportation paths through a path planning algorithm.

[0012] Preferably, the resource allocation strategy generation unit also includes a task queue dynamic adjustment module, which is used to dynamically adjust the generated task execution sequence according to the real-time resource status and strategy generation parameters; the task queue dynamic adjustment module includes a state perception submodule, which is used to obtain the inventory status of each storage location, the equipment operation status and the task completion progress in real time; an impact assessment submodule, which is used to analyze the impact of the current task queue state change on the task sequence execution efficiency; and a sequence reordering submodule, which is used to locally or globally reorder the task queue according to the impact assessment results.

[0013] Preferably, the local reordering or global reordering operation preserves the execution order of high priority tasks.

[0014] Compared with the prior art, the present invention has the following beneficial effects: The multi-timescale deep reinforcement learning-based adaptive collaborative warehouse resource allocation system of this invention uses a warehouse data acquisition unit to collect real-time operating data and environmental parameters from multiple types of equipment, and then standardizes this data to form structured data, providing an accurate and unified data foundation for subsequent operations. The warehouse resource status feature extraction unit utilizes structured data and a multi-dimensional resource status feature extraction method to comprehensively and accurately extract resource status features such as inventory turnover rate and shelf utilization, and generate abnormal features, enabling the system to promptly identify potential problems in the warehouse resource allocation process. The reinforcement learning state space construction unit analyzes and clusters resource state characteristics based on anomaly signatures, while also analyzing device interaction frequencies to form state space construction parameters. This provides input tailored to the warehouse's actual operating conditions for the deep reinforcement learning algorithm. The resource allocation strategy generation unit constructs a global strategy model, a local optimization model, and a multi-objective collaborative model. This model combines structured data to generate resource allocation strategies and classifies influencing factors, achieving comprehensive optimization of inventory scheduling, device collaboration, and path planning. This allows for the rational allocation of resources based on the actual needs of different storage locations and the overall system objectives. The policy adaptive optimization unit updates the parameters of the deep reinforcement learning algorithm through the experience replay mechanism and policy gradient optimization, so that the system can continuously adapt to the dynamic changes of the warehousing environment and continuously optimize the resource allocation strategy. In addition, the multi-source information fusion unit integrates equipment perception data, management system instruction data and external environment data, further improving the integrity and accuracy of the data; the task queue dynamic adjustment module adjusts the task execution sequence according to the real-time resource status, ensuring the execution of high-priority tasks and improving the overall efficiency and flexibility of warehousing operations. This system forms a complete closed loop from data collection, feature extraction, model construction, strategy generation to strategy optimization, effectively solving the problems of low efficiency and poor adaptability of traditional warehousing resource allocation, realizing adaptive and coordinated allocation of warehousing resources, significantly improving warehousing operation efficiency, reducing resource waste, and enhancing the ability of the warehousing system to cope with complex business scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a working principle diagram of the multi-time-scale deep reinforcement learning storage resource adaptive collaborative allocation system of the present invention; Figure 2 Flow chart of data collection and standardization processing for warehouse data collection unit; Figure 3 Flowchart for feature extraction and exception generation of the warehouse resource status feature extraction unit; Figure 4 Flowchart of state parameter parsing and space construction for reinforcement learning state space construction units. DETAILED DESCRIPTION

[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0017] See also Figures 1-4 The multi-time-scale deep reinforcement learning adaptive collaborative allocation system for storage resources of the present invention is specifically implemented as follows: During actual operation, the warehouse data collection unit collects real-time operating data and environmental parameters of multiple types of equipment in the warehouse scene. These data include, but are not limited to, the operating trajectory of forklifts, the working status of stackers, the temperature and humidity in the warehouse, and other information. After collecting the data, the data is standardized through the preset Z-score transformation, and the real-time operating data and environmental parameters of multiple types of equipment are converted into corresponding resource parameters. At the same time, the deletion method, interpolation method and statistical learning method are used to process the outliers of the standardized resource parameters to obtain non-interference observation data, and the degree of interaction between different devices is tracked from the non-interference observation data. After standardization, the dynamic load source information of the equipment that matches the dynamic load source of the storage location is obtained, and then the environmental impact analysis is performed to form structured data with a unified format. After the warehouse resource status feature extraction unit obtains structured data in a unified format, the edge computing node adaptation module pre-sets the edge computing node parameters to adapt it to the warehouse throughput and operation complexity, and establishes an adaptation statistical table. The warehouse management center module adjusts the extraction steps of the multi-dimensional resource status feature extraction method and presets the extraction parameters. The multi-storage location resource status feature extraction module performs periodic resource status feature extraction on different storage locations based on the multi-dimensional resource status feature extraction method, structured data, and adaptation statistical table, and generates abnormal resource status features for different storage locations; the resource status feature extraction module analyzes the execution status of the resource collaborative allocation task through the edge computing node, judges and analyzes the resource collaborative allocation task that discovers the resource status feature, and generates abnormal resource status features for the collaborative allocation resource. The two together form abnormal features. The reinforcement learning state space construction unit analyzes and clusters the resource collaborative allocation tasks and resource status characteristics of different storage locations within the storage space based on the abnormal characteristics of the unit time step. The storage resource status parameter analysis module locates the cargo location based on the abnormal characteristics of the collaborative allocation resource status and the abnormal characteristics of the resource status of different storage locations, including storage resource status parameter load analysis, storage resource status parameter environment analysis, and storage resource status parameter abnormality location. The resource status parameter classification and processing flow module classifies and forms the storage resource status parameter information in the abnormal characteristics, obtains the corresponding data processing flow, and forms the storage resource status parameter reference benchmark, the coefficients to be optimized of the storage resource status parameter calculation model, and the noise type. The dynamic load module detects the device interaction of the full time step dynamic load of different storage locations, statistically analyzes the source of the dynamic load of the equipment in different storage locations, and forms the state space construction parameters of the dynamic load of different storage locations. The resource allocation strategy generation unit models different storage locations within the storage space, constructing a global strategy model, a local optimization model, and a multi-objective collaborative model based on deep reinforcement learning. Based on the real-time operating data and environmental parameters of multiple types of equipment and various models, a preset state encoding mechanism and strategy generation algorithm are used to determine the resource allocation security of different storage locations, systems, and management terminals, obtaining allocation strategy parameters. Based on the resource allocation security determination parameters and the dynamic load source of equipment operation complexity, the operation difficulty of different storage locations is quantitatively assessed to obtain operation difficulty determination parameters. Based on the allocation strategy parameters and operation difficulty determination parameters, the resource allocation security of different storage locations, systems, and management terminals is automatically determined, forming a record of resource allocation influencing factors and visualizing the system resource allocation process. The strategy adaptive optimization unit continuously monitors the working status of different storage locations and the dynamic load sources of system equipment operation complexity, and generates records. It determines the resource allocation security of different storage locations, obtains resource allocation security test parameters, and evaluates the duration of different storage locations or allocation strategy levels, as well as the allocation accuracy of different storage locations for future resource allocation tasks. Based on the resource allocation security of different storage locations and the allocation accuracy of different storage locations for future resource allocation tasks, the deep reinforcement learning algorithm parameters are updated using the experience replay mechanism and policy gradient optimization.

[0018] Example 1: The warehouse data acquisition unit completes data collection and processing through the collaboration of multiple modules. The multi-source data acquisition module collects real-time operational data and environmental parameters generated by various types of equipment at different storage locations within the warehouse space at a fixed time step. At the warehouse operation site, various types of equipment, such as automated guided vehicles (AGVs), cranes, and forklifts, generate a large amount of data during their operation. AGV operational data includes driving speed, driving direction, current location coordinates, and cargo loading status; crane operational data includes lifting height, operating trajectory, hook status, and other data; and forklifts generate data such as cargo weight, mileage, and operating hours. Regarding environmental parameter collection, real-time information such as light intensity, air pressure, temperature, humidity, and air quality within the warehouse is collected. This data is aggregated to the multi-source data acquisition module from various sensors at different storage locations, built-in monitoring systems on the equipment, and other channels.

[0019] After collecting the raw data, the data standardization module processes it. Data standardization is based on preset rules and methods, with the goal of converting data of different types and magnitudes into a unified, comparable format. Each data dimension is processed according to a predetermined conversion method, so that data with different metrics and value ranges can be put under the same standard for subsequent analysis and use. For example, for equipment speed data, some equipment speeds may be in kilometers per hour, while others may be in meters per second. The data standardization module will convert them into the same unit and map the data to a reasonable range based on a specific mapping relationship.

[0020] The outlier processing module operates on standardized data. The deletion method is used to address data that significantly deviates from the normal range. In actual warehouse operations, extreme data may occur due to sensor failures, data transmission errors, and other factors. For example, if an AGV's speed record shows a value far exceeding its normal operating speed range, this data will be directly eliminated by the deletion method. Interpolation methods are suitable for handling situations where data is missing or there are minor anomalies. When equipment operating data within a certain time period is missing, or some data contains minor anomalies, interpolation methods use reasonable calculations based on the distribution of adjacent data to estimate and fill in the missing or abnormal data. For example, if a crane's lifting height data is missing for several time points within a certain period, interpolation methods use linear interpolation or other appropriate interpolation algorithms based on the lifting height data at the preceding and subsequent time points to calculate a reasonable estimate of the missing data. Statistical learning methods construct data distribution models and conduct in-depth data analysis to identify and address outliers. It will learn the distribution patterns of data and establish a mathematical model. When new data comes in, it will use the model to determine whether the data belongs to the normal distribution range. If not, it will be identified as an outlier and processed accordingly to obtain interference-free observation data.

[0021] The dynamic load source tracking module obtains information from non-interference observation data. It tracks the degree of interaction between different devices by analyzing information such as the strength and frequency of interaction signals between devices. During warehouse operations, devices frequently interact with each other, such as between AGVs and racks, or between cranes and forklifts. These interactions generate various signals. By capturing and analyzing these signals, the dynamic load source tracking module can understand the interactions between different devices and determine which interactions affect the dynamic load at a storage location, and to what extent. This data is then normalized based on the dynamic load sources at the storage location, generating device dynamic load source information that matches the dynamic load sources at that location. For example, if a storage location experiences a sudden increase in load during a specific time period, the dynamic load source tracking module may identify the cause as being from certain AGVs frequently handling goods at that location. The module then normalizes the relevant device interaction impact data to clarify the relationship between these devices and the dynamic load at that location.

[0022] The Environmental Impact Dynamic Load Information Parsing Module further analyzes the source information of equipment dynamic loads. It considers the impact of environmental factors on equipment operation, such as how temperature changes may affect equipment speed and performance, how humidity may affect cargo storage and metal parts of equipment, and how light intensity may affect operator efficiency and indirectly affect equipment operation. By analyzing the relationship between these environmental factors and equipment operation, the Environmental Impact Dynamic Load Information Parsing Module can analyze the environmental impact of equipment dynamic load source information, generating parameters for dynamic load sources in different environments. Ultimately, these parameters serve as structured data in a unified format, providing data support for subsequent tasks such as extracting warehouse resource status features. This enables the system to more rationally analyze and allocate warehouse resources based on this accurate and standardized data.

[0023] Example 2: During the actual operation of the warehouse resource status feature extraction unit, the edge computing node adaptation module begins to play a role during the system initialization phase. When the system starts, the staff presets the relevant parameters of the edge computing node based on the historical warehouse data accumulated in the past, covering information such as cargo throughput in different time periods, the complexity of various operations, and the load conditions of equipment operation, while combining rich operational experience. These parameters include but are not limited to the processing capacity of the edge computing node, such as the amount of data that can be processed per second; storage capacity, that is, the total amount of data that can be stored; and the priority of data processing. By reasonably setting these parameters, the edge computing node can adapt to the needs of warehouse operations of different scales and complexities. After the parameter preset is completed, the edge computing node adaptation module will create a detailed adaptation statistics table, which is used to record the processing effect of the edge computing node on various types of warehouse data under different parameter settings, such as processing time, processing accuracy, etc., so as to facilitate subsequent parameter adjustment and optimization.

[0024] The Warehouse Management Center module provides managers with a convenient human-computer interface. Within this interface, managers can flexibly adjust the extraction steps of the multi-dimensional resource status feature extraction method. For example, the feature extraction interval can be adjusted based on the actual warehouse operations. During periods of frequent goods inflow and outflow, the interval can be shortened to obtain resource status features more timely; during periods of relatively stable operations, the interval can be appropriately extended to reduce computing resource consumption. Managers can also set data filtering conditions, such as selecting only data from specific types of equipment or data within a specific time period, to improve the relevance and effectiveness of feature extraction. Furthermore, the module allows for preset extraction parameters, such as weight coefficients to determine the importance of different resource status features in the comprehensive evaluation, and thresholds to determine whether resource status is within the normal range. These settings enable the multi-dimensional resource status feature extraction method to better meet the actual needs of warehouse operations.

[0025] The multi-storage location resource status feature extraction module uses a multi-dimensional resource status feature extraction method, uniformly formatted structured data, and adaptive statistical tables to extract resource status features for each storage location according to a pre-set cycle. During the extraction process, several key indicators are calculated to assess the storage location's resource status. These include inventory turnover rate (the number of times inventory turns over within a certain period, reflecting the velocity of goods flow); location utilization rate (measuring the utilization of storage location space); equipment idle rate (the proportion of equipment that is not in use within a certain period); transport path length (recording the distance traveled during the transport process); and task wait time (the time between task issuance and execution). These calculated indicators are compared with normal ranges. If any indicator exceeds or falls below the normal range, the corresponding storage location's resource status is identified as abnormal, and different location resource status anomaly features are generated. For example, if the location utilization rate of a storage location is consistently below normal, indicating that the space is underutilized, the location is marked as abnormal.

[0026] The resource status feature extraction module leverages the powerful computing power of edge computing nodes to analyze the execution status of resource collaborative allocation tasks in real time. During task execution, a comprehensive assessment and analysis of the resource status features involved is performed. If a task execution time is excessively long, this may indicate resource misallocation or equipment failure. Irregular resource utilization, such as when some equipment is idle for extended periods while others are overly busy, generates a collaborative resource status anomaly feature. For example, in a cargo handling task, if improper path planning causes the AGV to travel an excessively long distance, causing the task execution time to exceed expectations, the resource status feature extraction module will identify this anomaly and generate the corresponding collaborative resource status anomaly feature. The resource status anomaly features for different storage locations and the collaborative resource status anomaly features together constitute a complete anomaly feature. These anomaly features provide important information for subsequent reinforcement learning state space construction and resource allocation strategy generation. This enables the system to make timely adjustments and optimizations to address issues arising in warehousing operations, achieving the rational allocation and efficient utilization of warehousing resources.

[0027] Example 3: The reinforcement learning state space construction unit, within the entire warehouse resource adaptive collaborative allocation system, undertakes the crucial task of converting warehouse resource state characteristics into state space construction parameters suitable for reinforcement learning processing. This unit is comprised of a warehouse resource state parameter parsing module, a resource state parameter classification and processing flow module, and a dynamic load module, which collaborate to accomplish this task.

[0028] The Warehouse Resource Status Parameter Parsing Module addresses anomalies in the status of collaboratively allocated resources and in the status of resources at different storage locations. In a real-world warehouse environment, various types of equipment are equipped with positioning devices. For example, automated guided vehicles (AGVs) are equipped with high-precision positioning sensors that enable real-time location information. When an anomaly is detected, the module first uses the equipment's positioning information, combined with task-related information such as the storage location of the goods and the starting and ending points of the handling task, to locate the anomalous location.

[0029] For load analysis of warehouse resource status parameters, such as AGVs, data such as load weight and operating time are analyzed. Load weight can be obtained using load cells installed on the AGV, while operating time is recorded by the system from the start to the completion of the AGV's mission. By analyzing this data, we can understand the load conditions of the equipment during different time periods and tasks. When analyzing the environment, we consider the impact of warehouse humidity on cargo storage. Humidity sensors are installed in the warehouse to monitor the ambient humidity in real time. Different types of cargo have different humidity requirements. When humidity exceeds the appropriate storage range, it may affect cargo quality and, in turn, the status of warehouse resources. By analyzing the relationship between humidity data and cargo storage status, we can assess the impact of environmental factors on resource status. During anomaly location, we combine device location information and timestamps to accurately determine the specific location and time of the anomaly. For example, if damaged cargo is detected in a storage location, the device's operating trajectory and time records can be used to determine whether the anomaly occurred during a specific handling process, pinpointing the specific device and time point.

[0030] The resource status parameter classification and processing flow module categorizes warehouse resource status parameter information from abnormal features. This is done by data type, such as numerical data (equipment operating speed, cargo weight, etc.), character data (equipment model, cargo name, etc.), and time data (task start time, equipment failure time, etc.). This is further subdivided based on data sources, such as sensor data and management system data. Develop corresponding data processing flows for different data types.

[0031] When forming a reference benchmark for storage resource status parameters, a statistical analysis method is used. Suppose that for the operating speed data of a certain device, a large number of data samples are collected over a period of time. , by calculating the average ,in represents the number of data samples, Indicates the Data samples are collected to obtain the average operating speed of the equipment, which is used as a reference benchmark. According to the changing trend of the data, the coefficients to be optimized in the warehouse resource status parameter calculation model are determined. For example, if it is found that the operating speed of the equipment gradually decreases with the increase of service life, a regression model is established to analyze the relationship between the two and determine the coefficients in the model in order to more accurately predict the operating status of the equipment. Identify the type of noise in the data. Common types of noise include random noise and periodic noise. Random noise is usually caused by slight errors in sensors, random interference in the environment, etc.; periodic noise may be related to the periodic operation of equipment, the periodic operation of the warehouse ventilation system, etc. By identifying the type of noise, appropriate filtering, denoising and other processing methods are adopted to improve the quality of the data.

[0032] The dynamic load module uses various sensors deployed throughout the warehouse space to monitor the dynamic load interactions of different storage locations in real time, over the entire time step. These sensors capture interaction signals between devices and record information such as the interaction time and content. In actual warehouse operations, frequent interactions occur between devices, such as cargo transfers between AGVs and stackers, and collaborative operations between cranes and forklifts. The dynamic load module analyzes this interaction information to statistically analyze the sources of dynamic load on devices at different storage locations. When a large volume of cargo enters and leaves a storage location within a short period of time, causing multiple devices to frequently operate there and increase the load, the dynamic load module can identify the uneven load caused by this concentrated task distribution. If a device failure disrupts the entire operation process, forcing other devices to take on additional tasks and thus causing load variations, the module can accurately identify the device failure as the source of the increased load. By processing and analyzing this information, state space parameters for the dynamic load at different storage locations are generated, providing key data support for the subsequent generation of resource allocation strategies, enabling the system to formulate more appropriate resource allocation strategies based on the actual load conditions at each location.

[0033] Example 4: The resource allocation strategy generation unit is responsible for generating a specific resource allocation strategy in the warehouse resource adaptive collaborative allocation system, and its multiple modules work together to achieve this goal.

[0034] The resource allocation model management module builds various models based on a deep learning framework. In the scenario of a large e-commerce company's smart warehousing center, the warehouse contains thousands of storage locations, storing goods of varying categories and specifications, and dozens of automated guided vehicles (AGVs), stacker cranes, and other equipment operating collaboratively. The global strategy model takes a comprehensive approach to the entire warehouse system, comprehensively considering all resources within the warehouse. For example, during a shopping festival, when a large number of orders arrive, the global strategy model comprehensively plans the storage arrangements for all storage locations, the task allocation of various equipment, and the scheduling of personnel. This enables the entire warehouse system to efficiently cope with the surge in business volume, comprehensively allocates resources, and ensures stable and efficient system operation. Local optimization models address specific needs. For example, if a storage location is dedicated to storing perishable electronics, the local optimization model will optimize the placement of goods, storage density, and equipment routing within that location based on the location's space, temperature and humidity control conditions, and the storage requirements and inbound and outbound frequency of the electronics, thereby improving operational efficiency and storage security at that specific location.

[0035] The multi-objective collaborative model includes sub-models for inventory balancing, equipment load balancing, and path overlap avoidance, which work together in practice. The inventory balancing sub-model dynamically adjusts the allocation ratio of shelves to different storage locations based on different sales seasons. For example, in the summer, when cooling products like air conditioners and fans are in high demand, the inventory balancing sub-model increases the inventory allocation ratio for these products in the corresponding storage locations while simultaneously reducing the inventory ratio for products like winter clothing. This ensures that inventory levels in each storage location approach a preset average, preventing overstocking in some locations while others frequently need to be restocked due to shortages. The equipment load balancing sub-model allocates tasks based on the performance and current operating status of different equipment when handling cargo handling tasks. When a batch of goods needs to be moved from the storage area to the sorting area, if an AGV has been operating continuously for an extended period and is nearing its load limit, the equipment load balancing sub-model will allocate some of the handling tasks to other idle or less-loaded AGVs. This ensures that the load difference between different equipment is below the threshold, preventing equipment failure due to excessive fatigue and improving overall equipment lifespan and operational stability. The path overlap avoidance sub-model plays a key role in warehouses with densely populated equipment. When multiple AGVs are performing handling tasks simultaneously, the path overlap avoidance sub-model uses a path planning algorithm to plan an optimal path for each AGV, reducing the probability of cross-overlapping of equipment handling paths, avoiding collisions between AGVs, and improving the safety and smoothness of operations.

[0036] The allocation strategy generation module operates by integrating real-time operational data from multiple types of equipment, environmental parameters, and various models. Suppose, at a certain moment, the temperature in the warehouse rises, affecting the operating speed of some equipment. Simultaneously, some AGVs are running low on battery power, and a batch of urgent orders need to be processed. The allocation strategy generation module first uses a preset state encoding mechanism to convert these warehouse system status information, such as equipment operating speed, battery power, and task urgency, into an input format acceptable to each model. Then, using a strategy generation algorithm, taking into account the overall planning of the global strategy model, the requirements of the local optimization model for specific areas, and the objectives of the multi-objective collaborative model, it determines the security of resource allocation across different storage locations, systems, and management terminals, and derives allocation strategy parameters. For example, AGVs with sufficient battery power will be prioritized for handling urgent orders, and appropriate routes will be planned to ensure that tasks are completed while ensuring the safety of equipment and goods.

[0037] The Operation Difficulty Assessment Module assesses the difficulty of an operation based on resource allocation safety assessment parameters and the dynamic load source of equipment operation complexity. For bulky and heavy goods stored on high-rise shelves, handling them not only requires specialized large equipment but is also difficult and poses safety risks. The Operation Difficulty Assessment Module comprehensively considers factors such as the urgency of the task, the complexity of equipment operation, and potential risks to quantitatively assess the operation difficulty of the storage location. The assessment parameter is then used to provide a reference for subsequent resource allocation.

[0038] The resource allocation influencing factor classification module automatically completes resource allocation security assessments for different storage locations, systems, and management terminals based on allocation strategy parameters and operation difficulty assessment parameters. In the above scenario, analysis can determine factors influencing resource allocation, including equipment performance (such as AGV power and operating speed), task priority (urgent orders take precedence), and cargo characteristics (volume, weight, storage requirements). This generates a record of factors influencing resource allocation. The allocation process visualization module graphically displays the system resource allocation process. Through the visual interface, managers can intuitively observe changes in the AGV's operating path and the transfer of goods between storage locations, facilitating real-time monitoring and management of warehouse operations.

[0039] Example 5: During the actual operation of the strategy adaptive optimization unit, the operating status recording module assumes the important responsibility of continuous monitoring and data recording. In large-scale automated warehousing scenarios, the warehouse is distributed with a large number of equipment such as automated guided vehicles (AGVs), stackers, and sorting robots, as well as numerous storage locations with different functions. The operating status recording module uses sensors throughout the warehouse, such as displacement sensors, speed sensors, and current sensors installed on equipment, as well as inventory sensors deployed at storage locations, to continuously monitor the operating status of different storage locations and the dynamic load sources of system equipment operation complexity.

[0040] Taking AGV as an example, displacement sensors track its driving trajectory in real time, speed sensors provide feedback on changes in operating speed, and current sensors monitor the operating current of the motor. These data can intuitively reflect the operating status of the AGV, including whether it is driving normally, whether there are any abnormal jams, and whether the load is overloaded. For storage locations, inventory sensors sense changes in the storage quantity of goods in real time. Combined with inbound and outbound time information, the usage dynamics of the storage locations can be clearly understood. At the same time, by analyzing the interaction data between devices, such as the time and number of times the AGV and stacker hand over goods, as well as the task allocation records in the task scheduling system, the source of the dynamic load of the system equipment operation complexity can be clearly identified, such as whether the task volume is too large due to a surge in orders, or whether the load is unbalanced due to equipment failure. This information is recorded in real time to form a detailed operating status record file.

[0041] The strategy effectiveness evaluation module assesses the resource allocation safety of different storage locations based on pre-defined criteria. These criteria cover multiple dimensions, including equipment operational safety, such as whether there is a risk of collision or exceeding safe load limits during task execution; cargo storage safety, such as whether cargo is in a suitable storage environment and whether there is any risk of damage due to improper storage; and task execution timeliness, namely, whether tasks are completed within the specified timeframe. By comparing and analyzing these criteria one by one, the resource allocation safety test parameters are derived.

[0042] For example, during a certain period of time, when evaluating a storage location for precision instruments, it was discovered that the shelf storing the instrument was at risk of tipping over due to the adjacent goods being stacked too high, which reduced the resource allocation safety score of the storage location. However, for another storage location for conventional goods, if all tasks can be completed on time, the equipment is operated in accordance with regulations, and the goods are stored in good condition, a higher safety score will be given. At the same time, the strategy effectiveness evaluation module will also analyze the duration of different storage locations or allocation strategy levels, and understand the stability of the current allocation strategy by counting the duration of different safety levels. In addition, combining historical data and current storage status, it evaluates the accuracy of the allocation of future resource allocation tasks to different storage locations, and determines whether the current resource allocation model can effectively respond to possible future task requirements.

[0043] The reinforcement learning network optimization module optimizes the deep reinforcement learning algorithm based on data provided by the operation status recording module and the strategy effectiveness evaluation module. If the system finds that the resource allocation security of a certain storage location is low and has remained unstable for a period of time, the reinforcement learning network optimization module will combine the allocation accuracy evaluation results of the storage location for future resource allocation tasks and use the experience replay mechanism to randomly extract relevant experience data from the historical operation status record file. This includes information such as the operation status, task execution status, and resource utilization efficiency under different resource allocation strategies.

[0044] This empirical data is reused to train the deep reinforcement learning algorithm, enabling the algorithm to learn from past experience and avoid repeating incorrect allocation strategies. At the same time, the policy gradient optimization method is used to adjust and update the parameters in the deep reinforcement learning algorithm based on the goals of current resource allocation security and future allocation accuracy. For example, the allocation weight parameters of different resource types in the global policy model are adjusted, or the optimization policy parameters for specific storage locations in the local optimization model are modified. This allows the resource allocation strategy generated by the deep reinforcement learning algorithm to better adapt to the ever-changing dynamic environment of the warehousing system, thereby improving the performance and efficiency of the entire adaptive collaborative allocation system for warehousing resources.

[0045] Furthermore, the system's multi-source information fusion unit also plays a crucial role. The time synchronization module aligns the timestamps of data from various channels, including device sensors, management systems, and external environmental monitoring devices. For example, the frequency at which device sensors collect data may differ from the frequency at which the management system records task times. The time synchronization module unifies this data to the same time base, ensuring temporal consistency and avoiding data confusion caused by time discrepancies. The spatial calibration module calibrates the spatial position of device sensor data based on the precise coordinate information of storage locations.

[0046] When there is a deviation between the AGV's positioning data and the actual location of the storage location, the spatial calibration module will correct the AGV's location information based on the warehouse layout and coordinate system to ensure the accuracy of the equipment's location data and provide a reliable spatial reference for resource allocation. When the conflict resolution module identifies conflicting information from different data sources, such as when the quantity of goods fed back by the equipment sensor is inconsistent with the quantity recorded by the management system, it will fuse the conflicting data through a voting mechanism or credibility weights. If the data from multiple sensors confirms each other, but the management system data contradicts it, the sensor data will be given a higher degree of credibility, and fusion will be performed based on the sensor data. The integrated fusion information set will be used as a supplementary input for structured data in a unified format, providing more comprehensive and accurate data support for each unit of the system.

[0047] The task queue dynamic adjustment module within the resource allocation strategy generation unit and the state perception submodule continuously monitor the inventory status, equipment operating status, and task completion progress of each storage location through a real-time communication interface. During a period of concentrated order processing, the state perception submodule can quickly obtain information such as the current location, remaining battery life, operational status, and changes in the inventory level at each storage location, as well as the number of pending tasks. Based on this real-time data, the impact assessment submodule applies a specific evaluation algorithm to analyze the impact of current task queue status changes on the efficiency of task sequence execution.

[0048] If a critical device suddenly fails, the impact assessment submodule will evaluate the scope and extent of the failure's impact on the efficiency of the entire task sequence execution based on the type of task the failed device is responsible for, the task priority, and the substitutability of other devices. The sequence rescheduling submodule will locally or globally reorder the task queue based on the impact assessment results. For high-priority tasks, such as cargo sorting tasks for urgent orders, the sequence rescheduling submodule will prioritize ensuring that their execution order is not affected. By adjusting the execution order of other low-priority tasks and replanning the equipment's operation path and task allocation, the module can optimize the warehouse operation process and ensure that the warehouse system can still operate as efficiently as possible in the event of an emergency.

[0049] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0050] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A multi-timescale deep reinforcement learning adaptive collaborative allocation system for warehouse resources, characterized by: include: The warehouse data collection unit is used to collect real-time operating data and environmental parameters of various types of equipment in the warehouse scene, and standardize the data to form structured data in a unified format; A warehouse resource status feature extraction unit is used to use the structured data in a unified format and the multi-dimensional resource status feature extraction method in the warehouse management center to complete the resource collaborative allocation task within the warehouse space and extract the resource status features of different warehouse locations and form abnormal features. The resource status features include: inventory turnover rate, storage space utilization rate, equipment idle rate, transportation path length, and task waiting time; The reinforcement learning state space construction unit is used to analyze and cluster the resource collaborative allocation tasks within the storage space and the resource state characteristics of different storage locations based on the abnormal characteristics of the unit time step. At the same time, it analyzes the device interaction frequency under the dynamic load of different storage locations to form the state space construction parameters of different storage location dynamic loads; A resource allocation strategy generation unit is used to model different storage locations in the storage space to form a resource allocation model, which includes a global strategy model based on deep reinforcement learning, a local optimization model, and a multi-objective collaborative model. Based on the resource allocation model and the structured data in a unified format, a resource allocation strategy is generated for the corresponding inventory scheduling, equipment collaboration, and path planning behaviors, and the factors affecting resource allocation are classified to achieve adaptive collaborative resource allocation in the storage space. The strategy adaptive optimization unit is used to model the resource allocation strategy of the storage space and identify the allocation strategy through the adaptive collaborative allocation parameters of the storage space resources using the deep reinforcement learning algorithm, and to update the parameters of the deep reinforcement learning algorithm using the experience replay mechanism and policy gradient optimization.

2. The adaptive collaborative allocation system for warehouse resources based on multi-time-scale deep reinforcement learning according to claim 1 is characterized in that: The warehouse data collection unit includes: Multi-source data acquisition module, used to collect real-time operating data and environmental parameters of multiple types of equipment generated by different storage locations in the storage space per unit time step; A data standardization module is used to standardize the real-time operation data and environmental parameters of the multiple types of equipment into corresponding resource parameters according to a preset Z-score transformation; The outlier processing module is used to process outliers on the resource parameters obtained by standardization to obtain interference-free observation data; outlier processing includes deletion, interpolation and statistical learning methods; A dynamic load source tracking module is used to track the interactive impact degree data of different devices from the non-interference observation data, and standardize the interactive impact degree data of different devices according to the dynamic load source of the storage location to obtain the dynamic load source information of the device that matches the dynamic load source of the storage location; The environment impact dynamic load information parsing module is used to perform environment impact analysis on the equipment dynamic load source information to form different environment dynamic load source parameter variables as the structured data with a unified format.

3. The multi-timescale deep reinforcement learning adaptive collaborative allocation system for warehouse resources according to claim 1 is characterized in that: The storage resource status feature extraction unit includes: The edge computing node adaptation module is used to preset edge computing node parameters, adapt to storage throughput and operation complexity, and establish adaptation statistics tables; The warehouse management center module is used to adjust the extraction steps and preset the extraction parameters of the multi-dimensional resource status feature extraction method in the warehouse management center; A multi-storage location resource status feature extraction module is used to extract periodic resource status features of different storage locations based on the multi-dimensional resource status feature extraction method, the structured data in a unified format, and the adaptation statistical table, and generate abnormal resource status features of different storage locations; The resource status feature extraction module is used for edge computing nodes to analyze the execution status of resource collaborative allocation tasks, to judge and analyze resource collaborative allocation tasks that have discovered resource status features, and to generate collaborative allocation resource status abnormal features. The collaborative allocation resource status abnormal features and the abnormal features of resource status at different storage locations form the abnormal features.

4. The multi-timescale deep reinforcement learning adaptive collaborative allocation system for warehouse resources according to claim 1 is characterized in that: The reinforcement learning state space construction unit includes: The storage resource status parameter parsing module is used to locate the storage location based on abnormal characteristics of the collaborative allocation resource status and abnormal characteristics of the resource status of different storage locations, including storage resource status parameter load parsing, storage resource status parameter environment parsing and storage resource status parameter abnormality location; The resource status parameter classification and processing flow module is used to classify and form the warehouse resource status parameter information in the abnormal characteristics of the collaborative allocation resource status and the abnormal characteristics of the resource status of different storage locations, and obtain the data processing flow corresponding to the warehouse resource status parameter information, forming the warehouse resource status parameter reference benchmark, the coefficient to be optimized of the warehouse resource status parameter calculation model, and the noise type; The dynamic load module is used to detect the equipment interaction of the full-time-step dynamic load of different storage locations, and statistically analyze the sources of the dynamic load of equipment in different storage locations to form the state space construction parameters of the dynamic load of different storage locations.

5. The multi-timescale deep reinforcement learning adaptive collaborative allocation system for warehouse resources according to claim 1 is characterized in that: The resource allocation strategy generating unit includes: Resource allocation model management module, used to build and manage global strategy models, local optimization models, and multi-objective collaborative models based on deep reinforcement learning; The allocation strategy generation module is used to determine the resource allocation security of different storage locations, systems, and management terminals based on the real-time operation data and environmental parameters of multiple types of equipment, the global strategy model, the local optimization model, and the multi-objective collaborative model, using a preset state coding mechanism and strategy generation algorithm to obtain allocation strategy parameters; An operation difficulty identification module is used to quantitatively evaluate the operation difficulty of different storage locations based on the resource allocation security identification parameters and the dynamic load source of the equipment operation complexity and obtain the operation difficulty identification parameters; A resource allocation influencing factor classification module is used to automatically complete the resource allocation security identification of different storage locations, systems and management terminals based on the allocation strategy parameters and operation difficulty identification parameters, and to form a record of resource allocation influencing factors; The allocation process visualization module is used to visualize the system resource allocation process.

6. The multi-timescale deep reinforcement learning adaptive collaborative allocation system for warehouse resources according to claim 1 is characterized in that: The strategy adaptive optimization unit includes: The operation status recording module is used to continuously monitor the working status of different storage locations and the dynamic load source of the system equipment operation complexity, forming a record of the status of different storage locations and the operation complexity of the storage system, providing a reference basis for subsequent resource allocation and coordinated scheduling; The strategy effectiveness evaluation module is used to determine the resource allocation security of different storage locations to obtain resource allocation security test parameters, evaluate the duration of different storage locations or allocation strategy levels, and evaluate the allocation accuracy of different storage locations for future resource allocation tasks; The reinforcement learning network optimization module is used to allocate resources based on the security of different storage locations, while combining the accuracy of different storage locations for future resource allocation tasks, and using the experience replay mechanism and policy gradient optimization to update the parameters of the deep reinforcement learning algorithm.

7. The multi-timescale deep reinforcement learning adaptive collaborative allocation system for warehouse resources according to claim 1 is characterized in that: It also includes a multi-source information fusion unit for integrating equipment perception data within the storage space, management system instruction data, and external environment data to form a fused information set; The multi-source information fusion unit includes a time synchronization module for performing timestamp alignment processing on data from different sources; The spatial calibration module is used to calibrate the spatial position of the device perception data based on the storage location coordinate information; The conflict resolution module is used to identify conflicting information from different data sources and perform fusion processing through a voting mechanism or credibility weights, wherein the fused information set serves as a supplementary input for the structured data in a unified format.

8. The multi-timescale deep reinforcement learning adaptive collaborative allocation system for warehouse resources according to claim 5 is characterized in that: The multi-objective collaborative model includes an inventory balancing sub-model, an equipment load balancing sub-model and a path overlap avoidance sub-model; the inventory balancing sub-model is used to dynamically adjust the cargo location allocation ratio so that the inventory level of each storage location approaches a preset mean; the equipment load balancing sub-model is used to make the difference in the operating load of different equipment less than a threshold through a task allocation strategy; the path overlap avoidance sub-model is used to reduce the probability of intersection and overlap of equipment transportation paths through a path planning algorithm.

9. The multi-timescale deep reinforcement learning adaptive collaborative allocation system for warehouse resources according to claim 5 is characterized in that: The resource allocation strategy generation unit also includes a task queue dynamic adjustment module for dynamically adjusting the generated task execution sequence according to the real-time resource status and strategy generation parameters; the task queue dynamic adjustment module includes a state perception submodule for obtaining the inventory status of each storage location, the equipment operation status and the task completion progress in real time; The impact assessment submodule is used to analyze the impact of the current task queue state change on the task sequence execution efficiency; the sequence reordering submodule is used to locally or globally reorder the task queue according to the impact assessment results.

10. The multi-timescale deep reinforcement learning adaptive collaborative allocation system for warehouse resources according to claim 9 is characterized in that: The local reordering or global reordering operation preserves the execution order of high priority tasks.

Citation Information

Patent Citations

  • Intelligent stockyard distribution, storage and transportation management and control system and method based on data analysis

    CN118761699A

  • Intelligent storage resource dynamic allocation method and system based on deep reinforcement learning

    CN119204589A

  • Scheduling automation system application state management method

    CN119292745A

  • Multi-robot collaborative scheduling system in automatic warehousing system

    CN120255517A

  • Space optimization management system for multi-source data fusion in warehouse management

    CN120338674A

Cited By

  • Deep learning warehouse location intelligent distribution method and system

    CN121436866A

  • Stacker multi-roadway orbital transfer scheduling method and system

    CN121504326A

  • Intelligent warehousing automatic control optimization method and system based on dynamic environment perception

    CN121742407A