Unmanned forklift dynamic task scheduling method and system based on deep reinforcement learning of cold chain warehouse
By using deep reinforcement learning and digital twin technology in cold chain warehouses, the path planning of unmanned forklifts is optimized, and the problem of inability to respond to dynamic environmental changes in the existing technology is solved, and efficient and intelligent unmanned forklift scheduling is achieved.
Patent Information
- Application Number
- CN202510725013.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The existing technology is difficult to respond to changes in the dynamic environment of cold chain warehouses in real time, and cannot fully consider a variety of optimization goals, resulting in low efficiency of unmanned forklift scheduling.
The dynamic task scheduling method of unmanned forklifts based on deep reinforcement learning is adopted. By arranging sensors in the cold chain warehouse to collect environmental data, a digital twin model is built, the initial unmanned forklift trajectory is generated, and the path is optimized using a multi-feature-coupled space-time joint path planning algorithm to achieve optimal scheduling.
Real-time response to the dynamic environment of cold chain warehouses is achieved, the total time consumption, energy consumption, collision risk and congestion of unmanned forklift paths is optimized, and the scheduling efficiency and intelligence level are improved.
Smart Images

Figure CN120235559A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned forklift scheduling, and particularly to a dynamic task scheduling method and system for unmanned forklifts based on deep reinforcement learning in cold chain warehouses. Background Art
[0002] With the rapid development of the cold chain logistics industry, traditional unmanned forklift scheduling methods have been difficult to meet the growing demands for high efficiency, precision, and intelligence. There are many deficiencies in the existing technology for the dynamic task scheduling of unmanned forklifts, making it difficult to respond to dynamic environmental changes in real time and unable to fully consider multiple optimization objectives. Summary of the Invention
[0003] One of the objectives of the present invention is to provide a dynamic task scheduling method for unmanned forklifts based on deep reinforcement learning in cold chain warehouses to solve the problem in the prior art that it is impossible to respond to dynamic environmental changes in real time. In response to the above problem, the present invention proposes a dynamic task scheduling method for unmanned forklifts in cold chain warehouses based on deep reinforcement learning. This method arranges various sensors in the target cold chain warehouse, collects environmental data and constructs a digital twin model, constructs a dynamic environmental state according to the data fed back by the sensors, and generates an initial unmanned forklift trajectory. Then, a path planning algorithm deployed on the cloud is used to optimize the initial trajectory to generate an optimal path and scheduling plan, realizing the optimization of the dynamic task scheduling of unmanned forklifts.
[0004] The present invention is realized through the following technical solutions. A dynamic task scheduling method for unmanned forklifts based on deep reinforcement learning in cold chain warehouses includes the following steps: arranging sensors in the target cold chain warehouse, sending the collected data to the cloud to build a sensor database, collecting the environmental data of the target cold chain warehouse, and building a digital twin model of the target cold chain warehouse on the cloud based on the collected environmental data using digital twin technology; constructing a dynamic environmental state of the target cold chain warehouse according to the data fed back by the sensors, and generating an initial unmanned forklift trajectory based on a path generation algorithm, and using a spatio-temporal joint path planning algorithm with multi-feature coupling deployed on the cloud to optimize the initial unmanned forklift trajectory to generate an optimal path and scheduling plan; the cloud sends instructions to the edge gateway, and the edge gateway decomposes the instructions and sends them to each unmanned forklift to implement the specific operation of dynamic task scheduling.
[0005] Furthermore, arranging sensors in the target cold chain warehouse includes: deployment of environmental detection sensors, UWB anchor deployment, and unmanned forklift sensor deployment; deployment of environmental detection sensors includes: installing detection sensors in the target cold chain warehouse for real-time detection of the position of obstacles; UWB anchor deployment includes: determining the number of UWB anchors according to the area and structure of the target cold chain warehouse and installing multiple UWB anchors to ensure signal coverage of the entire warehouse; unmanned forklift sensor deployment includes: installing anti-fog cameras at the front and rear of the unmanned forklift, installing lidar on the top of the unmanned forklift, and installing an IMU at the central position of the unmanned forklift.
[0006] Furthermore, data collected by the deployed sensors can be received through the edge gateway, and time synchronization and spatial alignment are performed on the collected data in the edge gateway to ensure the consistency of the data of each sensor.
[0007] Furthermore, collecting environmental data of the target cold chain warehouse includes: placing high-precision lidar and RGB-D camera devices on a mobile platform, planning a scanning path covering the entire target cold chain warehouse, including all channels, shelves, and important areas, moving along the predetermined path to ensure that the LiDAR and RGB-D camera fully cover every corner of the target cold chain warehouse, and collecting point cloud data of the entire warehouse; using the SLAM algorithm to convert the LiDAR data into high-precision three-dimensional point clouds and assigning color information to the point clouds in combination with the RGB-D data.
[0008] Furthermore, constructing the digital twin model of the target cold chain warehouse can also include marking key points of the target cold chain warehouse in the digital twin model. The key points can include: shelves, storage locations, temperature zone divisions, safety areas, and auxiliary facilities.
[0009] Furthermore, by numbering the shelves in the digital twin model and recording the specific position coordinates of each shelf on record, the marking of the shelves is realized; marking the coordinates and cargo types of each storage location to ensure clear classification of items and realize the marking of storage location information; marking the boundaries of different temperature zones in the cold storage and recording the temperature requirements of each temperature zone to realize the division of temperature zones; marking the safety fence area in the digital twin model and defining the no-go area for unmanned forklifts to realize the marking of safety areas; marking the positions of the charging piles and entrances / exits of unmanned forklifts in the digital twin model to ensure that these auxiliary facilities are accurately reflected in the digital twin model and realize the marking of auxiliary facilities.
[0010] Furthermore, the initial trajectory of the unmanned forklift, which is used to represent the state of the unmanned forklift at time t, includes position, heading, and speed in this state, and is represented by the following formula:
[0011] , where $x(t)$ is the abscissa position of the driverless forklift at time $t$; $y(t)$ is the ordinate position of the driverless forklift at time $t$; $\theta(t)$ is the heading angle of the driverless forklift at time $t$; $v(t)$ is the speed of the driverless forklift at time $t$.
[0012] Furthermore, the dynamic environmental state includes the set of dynamic obstacle positions and the congestion index of the storage location; the dynamic environmental state collects the position data of personnel and other forklifts in real time through sensors installed in the warehouse or factory, and obtains the task queue and space occupancy data of each storage location through the goods management system or the task scheduling system, processes the sensor data into the set of dynamic obstacle positions, and calculates the congestion index of each storage location according to the task queuing time and the space occupancy rate.
[0013] Furthermore, the status information of the set of obstacle positions and the congestion index is updated regularly to reflect the latest environmental changes, so as to obtain and update the dynamic environmental state in real time, and achieve effective path planning and scheduling management in a complex dynamic environment.
[0014] Furthermore, the specific expression of the set of positions $O(t)$ of dynamic obstacles is:
[0015] , where $(x_n(t), y_n(t))$ is the position of the $n$-th dynamic obstacle at time $t$. These positions are usually described by coordinates, such as in a two-dimensional space or
[0016] Furthermore, the congestion index $C_k(t)$ of the storage location has the following expression:
[0017] , where $\alpha$ and $\beta$ are weight coefficients used to balance the influence of the task queuing time and the space occupancy rate on the congestion index; $T_k(t)$ is the task queuing time of storage location $k$ at time $t$, indicating how many tasks are waiting to be processed; $O_k(t)$ is the space occupancy rate of storage location $k$ at time $t$, indicating the degree to which the storage location is occupied;
[0018] Furthermore, the task queuing time can be calculated by the following formula:
[0019] , where $N_k(t)$ is the number of tasks queued at storage location $k$ at time $t$; $p_j$ is the estimated processing time of the $j$-th task at storage location $k$.
[0020] Furthermore, the space occupancy rate can be calculated by the following formula:
[0021] , where is the occupied space of storage location k at time t, is the total space of storage location k.
[0022] Furthermore, the spatio-temporal joint path planning algorithm is constructed through the following steps: According to the main optimization objectives or constraint conditions of the unmanned forklift dynamic task scheduling, a main optimization objective sub-model is constructed. The main optimization objective sub-model includes: a path time-consuming sub-model that helps optimize the path of the unmanned forklift in the cold chain warehouse to minimize the total time-consuming; a path energy consumption sub-model that helps optimize the path of the forklift to minimize the energy consumption, extend the battery life, and reduce the operating cost; a collision risk sub-model that evaluates the collision risk on the path in real time and helps the forklift select the path with the lowest risk; a storage location congestion penalty sub-model that helps the forklift select the path with the lowest congestion level and avoid delays caused by congestion. Based on the constructed sub-models, the total objective function of the spatio-temporal joint path planning model is formed by weighted aggregation. The total objective function creates a comprehensive path cost minimization objective by combining the outputs of each sub-model. The total objective function is expressed by the following formula:
[0023] , where is the path time-consuming sub-model, is the weight coefficient of the path time-consuming sub-model; is the path energy consumption sub-model, is the weight coefficient of the path energy consumption sub-model; is the collision risk sub-model, is the weight coefficient of the collision risk probability sub-model; is the storage location congestion penalty sub-model, is the weight coefficient of the storage location congestion penalty sub-model.
[0024] Furthermore, the path time-consuming sub-model can be expressed by the following formula:
[0025] , and its constraint condition is: , where is the path time-consuming sub-model, which is used to calculate the time required for the unmanned forklift to travel on the path; is the start time of the unmanned forklift traveling on the path, is the end time of the unmanned forklift traveling on the path, is the speed of the unmanned forklift at time t, is the acceleration of the unmanned forklift at time t, is the maximum acceleration at the cold storage environment temperature under.
[0026] Furthermore, the path energy consumption sub-model can be expressed by the following formula:
[0027] , where is the path energy consumption sub-model, is the motion power, is the battery efficiency coefficient, is the battery efficiency coefficient at the ambient temperature .
[0028] Furthermore, the motion power can be calculated by the following formula:
[0029] , where the motion power consists of the square term of the speed and the square term of the acceleration, multiplied by the coefficients and respectively. These two coefficients are constants related to the motion resistance and the acceleration loss respectively.
[0030] Furthermore, the collision risk sub-model can be expressed by the following formula:
[0031] , where is the collision risk sub-model, is the exponential function, is the position of the obstacle at time t, is the safety radius, representing the influence range of the obstacle; when the passage is narrow, the safety radius will shrink.
[0032] Furthermore, the cargo location congestion penalty sub-model can be expressed by the following formula:
[0033] , where is the cargo location congestion penalty sub-model, is the set of cargo locations passed by the path of the driverless forklift, is the logarithm symbol, is the partial derivative symbol, and represent the weights of static congestion and dynamic congestion respectively; is the congestion change rate of cargo location k at time t, that is, the congestion trend, is the congestion index of the cargo location.
[0034] Furthermore, the spatio-temporal joint path planning algorithm further includes constraint conditions, which are used to ensure that the path planning not only minimizes the total objective function, but also guarantees safety, feasibility and task requirements. The constraint conditions include: a hard obstacle avoidance constraint that requires the distance between the driverless forklift and all dynamic obstacles at any time to be greater than or equal to the safety distance; a shelf aisle geometric constraint that is used to limit the position of the driverless forklift and must be within the boundary range of the shelf aisle; and a dynamic task scheduling coupling constraint that is used to ensure the timeliness of task scheduling.
[0035] Furthermore, the hard obstacle avoidance constraint can be expressed by the following formula:
[0036] , where is the safety distance, is the start time, is the end time.
[0037] Furthermore, the safety distance can be calculated by the following formula:
[0038] , where is the ground friction coefficient, which is set according to the change of the ground friction coefficient at the actual temperature; is the acceleration due to gravity.
[0039] Furthermore, the shelf aisle geometric constraint can be expressed by the following formula:
[0040] , where is the maximum abscissa position of the shelf aisle, is the minimum abscissa position of the shelf aisle, is the maximum ordinate position of the shelf aisle, is the minimum ordinate position of the shelf aisle. These positions jointly describe the boundary of the aisle and define the area where the forklift can move.
[0041] Furthermore, the dynamic task scheduling coupling constraint can be expressed by the following formula:
[0042] , where is the time when the driverless forklift i arrives at the picking position of task j, is the specified deadline, indicating that the time when the driverless forklift i arrives at the picking position of task j must be less than or equal to the specified deadline; is the time when the driverless forklift i arrives at the dropping position of task j, is the maximum tolerable delay time, indicating that the time when the driverless forklift i arrives at the dropping position of task j must be before the maximum tolerable delay time.
[0043] Furthermore, the edge gateway decomposes and issues the instructions, including the following steps: first, the edge gateway verifies the instruction data to ensure the integrity and correctness of the data; then it converts it into control instructions and task instruction sets that each unmanned forklift can understand, and distributes the path and task instructions of each unmanned forklift to the corresponding forklift control system; the unmanned forklift drives along the designated path according to the path instructions received from the edge gateway through its own navigation system; the edge gateway monitors the status and location of each unmanned forklift in real time, and feeds the monitoring data back to the digital twin model in the cloud.
[0044] On the other hand, the present invention provides a dynamic task scheduling system for unmanned forklifts based on deep reinforcement learning for cold chain warehouses, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the dynamic task scheduling method for unmanned forklifts based on deep reinforcement learning for cold chain warehouses as described above is implemented.
[0045] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0046] 1. The present invention can respond to the dynamic environmental changes of the cold chain warehouse in real time. By constructing a multi-feature coupled spatiotemporal joint path planning algorithm, it simultaneously optimizes multiple objectives such as the total time consumption, energy consumption, collision risk and congestion level of the unmanned forklift path, and timely adjusts the scheduling plan of the unmanned forklift to improve the scheduling efficiency.
[0047] 2. The present invention uses digital twin technology to accurately model the cold chain warehouse environment based on the collected environmental data, accurately describe the actual status of the warehouse, provide a reliable basis for scheduling decisions, and autonomously optimize the scheduling plan according to the dynamic environmental status, thereby improving the intelligence level of scheduling.
[0048] 3. The present invention leverages the powerful computing power of the cloud to efficiently process large amounts of sensor data, implement complex path planning calculations, and improve the optimization effect of the scheduling solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:
[0050] Figure 1 This is a flow chart of the method provided in Example 1 of the present invention. DETAILED DESCRIPTION
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0052] Embodiment 1
[0053] In the prior art, the shelves in cold chain warehouses are dense and the aisles are narrow, and they are in a low-temperature state all year round. Due to the special situation of cold chain warehouses, how to take into account possible obstacles such as personnel and equipment in the cold chain warehouse requires that the automated guided vehicle (AGV) needs to have good safety protection measures to avoid collisions and ensure the safety of personnel and equipment. At the same time, according to the current task requirements, how does the scheduling system dispatch the optimal AGV based on the picking point and dropping point of the task, the position information and loading information of each AGV, and the congestion degree information of the cargo openings of each cargo location; and realizing the precise positioning and navigation of the AGV in the cold chain warehouse is a common problem in the industry. It should be noted that the 'low temperature' in this application refers to the low temperature in the cold chain warehouse. Those skilled in the art can refer to the relevant regulations in the 'Classification and Basic Requirements for Cold Chain Logistics' with the national standard number: GB / T 28577-2021, and should understand that the 'low temperature' in this application means that frozen foods are stored at less than or equal to -18°C, and refrigerated foods need to be stored at 0~8°C. That is to say, those skilled in the art should understand that the low temperature range in this application should be at least: 8°C~-18°C.
[0054] A method for dynamically scheduling AGVs in the special environment of cold chain warehouses disclosed in this embodiment constructs a fine model of the cold chain warehouse in the cloud based on digital twin technology, and at the same time introduces multiple coupling characteristics considering the dynamic task scheduling of AGVs in the cold chain warehouse into the constructed fine model, thereby constructing a spatio-temporal joint path planning algorithm with multi-characteristic coupling, realizing efficient dynamic task scheduling of AGVs in the special environment of cold chain warehouses.
[0055] This embodiment includes two stages, namely the digital twin model construction stage for data collection and environment modeling. By deploying sensors at the cold chain warehouse site and scanning to construct a digital twin model in the cloud, a basic data perception layer that can reflect the cold chain warehouse site situation in real time in the cloud is constructed, providing basic data support for subsequent decision-making and scheduling.
[0056] The dynamic path planning and safety protection stage is used to perform path planning and obstacle avoidance based on the data collected in the first stage and through the constructed spatio-temporal joint path planning algorithm. Finally, a decision-making layer that can output dynamic scheduling instructions is constructed. The decision-making layer generates decisions and sends the instructions to the terminal. The terminal executes the dynamic scheduling instructions output by the decision-making layer to actually control the actions of the automated forklift, realizing the dynamic task scheduling of the automated forklift in the cold chain warehouse.
[0057] Figure 1 The flowchart of the method in this embodiment is shown. It can be seen from the figure that this embodiment includes the following steps:
[0058] The digital twin model construction stage includes two steps. The task in this stage is to build a refined cold chain warehouse digital twin model of the target cold chain warehouse in the cloud by deploying sensors in the target cold chain warehouse and collecting the environmental data of the target cold chain warehouse based on digital twin technology.
[0059] Step 1: First, deploy sensors in the target cold chain warehouse.
[0060] The sensor deployment specifically includes: the deployment of environmental detection sensors, the deployment of UWB anchors, and the deployment of automated forklift sensors.
[0061] The deployment of environmental detection sensors includes:
[0062] Install detection sensors in the cold chain warehouse for real-time detection of the positions of obstacles, which will be used to help analyze and check the positions of obstacles in the warehouse later.
[0063] The deployment of UWB anchors includes:
[0064] Determine the number of UWB anchors according to the area and structure of the cold storage to ensure signal coverage of the entire warehouse. Install the UWB anchors on the top or walls along the cold storage aisles to form a complete coverage network. The interval between each UWB anchor should be less than 20 meters to ensure the continuity and stability of signal coverage. When installing, pay attention to avoiding deploying UWB anchors near metal shelves to reduce signal interference. It is possible to consider installing the anchors on the top of the shelves or at positions far from metal structures. Each anchor needs to communicate wirelessly with the central control system (such as a UWB positioning server) to ensure real-time data transmission. Ensure that each UWB anchor has an independent power supply to avoid the positioning system failure due to power supply problems.
[0065] The deployment of automated forklift sensors includes:
[0066] Install anti-fog cameras at the front and rear of the driverless forklift. It is best to choose anti-fog cameras that can work properly in low-temperature environments. Install a low-temperature-resistant lidar on the top of the driverless forklift to ensure 360-degree dead-angle-free scanning. At the same time, install an IMU (Inertial Measurement Unit) at the center of the driverless forklift to accurately measure the posture and motion state of the forklift.
[0067] Receive the data collected by the deployed sensors through the edge gateway, and perform time synchronization and spatial alignment on the collected data in the edge gateway to ensure the consistency of the sensor data. Then send the data that has undergone time synchronization and spatial alignment to the cloud to build a sensor database for subsequent data calculations.
[0068] Step 2: Collect the environmental information of the cold chain warehouse, build a refined digital twin model of the target cold chain warehouse in the cloud based on the collected environmental information, and mark the positions of key points in the digital twin model after completion.
[0069] Specifically, the collection of the environmental information of the cold chain warehouse includes: Install high-precision lidar (LiDAR) and RGB-D camera installation equipment on a mobile platform, which can be a driverless forklift or a handheld device. Plan a scanning path that covers the entire target cold chain warehouse, including all aisles, shelves, and important areas. Move along the predetermined path to ensure that the LiDAR and RGB-D cameras can fully cover every corner of the target cold chain warehouse. Obtain the point cloud data of the entire warehouse, use the SLAM (Simultaneous Localization and Mapping) algorithm to convert the LiDAR data into high-precision three-dimensional point clouds, and combine the RGB-D data to assign color information to the point clouds to enhance the visualization effect of the map. Build a refined digital twin model of the target cold chain warehouse in the cloud based on the collected relevant data combined with digital twin technology. And mark the relevant information of key points in the digital twin model.
[0070] Specifically, the key points can include: shelves, storage locations, temperature zone divisions, safety areas, and auxiliary facilities. Number the shelves in the digital twin model, record the specific position coordinates of each shelf, and achieve the marking of the shelves. Mark the coordinates and cargo types (frozen / refrigerated) of each storage location to ensure clear item classification and achieve the marking of storage location information. Mark the boundaries of different temperature zones in the cold storage and record the temperature requirements of each temperature zone to achieve the division of temperature zones. Mark the safety fence area in the digital twin model, define the no-go area for the driverless forklift, and achieve the marking of the safety area. Mark the positions of the charging piles and entrances and exits of the driverless forklift in the digital twin model to ensure that these auxiliary facilities are accurately reflected in the digital twin model and achieve the marking of auxiliary facilities.
[0071] Step 3: After constructing the target cold chain warehouse digital twin model, construct the dynamic environmental state of the target cold chain warehouse at the current moment according to the data fed back by on-site sensors, and use the existing path generation algorithm to generate an initial unmanned forklift trajectory for the unmanned forklift. Use the spatio-temporal joint path planning mathematical model with multi-feature coupling deployed in the cloud to optimize the initial unmanned forklift trajectory, and finally generate an optimal path and scheduling plan, so that all unmanned forklifts can efficiently complete all assigned tasks on the premise of meeting safety, geometric, and task scheduling constraints.
[0072] Specifically, in this embodiment, the RRT* algorithm can be used to obtain an initial unmanned forklift trajectory P(t), which is used to represent the state of the unmanned forklift at time t, and this state includes position, heading, and speed. It can be expressed by the following formula:
[0073] ,
[0074] where, is the abscissa position of the unmanned forklift at time t; is the ordinate position of the unmanned forklift at time t; is the heading angle (the forward direction of the forklift) of the unmanned forklift at time t; is the speed of the unmanned forklift at time t.
[0075] In this embodiment, the dynamic environmental state can be divided into the position set of dynamic obstacles (personnel, other forklifts) at time t; and the congestion index (task queuing time + space occupancy rate) of storage location k at time t.
[0076] The position data of personnel and other forklifts are collected in real time through sensors installed in the warehouse or factory. And the task queue and space occupancy data of each storage location are obtained through the goods management system or task scheduling system. The sensor data is processed into an obstacle position set. At the same time, the congestion index of each storage location is calculated according to the task queuing time and space occupancy rate.
[0077] And the status information of the obstacle position set and congestion index is updated regularly to reflect the latest environmental changes. Thus, the dynamic environmental state is obtained and updated in real time, and effective path planning and scheduling management are realized in a complex dynamic environment.
[0078] Specifically, the position set of dynamic obstacles and the congestion index of each storage location can be expressed as:
[0079] The position set O(t) of dynamic obstacles can be obtained from the sensor data. The specific expression is:
[0080] ,
[0081] Among them, is the position of the nth dynamic obstacle at time t. These positions are usually described by coordinates, such as in a two-dimensional space or in a three-dimensional space . These positions can be obtained by the following methods: 1. Sensor data: Use sensors installed in the environment (such as lidar, cameras, ultrasonic sensors, etc.) to detect the positions of obstacles in real time. 2. Positioning system: Use a real-time positioning system (RTLS), such as a system based on ultra-wideband (UWB) or RFID technology, to track the positions of personnel and forklifts in real time. 3. Communication system: Obtain the position data of other forklifts through vehicle-to-vehicle (V2V) or vehicle-to-infrastructure (V2I) communication.
[0082] The congestion index of the storage location can be composed of the task queuing time and the space occupancy rate. The specific expression is:
[0083] ,
[0084] Among them, and are weight coefficients used to balance the influence of the task queuing time and the space occupancy rate on the congestion index.
[0085] is the task queuing time of storage location k at time t, indicating how many tasks are waiting to be processed; this task queuing time can obtain the task queue information of each storage location through the goods management system or the task scheduling system, and calculate the total time of the current queued tasks. Specifically, it can be calculated by the following formula:
[0086] ,
[0087] Among them, is the number of tasks queued at storage location k at time t; is the estimated processing time of the jth task at storage location k.
[0088] is the space occupancy rate of storage location k at time t, indicating the degree to which the storage location is occupied; the occupancy situation of the storage location can be tracked through sensor data or the goods management system, and the occupancy rate can be calculated. Specifically, it can be calculated by the following formula:
[0089] ,
[0090] Among them, is the occupied space of storage location k at time t, is the total space of storage location k.
[0091] Specifically, in this embodiment, the spatio-temporal joint path planning mathematical model is constructed through the following steps:
[0092] 1) First, considering the main optimization objectives or constraints of the unmanned forklift dynamic task scheduling, a main optimization objective sub-model is constructed.
[0093] In this embodiment, by considering the path time consumption, path energy consumption, collision risk, and cold storage location congestion penalty of the unmanned forklift, sub-models are constructed. These sub-models together constitute a mathematical model for the coupling requirements of dynamic path planning and multi-objective scheduling in the cold chain warehouse. Each sub-model can calculate different path attributes (time, energy consumption, risk, congestion) through time integration. Different parameters (such as speed v(t), acceleration a(t), trajectory P(t), etc.) are shared among the sub-models, and data and results are transmitted to each other through these parameters.
[0094] Specifically, in this embodiment, the sub-models may include: a path time consumption sub-model, a path energy consumption sub-model, a collision risk sub-model, and a location congestion penalty sub-model.
[0095] Among them, the path time consumption sub-model calculates the total time by integrating the reciprocal of the speed. This is because speed is the rate of change of distance with respect to time, and the reciprocal of speed is the rate of change of time with respect to distance. Summing up all the small time segments on the integral path gives the total path time consumption. This sub-model helps to optimize the path of the unmanned forklift in the cold chain warehouse to minimize the total time consumption. It can be expressed by the following formula:
[0096] ,
[0097] Its constraint conditions are: , where, is the path time consumption sub-model, which is used to calculate the time required for the unmanned forklift to travel on the path, considering the dynamic characteristics of the unmanned forklift and the influence of the ground friction coefficient in the cold storage. is the start time of the unmanned forklift traveling on the path, is the end time of the unmanned forklift traveling on the path, is the speed of the unmanned forklift at time t, is the acceleration of the unmanned forklift at time t, is the maximum acceleration at the cold storage environment temperature ; when the cold storage environment temperature is low, the ground freezes, the friction coefficient decreases, and it is easy to cause the maximum acceleration to decrease. is the acceleration constraint, which ensures the safe acceleration limit of the forklift in the low-temperature environment of the cold storage.
[0098] The path energy consumption sub-model takes into account the characteristics of battery efficiency and motion power. Battery efficiency is affected by ambient temperature, and low temperature will reduce battery efficiency. Motion power includes the energy consumption components of speed and acceleration. The greater the speed and acceleration, the greater the energy consumption. By dividing the motion power by the battery efficiency, the actual energy consumption of the unmanned forklift is reflected. This model helps optimize the path of the forklift to minimize energy consumption, extend battery life, and reduce operating costs. It can be expressed by the following formula:
[0099] ,
[0100] in, is the path energy consumption submodel, is the sports power, is the battery efficiency coefficient, For ambient temperature The battery efficiency coefficient under .
[0101] Specifically, in this embodiment, the exercise power can be calculated by the following formula:
[0102] ,
[0103] Among them, the motion power is composed of the square term of velocity and the square term of acceleration, which are multiplied by the coefficients and , these two coefficients are constants related to motion resistance and acceleration loss respectively.
[0104] The collision risk sub-model considers the dynamic obstacles existing in all time periods on the path, and maps the distance between the forklift position and the obstacle position into the collision risk through the Gaussian function form (exponential function). The closer the distance, the higher the risk index. This sub-model can evaluate the collision risk on the path in real time and help the forklift choose the path with the lowest risk. It can be expressed as follows:
[0105]
[0106] in, is the collision risk submodel, is an exponential function, is the position of the obstacle at time t, is the safety radius, which represents the impact range of obstacles; when the channel is narrow, the safety radius will be reduced.
[0107] It should be noted that in the collision risk sub-model formula shown, by summing, the calculation arrive The accumulation of negative exponents of the distance between all obstacles o and the unmanned forklift state P(t) within the time interval. The closer the obstacle is to the forklift, the larger the negative exponent value is, and the higher the collision risk is.
[0108] The cargo congestion penalty sub-model optimizes the path selection by considering the degree of congestion and its dynamic changes, helping the forklift to choose the path with the least congestion and avoid delays caused by congestion. It can be expressed by the following formula:
[0109] ,
[0110] in, is the cargo congestion penalty sub-model, is the collection of cargo spaces that the unmanned forklift passes through. is the logarithmic sign, is the partial differential symbol, and Represent the weights of static congestion and dynamic congestion respectively; is the congestion change rate of cargo space k at time t, that is, the congestion trend.
[0111] It should be noted that this formula calculates the cumulative congestion costs of all cargo locations k that the path passes through by summing up. represents the static congestion cost; Represents the dynamic congestion cost; the higher the congestion level or the greater the upward trend of congestion, the greater the penalty value.
[0112] 2) Based on the constructed sub-models, each sub-model is aggregated in a weighted manner to form the total objective function of the spatiotemporal joint path planning model. The total objective function creates a comprehensive path cost minimization goal by combining the outputs of each sub-model. Each sub-model optimizes path time, energy consumption, collision risk, and congestion penalty, which are finally expressed in the total objective function through weighted and synthetic methods. And by adjusting the weight coefficient, it can flexibly adapt to different operation strategies and needs.
[0113] The overall objective function can be expressed as follows:
[0114] ,
[0115] in, is the weight coefficient of the path time-consuming sub-model, is the weight coefficient of the path energy consumption sub-model, is the weight coefficient of the collision risk probability sub-model, is the weight coefficient of the cargo congestion penalty sub-model.
[0116] It should be noted that the weight coefficient is used to adjust the importance of each sub-model in the overall objective. These coefficients can be adjusted according to the cold storage operation strategy, specific needs and actual conditions. For example, if energy consumption cost is a key concern of cold storage operation, the value of the path energy consumption sub-model can be increased to make energy consumption occupy a more important position in the overall objective function.
[0117] 3) In order to ensure that the optimal path and scheduling plan output by the total objective function can achieve the minimum of the total objective function without collision, exceeding the channel range, and completing the task on time. Constraints can also be set for the total objective function to ensure safety, feasibility and task requirements. Ensure that the optimal path and scheduling plan outputted in the end can enable the unmanned forklift to operate efficiently and safely in the complex dynamic environment of the cold chain warehouse.
[0118] Constraints are restrictions or requirements on the variables in an optimization problem that must be satisfied during the optimization process. For the unmanned forklift path optimization problem, the constraints are to ensure that the path planning not only optimizes the overall objective function, but also guarantees safety, feasibility, and mission requirements.
[0119] Specifically, in this embodiment, obstacle avoidance hard constraints, shelf channel geometry constraints, and dynamic task scheduling coupling constraints may be included.
[0120] Among them, the obstacle avoidance hard constraint requires that the distance between the unmanned forklift and all dynamic obstacles at any time must be greater than or equal to the safety distance. It can be specifically expressed by the following formula:
[0121] ,
[0122] in, is the safety distance. In this embodiment, the safety distance can be calculated by the following formula:
[0123] ,in, is the ground friction coefficient, which is set according to the change of the ground friction coefficient under the actual temperature; is the acceleration due to gravity.
[0124] The rack channel geometric constraint is used to limit the position of the unmanned forklift and must be within the boundary of the rack channel. It can be expressed as follows:
[0125] ,
[0126] in, is the maximum position of the abscissa of the shelf channel, is the minimum horizontal coordinate position of the shelf channel, is the maximum vertical coordinate position of the shelf channel, is the minimum vertical coordinate position of the shelf aisle. These positions together describe the boundaries of the aisle and define the area where the forklift can move.
[0127] Dynamic task scheduling coupling constraints are used to ensure the timeliness of task scheduling. Specifically, it can be expressed as follows:
[0128] ,
[0129] in, is the time it takes for unmanned forklift i to arrive at the pickup location of task j, is the specified deadline, which means that the time for unmanned forklift i to arrive at the pickup location of task j must be less than or equal to the specified deadline.
[0130] is the time it takes for unmanned forklift i to arrive at the delivery location of task j, is the maximum tolerable delay time, which means that the time when unmanned forklift i arrives at the delivery location of task j must be before the maximum tolerable delay time.
[0131] It should be noted that the feasibility of the path planning scheme is guaranteed by the constraints, that is, the unmanned forklift will not collide, will not exceed the channel range during driving, and can complete the task on time. Under the premise of meeting the constraints, the path planning model needs to weigh multiple goals (such as path time, energy consumption, collision risk and congestion penalty) and finally find a comprehensive optimal path. Ensure that the unmanned forklift will not collide, exceed the channel range, and complete the task on time while minimizing the total objective function.
[0132] Step 4: The cloud sends the generated optimal path and scheduling plan to the edge gateway, which decomposes the instructions and sends them to each unmanned forklift to implement the specific operations of dynamic task scheduling.
[0133] Specifically, after receiving the optimal path and scheduling plan, the edge gateway will first verify the data to ensure its integrity and correctness.
[0134] It is then converted into control instructions and task instruction sets that each unmanned forklift can understand. Based on the parsed data, the path and task instructions of each unmanned forklift are distributed to the corresponding forklift control system.
[0135] The unmanned forklift drives along the designated path through its own navigation system according to the path instructions received from the edge gateway. The navigation system uses sensors on the forklift (such as lidar, camera, ultrasonic sensor, etc.) for real-time environmental perception and path tracking. According to the scheduling plan, the unmanned forklift will execute specific task instructions during driving, such as loading goods, transporting to designated cargo locations, and unloading goods.
[0136] The edge gateway monitors the status and location of each unmanned forklift in real time, and feeds the monitoring data back to the digital twin model in the cloud. Through the feedback data from on-site sensors and forklifts, it ensures the safety and efficiency of task execution. If an emergency occurs during task execution (such as new obstacles, changes in task priority, etc.), the edge gateway will adjust the path and task instructions of the unmanned forklift in real time, and send the adjusted instructions to the corresponding forklift to ensure that the task can be completed smoothly.
[0137] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A dynamic task scheduling method for unmanned forklifts based on deep reinforcement learning in cold chain warehouses, characterized in that The described dynamic task scheduling method includes: Deploy sensors in the target cold chain warehouse and send the collected data to the cloud to build a sensor database. Collect the environmental data of the target cold chain warehouse and build a digital twin model of the target cold chain warehouse in the cloud based on the collected environmental data using digital twin technology. Build the dynamic environmental state of the target cold chain warehouse according to the data fed back by the sensors and generate an initial unmanned forklift trajectory based on the path generation algorithm. Use the spatio-temporal joint path planning algorithm with multi-feature coupling deployed in the cloud to optimize the initial unmanned forklift trajectory and generate the optimal path and scheduling plan. The cloud sends instructions to the edge gateway, and the edge gateway decomposes the instructions and sends them to each unmanned forklift to implement the specific operation of dynamic task scheduling.
2. The dynamic task scheduling method for an unmanned forklift based on deep reinforcement learning in a cold chain warehouse according to claim 1, wherein The deployment of sensors in the target cold chain warehouse includes: the deployment of environmental detection sensors, the deployment of UWB anchors, and the deployment of unmanned forklift sensors. The deployment of the environmental detection sensors includes: installing detection sensors in the target cold chain warehouse for real-time detection of the position of obstacles. The deployment of the UWB anchors includes: determining the number of UWB anchors according to the area and structure of the target cold chain warehouse and installing multiple UWB anchors to ensure signal coverage of the entire warehouse. The deployment of the unmanned forklift sensors includes: installing anti-fog cameras at the front and rear of the unmanned forklift, installing lidar on the top of the unmanned forklift, and installing an IMU at the central position of the unmanned forklift.
3. The dynamic task scheduling method for an unmanned forklift based on deep reinforcement learning in a cold chain warehouse according to claim 1, wherein The collection of the environmental data of the target cold chain warehouse includes: Mount high-precision lidar and RGB-D camera devices on a mobile platform and plan a scanning path covering the entire target cold chain warehouse, including all aisles, shelves, and important areas. Move along the predetermined path to ensure that the LiDAR and RGB-D camera fully cover every corner of the target cold chain warehouse and collect the point cloud data of the entire warehouse. Use the SLAM algorithm to convert the LiDAR data into high-precision three-dimensional point clouds and assign color information to the point clouds in combination with the RGB-D data.
4. The dynamic task scheduling method for an unmanned forklift based on deep reinforcement learning in a cold chain warehouse according to claim 1, wherein The dynamic environmental state includes a set of dynamic obstacle positions and a cargo location congestion index. The dynamic environmental state collects the position data of personnel and other forklifts in real time through sensors installed in the warehouse or factory, and obtains the task queue and space occupancy data of each cargo location through the cargo management system or task scheduling system. Process the sensor data into a set of dynamic obstacle positions, and calculate the congestion index of each cargo location according to the task queuing time and space occupancy rate.
5. The dynamic task scheduling method for an unmanned forklift based on deep reinforcement learning in a cold chain warehouse according to claim 1, wherein The spatio-temporal joint path planning algorithm is constructed through the following steps: According to the main optimization objectives or constraints of the unmanned forklift dynamic task scheduling, construct a main optimization objective sub-model, and the main optimization objective sub-model includes: A path time-consuming sub-model that helps optimize the path of the unmanned forklift in the cold chain warehouse to make the total time-consuming shortest. A path energy consumption sub-model that helps optimize the path of the forklift to minimize energy consumption, extend the battery life, and reduce the operating cost. A collision risk sub-model that evaluates the collision risk on the path in real time and helps the forklift select the path with the lowest risk. A sub-model for forklift congestion penalty that helps the forklift select the path with the lowest congestion level and avoid delays caused by congestion; Based on the constructed sub-models, the overall objective function of the spatio-temporal joint path planning model is aggregated by weighting. The overall objective function creates a comprehensive objective of minimizing the path cost by combining the outputs of each sub-model. The overall objective function is expressed by the following formula: , Among them, is the path time-consuming sub-model, is the weight coefficient of the path time-consuming sub-model; is the path energy consumption sub-model, is the weight coefficient of the path energy consumption sub-model; is the collision risk sub-model, is the weight coefficient of the collision risk probability sub-model; is the storage location congestion penalty sub-model, is the weight coefficient of the storage location congestion penalty sub-model.
6. The dynamic task scheduling method for an unmanned forklift based on deep reinforcement learning in a cold chain warehouse according to claim 5, wherein The spatio-temporal joint path planning algorithm also includes constraint conditions, which are used to ensure that the path planning not only minimizes the overall objective function but also guarantees safety, feasibility, and task requirements. The constraint conditions include: A hard obstacle avoidance constraint that requires the distance between the driverless forklift and all dynamic obstacles at any time to be greater than or equal to the safety distance; A geometric constraint of the shelf aisle that is used to limit the position of the driverless forklift and must be within the boundary of the shelf aisle; And a dynamic task scheduling coupling constraint for ensuring the timeliness of task scheduling.
7. The dynamic task scheduling method for the driverless forklift based on deep reinforcement learning in the cold chain warehouse according to claim 6, wherein, The hard obstacle avoidance constraint is expressed by the following formula: , Among them, O(t) is the set of positions of dynamic obstacles; P(t) is the initial trajectory of the driverless forklift; is the safety distance, is the position of the obstacle at time t, is the start time, is the end time.
8. The dynamic task scheduling method for the driverless forklift based on deep reinforcement learning in the cold chain warehouse according to claim 7, wherein The safety distance is calculated by the following formula: , Among them, is the ground friction coefficient, is the acceleration due to gravity, is the speed of the driverless forklift at time t.
9. The dynamic task scheduling method for an unmanned forklift based on deep reinforcement learning in a cold chain warehouse according to claim 1, characterized in that, The edge gateway decomposes and distributes the instructions, including the following steps: First, the edge gateway checks the data of the instructions to ensure the integrity and correctness of the data; Then it converts them into control instructions and task instruction sets that each driverless forklift can understand, and distributes the path and task instructions of each driverless forklift to the corresponding forklift control system; The driverless forklift travels along the specified path through its own navigation system according to the path instructions received from the edge gateway; The edge gateway monitors the status and position of each driverless forklift in real time and feeds the monitoring data back to the digital twin model in the cloud.
10. An unmanned forklift dynamic task scheduling system based on deep reinforcement learning in a cold chain warehouse, characterized in that, The dynamic task scheduling system includes: A processor; A memory storing a computer program, which when executed by the processor, implements the dynamic task scheduling method for driverless forklifts based on deep reinforcement learning in a cold chain warehouse as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Automatic guided vehicle task allocation and path planning method based on deep reinforcement learning
CN117055563A
Multi-level low-altitude air route network construction method in complex urban environment
CN119516846A
Cold chain warehouse unmanned aerial vehicle obstacle avoidance and path planning checking method and system based on AI intelligence
CN120010517A
Motorcade multi-target dynamic scheduling method and system based on edge calculation
CN120069722A
Method and apparatus for coordinating railway line of road and yard planners
US20060212184A1
Cited By
Warehousing path planning method, system and equipment based on industrial Internet of Things, and medium
CN120409872A
Warehouse path planning method, system, equipment and medium based on industrial Internet of Things
CN120409872B
Method and system for realizing AR navigation based on cold chain warehouse AI remote control
CN120445229A
Intelligent monitoring and self-adaptive scheduling management platform for safe operation of parking lot vehicle based on cloud-side collaborative AI (artificial intelligence)
CN121638581A
Cold chain sorting optimization system and method based on digital twinning and reinforcement learning
CN121764006A