Warehouse scheduling decision method and device, terminal equipment and storage medium
By training and applying a deep belief network model, the problem of low efficiency in equipment scheduling and storage location allocation in the warehousing system was solved, enabling efficient scheduling decisions in the warehousing system and meeting the demand for high throughput.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TONGJI UNIV
- Filing Date
- 2022-08-31
- Publication Date
- 2026-05-29
AI Technical Summary
When faced with urgent orders and a surge in order volume, the existing warehousing system suffers from low efficiency in equipment scheduling and storage allocation, resulting in long total system operation time and difficulty in meeting high throughput demands.
A deep belief network model is adopted, and by training the lift selection learning model, shuttle selection learning model and cargo location priority learning model, a scheduling decision scheme is generated, system operation information is obtained in real time, potential patterns are learned, and on-site operations are guided.
This enabled timely and efficient scheduling decisions for the warehousing system, reduced total operation time, improved equipment utilization and operational efficiency, and met high throughput requirements.
Smart Images

Figure CN115409448B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of warehouse system scheduling and optimization, and in particular to a warehouse scheduling decision-making method, apparatus, terminal equipment, and storage medium. Background Technology
[0002] Currently, the order tasks in the e-commerce logistics industry are characterized by large scale, high timeliness, and high volatility, which places demands on the development of modern logistics warehousing technology towards intensification, automation, integration, and intelligence. Automated intensive warehousing systems are a new type of logistics storage system that integrates high-density automated racking, conveyor belts, multi-level shuttles, elevators, automatic barcode identification systems, and warehouse management systems. They offer advantages such as high space utilization, large storage capacity, high operational efficiency, high throughput, and fast response speed, making them an ideal choice for building smart warehousing centers.
[0003] However, the order processing scenarios in the e-commerce logistics industry are prone to sudden requests such as urgent orders, order insertions, and surges in order volume, placing certain demands on the equipment scheduling and cargo allocation of warehousing systems. Currently, the equipment scheduling and cargo allocation problems in warehousing systems are mostly solved using queuing theory and mixed-integer programming models with optimization algorithms. These models are complex and have poor real-time performance, resulting in slow system response and difficulty in meeting the demands of frequent task changes and system responses. Traditional classical scheduling rules offer better real-time performance, but their operational efficiency is low, equipment utilization is low, leading to a longer total system operation time and making it difficult to meet high throughput requirements.
[0004] Therefore, it is necessary to propose a solution to save the total operation time of warehouse scheduling. Summary of the Invention
[0005] The main purpose of this application is to provide a warehouse scheduling decision-making method, device, terminal equipment, and storage medium, which aims to solve the problems of low efficiency in warehouse scheduling operations, low equipment utilization, and long total system operation time, so as to achieve timely and efficient warehouse system scheduling decisions.
[0006] To achieve the above objectives, this application provides a warehouse scheduling decision-making method, which includes:
[0007] When a goods outbound request is detected, the current attribute characteristic data of the intensive warehousing system is obtained;
[0008] The attribute feature data is input into a pre-trained deep belief network model for scheduling decision-making, generating a scheduling decision scheme. The deep belief network model consists of one or more of the following: a hoist selection learning model, a shuttle selection learning model, and a cargo location priority learning model.
[0009] Optionally, before the step of inputting the attribute feature data into a pre-trained deep belief network model for scheduling decision-making and generating a scheduling decision scheme, the method further includes:
[0010] The deep belief network model obtained through training specifically includes:
[0011] The deep belief network model is trained offline to obtain the offline trained deep belief network model;
[0012] The offline-trained deep belief network model is then trained online to obtain a well-trained deep belief network model.
[0013] Optionally, the step of training the deep belief network model offline to obtain the offline-trained deep belief network model includes:
[0014] An integrated optimization mathematical model is established, and an optimization algorithm is used to solve the integrated optimization mathematical model to obtain a simulated decision scheme.
[0015] The simulated decision-making scheme is imported into a pre-built warehousing system simulation model for partitioning, resulting in simulated label data;
[0016] The pre-established outbound order plan and the simulated decision-making scheme are imported into the warehousing system simulation model to simulate operations and obtain simulated operational attribute status data.
[0017] Simulated attribute feature data is generated based on the aforementioned operational attribute status data;
[0018] Acquire historical interaction data of the intensive warehousing system to generate historical attribute feature data and historical tag data;
[0019] By combining the simulated attribute feature data, the historical attribute feature data, the simulated label data, and the historical label data, the deep belief network model is trained offline to obtain the offline-trained deep belief network model.
[0020] Optionally, the step of training the offline-trained deep belief network model online to obtain a trained deep belief network model includes:
[0021] Acquire online interactive data between the warehouse management system and the actual operation site of the intensive warehousing system;
[0022] The offline-trained deep belief network model is trained online based on the online interaction data to obtain the trained deep belief network model.
[0023] Optionally, the deep belief network model consists of the hoist selection learning model, the shuttle selection learning model, and the cargo location priority learning model. The step of combining the simulated attribute feature data, the historical attribute feature data, the simulated label data, and the historical label data to perform offline training on the deep belief network model to obtain the offline-trained deep belief network model includes:
[0024] Based on the simulated attribute feature data and the historical attribute feature data, first training attribute feature data is generated for offline training of the lift machine selection learning model;
[0025] Based on the simulated label data and the historical label data, first training label data is generated for offline training of the booster selection learning model;
[0026] The hoist selection learning model is trained offline by combining the attribute feature data and the label data used for training, generating a decision scheme for the selected hoist, and obtaining the hoist selection learning model after offline training.
[0027] Combining the selected elevator decision scheme, the simulated attribute feature data, and the historical attribute feature data, second training attribute feature data is generated for offline training of the shuttle selection learning model.
[0028] Combining the selected elevator decision scheme, the simulated label data, and the historical label data, second training label data is generated for offline training of the shuttle selection learning model;
[0029] The shuttle selection learning model is trained offline by combining the attribute feature data and the label data used for the second training, generating a selected hoist-shuttle decision scheme, and obtaining the offline trained shuttle selection learning model.
[0030] Combining the selected hoist-shuttle decision scheme, the simulated attribute feature data, and the historical attribute feature data, a third set of training attribute feature data is generated for offline training of the cargo location priority learning model.
[0031] Combining the selected hoist-shuttle decision scheme with the simulated label data and the historical label data, a third set of training label data is generated for offline training of the cargo location priority learning model.
[0032] The location priority learning model is trained offline by combining the attribute feature data and the label data used for the third training, generating a selected hoist-shuttle-location decision scheme, obtaining the shuttle selection learning model after offline training, and obtaining the deep belief network model after offline training.
[0033] Optionally, the third training attribute feature data includes: attribute features of the selected roadway, attribute features of the selected hoist, attribute features of the selected shuttle, location attribute features of the goods to be shipped, and location priority attribute features generated by comparing locations pairwise.
[0034] Optionally, after the step of inputting the attribute features into a pre-trained deep belief network model for scheduling decision-making and generating a scheduling decision scheme, the method further includes:
[0035] Executing warehouse scheduling tasks according to the aforementioned scheduling decision scheme specifically includes:
[0036] When the selected shuttle car performs a cross-level pickup task, the working status information of the shuttle car on the target level where the target goods are located is detected.
[0037] When it is detected that the shuttle in the target layer is performing a work task, a task transfer strategy is executed, specifically including:
[0038] Cancel the current cross-level pickup task of the selected shuttle vehicle;
[0039] When the shuttle in the target layer is detected to have finished its task, the cross-layer pickup task is executed.
[0040] This application also proposes a warehouse scheduling decision-making device, which includes:
[0041] The data acquisition module is used to acquire the current attribute feature data of the intensive warehousing system when a goods outbound request is detected;
[0042] The scheduling decision module is used to input the attribute feature data into a pre-trained deep belief network model to make scheduling decisions and generate scheduling decision schemes. The deep belief network model is composed of one or more of the following: hoist selection learning model, shuttle selection learning model, and cargo location priority learning model.
[0043] This application also proposes a terminal device, which includes a memory, a processor, and a warehouse scheduling decision program stored in the memory and executable on the processor. When the warehouse scheduling decision program is executed by the processor, it implements the steps of the warehouse scheduling decision method as described above.
[0044] This application also proposes a computer-readable storage medium storing a warehouse scheduling decision program, which, when executed by a processor, implements the steps of the warehouse scheduling decision method described above.
[0045] The warehouse scheduling decision-making method, apparatus, terminal equipment, and storage medium proposed in this application acquire current attribute feature data of the intensive warehousing system when a goods outbound request is detected. This attribute feature data is then input into a pre-trained deep belief network model for scheduling decision-making, generating a scheduling decision scheme. The deep belief network model is composed of one or more of a hoist selection learning model, a shuttle selection learning model, and a storage location priority learning model. By using the trained deep belief network model to schedule goods outbound requests in the intensive warehousing system, the problem of low operational efficiency and low equipment utilization, resulting in long total system operation time, can be solved, achieving timely and efficient scheduling decisions for the warehousing system. This application's solution does not require the establishment of complex mathematical models and assumptions; instead, it learns the potential patterns of intensive warehousing system scheduling from the massive amounts of data accumulated during system operation, providing decision guidance for on-site operations. Based on the goal of minimizing the total system operation time, this application's solution achieves timely and efficient warehouse system scheduling decisions by acquiring real-time system operation information. Attached Figure Description
[0046] Figure 1 This is a schematic diagram of the functional modules of the terminal equipment to which the warehouse scheduling decision-making device belongs in this application;
[0047] Figure 2 This is a flowchart illustrating a first exemplary embodiment of the warehouse scheduling decision-making method of this application;
[0048] Figure 3 This is a schematic diagram of the deep belief network model involved in the embodiment of the warehouse scheduling decision-making method of this application;
[0049] Figure 4 This is a flowchart illustrating a second exemplary embodiment of the warehouse scheduling decision-making method of this application;
[0050] Figure 5 This is a flowchart illustrating a third exemplary embodiment of the warehouse scheduling decision-making method of this application;
[0051] Figure 6 This is a flowchart illustrating the fourth exemplary embodiment of the warehouse scheduling decision method of this application;
[0052] Figure 7 This is a flowchart illustrating the fifth exemplary embodiment of the warehouse scheduling decision-making method of this application;
[0053] Figure 8 This is a flowchart illustrating the seventh exemplary embodiment of the warehouse scheduling decision-making method of this application;
[0054] Figure 9 This is a schematic diagram illustrating the learning and real-time decision-making of an automated, intensive warehousing system scheme involved in the embodiment of the warehousing scheduling decision-making method of this application;
[0055] Figure 10 This is a flowchart illustrating the real-time decision-making process for equipment scheduling and cargo location allocation involved in the embodiment of the warehouse scheduling decision-making method of this application.
[0056] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0057] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0058] The main solution of this application embodiment is as follows: A deep belief network model is trained offline to obtain an offline-trained deep belief network model; the offline-trained deep belief network model is then trained online to obtain a trained deep belief network model; when a goods outbound request is detected, the current attribute feature data of the intensive warehousing system is obtained; the attribute feature data is input into the pre-trained deep belief network model for scheduling decisions, generating a scheduling decision scheme. The deep belief network model is composed of one or more of a hoist selection learning model, a shuttle selection learning model, and a storage location priority learning model. By using the trained deep belief network model to make scheduling decisions for goods outbound requests in the intensive warehousing system, the problems of low operational efficiency, low equipment utilization, and long total system operation time can be solved, achieving timely and efficient scheduling decisions for the warehousing system. This application solution does not require the establishment of complex mathematical models and assumptions; instead, it learns the potential patterns of intensive warehousing system scheduling from the massive amounts of data accumulated during system operation and provides decision guidance for on-site operations. Based on the proposed solution, which aims to minimize the total system operation time, the timely and efficient scheduling decision-making of the warehousing system is achieved by acquiring real-time information on the system's operation process.
[0059] Specifically, refer to Figure 1 , Figure 1This is a schematic diagram of the functional modules of the terminal equipment to which the warehouse scheduling decision-making device of this application belongs. The warehouse scheduling decision-making device can be an independent device capable of warehouse scheduling decisions and network model training, and can be implemented on the terminal equipment in hardware or software form. The terminal equipment can be a smart mobile terminal with data processing capabilities, such as a mobile phone or computer, or a fixed terminal device or server with data processing capabilities.
[0060] In this embodiment, the terminal equipment of the warehouse scheduling decision-making device includes at least an output module 110, a processor 120, a memory 130, and a communication module 140.
[0061] The memory 130 stores the operating system and the warehouse scheduling decision program. The warehouse scheduling decision device can store the following information in the memory 130: the current attribute feature data of the acquired intensive warehousing system; the scheduling decision scheme generated after inputting the attribute feature data into a pre-trained deep belief network model for scheduling decision-making; the established integrated optimization mathematical model; the simulated decision scheme obtained by solving the integrated optimization mathematical model using an optimization algorithm; the simulated label data obtained by importing the simulated decision scheme into a pre-constructed warehouse system simulation model for partitioning; the established outbound order plan; the simulated operation attribute status data obtained by importing the outbound order plan and the simulated decision scheme into the warehouse system simulation model for simulated operation; the simulated attribute feature data generated based on the operation attribute status data; the historical interaction data of the acquired intensive warehousing system; the generated historical attribute feature data and historical label data; and the online interaction data between the acquired warehouse management system and the actual operation site of the intensive warehousing system. The output module 110 can be a display screen, etc. The communication module 140 may include a WIFI module, a mobile communication module, and a Bluetooth module, and can communicate with external devices or servers through the communication module 140.
[0062] When the warehouse scheduling decision program in memory 130 is executed by the processor, it performs the following steps:
[0063] When a goods outbound request is detected, the current attribute characteristic data of the intensive warehousing system is obtained;
[0064] The attribute feature data is input into a pre-trained deep belief network model for scheduling decision-making, generating a scheduling decision scheme. The deep belief network model consists of one or more of the following: a hoist selection learning model, a shuttle selection learning model, and a cargo location priority learning model.
[0065] Furthermore, when the warehouse scheduling decision program in memory 130 is executed by the processor, it also performs the following steps:
[0066] The deep belief network model obtained through training specifically includes:
[0067] The deep belief network model is trained offline to obtain the offline trained deep belief network model;
[0068] The offline-trained deep belief network model is then trained online to obtain a well-trained deep belief network model.
[0069] Furthermore, when the warehouse scheduling decision program in memory 130 is executed by the processor, it also performs the following steps:
[0070] An integrated optimization mathematical model is established, and an optimization algorithm is used to solve the integrated optimization mathematical model to obtain a simulated decision scheme.
[0071] The simulated decision-making scheme is imported into a pre-built warehousing system simulation model for partitioning, resulting in simulated label data;
[0072] The pre-established outbound order plan and the simulated decision-making scheme are imported into the warehousing system simulation model to simulate operations and obtain simulated operational attribute status data.
[0073] Simulated attribute feature data is generated based on the aforementioned operational attribute status data;
[0074] Acquire historical interaction data of the intensive warehousing system to generate historical attribute feature data and historical tag data;
[0075] By combining the simulated attribute feature data, the historical attribute feature data, the simulated label data, and the historical label data, the deep belief network model is trained offline to obtain the offline-trained deep belief network model.
[0076] Furthermore, when the warehouse scheduling decision program in memory 130 is executed by the processor, it also performs the following steps:
[0077] Acquire online interactive data between the warehouse management system and the actual operation site of the intensive warehousing system;
[0078] The offline-trained deep belief network model is trained online based on the online interaction data to obtain the trained deep belief network model.
[0079] Furthermore, when the warehouse scheduling decision program in memory 130 is executed by the processor, it also performs the following steps:
[0080] Based on the simulated attribute feature data and the historical attribute feature data, first training attribute feature data is generated for offline training of the lift machine selection learning model;
[0081] Based on the simulated label data and the historical label data, first training label data is generated for offline training of the booster selection learning model;
[0082] The hoist selection learning model is trained offline by combining the attribute feature data and the label data used for training, generating a decision scheme for the selected hoist, and obtaining the hoist selection learning model after offline training.
[0083] Combining the selected elevator decision scheme, the simulated attribute feature data, and the historical attribute feature data, second training attribute feature data is generated for offline training of the shuttle selection learning model.
[0084] Combining the selected elevator decision scheme, the simulated label data, and the historical label data, second training label data is generated for offline training of the shuttle selection learning model;
[0085] The shuttle selection learning model is trained offline by combining the attribute feature data and the label data used for the second training, generating a selected hoist-shuttle decision scheme, and obtaining the offline trained shuttle selection learning model.
[0086] Combining the selected hoist-shuttle decision scheme, the simulated attribute feature data, and the historical attribute feature data, a third set of training attribute feature data is generated for offline training of the cargo location priority learning model.
[0087] Combining the selected hoist-shuttle decision scheme with the simulated label data and the historical label data, a third set of training label data is generated for offline training of the cargo location priority learning model.
[0088] The location priority learning model is trained offline by combining the attribute feature data and the label data used for the third training, generating a selected hoist-shuttle-location decision scheme, obtaining the shuttle selection learning model after offline training, and obtaining the deep belief network model after offline training.
[0089] Furthermore, when the warehouse scheduling decision program in memory 130 is executed by the processor, it also performs the following steps:
[0090] Executing warehouse scheduling tasks according to the aforementioned scheduling decision scheme specifically includes:
[0091] When the selected shuttle car performs a cross-level pickup task, the working status information of the shuttle car on the target level where the target goods are located is detected.
[0092] When it is detected that the shuttle in the target layer is performing a work task, a task transfer strategy is executed, specifically including:
[0093] Cancel the current cross-level pickup task of the selected shuttle vehicle;
[0094] When the shuttle in the target layer is detected to have finished its task, the cross-layer pickup task is executed.
[0095] This embodiment, through the above-described scheme, specifically acquires the current attribute feature data of the intensive warehousing system when a goods outbound request is detected; inputs the attribute feature data into a pre-trained deep belief network model for scheduling decision-making, generating a scheduling decision scheme. The deep belief network model is composed of one or more of a hoist selection learning model, a shuttle selection learning model, and a storage location priority learning model. By using the trained deep belief network model to make scheduling decisions for goods outbound requests in the intensive warehousing system, the problems of low operational efficiency, low equipment utilization, and long total system operation time can be solved, achieving the goal of timely and efficient scheduling decisions for the warehousing system. This application scheme does not require the establishment of complex mathematical models and assumptions; instead, it learns the potential patterns of intensive warehousing system scheduling from the massive data accumulated during system operation and provides decision-making guidance for on-site operations. Based on the goal of minimizing the total system operation time, this application scheme achieves timely and efficient scheduling decisions for the warehousing system by acquiring real-time system operation information.
[0096] Based on, but not limited to, the terminal device architecture described above, this application proposes method embodiments.
[0097] Reference Figure 2 , Figure 2 This is a flowchart illustrating a first exemplary embodiment of the warehouse scheduling decision-making method of this application. The executing entity of this embodiment can be a warehouse scheduling decision-making device, a warehouse scheduling decision-making terminal device, or a server. This embodiment uses a warehouse scheduling decision-making device as an example, which can be integrated into terminal devices such as smartphones and tablets with data processing capabilities. The warehouse scheduling decision-making method includes:
[0098] Step S1001: When a goods outbound request is detected, obtain the current attribute feature data of the intensive warehousing system.
[0099] Specifically, when the system detects a goods outbound request, it acquires the current attribute feature data of the intensive warehousing system. This attribute feature data is related to the operational status information of the intensive warehousing system and can be extracted from the system's operational status information. In this embodiment, the intensive warehousing system can be a logistics storage system integrating high-density automated racking, conveyor belts, multi-level shuttles, elevators, barcode automatic identification systems, and warehouse management systems. Generally, the structural distribution of an intensive warehousing system includes: every two rows of racks form an aisle; each aisle entrance is equipped with a shuttle elevator and a cargo elevator; each type of goods is stored in multiple locations on different racks. Each aisle is equipped with several shuttles, which can retrieve goods across levels via the shuttle elevator, but can only move within fixed aisles. Goods can be outbound via location elevators. Each level of the racking has a cargo buffer area at its entrance. When an aisle is selected, the corresponding shuttle elevator and cargo elevator are simultaneously selected.
[0100] Step S1002: Input the attribute feature data into a pre-trained deep belief network model for scheduling decision-making and generate a scheduling decision scheme. The deep belief network model is composed of one or more of the following: hoist selection learning model, shuttle selection learning model, and cargo location priority learning model.
[0101] Specifically, the acquired attribute feature data is input into a pre-trained deep belief network model for scheduling decisions, generating a scheduling decision scheme. For example... Figure 3 As shown, Figure 3 This is a schematic diagram of the deep belief network model involved in the warehouse scheduling decision-making method of this application. The structure of the deep belief network model is constructed by sequentially stacking Restricted Boltzmann Machines (RBMs), that is, the output of the previous RBM is the input of the next RBM. The learning process of this network structure is divided into two stages: first, unsupervised pre-training is performed layer by layer on the RBMs, and then the entire network is tuned in a supervised manner using the backpropagation algorithm (BP).
[0102] RBM is a generative model with a two-layer neural network structure. The first layer is the visible layer v, which receives input and consists of n visible units v = (v1, v2, ..., vn). n The system is composed of m hidden units, generally following a Bernoulli or Gaussian distribution. The second layer is the hidden layer h, consisting of m hidden units h = (h1, h2, ..., h...). m The structure generally follows a Bernoulli distribution. All visible and hidden units are fully connected, while units within the visible and hidden layers themselves are not connected to each other; that is, there is full connectivity between layers and no connectivity within layers. The weights between connected neurons are w = {w...} ij}∈R n×m, i=1,2,…,n, j=1,2,…,m. a={a i}∈R n and b = {b j}∈R m , where represent the biases of the i-th visible unit and the j-th hidden unit, respectively. As shown in Equation 1 below, for an RBM where both v and h follow a Bernoulli distribution, its energy function is:
[0103]
[0104] In the formula, v i and h j Let w represent the binary states of the i-th visible unit and the j-th hidden unit, respectively. ij Let represent the weight between the i-th visible unit and the j-th hidden unit. Lower energy indicates that the network is in a more ideal state, i.e., with the lowest learning error. After regularizing and exponentializing this energy function, we obtain the joint probability distribution formula for a set of states of visible and hidden nodes, as shown in Equations 2 and 3 below:
[0105]
[0106]
[0107] The conditional probability distributions of visible neurons and hidden neurons are shown in Equations 4 and 5 below:
[0108]
[0109]
[0110] Given v and h, the hidden layer unit h can be obtained. j and visible layer unit v i The activation state probability is calculated. A contrastive divergence algorithm is used to solve the problem layer by layer to obtain the optimal solution for the RBM weights at each layer. By calculating the gradient of the log-likelihood function logP(v,h|θ), the RBM weight update formula can be obtained, as shown in equations 6 and 7 below:
[0111]
[0112]
[0113] In the formula, τ and η represent the number of iterations and the learning rate of the RBM, respectively, and E data (v i h j ) and E model (v i h j ) represent the expectation of the observed data in the training set and the expectation of the distribution determined by the model, respectively.
[0114] Forward stacked RBM learning is an unsupervised learning method. After greedily training each layer of the RBM, the initial weights W = {w1, w2, ... w} are obtained. l This essentially provides prior knowledge of the input data for supervised learning. Backward fine-tuning learning starts from the output layer of the DBN network and uses known labels to progressively fine-tune the model parameters towards the input layer. That is, it uses the BP algorithm to perform supervised training of the network, fine-tuning the parameters from the output layer to the input layer to reduce gradients. The final layer of the deep belief network model uses a softmax classifier model, and the model's output is presented as a probability. The class corresponding to the highest probability is selected as the model's decision.
[0115] Based on the above network structure, the deep belief network model can be composed of one or more of the following: hoist selection learning model, shuttle selection learning model, and cargo location priority learning model.
[0116] This embodiment, through the above-described scheme, specifically acquires the current attribute feature data of the intensive warehousing system when a goods outbound request is detected; inputs the attribute feature data into a pre-trained deep belief network model for scheduling decision-making, generating a scheduling decision scheme. The deep belief network model is composed of one or more of a hoist selection learning model, a shuttle selection learning model, and a storage location priority learning model. By using the trained deep belief network model to make scheduling decisions for goods outbound requests in the intensive warehousing system, the problems of low operational efficiency, low equipment utilization, and long total system operation time can be solved, achieving the goal of timely and efficient scheduling decisions for the warehousing system. This application scheme does not require the establishment of complex mathematical models and assumptions; instead, it learns the potential patterns of intensive warehousing system scheduling from the massive data accumulated during system operation and provides decision-making guidance for on-site operations. Based on the goal of minimizing the total system operation time, this application scheme achieves timely and efficient scheduling decisions for the warehousing system by acquiring real-time system operation information.
[0117] Reference Figure 4 , Figure 4 This is a flowchart illustrating a second exemplary embodiment of the warehouse scheduling decision method of this application. Based on the above... Figure 2 In the embodiment shown, prior to the step of obtaining the current attribute characteristic data of the intensive warehousing system when a goods outbound request is detected, the warehousing scheduling decision method further includes:
[0118] Step S1000: The deep belief network model is trained. In this embodiment, step S1000 is performed before step S1001. In other embodiments, step S1000 can also be performed between step S1001 and step S1002.
[0119] Compared to the above Figure 2 The embodiment shown also includes a scheme for training the deep belief network model.
[0120] Specifically, the steps for training the deep belief network model may include:
[0121] Step S1100: The deep belief network model is trained offline to obtain the offline trained deep belief network model.
[0122] Step S1200: The offline-trained deep belief network model is trained online to obtain a trained deep belief network model.
[0123] More specifically, in this embodiment, the constructed deep belief network model is trained offline. After offline training, an offline-trained deep belief network model is obtained. Offline training refers to training the deep belief network model in a simulated or analog environment using pre-acquired training sample data. Then, the offline-trained deep belief network model is trained online. After online training, an online-trained deep belief network model is obtained. Online training refers to training the deep belief network model in a real environment using incremental sample data generated from actual field operations. Finally, after training is complete, a trained deep belief network model is obtained.
[0124] Then, the intensive warehousing system can be scheduled and decided using the trained deep belief network model.
[0125] This embodiment, through the above-described scheme, specifically obtains the deep belief network model through training; when a goods outbound request is detected, it acquires the current attribute feature data of the intensive warehousing system; the attribute feature data is input into the pre-trained deep belief network model for scheduling decision-making, generating a scheduling decision scheme. The deep belief network model is composed of one or more of the following: a hoist selection learning model, a shuttle selection learning model, and a storage location priority learning model. By using the trained deep belief network model to make scheduling decisions for goods outbound requests in the intensive warehousing system, the problems of low operational efficiency, low equipment utilization, and long total system operation time can be solved, achieving the goal of timely and efficient scheduling decisions for the warehousing system. This application scheme eliminates the need for complex mathematical models and assumptions; instead, it learns the potential patterns of intensive warehousing system scheduling from the massive amounts of data accumulated during system operation, providing decision guidance for on-site operations. Based on the goal of minimizing the total system operation time, this application scheme achieves timely and efficient scheduling decisions for the warehousing system by acquiring real-time system operation information.
[0126] Furthermore, referring to Figure 5 , Figure 5 This is a flowchart illustrating a third exemplary embodiment of the warehouse scheduling decision-making method of this application. Based on the above... Figure 4 In the embodiment shown, step S1100, which involves offline training of the deep belief network model to obtain the offline-trained deep belief network model, may include:
[0127] Step S1110: Establish an integrated optimization mathematical model, and use an optimization algorithm to solve the integrated optimization mathematical model to obtain a simulated decision scheme.
[0128] Specifically, based on reasonable assumptions and constraints of the system operation process, an integrated optimization mathematical model is established with the goal of minimizing the total operation time. Then, the established integrated optimization mathematical model is solved using an optimization algorithm to obtain a simulated decision scheme. The simulated decision scheme is an approximately optimal warehouse scheduling decision scheme, and the content of the decision scheme may include equipment scheduling and storage location allocation.
[0129] Step S1120: Import the simulated decision scheme into the pre-built warehousing system simulation model for partitioning to obtain simulated label data.
[0130] Specifically, the obtained simulated decision scheme is imported into a pre-built warehousing system simulation model to divide the scheduling scheme instructions. The scheduling scheme instructions can be divided into equipment scheduling instructions and storage location allocation instructions. Then, based on the obtained scheduling scheme instructions, simulated label data for training is generated.
[0131] Step S1130: Import the pre-established outbound order plan and the simulated decision scheme into the warehousing system simulation model to perform simulation operations and obtain simulated operation attribute status data.
[0132] Specifically, an outbound order plan is established based on the problem example. The pre-established outbound order plan and the simulated decision-making scheme are imported into the pre-constructed warehouse system simulation model. The operation process is simulated through the warehouse system simulation model, and the simulated operational attribute status data is extracted. This operational attribute status data is related to the operation of the warehouse system, such as equipment scheduling and storage location allocation.
[0133] Step S1140: Generate simulated attribute feature data based on the running attribute status data.
[0134] Specifically, attribute feature data for training simulations is generated based on the obtained running attribute state data.
[0135] Step S1150: Obtain historical interaction data of the intensive warehousing system, and generate historical attribute feature data and historical tag data.
[0136] Specifically, historical interaction data of the intensive warehousing system is acquired. This historical interaction data is generated by the warehousing management system and the intensive warehousing system during historical operations. The content of the historical interaction data may include, but is not limited to, historical operational attribute status data and historical scheduling scheme instructions. Then, historical attribute feature data and historical tag data are generated based on the historical interaction data.
[0137] Step S1160: Combine the simulated attribute feature data, the historical attribute feature data, the simulated label data, and the historical label data to perform offline training on the deep belief network model, thereby obtaining the offline-trained deep belief network model.
[0138] Specifically, the simulated attribute feature data, the historical attribute feature data, the simulated label data, and the historical label data are combined, that is, the simulated data and the historical data are merged to generate attribute feature data and label data for training. The deep belief network model is then trained offline, and after training is completed, the offline trained deep belief network model is obtained.
[0139] This embodiment constructs operational status attribute features of an intensive warehousing system based on system operation data. By extracting key attributes that affect equipment scheduling and storage location allocation during system operation, it establishes attribute features within corresponding ranges for different model learning objectives, thereby reducing the dimensionality of model input and training time.
[0140] This embodiment, following the above-described scheme, specifically obtains the deep belief network model through training. When a goods outbound request is detected, the current attribute feature data of the intensive warehousing system is acquired. This attribute feature data is input into the pre-trained deep belief network model for scheduling decision-making, generating a scheduling decision scheme. The deep belief network model is composed of one or more of the following: a hoist selection learning model, a shuttle selection learning model, and a storage location priority learning model. By using the trained deep belief network model to make scheduling decisions for goods outbound requests in the intensive warehousing system, the problems of low operational efficiency, low equipment utilization, and long total system operation time can be solved, achieving timely and efficient scheduling decisions for the warehousing system. This application scheme eliminates the need for complex mathematical models and assumptions; instead, it learns the potential patterns of intensive warehousing system scheduling from the massive amounts of data accumulated during system operation, providing decision guidance for on-site operations. Based on the goal of minimizing the total system operation time, this application scheme achieves timely and efficient scheduling decisions for the warehousing system by acquiring real-time system operation information.
[0141] Furthermore, referring to Figure 6 , Figure 6 This is a flowchart illustrating a fourth exemplary embodiment of the warehouse scheduling decision-making method of this application. Based on the above... Figure 5 In the embodiment shown, step S1200, which involves training the offline-trained deep belief network model online to obtain a trained deep belief network model, may include:
[0142] Step S1210: Obtain online interactive data between the warehouse management system and the actual operation site of the intensive warehousing system.
[0143] Specifically, the system acquires online interaction data generated by the warehouse management system and the intensive warehousing system at the actual operational site. This online interaction data may include, but is not limited to, online operational attribute status data and online scheduling instructions. Based on this online interaction data, online attribute feature data and online tag data can be generated.
[0144] Step S1220: Train the offline-trained deep belief network model online based on the online interaction data to obtain the trained deep belief network model.
[0145] Specifically, the offline-trained deep belief network model is trained online based on the online interaction data. That is, the online attribute feature data and online label data generated from the online interaction data are input into the offline-trained deep belief network model for online training. After training is completed, the online-trained deep belief network model is obtained. At this point, the trained deep belief network model is obtained.
[0146] Then, the intensive warehousing system can be scheduled and decided using the trained deep belief network model.
[0147] This embodiment, through the above-described scheme, specifically obtains the deep belief network model through training; when a goods outbound request is detected, it acquires the current attribute feature data of the intensive warehousing system; the attribute feature data is input into the pre-trained deep belief network model for scheduling decision-making, generating a scheduling decision scheme. The deep belief network model is composed of one or more of the following: a hoist selection learning model, a shuttle selection learning model, and a storage location priority learning model. By using the trained deep belief network model to make scheduling decisions for goods outbound requests in the intensive warehousing system, the problems of low operational efficiency, low equipment utilization, and long total system operation time can be solved, achieving the goal of timely and efficient scheduling decisions for the warehousing system. This application scheme eliminates the need for complex mathematical models and assumptions; instead, it learns the potential patterns of intensive warehousing system scheduling from the massive amounts of data accumulated during system operation, providing decision guidance for on-site operations. Based on the goal of minimizing the total system operation time, this application scheme achieves timely and efficient scheduling decisions for the warehousing system by acquiring real-time system operation information.
[0148] Reference Figure 7 , Figure 7 This is a flowchart illustrating a fifth exemplary embodiment of the warehouse scheduling decision-making method of this application. Based on the above... Figure 6 In the embodiment shown, the deep belief network model consists of the hoist selection learning model, the shuttle selection learning model, and the cargo location priority learning model. Step S1160 involves combining the simulated attribute feature data, the historical attribute feature data, the simulated label data, and the historical label data to perform offline training on the deep belief network model. The resulting offline-trained deep belief network model may include:
[0149] Step S1161: Based on the simulated attribute feature data and the historical attribute feature data, generate first training attribute feature data for offline training of the lift machine selection learning model.
[0150] Specifically, based on the obtained simulated attribute feature data and the historical attribute feature data, training attribute feature data is generated for offline training of the hoist selection learning model, and this data serves as the first training attribute feature data. The hoist selection learning model is used to generate scheduling decision schemes for selected lanes and selected hoists. The attribute feature data for the first training includes continuous and Boolean types, and its content includes characteristics of all lanes and equipped hoists within the warehousing system, such as the number of idle shuttles in each lane, the number of shuttles reaching the target layer in each lane, the lane's storage capacity, whether the shuttle hoist is idle, and the target layer of the shuttle hoist. Assuming the warehousing system's automated storage and retrieval system has m lanes, then it is equipped with m shuttle hoists and m cargo hoists. For each lane and hoist's attributes, m corresponding attribute features need to be established.
[0151] The calculation of the lane storage rate SR_lane, a characteristic attribute, is shown in Formula 8 below:
[0152]
[0153] Where, n r1 and n r2 N refers to the number of occupied shelving units on both sides of the aisle. r1 and N r2 These refer to the total number of units on both sides of the aisle.
[0154] Step S1162: Generate training label data for offline training of the booster selection learning model based on the simulated label data and the historical label data.
[0155] Specifically, based on the obtained simulated label data and the historical label data, training label data for offline training of the lift machine selection learning model is generated, and this data is used as the first training label data, wherein the first training label data is the lift machine number.
[0156] Step S1163: Combine the first training attribute feature data and the first training label data to perform offline training on the lift selection learning model, generate the selected lift decision scheme, and obtain the offline trained lift selection learning model.
[0157] Specifically, the hoist selection learning model is trained offline by combining the obtained first training attribute feature data and the first training label data to generate a selected hoist decision scheme. The content of the selected hoist decision scheme includes the selected roadway and the selected hoist, as well as their attribute feature data and label data. At this time, the hoist selection learning model after offline training is obtained.
[0158] Step S1164: Combining the selected elevator decision scheme, the simulated attribute feature data, and the historical attribute feature data, generate second training attribute feature data for offline training of the shuttle selection learning model.
[0159] Specifically, in conjunction with the selected hoist decision-making scheme, which includes attribute feature data of the selected roadway and the selected hoist, as well as the simulated attribute feature data and the historical attribute feature data, training attribute feature data for offline training of the shuttle selection learning model is generated, and this data serves as the second training attribute feature data. The shuttle selection learning model is used to generate scheduling decision schemes for the selected roadway, the selected hoist, and the selected shuttle. The attribute types of the second training attribute feature data include continuous and Boolean types, and its content includes the attribute features of the selected roadway, the attribute features of the selected hoist, and the attribute features of all shuttles equipped in the selected roadway, such as whether the shuttle is idle, whether the shuttle is in the target layer, the shuttle's target layer, the number of tasks to be completed by the shuttle, and the shuttle's task completion rate. Assuming each roadway has k shuttles, for each shuttle's attribute, k corresponding attribute features need to be established.
[0160] The calculation of the attribute feature shuttle task completion degree S_comp is shown in Formula 9 below:
[0161]
[0162] Among them, T comp T refers to the number of tasks completed by the shuttle. total This refers to the total number of tasks completed by the shuttle up to the current moment.
[0163] Step S1165: Combining the selected elevator decision scheme, the simulated label data, and the historical label data, generate second training label data for offline training of the shuttle selection learning model.
[0164] Specifically, in conjunction with the selected hoist decision scheme, including the label data of the selected roadway and the selected hoist, as well as the simulated label data and the historical label data, training label data for offline training of the shuttle selection learning model is generated, and used as the second training label data, wherein the second training label data is the shuttle number.
[0165] Step S1166: Combine the second training attribute feature data and the second training label data to perform offline training on the shuttle selection learning model, generate the selected elevator-shuttle decision scheme, and obtain the offline trained shuttle selection learning model.
[0166] Specifically, the shuttle selection learning model is trained offline by combining the obtained second training attribute feature data and the second training label data to generate a selected hoist-shuttle decision scheme. The selected hoist-shuttle decision scheme includes the selected lane, the selected hoist and the selected shuttle, as well as their attribute feature data and label data. At this time, the offline trained shuttle selection learning model is obtained.
[0167] Step S1167: Combining the selected hoist-shuttle decision scheme, the simulated attribute feature data, and the historical attribute feature data, generate third training attribute feature data for offline training of the cargo location priority learning model.
[0168] Specifically, in conjunction with the selected elevator-shuttle decision scheme, which includes attribute feature data of the selected aisle, selected elevator, and selected shuttle, as well as the simulated attribute feature data and the historical attribute feature data, training attribute feature data is generated for offline training of the location priority learning model, and serves as the third training attribute feature data. The location priority learning model is used to generate scheduling decision schemes for the selected aisle, selected elevator, selected shuttle, and selected location. The attribute types of the third training attribute feature data include continuous and Boolean types, and its content includes attribute features of the selected aisle, the selected elevator, the selected shuttle, and all location locations in the selected aisle that currently require outbound shipment, as well as other location-related features, such as whether there is an unselected shuttle on the location's floor, the shelf where the location is located, the floor where the location is located, the storage capacity of the shelf where the location is located, and the time required for the shuttle to reach the location.
[0169] The location rate SR_rack of the shelf where the attribute-characteristic storage location is located is calculated as shown in Formula 10 below:
[0170]
[0171] Where, n r N refers to the number of shelving units currently in use. r This refers to the total number of units on the shelf.
[0172] Among them, the time required for the shuttle vehicle to reach the cargo location is characterized by its attributes. The calculation is shown in Formula 11 below:
[0173]
[0174] Among them, T load_1 The time it takes for a cargo container to be loaded from the rack to the shuttle, T load_2 The time it takes for the shuttle car to move from the track into the hoist, T unload_1 The time it takes for the cargo container to be unloaded from the shuttle to the buffer area, T unload_2 The time it takes for the shuttle car to move from the hoist onto the track, z shuttle The floor where the shuttle is located, z cargo This refers to the floor where the target storage location is located. This represents the travel time of the horizontally upward shuttle from the standby point to the cargo location, calculated as shown in Formula 12 below:
[0175]
[0176] in, y1 and y2 represent the locations of the standby point and the storage location, respectively; w0 represents the width of the shelving unit; a s The acceleration of the shuttle, v s 'c' represents the maximum speed of the shuttle, and 'c' represents the total number of columns on the rack.
[0177] This represents the travel time of the shuttle car across floors via the elevator in the vertical direction, and the calculation process is shown in Formula 13 below:
[0178]
[0179] Where z1 and z2 represent the floor where the shuttle is located and the floor where the target storage location is located, respectively, h0 represents the height of the racking unit, and a l v represents the acceleration of the hoist. l This indicates the maximum travel speed of the elevator, and n-tier indicates the total number of shelves.
[0180] Step S1168: Combining the selected hoist-shuttle decision scheme, the simulated label data, and the historical label data, generate third training label data for offline training of the cargo location priority learning model.
[0181] Specifically, in conjunction with the selected hoist-shuttle decision scheme, which includes label data of the selected lane, the selected hoist, and the selected shuttle, as well as the simulated label data and the historical label data, training label data is generated for offline training of the cargo location priority learning model, and is used as the third training label data, wherein the third training label data is cargo location priority "high" and "low".
[0182] Step S1169: Combine the attribute feature data and label data used for the third training to perform offline training on the cargo location priority learning model, generate the selected hoist-shuttle-cargo location decision scheme, obtain the shuttle selection learning model after offline training, and obtain the deep belief network model after offline training.
[0183] Specifically, by combining the obtained third training attribute feature data and the third training label data, the cargo location priority learning model is trained offline to generate a selected hoist-shuttle-cargo location decision scheme. The selected hoist-shuttle-cargo location decision scheme includes the selected lane, the selected hoist, the selected shuttle, and the selected cargo location, as well as their attribute feature data and label data. At this time, the offline-trained shuttle selection learning model is obtained, and the offline-trained deep belief network model is also obtained.
[0184] In this embodiment, the attribute feature data and construction scope of the same object source differ for different learning objectives, as shown in Table 1 below. The detailed description of the attribute features established for learning hoist selection, shuttle selection, and cargo location priority is shown in Table 2 below:
[0185]
[0186]
[0187] Table 1: Attribute characteristics and construction scope of the same object source
[0188]
[0189] Table 2: Attributes and features used for learning elevator selection, shuttle selection, and storage location priority.
[0190] This embodiment establishes a three-stage deep belief network model for hoist selection, shuttle selection, and cargo location priority, aiming to minimize the system's operation time and achieve integrated learning of optimal scheduling solutions.
[0191] This embodiment, through the above-described scheme, specifically obtains the deep belief network model through training; when a goods outbound request is detected, it acquires the current attribute feature data of the intensive warehousing system; the attribute feature data is input into the pre-trained deep belief network model for scheduling decision-making, generating a scheduling decision scheme. The deep belief network model is composed of one or more of the following: a hoist selection learning model, a shuttle selection learning model, and a storage location priority learning model. By using the trained deep belief network model to make scheduling decisions for goods outbound requests in the intensive warehousing system, the problems of low operational efficiency, low equipment utilization, and long total system operation time can be solved, achieving the goal of timely and efficient scheduling decisions for the warehousing system. This application scheme eliminates the need for complex mathematical models and assumptions; instead, it learns the potential patterns of intensive warehousing system scheduling from the massive amounts of data accumulated during system operation, providing decision guidance for on-site operations. Based on the goal of minimizing the total system operation time, this application scheme achieves timely and efficient scheduling decisions for the warehousing system by acquiring real-time system operation information.
[0192] Furthermore, based on the above embodiments, in this embodiment, the attribute feature data used for the third training includes: attribute features of the selected roadway, attribute features of the selected hoist, attribute features of the selected shuttle, location attribute features of the goods to be shipped, and location priority attribute features generated by comparing locations pairwise.
[0193] Specifically, the third set of training attribute feature data used to train the storage location priority learning model includes: attribute features of the selected aisle, attribute features of the selected hoist, attribute features of the selected shuttle, storage location attribute features of the goods to be shipped, and attribute features of storage location priority generated through pairwise comparison of storage locations. The pairwise comparison of storage locations is used to construct extended attribute features for storage location priority. Specifically, storage location cl1 is set as A, and storage location cl2 as B. Taking A as the object, B is compared with A. For example, for the attribute feature "number of storage location layers", a Boolean type feature "number of storage location layers (A) < number of storage location layers (B)" is constructed, and so on. If the selection result is A priority over B, the priority output is "high". Next, storage location cl2 is set as A, and storage location cl1 as B, and the same method is used; the priority output is "low". Assuming that a certain type of goods has n storage locations in a certain aisle, the above method needs to be used to establish... For attribute characteristics.
[0194] This embodiment proposes a pairwise comparison method for the storage location priority learning model, which expands the Boolean type attribute features and improves the learning accuracy of the model.
[0195] like Figure 8 As shown, Figure 8This is a flowchart illustrating the seventh exemplary embodiment of the warehouse scheduling decision method of this application. Based on the above embodiment, in this implementation, after step S1002, where the attribute features are input into a pre-trained deep belief network model for scheduling decision-making and a scheduling decision scheme is generated, the method further includes: step S1003, where the warehouse scheduling task is executed according to the scheduling decision scheme.
[0196] Specifically, the steps for executing warehouse scheduling tasks according to the scheduling decision scheme may include:
[0197] When the selected shuttle car performs a cross-level pickup task, the working status information of the shuttle car on the target level where the target goods are located is detected.
[0198] When it is detected that the shuttle in the target layer is performing a work task, the task transfer strategy is executed.
[0199] Specifically, during the selected shuttle's task execution, when a cross-level cargo retrieval task needs to be achieved through the selected elevator, the working status information of the shuttle on the target level where the target cargo is located is first detected; when it is detected that the shuttle on the target level is performing a work task on that target level, the selected shuttle executes a task transfer strategy to avoid shuttle conflicts during task execution.
[0200] Furthermore, the step of executing a task transfer strategy when it is detected that the shuttle in the target layer is performing a work task may include:
[0201] Cancel the current cross-level pickup task of the selected shuttle vehicle;
[0202] When the shuttle in the target layer is detected to have finished its task, the cross-layer pickup task is executed.
[0203] Specifically, when the selected shuttle is preparing to perform a cross-level retrieval task, if it is detected that a shuttle on the target level is already performing a task on that level, the selected shuttle's current cross-level retrieval task is canceled, and the execution time of that task is transferred to the next moment of the current shuttle performing the task on the target level. In other words, on the target level, after the shuttle performing its current task completes its task, the selected shuttle will begin performing the cross-level retrieval task. Therefore, when it is detected that the shuttle on the target level has finished performing its task, the selected shuttle begins performing the cross-level retrieval task.
[0204] This embodiment, through the above-described scheme, specifically detects the working status information of the shuttle on the target layer where the target goods are located when the selected shuttle is performing a cross-layer pickup task; when it is detected that the shuttle on the target layer is performing a work task, a task transfer strategy is executed, which can effectively avoid shuttle conflicts during task execution.
[0205] like Figure 9 As shown, Figure 9 This is a schematic diagram illustrating the scheme learning and real-time decision-making of an automated high-density warehousing system, as described in the embodiment of the warehousing scheduling decision-making method of this application. In this embodiment, the process of scheme learning and real-time decision-making for the automated high-density warehousing system may include: real-time decision-making based on deep belief networks and on-site operation of the automated high-density warehousing system. The automated high-density warehousing system provides real-time system operation information to the deep belief network, while the deep belief network provides real-time decision-making guidance for the on-site operation of the automated high-density warehousing system.
[0206] The real-time decision-making process based on deep belief networks includes:
[0207] First, based on real-time system operation information, the attribute status information of the roadway and the hoist is obtained, and the first attribute feature data is extracted and generated. The first attribute feature data is then input into the hoist selection learning model for scheduling decision-making, and a decision scheme for the selected hoist is generated.
[0208] Then, combining the selected hoist decision scheme, the attribute status information of the selected roadway and hoist, as well as the attribute status information of the shuttle, are obtained, and second attribute feature data is extracted and generated. The second attribute feature data is then input into the shuttle selection learning model for scheduling decision-making, thereby generating the selected hoist-shuttle decision scheme.
[0209] Finally, based on the selected hoist-shuttle decision scheme, the attribute status information of the selected roadway and hoist, the attribute status information of the selected shuttle, and the attribute status information of the cargo location are obtained. The third attribute feature data is extracted and generated. The third attribute feature data is input into the cargo location priority learning model for scheduling decision-making, and the selected hoist-shuttle-cargo location decision scheme is generated.
[0210] Guided by the selected hoist-shuttle-location decision scheme generated by the deep belief network, the automated high-density warehousing system executes relevant warehousing scheduling tasks and provides the deep belief network with the real-time system operation information generated during the task execution process.
[0211] This embodiment integrates the trained three-stage deep belief network model into the warehouse management system using the above-described scheme. By acquiring real-time system operation attribute status information, it enables immediate and efficient decision-making regarding hoist and shuttle scheduling and cargo location allocation. Compared with traditional classical scheduling rules, it achieves a shorter total system operation time under different task scales.
[0212] like Figure 10 As shown, Figure 10This is a flowchart illustrating the real-time decision-making process for equipment scheduling and storage location allocation involved in an embodiment of the warehouse scheduling decision-making method of this application. In this embodiment, the real-time decision-making process for equipment scheduling and storage location allocation of the warehouse scheduling decision-making method may include:
[0213] When a request for outbound goods is received from the system at a certain time t, the system first constructs attribute features of the current system operation site for learning the selection of roadways and elevators based on the system's operational attribute status information. Then, through the trained elevator selection learning model, the system selects the roadway and elevator for the requested outbound goods.
[0214] Then, based on the above steps, the attribute features of the current system operation site are constructed for shuttle selection learning. The trained shuttle selection learning model is used to select a shuttle for goods requesting to be dispatched.
[0215] Then, based on the above steps, the attribute status information of the current system operation site for learning the storage location priority is obtained, all storage locations of the requested outbound goods in the selected aisle are retrieved, attribute features for storage location priority learning are constructed by comparing storage locations pairwise, and then the highest priority storage location is selected for the requested outbound storage location through the trained storage location priority learning model.
[0216] Finally, based on the above steps, a decision scheme for the selected elevator-shuttle-cargo location is generated, and the warehouse scheduling task is executed according to the scheme to complete the outbound plan for this batch of orders.
[0217] Next, determine whether the outbound plan for this batch of orders has been completed. If it has, end the current warehouse scheduling task; if not, wait for the next outbound request.
[0218] This embodiment, through the above-described scheme, specifically obtains the deep belief network model through training; when a goods outbound request is detected, it acquires the current attribute feature data of the intensive warehousing system; the attribute feature data is input into the pre-trained deep belief network model for scheduling decision-making, generating a scheduling decision scheme. The deep belief network model is composed of one or more of the following: a hoist selection learning model, a shuttle selection learning model, and a storage location priority learning model. By using the trained deep belief network model to make scheduling decisions for goods outbound requests in the intensive warehousing system, the problems of low operational efficiency, low equipment utilization, and long total system operation time can be solved, achieving the goal of timely and efficient scheduling decisions for the warehousing system. This application scheme eliminates the need for complex mathematical models and assumptions; instead, it learns the potential patterns of intensive warehousing system scheduling from the massive amounts of data accumulated during system operation, providing decision guidance for on-site operations. Based on the goal of minimizing the total system operation time, this application scheme achieves timely and efficient scheduling decisions for the warehousing system by acquiring real-time system operation information.
[0219] Furthermore, this embodiment provides a specific implementation case and performance verification case of a warehouse scheduling decision-making method.
[0220] Taking an automated high-density storage center's three-dimensional warehouse as an example, the storage system has four rows of aisles. Each row of aisles has one single-depth high-density rack on each side, with each rack having 20 layers and 40 columns. Each layer of the rack has a buffer area for buffering boxes, with a capacity of 1. Each row of aisles is equipped with one shuttle elevator, one cargo elevator, and three multi-level shuttles. Taking the system's outbound operations as an example, the initial positions of the goods in the storage system and the historical outbound order list are known. Experimental analysis is conducted based on a simulation case of this real storage system. A simulation scenario of the storage system is established using the professional discrete-time simulation software platform Siemens Tecnomatix.
[0221] To implement the real-time decision-making method proposed in this invention, a DBN program was developed and integrated with a simulation platform. The DBN program was developed using Python on TensorFlow, while the system simulation program was developed on Tecnomatix using the built-in SimTalk programming language. The simulation program includes the following subroutines: a warehouse system controller, a system operation status controller, a task controller, a scheduling instruction generator, a communicator, and a scheduler. The warehouse system controller simulates the system operation process, the task controller manages the outbound order plan, and the communicator establishes real-time information interaction with the simulation program and the DBN program through a COM interface.
[0222] The decision-making process is as follows: An integrated optimization model for equipment scheduling and location allocation is established based on the problem instance, and an optimization algorithm is used to solve for the optimal decision-making scheme. The outbound order plan and decision-making scheme of the problem instance are imported into the task controller, and the operation process is simulated through the warehouse system controller. When the task controller triggers the outbound task sequentially according to the outbound order plan, the real-time operating attribute status information of the system is sent to the system operation status controller. The operation status controller records and processes the information, generating feature data for model training. The scheduling instruction generator divides the equipment scheduling and location allocation instructions according to the imported decision-making scheme, and establishes label data for model training. The generated system status feature data and label data are used to train the DBN model, resulting in the trained hoist, shuttle selection learning, and location priority learning models.
[0223] The real-time decision-making process is as follows: The task controller triggers outbound requests sequentially based on the outbound order plan. The system operation status controller generates real-time system attribute characteristic data and sends it to the communicator. The communicator sends the data to the trained DBN program, outputting the equipment scheduling and storage location allocation scheme, which is then sent to the scheduler. The scheduler sends real-time scheduling instructions to the warehouse system controller via the communicator for execution. The complete task execution process information is saved in the simulation platform's database.
[0224] Historical order data was processed using the above method, resulting in 11,880 samples for hoist selection learning, 11,880 samples for shuttle selection learning, and 29,466 samples for location priority learning. 80% of the samples were used as the training set, and 20% as the test set. According to the feature data construction method proposed in this invention, the number of attribute features used for lane / hoist selection learning is (4+7)*4+4*4=60, the number of attribute features used for shuttle selection learning is (3+5)*3+4+7+4=39, and the number of attribute features used for location priority learning is (2+6)*2+7+3+5+4+7+4=46.
[0225] The parameters of an algorithm have a crucial impact on its performance. The Taguchi method can select the optimal parameters with fewer experiments, significantly reducing computational costs. Deep belief networks include the following important parameters: (1) the number of network nodes N n (2) Number of hidden layers N h (3) RBM learning rate r1; (4) BP learning rate r2; (5) Batch size bs. Four factor levels are considered for each parameter, as shown in Table 3 below. An orthogonal matrix L is used. 16 (4 5Experimental design: The number of RBM iterations was set to 20 and the number of BP iterations to 300 for all experiments, with the ReLU activation function. Parameter selection experiments were conducted using the above dataset. Each case was run 30 times. The mean model learning accuracy for hoist selection (LS), shuttle selection (SS), and location priority (LP) is shown in Table 4 below.
[0226]
[0227] Table 3: Factor Levels of Key Parameters in Deep Belief Networks
[0228]
[0229]
[0230] Table 4: Mean Model Learning Accuracy for Hoist Selection (LS), Shuttle Selection (SS), and Location Priority (LP)
[0231] The signal-to-noise ratio response values and the importance ranking of each parameter for the LS, SS, and LP models are shown in Tables 5, 6, and 7 below, respectively.
[0232] Level <![CDATA[N n ]]> <![CDATA[N h ]]> <![CDATA[r1]]> <![CDATA[r2]]> bs 1 -1.2604 -0.6535 -0.5608 -0.8024 -0.5410 2 -0.5895 -0.6042 -0.7203 -0.5996 -0.5801 3 -0.8374 -0.5829 -0.6369 -0.5760 -0.7241 4 -0.5693 -1.4161 -1.3387 -1.2787 -1.4115 Delta 0.6911 0.8332 0.7779 0.7027 0.8706 Rank 5 2 3 4 1
[0233] Table 5
[0234] Level <![CDATA[N n ]]> <![CDATA[N h ]]> <![CDATA[r1]]> <![CDATA[r2]]> bs 1 -0.5357 -0.3567 -0.3769 -0.4709 -0.2816 2 -0.4875 -0.3419 -0.3481 -0.4393 -0.4056 3 -0.3300 -0.5506 -0.3819 -0.3697 -0.2855 4 -0.3348 -0.4387 -0.5811 -0.4080 -0.7153 Delta 0.2057 0.2087 0.2330 0.1012 0.4337 Rank 4 3 2 5 1
[0235] Table 6
[0236]
[0237]
[0238] Table 7
[0239] Delta represents the ranking of parameter importance. It can be seen that the order of parameter influence differs across models. Therefore, the parameters for all three models need careful selection to prevent underfitting or overfitting. The parameter set with the highest signal-to-noise ratio response is selected as the optimal parameter set, i.e., LS:N. n =256, N h =3, r1=0.02, r2=0.1, bs=128; SS: N n =128, N h =2, r1=0.05, r2=0.1, bs=128; LP: N n =256, N h =4, r1=0.05, r2=0.2, bs=256.
[0240] To verify the performance of the proposed real-time decision-making method, it is compared with traditional classical scheduling rules. Based on the case study, the operating parameters of the intensive warehousing system are set as shown in Table 8. Classical scheduling rules can be divided into three categories according to their purpose: hoist scheduling, shuttle scheduling, and location allocation. Detailed information on classical rules is shown in Table 9. These three categories are combined to form eight scheduling rule groups for warehousing systems, as shown in Table 10. For example, scheduling rule group s1 indicates that hoist scheduling uses the shortest task queue rule, shuttle scheduling uses the longest waiting time rule, and location allocation uses the shortest handling distance rule. Experiments were conducted on six task sizes (Ts), using the total system operation time as the indicator. P in the table represents the proposed method. The experimental results are shown in Table 11. The reduction rate of total operation time compared to the classical scheduling rule group is shown in Table 12. Here, MI represents the minimum reduction rate of total operation time for the same task size, and AV represents the average reduction rate of total operation time for the same scheduling rule group under different task sizes.
[0241] <![CDATA[w0]]> <![CDATA[h0]]> <![CDATA[v s ]]> <![CDATA[v l ]]> <![CDATA[a s ]]> <![CDATA[a l ]]> <![CDATA[T load_1 ]]> <![CDATA[T load_2 ]]> <![CDATA[T unload_1 ]]> <![CDATA[T unload_2 ]]> 0.75 0.75 4 2 1 1.5 6 7 6 7
[0242] Table 8: Operating Parameter Settings for Intensive Storage Systems
[0243]
[0244]
[0245] Table 9: Detailed Information on Classic Rules
[0246] No. Scheduling rule group No. Scheduling rule group S1 STQ-LWT-MHD S5 MNIS-LWT-MHD S2 STQ-LWT-STT S6 MNIS-LWT-STT S3 STQ-STQ-MHD S7 MNIS-STQ-MHD S4 STQ-STQ-STT S8 MNIS-STQ-STT
[0247] Table 10: Scheduling Rules Group for the Warehouse System
[0248]
[0249] Table 11: Experimental Results
[0250]
[0251]
[0252] Table 12: Total Job Time Reduction Rate of This Method Compared with Classic Scheduling Rules
[0253] Table 11 shows that, for different task sizes, the total operation time of the proposed method is shorter than that of the classical scheduling rule group. Table 12 shows that, compared with the classical scheduling rule group, the proposed method reduces the average total system operation time by a minimum of 6.54% and a maximum of 17.22%. As the task size increases, the minimum reduction rate of the total system operation time gradually increases, reaching a minimum reduction of 7.4% when the batch task size reaches 100. Therefore, this verifies the superiority of the proposed real-time decision-making method in intensive warehousing system operations.
[0254] This embodiment, through the above-described scheme, specifically acquires the current attribute feature data of the intensive warehousing system when a goods outbound request is detected; inputs the attribute feature data into a pre-trained deep belief network model for scheduling decision-making, generating a scheduling decision scheme. The deep belief network model is composed of one or more of a hoist selection learning model, a shuttle selection learning model, and a storage location priority learning model. By using the trained deep belief network model to make scheduling decisions for goods outbound requests in the intensive warehousing system, the problems of low operational efficiency, low equipment utilization, and long total system operation time can be solved, achieving the goal of timely and efficient scheduling decisions for the warehousing system. This application scheme does not require the establishment of complex mathematical models and assumptions; instead, it learns the potential patterns of intensive warehousing system scheduling from the massive data accumulated during system operation and provides decision-making guidance for on-site operations. Based on the goal of minimizing the total system operation time, this application scheme achieves timely and efficient scheduling decisions for the warehousing system by acquiring real-time system operation information.
[0255] Furthermore, this application also proposes a warehouse scheduling decision-making device, which includes:
[0256] The data acquisition module is used to acquire the current attribute feature data of the intensive warehousing system when a goods outbound request is detected;
[0257] The scheduling decision module is used to input the attribute feature data into a pre-trained deep belief network model to make scheduling decisions and generate scheduling decision schemes. The deep belief network model is composed of one or more of the following: hoist selection learning model, shuttle selection learning model, and cargo location priority learning model.
[0258] The principle and implementation process of warehouse scheduling decision-making in this embodiment are explained in the above embodiments, and will not be repeated here.
[0259] Furthermore, this application also proposes a terminal device, which includes a memory, a processor, and a warehouse scheduling decision program stored in the memory and executable on the processor. When the warehouse scheduling decision program is executed by the processor, it implements the steps of the warehouse scheduling decision method as described above.
[0260] Since this warehouse scheduling decision program adopts all the technical solutions of all the aforementioned embodiments when it is executed by the processor, it has at least all the beneficial effects brought about by all the technical solutions of all the aforementioned embodiments, which will not be repeated here.
[0261] Furthermore, embodiments of this application also propose a computer-readable storage medium storing a warehouse scheduling decision program, which, when executed by a processor, implements the steps of the warehouse scheduling decision method as described above.
[0262] Since this warehouse scheduling decision program adopts all the technical solutions of all the aforementioned embodiments when it is executed by the processor, it has at least all the beneficial effects brought about by all the technical solutions of all the aforementioned embodiments, which will not be repeated here.
[0263] Compared to existing technologies, the warehouse scheduling decision-making method, apparatus, terminal equipment, and storage medium proposed in this application obtain current attribute feature data of the intensive warehousing system when a goods outbound request is detected. This attribute feature data is then input into a pre-trained deep belief network model for scheduling decision-making, generating a scheduling decision scheme. The deep belief network model is composed of one or more of a hoist selection learning model, a shuttle selection learning model, and a storage location priority learning model. By using the trained deep belief network model to schedule goods outbound requests in the intensive warehousing system, the problem of low operational efficiency and low equipment utilization, resulting in long total system operation time, can be solved, achieving timely and efficient scheduling decisions for the warehousing system. This application's solution does not require the establishment of complex mathematical models and assumptions; instead, it learns the potential patterns of intensive warehousing system scheduling from the massive amounts of data accumulated during system operation, providing decision guidance for on-site operations. Based on the goal of minimizing the total system operation time, this application's solution achieves timely and efficient warehousing system scheduling decisions by acquiring real-time system operation information.
[0264] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0265] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0266] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of this application.
[0267] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A warehouse scheduling decision-making method, characterized in that, The warehouse scheduling decision-making method includes: When a goods outbound request is detected, the current attribute feature data of the intensive warehousing system is obtained; wherein, the attribute feature data refers to the system state attribute feature data at the time of task arrival, including one or more state attribute features of aisles, elevators, shuttle cars and storage locations; The attribute feature data is input into a pre-trained deep belief network model for scheduling decision-making, generating a scheduling decision scheme. The structure of the deep belief network model is constructed by sequentially stacking Restricted Boltzmann Machines (RBMs). The learning process of the structure includes unsupervised pre-training of the RBMs layer by layer, followed by supervised tuning of the entire network using the backpropagation algorithm (BP). The deep belief network model consists of one or more of the following: a hoist selection learning model, a shuttle selection learning model, and a cargo location priority learning model. Before the step of obtaining the current attribute feature data of the intensive warehousing system when a goods outbound request is detected, the offline training of the deep belief network model is also included: Based on the simulated attribute feature data and historical attribute feature data, generate the first training attribute feature data for offline training of the booster machine to select the learning model; Based on the simulated label data and historical label data, generate first training label data for offline training of the booster selection learning model; The hoist selection learning model is trained offline by combining the attribute feature data and the label data used for training, generating a decision scheme for the selected hoist, and obtaining the hoist selection learning model after offline training. Combining the selected elevator decision scheme, the simulated attribute feature data, and the historical attribute feature data, second training attribute feature data is generated for offline training of the shuttle selection learning model. Combining the selected elevator decision scheme, the simulated label data, and the historical label data, second training label data is generated for offline training of the shuttle selection learning model; The shuttle selection learning model is trained offline by combining the attribute feature data and the label data used for the second training, generating a selected hoist-shuttle decision scheme, and obtaining the offline trained shuttle selection learning model. Combining the selected elevator-shuttle decision scheme, the simulated attribute feature data, and the historical attribute feature data, a third set of training attribute feature data is generated for offline training of the cargo location priority learning model. The third set of training attribute feature data includes: attribute features of the selected aisle, attribute features of the selected elevator, attribute features of the selected shuttle, cargo location attribute features of the goods to be shipped, and attribute features of cargo location priority generated by pairwise comparison of cargo locations. Combining the selected hoist-shuttle decision scheme with the simulated label data and the historical label data, a third set of training label data is generated for offline training of the cargo location priority learning model. The cargo location priority learning model is trained offline by combining the attribute feature data and the label data used for the third training, generating a selected hoist-shuttle-cargo location decision scheme, obtaining the shuttle selection learning model after offline training, and obtaining the deep belief network model after offline training. After the step of inputting the attribute feature data into a pre-trained deep belief network model for scheduling decision-making and generating a scheduling decision scheme, the method further includes: Execute warehouse scheduling tasks according to the scheduling decision plan, specifically including: When the selected shuttle car performs a cross-level pickup task, the working status information of the shuttle car on the target level where the target goods are located is detected. When it is detected that the shuttle in the target layer is performing a work task, a task transfer strategy is executed, specifically including: Cancel the current cross-level pickup task of the selected shuttle vehicle; When the shuttle in the target layer is detected to have finished its task, the cross-layer pickup task is executed.
2. The warehouse scheduling decision-making method according to claim 1, characterized in that, Before the step of obtaining the current attribute feature data of the intensive warehousing system when a goods outbound request is detected, the method further includes: The deep belief network model obtained through training specifically includes: The deep belief network model is trained offline to obtain the offline trained deep belief network model; The offline-trained deep belief network model is then trained online to obtain a well-trained deep belief network model.
3. The warehouse scheduling decision-making method according to claim 2, characterized in that, The step of training the deep belief network model offline to obtain the offline-trained deep belief network model includes: An integrated optimization mathematical model is established, and an optimization algorithm is used to solve the integrated optimization mathematical model to obtain a simulated decision scheme. The simulated decision-making scheme is imported into a pre-built warehousing system simulation model for partitioning, resulting in simulated label data; The pre-established outbound order plan and the simulated decision-making scheme are imported into the warehousing system simulation model to simulate operations and obtain simulated operational attribute status data. Simulated attribute feature data is generated based on the aforementioned operational attribute status data; Acquire historical interaction data of the intensive warehousing system to generate historical attribute feature data and historical tag data; By combining the simulated attribute feature data, the historical attribute feature data, the simulated label data, and the historical label data, the deep belief network model is trained offline to obtain the offline-trained deep belief network model.
4. The warehouse scheduling decision-making method according to claim 3, characterized in that, The step of training the offline-trained deep belief network model online to obtain a trained deep belief network model includes: Acquire online interactive data between the warehouse management system and the actual operation site of the intensive warehousing system; The offline-trained deep belief network model is trained online based on the online interaction data to obtain the trained deep belief network model.
5. A warehouse scheduling decision-making device, characterized in that, The warehouse scheduling decision-making device includes: The data acquisition module is used to acquire the current attribute feature data of the intensive warehousing system when a goods outbound request is detected; wherein, the attribute feature data refers to the system state attribute feature data at the time of task arrival, including one or more state attribute features of aisles, elevators, shuttle cars and storage locations; The data acquisition module, prior to the step of acquiring the current attribute feature data of the intensive warehousing system when a goods outbound request is detected, also includes offline training of the deep belief network model: Based on the simulated attribute feature data and historical attribute feature data, generate the first training attribute feature data for offline training of the booster machine to select the learning model; Based on the simulated label data and historical label data, generate the first training label data for offline training of the booster machine to select the learning model; The hoist selection learning model is trained offline by combining the attribute feature data and the label data used for training, generating a decision scheme for the selected hoist, and obtaining the hoist selection learning model after offline training. Combining the selected elevator decision scheme, the simulated attribute feature data, and the historical attribute feature data, second training attribute feature data is generated for offline training of the shuttle selection learning model. Combining the selected elevator decision scheme, the simulated label data, and the historical label data, second training label data is generated for offline training of the shuttle selection learning model; The shuttle selection learning model is trained offline by combining the attribute feature data and the label data used for the second training, generating a selected hoist-shuttle decision scheme, and obtaining the offline trained shuttle selection learning model. Combining the selected elevator-shuttle decision scheme, the simulated attribute feature data, and the historical attribute feature data, a third set of training attribute feature data is generated for offline training of the cargo location priority learning model. The third set of training attribute feature data includes: attribute features of the selected aisle, attribute features of the selected elevator, attribute features of the selected shuttle, cargo location attribute features of the goods to be shipped, and attribute features of cargo location priority generated by pairwise comparison of cargo locations. Combining the selected hoist-shuttle decision scheme with the simulated label data and the historical label data, a third set of training label data is generated for offline training of the cargo location priority learning model. The cargo location priority learning model is trained offline by combining the attribute feature data and the label data used for the third training, generating a selected hoist-shuttle-cargo location decision scheme, obtaining the shuttle selection learning model after offline training, and obtaining the deep belief network model after offline training. The scheduling decision module is used to input the attribute feature data into a pre-trained deep belief network model for scheduling decisions and to generate scheduling decision schemes. The deep belief network model is constructed by sequentially stacking Restricted Boltzmann Machines (RBMs). The learning process of the structure includes unsupervised pre-training of the RBMs layer by layer, followed by supervised optimization of the entire network using the backpropagation (BP) algorithm. The deep belief network model consists of one or more of the following: a hoist selection learning model, a shuttle selection learning model, and a cargo location priority learning model. The scheduling decision module is also used to execute warehouse scheduling tasks according to the scheduling decision plan, specifically including: When the selected shuttle car performs a cross-level pickup task, the working status information of the shuttle car on the target level where the target goods are located is detected. When it is detected that the shuttle in the target layer is performing a work task, a task transfer strategy is executed, specifically including: Cancel the current cross-level pickup task of the selected shuttle vehicle; When the shuttle in the target layer is detected to have finished its task, the cross-layer pickup task is executed.
6. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a warehouse scheduling decision program stored in the memory and executable on the processor. When the warehouse scheduling decision program is executed by the processor, it implements the steps of the warehouse scheduling decision method as described in any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a warehouse scheduling decision program, which, when executed by a processor, implements the steps of the warehouse scheduling decision method as described in any one of claims 1-4.