A cloth delivery scheduling method and system based on multi-strategy cooperation
By employing a multi-strategy collaborative fabric outbound scheduling method, which combines the analytic hierarchy process (AHP), fuzzy comprehensive evaluation, and reinforcement learning models, the problems of priority assessment, resource scheduling, and path planning in fabric outbound scheduling are solved. This enables efficient and flexible order processing and resource utilization, thereby improving enterprise operational efficiency and customer satisfaction.
Patent Information
- Application Number
- CN202511274574.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-08
AI Technical Summary
Existing fabric outbound scheduling methods suffer from one-sided priority assessment, rigid resource scheduling, redundant paths, and insufficient dynamic adaptability, resulting in delayed processing of high-value customer orders, insufficient resource utilization, and low picking efficiency.
A multi-strategy collaborative fabric outbound scheduling method is adopted. The order priority is calculated by using the analytic hierarchy process and the fuzzy comprehensive evaluation model. The scheduling strategies of busy period and idle period are combined. The shortest path algorithm is used to plan the picking route. During the picking process, the order merging decision is made by the reinforcement learning model to realize the dynamic matching of resources and orders.
It improves fabric outbound efficiency, ensures high-value orders are prioritized, reduces path redundancy, improves resource utilization and picking operation continuity, adapts to changes in order structure, and optimizes operating costs.
Smart Images

Figure CN120782088B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a cloth warehouse-out scheduling method based on multi-strategy cooperation, and belongs to the technical field of intelligent scheduling of warehouse logistics. BACKGROUND
[0002] In the operation practice of manufacturing and logistics industry, the pros and cons of the order warehouse-out scheduling strategy directly affect the enterprise operation efficiency and cost control level. With the rapid development of e-commerce and manufacturing industry, cloth orders show the characteristics of multi-category, small batch, high frequency and significant difference in urgency. The traditional warehouse-out scheduling method gradually exposes the problem of insufficient adaptability.
[0003] First, the priority evaluation system is missing. The traditional method mostly adopts the first-come-first-served (FCFS) mode, which only sorts orders according to their arrival time, without considering multi-dimensional factors such as customer level, order urgency, and profit contribution. During the peak period of orders, the urgent orders of high-value customers may be delayed due to late sorting, resulting in a decline in customer experience and even loss of orders.
[0004] Second, the order processing and resource scheduling are disconnected. Cloth orders often contain multiple SKUs (inventory units), and the number of single SKUs (rolls) varies greatly. The existing scheduling does not establish a dynamic matching mechanism with the capacity of picking equipment (such as forklift load), which may lead to extreme situations such as "insufficient capacity for one-time picking" or "high empty load rate of equipment".
[0005] Third, the path planning and order consolidation are not intelligent enough. The picking path mostly relies on manual experience or simple sequential planning, resulting in redundant and repeated paths. In addition, there is a lack of dynamic consolidation mechanism for similar orders, and the picking efficiency is limited by the single order processing mode, resulting in a long overall fulfillment cycle.
[0006] Fourth, the dynamic adaptability is poor. When the order structure such as category, quantity, and urgency changes suddenly, the traditional method is difficult to quickly adjust the scheduling strategy, which may lead to systematic delay. Especially in the cloth industry, seasonal demand fluctuates greatly, and fixed processes cannot meet the requirements of flexible production and distribution.
[0007] Therefore, it is urgent to develop a multi-strategy cooperation method that integrates dynamic evaluation of order priority, elastic scheduling of resources, intelligent path planning, and dynamic consolidation decision-making, to improve the efficiency and intelligence level of cloth warehouse-out, and realize the collaborative optimization of order allocation and job scheduling under resource constraints, so as to achieve the purpose of cost reduction and efficiency improvement. SUMMARY
[0008] The technical problems to be solved by the present application are: the existing cloth warehouse scheduling method has one-sided priority evaluation, rigid resource scheduling, redundant path, and insufficient dynamic adaptability, how to realize dynamic matching of resources and orders, improve warehouse efficiency, reduce operation cost and guarantee priority performance of high value orders through the fusion of flexible scheduling, double mode scheduling, order priority evaluation, intelligent path planning and order consolidation decision driven by reinforcement learning.
[0009] To solve the above technical problems, the technical scheme adopted by the present application is:
[0010] A cloth warehouse scheduling method based on multi-strategy cooperation, comprising the following steps: step 1: establishing a warehouse space coordinate model, and establishing a mapping relationship between the warehouse position code and the three-dimensional coordinates; step 2: using an analytic hierarchy process-fuzzy comprehensive evaluation model to calculate the order priority for newly arrived orders; step 3: according to the current job resource state and the order priority calculation result, triggering the busy period scheduling strategy or the idle period scheduling strategy to complete the allocation of orders and job resources; step 4: for the allocated orders, using the shortest path algorithm to plan the picking path; step 5: during the picking execution process, combining real-time state information, and using a reinforcement learning model to consolidate orders; step 6: repeating steps 2 to 5 to calculate the order scheduling and picking path of each period, and updating the personnel and equipment usage according to the solution result.
[0011] The foregoing cloth warehouse scheduling method based on multi-strategy cooperation, in the step 1, comprises:
[0012] Step 1.1: taking the lower left corner of the order consolidation area as the coordinate origin , constructing a three-dimensional Cartesian coordinate system, X axis, Y axis and Z axis representing horizontal, vertical and height directions respectively, and dividing the warehouse space into a three-dimensional grid;
[0013] Step 1.2: dividing the warehouse functional area, including the order consolidation area, the corridor and the fabric storage area, the corridor includes the main corridor, the auxiliary corridor and the horizontal corridor, and defining the coordinate range of each area: let be the width of the order consolidation area in the X axis direction, be the length of the order consolidation area in the Y axis direction, be the width of the main corridor in the X axis direction, be the width of the auxiliary corridor in the X axis direction, be the width of the right corridor in the X axis direction, be the total length of the warehouse in the Y axis direction, be the total length of the warehouse in the X axis direction, be the length of the left main shelf in the X axis direction, be the length of the right auxiliary shelf in the X axis direction, then the X axis range of the order consolidation area is , the Y-axis range of the main corridor is ; the X-axis range of the main corridor is , the Y-axis range of the main corridor is ; the X-axis range of the secondary corridor is , the Y-axis range of the secondary corridor is ; the X-axis range of the right corridor is , the Y-axis range of the right corridor is ; the X-axis range of the fabric storage area is , the Y-axis range of the fabric storage area is , containing the left main shelf and the right secondary shelf ;
[0014] Step 1.3: According to the actual situation of the SKU storage position, a mapping formula of coding and three-dimensional coordinates is established, which uniquely maps the storage position coding to the coordinate point in the three-dimensional Cartesian coordinate system , realizing the digital modeling of the warehouse space; let be the column number in the storage position coding, be the row number in the storage position coding, be the layer number in the storage position coding, be the column spacing, be the maximum number of columns of the left main shelf, be the row spacing, be the layer spacing, and the mapping formula is as follows:
[0015] ,
[0016] ,
[0017] .
[0018] The foregoing fabric warehouse out-of-warehouse scheduling method based on multi-strategy cooperation, in step 2, comprises:
[0019] Step 2.1: Determine the evaluation index and establish a hierarchical system, with the order priority as the target layer; the customer importance, order urgency, and order amount as the criterion layer; and the order fabric amount as the positive index, the order arrival time as the reverse index, the order completion time as the reverse index, and the order profit rate as the positive index;
[0020] Step 2.2: The weights are determined by using the analytic hierarchy process, the local weights of each index in the index layer relative to the corresponding criterion layer are calculated, and the weights of each criterion in the criterion layer relative to the target layer are calculated; the expert scoring method is used, and the elements in the same level are compared with each other according to the analytic hierarchy process scale rule, the index layer judgment matrix and the criterion layer judgment matrix are constructed;
[0021] The maximum eigenvalue of the index layer judgment matrix and the criterion layer judgment matrix is solved by the eigenvalue method and the corresponding normalized eigenvector , that is, the relative weight of each index.
[0022] The maximum eigenvalue of the index layer and the corresponding eigenvector are calculated, first, the sum of each column of the judgment matrix is calculated to obtain the column sum vector, and the judgment matrix is divided by the column sum of the corresponding column to obtain the normalized matrix .
[0023] The average value of each row is calculated to obtain the weight vector of the index layer .
[0024] The weight vector of the index layer is adjusted and normalized to obtain the eigenvector one .
[0025] The weighted sum vector is calculated , each weighted value is divided by the corresponding weight component, and the average value is taken to obtain the maximum eigenvalue .
[0026] The maximum eigenvalue of the criterion layer and the corresponding eigenvector are calculated, the judgment matrix is divided into two 2x2 sub-matrices , which are processed respectively.
[0027] For the sub-matrix , first, the sum of each column of the sub-matrix is calculated to obtain the column sum vector, and the sub-matrix is divided by the column sum of the corresponding column to obtain the normalized matrix .
[0028] The average value of each row is calculated to obtain the weight vector one.
[0029] The sub-matrix is processed again:
[0030] First, the sum of each column of the sub-matrix is calculated to obtain the column sum vector, and the sub-matrix is divided by the column sum of the corresponding column to obtain the normalized matrix .
[0031] Calculate the average value of each row to obtain weight vector two;
[0032] Merge weight vector one and weight vector two according to the order of the sub-matrix to obtain the original quaternion vector ;
[0033] Normalize the original quaternion vector to obtain the adjusted feature vector two ;
[0034] Conduct consistency test on the judgment matrix and calculate the consistency ratio , , wherein, is the consistency index, is the order of the judgment matrix, is the average random consistency index;
[0035] When CR < 0.1, the consistency test is passed;
[0036] Calculate the weight of each index relative to the target layer , including the order fabric quantity , order arrival time , order completion time , order profit rate , the calculation method is as follows:
[0037] = index item weight × corresponding criterion item weight;
[0038] Obtain the composite weight original vector of each index , and finally normalize the composite weight original vector to obtain the weight vector W;
[0039] Step 2.3: Use fuzzy comprehensive evaluation method to evaluate the priority, and establish fuzzy comment set = (excellent, good, medium, poor) = (7, 5, 3, 1); For order fabric quantity , order arrival time , order completion time , order profit rate , develop grade evaluation criteria respectively, determine the index value or time interval range corresponding to each grade, distinguish between positive and negative indicators, and use the range transformation method to determine the membership matrix :
[0040] Among them, for positive indicators , when the estimated value of the index is If the estimated value is greater than the standard of the grade "excellent" in the evaluation index grade, the corresponding grade excellent membership is 1, and the rest is 0; when the estimated value If the estimated value is less than the standard of the grade "poor" in the evaluation index grade, the corresponding grade excellent membership is 0, and the rest is 1; when the estimated value If the estimated value is between the standard of the grade "excellent" and the standard of the grade "good", the first The membership calculation method of the first
[0041] ,
[0042] The lower limit value of the jth grade is represented by j=1, corresponding to the lower limit of "excellent", j=2 corresponding to the lower limit of "good", and j=3 corresponding to the lower limit of "medium";
[0043] The membership calculation method of the first ;
[0044] The membership of other grades is 0;
[0045] For the reverse index When the estimated value of the index is greater than the standard of the grade "excellent" in the evaluation index grade, the corresponding grade excellent membership is 1, and the rest is 0; when the estimated value If the estimated value is less than the standard of the grade "poor" in the evaluation index grade, the corresponding grade excellent membership is 0, and the rest is 1; when the estimated value If the estimated value is between the standard of the grade "excellent" and the standard of the grade "good", the first The membership calculation formula of the first
[0046] ,
[0047] The membership calculation formula of the first ;
[0048] The membership of other grades is 0;
[0049] Step 2.4: Construct the membership matrix , the element represents the membership of the first The index C1-C4 belongs to the first The degree of association between the index and the grade is quantified, and the priority score is calculated , is the index weight vector, R is the fuzzy comment set, representing the quantified score of each grade, is the membership matrix, and the superscript represents transposition.
[0050] The aforementioned fabric outbound scheduling method based on multi-strategy collaboration includes, in step 3: Step 3.1: Define scheduling trigger conditions, assuming the number of idle job resources is... The number of pending orders is ,when When idle period scheduling is triggered, When busy period scheduling is triggered; Step 3.2: When At that time, the workload-based lottery scheduling algorithm triggers idle period scheduling: Step 3.3: When When busy period scheduling is triggered, it is a priority-oriented skill-based tiered scheduling.
[0051] The aforementioned fabric outbound scheduling method based on multi-strategy collaboration includes, in step 3.2:
[0052] Step 3.2.1: Calculate the relative workload of employees The calculation formula is:
[0053] ,
[0054] in, For the first Total working hours of employees Total shift hours for employees;
[0055] Step 3.2.2: Allocate lottery tickets inversely proportional to the relative workload. The calculation formula is:
[0056] ,
[0057] in, As the base lottery number, As an adjustment factor, the lower the workload, the more lottery tickets are drawn; For the first The number of lottery tickets won by employees For the first The relative workload of employees;
[0058] Step 3.2.3: Construct a lottery pool of idle personnel, randomly select lottery tickets to allocate orders, and achieve resource balance.
[0059] The aforementioned fabric outbound scheduling method based on multi-strategy collaboration includes, in step 3.3: Step 3.3.1: Scoring by priority. Arrange the waiting orders in descending order to obtain the sorted results. ; Indicates the first One order; Step 3.3.2: Divide the workers into two skill levels, including skilled workers and inexperienced workers, the first The lottery distribution coefficient of the hierarchical personnel is ; step 3.3.3: the lottery quota allocated to the high-skilled workers is greater than the lottery quota allocated to the low-skilled workers, and a differentiated lottery pool is constructed; step 3.3.4: the orders sorted , are randomly distributed by lottery from the lottery pool, and high-priority orders are preferentially guaranteed to be processed by skilled workers.
[0060] The foregoing method for scheduling cloth delivery based on multi-strategy cooperation, in step 4, comprises:
[0061] Step 4.1: screening valid SKUs, denoted as set ;
[0062] Step 4.2: converting valid SKUs to three-dimensional coordinates through warehouse layout mapping ;
[0063] Step 4.3: constructing a picking path graph model , node set , where is an order merging area, is a target warehouse coordinate; edge set , representing feasible paths that meet warehouse corridor constraints; edge weight is the passing distance, which is calculated using Manhattan distance, and the calculation formula is ; the Z-axis difference is ignored; , , are the coordinates of the target warehouse of the first valid SKU in the horizontal direction of the warehouse, the vertical direction of the warehouse, and the vertical direction of the warehouse height, respectively; , are the coordinates of the target warehouse of the first valid SKU in the horizontal direction of the warehouse and the vertical direction of the warehouse, respectively; they are determined through the mapping formula of step 1;
[0064] Step 4.4: using Dijkstra's algorithm to solve the shortest path:
[0065] Step 4.4.1: initializing the distance array, , representing the current shortest distance estimate from the starting point to node , , , representing the shortest distance from the starting point to itself as 0, , +∞ representing positive infinity, indicating that the distance from the starting point to other nodes is unknown in the initial state; initializing the priority queue , the unvisited node set denotes set difference operation, i.e. removing nodes from set to obtain the set of unvisited nodes;
[0066] Step 4.4.2: Take out node from the priority queue , and update the current shortest distance estimate to be the minimum, if , then skip; for each adjacent node , calculate the candidate distance , if , then update , and add to the priority queue ; remove node from , repeat the above operations of this step until , denotes empty set, i.e. all nodes have been processed; Synchronize with step 4.3 , and calculate the travel distance of node and ;
[0067] Step 4.4.3: Backtrack the predecessor nodes from the end point to generate the shortest path , the corresponding picking order is obtained by sorting the elements in set according to the above shortest path , where is the th picked SKU in set .
[0068] The foregoing cloth warehouse-out scheduling method based on multi-strategy cooperation, in step 5, comprises:
[0069] Step 5.1: Perform the picking operation according to the planned path in turn, and update the current load and the number of picked SKUs;
[0070] Step 5.2: Determine whether all SKUs of the order / batch are picked, if not, return to step 5.1 to continue moving to the next SKU position; if yes, trigger the order merging decision based on reinforcement learning; the order refers to the complete demand instruction issued by the customer, including a plurality of SKUs and a plurality of volumes; the batch refers to when the total volume of SKUs contained in the order exceeds the maximum load capacity of the picking device, the order is split into a plurality of batches, and the total number of SKUs of a single batch needs to meet the device capacity constraint.
[0071] Step 5.3: Construct a 24-dimensional state vector , including four-dimensional picker state, normalized warehouse-out frequency and 20 features of orders to be merged; is a real number field;
[0072] The four-dimensional picker state includes position coordinates , current load ratio , is the current number of loads, is the maximum capacity of the picking device;
[0073] Select the top 5 orders to be processed, each order contains 4 features, including:
[0074] type , town = 1, town = 0;
[0075] standardized quantity , is the total number of orders;
[0076] urgency , is the remaining processing time, is the time limit;
[0077] position similarity , is the average distance between the current path and the order location, the smaller the distance, the higher the similarity, and the remaining dimensions are filled with 0 when there are less than 5 orders;
[0078] Step 5.4: Select the merging action using the DQN algorithm, the action space corresponds to three merging strategies, all of which need to meet the forklift capacity constraint, where is not merged, maintain the current task; is to merge the entire order, select the highest priority order to be merged and include all remaining SKUs; is to merge part of the complete SKU, select the order to be merged with the highest position similarity and extract the SKU that can be completely accommodated; complete accommodation means that the current picking device has enough capacity to load all the volumes of the SKU at once without overloading;
[0079] Step 5.5: Execute the selected merging action and calculate the reward, design the reward function where, is the basic merging reward, , is the number of merged volumes, is the basic merging reward coefficient, encouraging merging behavior, giving rewards according to the number of merged SKUs, and higher rewards for complete order merging; is the path saving reward, , is the individual processing distance, is the distance after merging, is the reward coefficient for path saving, and the reward for the reduced moving distance after merging; is the capacity utilization reward, , is the load rate after merging, is the capacity utilization reward coefficient, and the forklift is encouraged to be fully loaded; is the order completion reward, if the order is completely merged , otherwise , is the order completion reward coefficient, and the order splitting is reduced; is the outbound frequency penalty, , is the unloading frequency, is the outbound frequency penalty coefficient;
[0080] Step 5.6: if the merging action is selected, or , the path planning algorithm of step 4 is re-executed based on the merged SKU to generate a new picking path and return to step 5.1 to continue picking; if the non-merging action is selected, , the current batch of picking is ended.
[0081] The aforementioned cloth outbound scheduling method based on multi-strategy cooperation, in step 5.4, comprises:
[0082] Step 5.4.1: a 3-layer fully connected neural network is constructed, the input state vector is input, and the 3-dimensional action value is output, and the hidden layer adopts the ReLU activation function;
[0083] Step 5.4.2: an experience replay mechanism is adopted to store samples into the experience pool, and 64 samples are randomly sampled each time for training;
[0084] Step 5.4.3: the reward function value , is updated, wherein is the value estimation of the action executed in the current state , is the immediate reward, is the discount factor, is the new state after executing the action, is the next action selected in the state , is the maximum value of each action in the new state, is the conditional symbol, represents expectation, and is the value of the next state . Performing an action Later, the mathematical expression of the expected sum of all possible rewards and values thereafter;
[0085] Step 5.4.4: Adopt - Greedy policy balances exploration and exploitation, sets the initial exploration rate , decays to 0.01 with training rounds, , extensive exploration at the beginning, , then use the optimal strategy for exploitation;
[0086] Step 5.4.5: Set the number of iterations for the target network to synchronize the main network parameters.
[0087] A computer system includes a memory, a processor, and a computer program stored on the memory, the processor executing the computer program to implement the steps of a fabric delivery scheduling method based on multi-strategy cooperation.
[0088] The present application has the beneficial effects: the fabric delivery scheduling method based on multi-strategy cooperation provided by the present application, aiming at the problems of one-sided priority evaluation, rigid resource scheduling, redundant path planning and insufficient dynamic adaptability in current fabric delivery scheduling, through the fusion of order priority dynamic evaluation, resource elastic scheduling, intelligent path planning and reinforcement learning driven order merging decision, a set of multi-strategy cooperative intelligent scheduling mechanism is established.
[0089] Based on the dynamic characteristics of the order life cycle, the present application realizes the accurate quantification of order priority through hierarchical index system and fuzzy comprehensive evaluation, solves the resource mismatch problem caused by the traditional method of relying on single dimension sorting only, ensures that high value and high urgency orders are given priority, and improves customer service quality and cooperation stability. At the same time, through elastic scheduling and double mode scheduling strategy, the dynamic adaptation of operation personnel and order quantity is realized, avoiding resource idling or overload under fixed scheduling mode, and improving the overall utilization efficiency of human and equipment resources.
[0090] In addition, the present application optimizes the picking path through weighted undirected graph construction and shortest path algorithm, combines the intelligent adjustment of order merging decision, reduces the invalid movement and repeated work in the warehouse, and reduces the path redundancy; through reinforcement learning for dynamic merging of similar orders, the limitation of single order processing mode is broken through, and the coherence of picking operation and equipment full load rate are improved.
[0091] The method proposed by the present application has strong dynamic adaptability, can adjust the scheduling strategy in real time according to the change of order structure, adapts to the seasonal demand fluctuation and order structure diversification scene of fabric industry, and finally realizes the cooperative goals of improving delivery efficiency, optimizing operation cost and improving customer satisfaction. Attached Figure Description
[0092] Figure 1 This is a flowchart of the fabric outbound scheduling method based on multi-strategy collaboration in Embodiment 1 of the present invention;
[0093] Figure 2 This is the hierarchical structure of the AHP-FCE evaluation model in Embodiment 1 of the present invention;
[0094] Figure 3 This is a flowchart of the dual-mode scheduling mechanism in Embodiment 1 of the present invention;
[0095] Figure 4 This is a flowchart of the Dijkstra algorithm path planning in Embodiment 1 of the present invention;
[0096] Figure 5 This is a flowchart of the order merging decision-making process of the DQN model in Embodiment 1 of the present invention;
[0097] Figure 6 These are the warehouse layout modeling parameters of Embodiment 1 of the present invention. Detailed Implementation
[0098] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0099] The present invention will be further described below with reference to specific embodiments.
[0100] Example 1
[0101] like Figure 1 As shown, this embodiment provides a fabric outbound scheduling method based on multi-strategy collaboration, including the following steps: Step 1: Establish a warehouse spatial coordinate model and establish the mapping relationship between warehouse location codes and three-dimensional coordinates; Step 2: Calculate the order priority for newly arriving orders using the Analytic Hierarchy Process-Fuzzy Comprehensive Evaluation (AHP-FCE) model; Step 3: Trigger busy period scheduling or idle period scheduling based on the current operational resource status and order priority calculation results to complete the allocation of orders and operational resources; Step 4: For allocated orders, plan the picking route using the shortest path algorithm; Step 5: During the picking process, combine real-time status information and use a reinforcement learning model to merge orders; Step 6: Repeat steps 2 to 5 to calculate the order scheduling and picking route for each subsequent time period, and update the personnel and equipment usage based on the solution results.
[0102] Step 1 includes:
[0103] Step 1.1: As Figure 6 As shown, the origin of the coordinate system is the bottom left corner of the order merging area. A three-dimensional Cartesian coordinate system is constructed, with the X-axis, Y-axis, and Z-axis representing the horizontal, vertical, and height directions, respectively. The X-axis is 100 meters long, the Y-axis is 70 meters long, and the Z-axis is 4 meters long, dividing the warehouse space into a three-dimensional grid.
[0104] Step 1.2: Divide the warehouse into functional areas, including an order merging area, corridors, and fabric storage areas. The corridors include main corridors, secondary corridors, and transverse corridors. Define the coordinate range of each area: Let... This represents the width of the order merging area along the X-axis. The length of the order merging area in the Y-axis direction. The width of the main corridor along the X-axis. Let X be the width of the secondary corridor along the X-axis. Let be the width of the right corridor in the X-axis direction. Let Y be the total longitudinal length of the warehouse along the Y-axis. Let X be the total horizontal length of the warehouse along the X-axis. This represents the length of the left-side main shelf along the X-axis. Let the length of the right-side auxiliary shelf be in the X-axis direction. Then the X-axis range of the order merging area is: The Y-axis range is The X-axis range of the main corridor is... The Y-axis range is The X-axis range of the secondary corridor is The Y-axis range is The X-axis range of the right corridor is... The Y-axis range is The X-axis range of the fabric storage area is... The Y-axis range is Including the main shelving on the left (35 columns). With the right-hand auxiliary shelf (15 columns) "Column" refers to the number of partition units along the X-axis of the warehouse shelf, with the width of each unit equal to the column spacing; "row" refers to the number of partition units along the Y-axis, with each row containing several columns; "layer" refers to the number of partition units along the Z-axis.
[0105] In this embodiment, =8m, =28m, =8m, =8m, =6m, =70m, =100m, = 49m, = 21m, the order merging area X-axis range is [0, 8], Y-axis range is [0, 28]; the main corridor X-axis range is [8, 16], Y-axis range is [0, 70]; the secondary corridor X-axis range is [64, 72], Y-axis range is [0, 70]; the right corridor X-axis range is [94, 100], Y-axis range is [0, 70]; the fabric storage area X-axis range is [16, 94], Y-axis range is [0, 70], containing the left main rack (35 columns) and the right secondary rack (15 columns) ; the transverse corridor Y-axis coordinate is , a total of 35;
[0106] Step 1.3: According to the actual situation of the SKU storage position, a mapping formula of coding and three-dimensional coordinates is established, which uniquely maps the storage position coding to the coordinate point in the three-dimensional Cartesian coordinate system , realizing the digital modeling of the warehouse space. Let be the column number in the storage position coding, be the row number in the storage position coding, be the layer number in the storage position coding, be the column spacing, be the maximum column number of the left main rack, be the row spacing, be the layer spacing, and the mapping formula is as follows:
[0107]
[0108]
[0109] .
[0110] According to the actual situation of the warehouse, = 8m, = 8m, = 8m, = 1.4m, = 35m, = 49m, = 2m, = 1m, the formula after substitution is as follows:
[0111]
[0112]
[0113] .
[0114] In the step 2, it includes:
[0115] Step 2.1: Determine evaluation indexes and establish a hierarchy system as shown in the following table. Figure 2 As shown, the target layer is order priority; the criterion layer is customer importance, order urgency, and order amount; and the index layer includes order cloth amount , which is a positive index, order arrival time , which is a negative index, order completion time , which is a negative index, order profit rate , which is a positive index.
[0116] Step 2.2: Determine weights using the analytic hierarchy process (AHP), calculate the local weights of each index in the index layer relative to the criterion layer to which it belongs, and the weights of each criterion in the criterion layer relative to the target layer; use expert scoring method, according to the analytic hierarchy process scale rules shown in Table 1, compare each element in the same hierarchy two by two, and construct the index layer judgment matrix and the criterion layer judgment matrix.
[0117] Table 1: Analytic hierarchy process scale table
[0118]
[0119] The index layer judgment matrix is
[0120]
[0121] The criterion layer judgment matrix is
[0122]
[0123] Some elements in the criterion layer matrix are 0 because some indexes are not comparable, and the matrix structure is a block matrix, i.e., unrelated factors are treated separately.
[0124] The maximum eigenvalue of the index layer judgment matrix and the criterion layer judgment matrix is solved by the eigenvalue method and the corresponding normalized eigenvector , which is the relative weight of each index.
[0125] To calculate the maximum eigenvalue and its corresponding eigenvector of the index layer, first calculate the sum of each column of the judgment matrix to obtain the column sum vector , divide each element of the judgment matrix by the column sum of the corresponding column to obtain the normalized matrix ,
[0126] ,
[0127] Calculate the average value of each row to obtain the weight vector of the index layer .
[0128]
[0129] The index layer weight vector Adjust and normalize to get the feature vector one ;
[0130] Calculate the weighted sum vector , each weighted value is divided by the corresponding weight component, and the average is obtained to get the maximum feature value ;
[0131]
[0132] Calculate the maximum feature value of the criterion layer and the corresponding feature vector, since the judgment matrix is a block diagonal matrix, which is divided into two 2x2 submatrices , Process respectively.
[0133] For the submatrix
[0134] , first calculate the sum of each column of the submatrix , get the column sum vector [3.0084, 1.4979], divide each element of the submatrix by the column sum of the corresponding column to get the normalized matrix ;
[0135]
[0136] Calculate the average value of each row to get the weight vector one [0.3324, 0.6676].
[0137] Process the submatrix again:
[0138]
[0139] First, calculate the sum of each column of the submatrix , get the column sum vector [1.9512, 2.0512], divide each element of the submatrix by the column sum of the corresponding column to get the normalized matrix ;
[0140]
[0141] Calculate the average value of each row to get the weight vector two [0.5125, 0.4875]
[0142] Merge the weight vector one and the weight vector two according to the order of the submatrix to get the original four-element vector , expressed as:
[0143]
[0144] The original four-element vector is normalized to obtain The corresponding , , , The adjusted feature vector two is obtained.
[0145] The results are shown in Table 2:
[0146] Table 2 Maximum eigenvalue of judgment matrix and its corresponding eigenvector
[0147]
[0148] The consistency of the judgment matrix is checked, and the consistency ratio , , wherein is the consistency index, is the order of the judgment matrix, is the average random consistency index, as shown in Table 3:
[0149] Table 3 RI table of AHP
[0150]
[0151] When CR<0.1, the consistency check is passed.
[0152] For the index layer, For the criterion layer, , the consistency ratio values are all less than 0.1, passing the consistency check;
[0153] The weight of each index relative to the target layer is calculated , including the order fabric quantity , order arrival time , order completion time , order profit margin , the calculation method is as follows:
[0154] = index item weight × corresponding criterion item weight;
[0155] That is
[0156]
[0157]
[0158]
[0159]
[0160] Obtain the original vector of composite weights for each indicator. ,right Normalization yields the weight vector W as
[0161]
[0162] Step 2.3: Use the fuzzy comprehensive evaluation method (FCE) to evaluate priorities and establish a fuzzy evaluation set. =(Excellent, Good, Average, Poor)=(7,5,3,1); Regarding the amount of fabric in the order. Order arrival time Order completion time Order Profit Margin Separate evaluation criteria were established, and the corresponding indicator values or time intervals for each level were determined, as shown in Table 4:
[0163] Table 4. Order Priority Index Rating Criteria
[0164]
[0165] To distinguish between positive and negative indicators, the membership matrix is determined using the range transformation method. :
[0166] Among them, for positive indicators When the estimated value of the indicator When the value exceeds the "Excellent" standard in the evaluation index, the membership degree for the "Excellent" level is 1, while for other levels it is 0; when the estimated value... When the value is less than the "poor" standard in the evaluation index level, the corresponding level of "excellent" membership is 0, and the membership of other levels is 1; when the estimated value In between, the first The membership degree is calculated as follows: where, This represents the lower limit of the interval for the j-th grade. j=1 corresponds to the lower limit for "excellent", j=2 corresponds to the lower limit for "good", and so on.
[0167] ,
[0168] No. The membership degree is calculated as follows:
[0169] ,
[0170] Other levels have a membership degree of 0;
[0171] For contrarian indicators When the estimated value of the indicator When the value is less than the "Excellent" standard in the evaluation index, the membership degree of the corresponding "Excellent" level is 1, and the membership degree of the other levels is 0; when the estimated value When the value exceeds the "difference" standard in the evaluation index level, the corresponding level difference membership degree is 0; for other levels, it is 1. In between, the first The formula for calculating the membership degree is:
[0172] ,
[0173] The membership degree of level j is calculated as follows:
[0174] ,
[0175] Other levels have a membership degree of 0;
[0176] Step 2.4: Construct the membership matrix Its elements Indicates the first The indicators (C1-C4) belong to the first The membership degree of each grade (Excellent, Good, Average, Poor) is used to quantify the correlation between the indicator and the grade, and to calculate the priority score. , Let R be the indicator weight vector, and R be the fuzzy comment set. This represents the quantitative score for each level. This is the membership matrix, with superscripts... This indicates transpose.
[0177] Taking orders 0101046 and 0101047 as examples, the priorities are calculated using the AHP-FCE model as follows:
[0178] For order 0101046, an order placed outside the town at 15:20, the total amount of fabric shipped was 39 rolls, with SKU shipments of [5,2,1,3,13,1,3,11], totaling 8 SKUs.
[0179] Order fabric quantity Volume 39 belongs to the interval [20, 40), a positive index. According to the range transformation method:
[0180]
[0181] Obtain the membership vector ;
[0182] Order arrival time The value is 15:20, falling within the range [15:00, 18:00), making it a reverse indicator. According to the range transformation method:
[0183]
[0184] Get the membership vector ;
[0185] Order completion time 2 hours, corresponding to the completion time level is "poor", get the membership vector ;
[0186] Order profit margin calculation C4, order total amount = 39 x 100 = 3900, total cost = 8 x 50 = 400, profit margin = (3900-400) / 3900≈89.74%, belongs to the interval [60,80), positive index. According to the range transformation method:
[0187]
[0188] Out of range, take 1, get the membership vector ;
[0189] The membership matrix is
[0190]
[0191] The priority score of this order is
[0192] ;
[0193] For order 0101047, the order is in town, the order time is 15:22, the total fabric out of stock is 87 rolls, and the SKU out of stock is [4,60,5,15,3], a total of 5 SKUs.
[0194] Order fabric quantity 87 rolls ≥ 60, positive index, get the membership vector ;
[0195] Order arrival time 15:22, belongs to the interval [15:00,18:00), reverse index. According to the range transformation method:
[0196]
[0197] Get the membership vector ;
[0198] Order completion time 1 hour, corresponding to the completion time level is "medium", get the membership vector ;
[0199] Order profit rate calculation C4, order total amount = 87x100 = 8700, total cost = 5x50 = 250, profit rate = (8700-250) / 8700≈97.13%, belongs to [60,80) interval, positive direction indicator. According to the range transformation method:
[0200] ;
[0201] Out of range, take 1, get membership degree vector ;
[0202] Then the membership matrix is
[0203]
[0204] The priority score of this order is
[0205] ;
[0206] Order 0101046 priority score is 2.84, order 0101047 priority score is 4.19. The difference between the two mainly comes from the significant difference in the amount of cloth C1, the large order of 87 rolls occupies an advantage in the positive direction indicator, which meets the judgment standard of high priority order, and provides a quantitative basis for subsequent busy period scheduling of resource inclination to high value order.
[0207] In the step 3, as shown in Figure 3 , including:
[0208] Step 3.1: define the scheduling trigger condition, let the number of idle job resources be , the number of waiting orders be , trigger idle period scheduling when , trigger busy period scheduling when ;
[0209] Step 3.2: when , the workload-based lottery scheduling algorithm is triggered to trigger idle period scheduling:
[0210] Step 3.2.1: calculate the relative workload of employees , the calculation formula is:
[0211] ,
[0212] Among them, is the cumulative working time of the th employee, is the total length of the employee's shift;
[0213] Step 3.2.2: allocate lottery numbers in inverse proportion to the relative workload , the calculation formula is:
[0214] ,
[0215] wherein, is the benchmark lottery number, is the adjustment coefficient, the lower the workload, the more lottery numbers; is the lottery number obtained by the kth employee, is the relative workload of the kth employee;
[0216] Step 3.2.3: Build an idle staff lottery pool, randomly draw lottery allocation orders, and achieve resource balance;
[0217] Step 3.3: When , trigger busy period scheduling, i.e. priority-oriented skill stratification scheduling:
[0218] Step 3.3.1: Sort the waiting orders in descending order of priority score , get the sorted result ; denotes the mth order;
[0219] Step 3.3.2: Divide the workers into 2 skill levels, including skilled workers and unskilled workers, and the lottery allocation coefficient of the kth level worker is , ;
[0220] Step 3.3.3: Assign more lottery quotas to high-skilled employees and fewer lottery quotas to low-skilled employees, and build a differentiated lottery pool. For example, when 1 skilled worker and 1 unskilled worker are idle at the same time, the skilled worker gets 10 lotteries and the unskilled worker gets 3 lotteries, and the total lottery pool is 13. The probability of the skilled worker being drawn is 10 / 13, which is significantly higher than that of the unskilled worker.
[0221] Step 3.3.4: For the sorted orders , randomly draw and allocate from the lottery pool, and give priority to high-priority orders handled by skilled workers.
[0222] For example, at 15:20, order 0101046 arrives, and at 15:22, orders 0101047, 0101048, and 0101049 arrive at the same time. At this time, all workers are picking goods, and there is no idle staff, so , trigger busy period scheduling.
[0223] Sort the waiting orders by priority score in descending order, and the result is as follows (order number, type, total volume, priority value):
[0224] (0101047, in-town, 87 volumes, 4.1939) > (0101046, out-of-town, 35 volumes, 2.8420) > (0101048, out-of-town, 26 volumes, 2.5177) > (0101049, out-of-town, 19 volumes, 2.3406).
[0225] When there are idle workers, high-priority order 0101047 is assigned to skilled workers to ensure that in-town orders are shipped within 1 hour.
[0226] Take order 0101044 (arrived at 15:15, in-town order, 39 volumes, containing 8 SKUs: 1930, 2834, 3907, 1337, 3946, 5225, 2021, 1430) as an example, in step 4, as shown, including: Figure 4
[0227] Step 4.1: Filter valid SKUs, represented as set , where the number of valid SKUs ;
[0228] Step 4.2: Convert valid SKUs to three-dimensional coordinates through warehouse layout mapping ;
[0229] Step 4.3: Build a picking path graph model , node set , where is the order consolidation area with coordinates , is the target bin coordinates; edge set , representing feasible paths that meet the warehouse corridor constraints; edge weight is the travel distance, calculated using Manhattan distance, with the formula ; ignore the Z-axis difference;
[0230] Step 4.4: Solve the shortest path using Dijkstra's algorithm:
[0231] Step 4.4.1: Initialize the distance array, representing the current shortest distance estimate from the starting point to node , , , indicating that the shortest distance from the starting point to itself is 0, is the order consolidation area, which is the starting point of path planning, , +∞ represents positive infinity, indicating that the distance from the starting point to other nodes is unknown in the initial state; initialize the priority queue , the set of unvisited nodes ; \ represents the set difference operation, that is, remove nodes from the set to obtain the set of unvisited nodes;
[0232] Step 4.4.2: Take out the node from the priority queue with the minimum current shortest distance estimate value , if , skip; for each adjacent node , calculate the candidate distance , if , update , and add to the priority queue ; remove node from , repeat the above operations in this step until , represents an empty set, that is, all nodes have been processed;
[0233] Step 4.4.3: Backtrack the predecessor nodes from the end point to generate the shortest path , the corresponding picking order is the shortest path set obtained by sorting the elements in the set according to the above shortest path , where is the SKU selected in the set . In the example of order 0101044, the picking order is specifically [1337, 1430, 1930, 2021, 2834, 5225, 3907, 3946]. The total distance is . If the order is picked directly according to the original SKU order, the calculated distance is 514.00m, and the method reduces the distance by 37.9% through path optimization.
[0234] In the step 5, as shown in Figure 5 , comprising:
[0235] Step 5.1: Perform the picking operation according to the planned path in turn, and update the current load and the number of selected SKUs;
[0236] Step 5.2: Determine whether all SKUs of the order / batch have been picked, if not, return to step 5.1 and move to the next SKU position; if yes, trigger the order merging decision based on reinforcement learning; the order refers to the complete demand instruction issued by the customer, including several SKUs and several volumes; the batch refers to when the total volume of SKUs contained in the order exceeds the maximum load capacity of the picking device, the order is split into multiple batches, and the total number of SKUs in a single batch needs to meet the device capacity constraint.
[0237] Step 5.3: Construct a 24-dimensional state vector , including a four-dimensional picker state, a standardized outbound frequency and 20-dimensional features of the orders to be merged; are real numbers;
[0238] The four-dimensional picker state includes position coordinates , current load ratio , current load volume, maximum capacity of the picking device.
[0239] Select the top 5 orders to be processed, each order contains 4 features, including:
[0240] type , town = 1, town = 0;
[0241] standardized quantity , total volume of the order;
[0242] urgency , remaining processing time, time limit;
[0243] position similarity , the average distance between the current path and the order location, the smaller the distance, the higher the similarity, and the remaining dimensions are filled with 0 when there are less than 5 orders;
[0244] Step 5.4: Select the merging action using the DQN algorithm, the action space corresponds to three merging strategies, all of which need to meet the forklift capacity constraint, where do not merge, maintain the current task; merge the entire order, select the highest priority order to be merged and include all remaining SKUs; merge part of the complete SKU, select the order to be merged with the highest position similarity and extract the SKU that can be completely accommodated; complete accommodation means that the remaining capacity of the current picking device can load the entire volume of the SKU at once without overloading.
[0245] Step 5.4.1: Construct a 3-layer fully connected neural network, input state vector, output 3-dimensional action value, hidden layer uses ReLU activation function;
[0246] Step 5.4.2: Use the experience replay mechanism to store samples into the experience pool, and train 64 samples randomly sampled each time;
[0247] Step 5.4.3: Update the reward function value , where is the value estimate of the current state performing action , is the immediate reward, is the discount factor, is the new state after performing the action, is the next action selected in state , is the maximum value of each action in the new state, is the conditional symbol, represents the expectation, which is a mathematical expression of "the sum of the expected rewards and values of all possible subsequent rewards and values after performing action in state ", reflecting the average expectation of long-term cumulative rewards;
[0248] Step 5.4.4: Use - Greedy policy to balance exploration and utilization, set the initial exploration rate = 1.0, decay to 0.01 with training rounds, when it is early, extensive exploration is performed, when it is later, the optimal strategy is utilized;
[0249] Step 5.4.5: Synchronize the main network parameters every 10 iterations of the target network to avoid training shock. The synchronization period of the target network can be adjusted to 5-20 rounds according to the training stability, and the embodiment takes 10 rounds.
[0250] Step 5.5: Perform the selected merge action and calculate the reward, design the reward function where, is the basic merge reward, , is the number of merged volumes, the basic merge reward coefficient = 0.5, encouraging merging behavior, giving rewards according to the number of merged SKUs, and higher rewards when the order is completely merged; is the path saving reward, , is a single processing distance, is a combined distance, path saving reward coefficient =0.3, reward for reduced movement distance after merging; is a capacity utilization reward, , is a combined load rate, capacity utilization reward =10, encourage forklifts to be fully loaded, promote efficient use of resources; is an order completion reward, if the order is completely merged , otherwise , order completion reward coefficient =10, reduce order splitting; is a warehouse-out frequency penalty, , is a warehouse-out frequency, warehouse-out frequency penalty coefficient =0.2;
[0251] Step 5.6: if the merging action is selected, or , then the path planning algorithm of step 4 is re-executed based on the merged SKU to generate a new picking path and return to step 5.1 to continue picking; if the non-merging action is selected, , then the current batch of picking is ended.
[0252] In the step 6, it includes:
[0253] Set the fixed time period length minutes, divide the entire job into consecutive time periods, initialize the time period counter ; the initial state of the first time period is inherited from the end state of the first time period, including the current position of the personnel, the remaining load of the equipment, the unfinished order information and the newly arrived orders; for the order set of the first time period, steps 2 to 5 are repeatedly executed.
[0254] The present application significantly improves the warehouse-out efficiency of the cloth, optimizes the resource allocation, and reduces the operating cost. The implementation effect is shown in Table 5:
[0255] Table 5 Result Comparison
[0256]
[0257] The embodiment takes the actual operation scene of a certain F cloth store as the basis, and compares and verifies the traditional scheduling method and the multi-strategy cooperative scheduling method described in the application. F cloth store mainly wholesales children's clothing fabrics, the picking equipment is a forklift, the maximum capacity is 50 rolls, the operation personnel include 7 skilled workers and 2 unskilled workers, the working period is 10:00-22:00, including lunch break time, the warehouse position coding rule is “RRCCLL”, the first two digits of the SKU code are fabric varieties (01-80), and the last two digits are color numbers (01-50), the town order requires to be shipped out within 1 hour, and the town order requires to be shipped out within 2 hours. The following combines order data and warehouse position information, as shown in Table 6 and Table 7, to explain the implementation process and effect of the application in detail.
[0258] Table 6 Part of order data
[0259]
[0260] Explanation: Each number in the “Outbound quantity (rolls) corresponding to each SKU” column corresponds to the outbound quantity of each SKU in the “SKU code” column, for example, in 1930, 2834, 3907, 1337, 3946, 5225, 2021, 1430 and [3, 7, 1, 6, 2, 1, 14, 5] of the first row, the outbound quantity of the SKU with the number 1930 is 3 rolls, the outbound quantity of the SKU with the number 2834 is 7 rolls, and so on.
[0261] Table 7 Warehouse position information
[0262]
[0263] Table 7 (continued)
[0264]
[0265] Explanation: The first two digits of the SKU code represent the fabric variety, 01 to 80 represent 80 varieties, and the last two digits represent the fabric color number, 01 to 50 represent 50 color numbers.
[0266] Example 2
[0267] A computer system includes a memory, a processor, and a computer program stored on the memory, the processor executes the computer program to implement the steps of the method as described in Example 1.
[0268] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks. Figure 1 one or more flow or blocks.
[0269] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure 1 one or more flow or blocks. Figure 1 one or more flow or blocks.
[0270] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks. Figure 1 one or more flow or blocks.
[0271] The above only is the preferred embodiment of the present application, it should be pointed out that, for those skilled in the technical field, without departing from the technical principles of the present application, can also make a number of improvements and variations, these improvements and variations should also be considered as the protection scope of the present application.
Claims
1. A multi-strategy cooperation-based material delivery scheduling method, characterized in that, The method comprises the following steps: Step 1: establishing a warehouse space coordinate model, and establishing a mapping relationship between a warehouse position code and three-dimensional coordinates; Step 2: calculating an order priority of a newly arrived order by using an analytic hierarchy process-fuzzy comprehensive evaluation model; Step 3: triggering a busy period scheduling strategy or an idle period scheduling strategy according to a current job resource state and a calculation result of the order priority, and completing allocation of the order and the job resource; Step 4: planning a picking path for the allocated order by using a shortest path algorithm; Step 5: in a picking execution process, combining real-time state information, and using a reinforcement learning model to perform order merging; Step 6: repeatedly executing steps 2 to 5 to calculate order scheduling and a picking path in each time period, and updating personnel and equipment usage according to a solution result; In the step 1, comprising: Step 1.1: Taking the lower left corner of the order merging area as the coordinate origin , a three-dimensional Cartesian coordinate system is constructed, and the X-axis, Y-axis and Z-axis represent the horizontal direction, vertical direction and height direction respectively, and the warehouse space is divided into a three-dimensional grid. Step 1.2: Divide the warehouse functional areas, including the order merging area, the corridors including the main corridor, the sub-corridor, the transverse corridor, and the fabric storage area, and define the coordinate ranges of each area: Let be the width of the order merging area in the X-axis direction, be the length of the order merging area in the Y-axis direction, be the width of the main corridor in the X-axis direction, be the width of the sub-corridor in the X-axis direction, be the width of the right corridor in the X-axis direction, be the total length of the warehouse in the Y-axis direction, be the total length of the warehouse in the X-axis direction, be the length of the left main rack in the X-axis direction, be the length of the right sub-rack in the X-axis direction, then the X-axis range of the order merging area is , the Y-axis range is ; the X-axis range of the main corridor is , the Y-axis range is ; the X-axis range of the sub-corridor is , the Y-axis range is ; the X-axis range of the right corridor is , the Y-axis range is ; the X-axis range of the fabric storage area is , the Y-axis range is , including the left main rack and the right sub-rack ; Step 1.3: Establish codes and three-dimensional coordinates based on the actual situation of SKU warehouse locations. The mapping formula uniquely maps the warehouse code to a coordinate point in the three-dimensional Cartesian coordinate system. To achieve digital modeling of warehouse space; This refers to the column number in the position code. This refers to the rank number in the warehouse code. This refers to the layer number in the warehouse code. For column spacing, This represents the maximum number of columns on the left-hand main shelf. For row spacing, The interlayer spacing is mapped using the following formula: ; ; ; In the step 3, comprising: Step 3.1: Define the dispatch trigger condition, let the number of idle job resources be , the number of waiting orders be , the idle period dispatch be triggered when , and the busy period dispatch be triggered when . Step 3.2: When the workload-based lottery scheduling algorithm, triggers the idle period scheduling: Step 3.3: When the busy period schedule, i.e. the priority-oriented skill tiered schedule, is triggered.
2. The material delivery scheduling method based on multi-strategy cooperation according to claim 1, characterized in that, In the step 2, comprising: Step 2.1: Determine evaluation indexes and establish a hierarchy system, the target layer is order priority; the criterion layer is customer importance, order urgency, order amount; the index layer includes order cloth amount , as a positive index, order arrival time , as a negative index, order completion time , as a negative index, order profit rate , as a positive index; Step 2.2: determining a weight by using an analytic hierarchy process, calculating a local weight of each index in an index layer relative to a corresponding criterion layer, and a weight of each criterion in the criterion layer relative to a target layer; comparing each element in the same level two by two according to an analytic hierarchy process scale rule to construct an index layer judgment matrix and a criterion layer judgment matrix by using an expert scoring method; The maximum characteristic value of the index layer judgment matrix and the criterion layer judgment matrix is solved by the characteristic value method respectively and the corresponding normalized characteristic vector , that is, the relative weight of each index. The maximum characteristic value and the corresponding characteristic vector of the index layer are calculated, and the judgment matrix is calculated The sum of each column is obtained, and a column sum vector is obtained Each element is divided by the column sum of the corresponding column to obtain a normalized matrix ; The average value of each row is calculated to obtain the index layer weight vector ; weight vector of the index layer adjusting and normalizing to obtain the feature vector ; Computing a weighted sum vector Each weighted value is divided by the corresponding weight component and averaged to obtain the maximum eigenvalue ; The maximum eigenvalue and the corresponding eigenvector of the criterion layer are calculated, and the judgment matrix is Split into two 2x2 sub-matrix Respectively handle; For a submatrix , first compute the sum of each column of the submatrix , get the column sum vector, divide each element of the submatrix by the column sum of the corresponding column, get the normalized matrix ; Step 2.3: calculating an average value of each row to obtain a weight vector one; Post-processing sub-matrix : First compute the sub-matrix The sum of each column, resulting in a column sum vector, is computed by Each element is divided by the column sum of its corresponding column, resulting in a normalized matrix ; Step 2.4: calculating an average value of each row to obtain a weight vector two; combining weight vector one and weight vector two in submatrix order to obtain original quaternion vector ; to the original quaternion normalized, resulting in the feature vector ; A consistency ratio is calculated for the judgment matrix , , wherein, is a consistency index, is the order of the judgment matrix, is an average random consistency index; When CR < 0.1, the consistency test passes; Calculating each index The indexes include the order cloth amount , order arrival time 2, order completion time 3, order profit rate 4, the calculation method is as follows: ; The synthetic weight original vector of each index is obtained The synthetic weight original vector is normalized The weight vector W is finally normalized Step 2.3: Priority evaluation is carried out by using fuzzy comprehensive evaluation method, and a fuzzy comment set is established Excellent, good, medium, poor ; For the order cloth quantity , order arrival time 2, order completion time 3, order profit rate 4, respectively, establish grade evaluation criteria, determine the corresponding index value or time interval range of each grade, distinguish between positive and negative indicators, and determine the membership matrix by using the range transformation method : Wherein, for the positive indicators When the estimated value of the indicator is greater than the standard of the grade "excellent" in the evaluation indicator grade, then the corresponding grade excellent membership is 1, and the rest is 0; when the estimated value is less than the standard of the grade "poor" in the evaluation indicator grade, then the corresponding grade excellent membership is 0, and the rest is 1; when the estimated value is between the standards of the grades "good" and "excellent", the first grade membership is calculated as follows: ; represents the lower limit value of the jth level, j = 1 corresponds to the lower limit of "excellent", j = 2 corresponds to the lower limit of "good", and j = 3 corresponds to the lower limit of "fair"; The first The level membership degree is calculated as follows: ; Other level membership is 0; For the reverse index When the estimated value of the index is less than the standard of the grade "excellent" in the evaluation index grade, the corresponding grade excellent membership is 1, and the rest is 0; when the estimated value is greater than the standard of the grade "poor" in the evaluation index grade, the corresponding grade poor membership is 0, and the rest is 1; when the estimated value is between the standards of the grades "good" and "excellent", the first grade membership calculation formula is: ; No. The formula for calculating the membership degree is: ; Other level membership is 0; Step 2.4: Constructing the membership matrix , whose elements represent the membership of the th indicator C1-C4 to the th grade, which is used to quantify the degree of association between the indicator and the grade, and the priority score is calculated by , , where W is the indicator weight vector, R is the fuzzy comment set, representing the quantitative score of each grade, is the membership matrix, and the superscript represents the transpose.
3. The material delivery scheduling method based on multi-strategy cooperation according to claim 1, characterized in that, In step 3.2, comprising: Step 3.2.1: Calculate the relative work load of the employee The formula is: ; wherein, is the first cumulative work time for the employee, is the total on-duty time for the employee; Step 3.2.2: Distribute the number of lotteries in inverse proportion to the relative work load The formula is: ; wherein, is the reference number of lottery tickets, is the adjustment factor, the lower the workload, the more lottery tickets; is the number of lottery tickets obtained by the first employee, is the relative workload of the first employee; Step 3.2.3: constructing an idle personnel lottery pool, randomly drawing a lottery to allocate an order, and realizing resource balance.
4. The material delivery scheduling method based on multi-strategy cooperation according to claim 1, characterized in that, In step 3.3, comprising: Step 3.3.1: Priority Score Waiting orders are ranked in descending order to get the ranking result ; represents the th order; Step 3.3.2: divide the workers into skill levels, including skilled workers and unskilled workers, the first level workers having a lottery distribution coefficient of ; Step 3.3.3: the lottery quota allocated to a high-skilled worker is greater than the lottery quota allocated to a low-skilled worker, and a differentiated lottery pool is constructed; Step 3.3.4: Order sorting From the lottery pool, the lottery is randomly assigned, and high-priority orders are guaranteed to be processed by skilled workers.
5. The material dispatch scheduling method based on multi-strategy cooperation according to claim 1, characterized in that, In the step 4, comprising: Step 4.1: Filter for valid SKUs, represented as a set ; Step 4.2: Convert valid SKUs to 3D coordinates through warehouse layout mapping ; Step 4.3: Constructing the picking path graph model , a node set , wherein is an order merging area, is a target bin coordinate; an edge set , representing a feasible path meeting the warehouse corridor constraint; an edge weight is a passing distance, calculated using Manhattan distance, and the calculation formula is ; , , are the coordinates of the target bin in the horizontal direction, the vertical direction and the vertical direction of the warehouse height corresponding to the first effective SKU, respectively; , are the coordinates of the target bin in the horizontal direction and the vertical direction of the warehouse corresponding to the first effective SKU, respectively; Step 4.4: solving a shortest path by using a Dijkstra algorithm: Step 4.4.1: initialize distance array, denotes the start node the current shortest distance estimate from the start node to node , denotes the start node the shortest distance from the start node to itself is 0, , +∞ denotes positive infinity, indicating that the distance from the start node to other nodes is unknown in the initial state; initialize priority queue , the set of unvisited nodes ; denotes the set difference operation, i.e., removing node from the set results in the set of unvisited nodes; Step 4.4.2: From the priority queue Extracting nodes This makes the current shortest distance estimate Minimum if If the adjacent node is not found, skip it; for each adjacent node... Calculate candidate distance ,like Then update and will Add to priority queue ; will node from Remove from the middle, and repeat the above steps until... , This indicates an empty set, meaning all nodes have been processed. Same as step 4.3 , for nodes and The passage distance; Step 4.4.3: generating the shortest path by backtracking from the end node to the predecessor node The corresponding picking sequence is obtained by sorting the elements in the set according to the shortest path above , wherein is the th picked SKU in the set .
6. The material dispatch scheduling method based on multi-strategy cooperation according to claim 1, characterized in that, In the step 5, comprising: Step 5.1: sequentially performing a picking operation according to a planned path, and updating a current load and a picked SKU quantity; Step 5.2: judging whether all SKUs of the order / batch are picked, if not, returning to step 5.1 to continue moving to a next SKU position; if yes, triggering an order merging decision based on reinforcement learning; the order refers to a complete demand instruction of a customer, including a plurality of SKUs and a plurality of volumes; the batch refers to that when the total volume of SKUs of an order exceeds the maximum load capacity of a picking device, the order is split into a plurality of batches, and the total number of SKUs of a single batch needs to meet the device capacity constraint; Step 5.3: Constructing the 24-dimensional state vector including a four-dimensional picker state, a normalized number of times the bin is picked and 20-dimensional features of the orders to be merged; is a real number field; The four-dimensional picker state comprises a position coordinate , a current load ratio , , a current load volume, , a maximum capacity of the picking device; Selecting the top 5 orders to be processed, each order containing 4 features, including: Type , town = 1, outside town = 0; Standardized quantity , is the total number of volumes ordered; urgency , as the remaining processing time, as the time limit; Position similarity , is the average distance of the current path to the order positions, the smaller the distance, the higher the similarity, and the remaining dimensions are filled with 0 when there are less than 5 orders; Step 5.4: Select the merge action using the DQN algorithm, action space For the three merge strategies, the forklift capacity constraint needs to be met, wherein For no merge, maintain the current task; 1 For merging the entire order, select the highest-priority order to be merged and include all remaining SKUs; 2 For merging part of the complete SKU, select the order to be merged with the highest location similarity and extract the SKU that can be completely accommodated; complete accommodation means that the remaining capacity of the current picking device can load all the volumes of the SKU at one time without overloading; Step 5.5: Perform selected merge action and calculate reward, design reward function wherein, is the base merge reward, , is the number of merged orders, is the base merge reward coefficient, encouraging merge behavior, giving reward by the number of SKU merged, giving higher reward when order is completely merged; is the path saving reward, , is the distance of separate handling, is the distance after merging, is the path saving reward coefficient, rewarding the reduced moving distance after merging; is the capacity utilization reward, , is the load rate after merging, is the capacity utilization reward coefficient, encouraging the forklift to be fully loaded; is the order completion reward, if the order is completely merged , otherwise , is the order completion reward coefficient, reducing order splitting; is the put-out frequency penalty, , is the put-out frequency, is the put-out frequency penalty coefficient; Step 5.6: If the merge action is selected, 1 or the path planning algorithm of Step 4 is re-executed based on the merged SKU, a new picking path is generated and the process returns to Step 5.1 to continue picking; if the non-merge action is selected, 0, the current batch of picking is ended.
7. The material dispatch scheduling method based on multi-strategy cooperation according to claim 6, characterized in that, In the step 5.4, comprising: Step 5.4.1: constructing a 3-layer fully connected neural network, inputting a state vector, and outputting a 3-dimensional action value, and using a ReLU activation function for a hidden layer; Step 5.4.2: Store samples with experience replay mechanism to the experience pool, train with 64 samples randomly sampled each time; Step 5.4.3: Update the reward function value , where is the value estimate for the current state performing action , is the immediate reward, is the discount factor, is the new state after performing the action, is the next action chosen in state , is the maximum value of each action in the new state, is the conditional symbol, denotes expectation, and is the mathematical expression for the expected sum of all possible rewards and values following the performance of action in state . Step 5.4.4: Use - A greedy strategy balances exploration and exploitation, setting an initial exploration rate. According to the training rounds Decay to 0.01, In the early stages Conduct extensive exploration, Later Then, the optimal strategy will be utilized; Step 5.4.5: synchronizing main network parameters every iteration setting number of rounds for a target network.
8. A computer system comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program comprises instructions that, when executed by the processor, cause the processor to perform the method of any one of claims 1 to 7. The processor executes the computer program to implement the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Textile order information storage management system
CN111784230A
Dynamic goods picking method and system considering goods picking list relevance
CN113343570A