Automatic guided vehicle scheduling optimization method and device, medium and terminal
By generating fleet heterogeneous data in the intelligent warehouse and integrating undirected graph model, dynamically adjusting the automatic guided vehicle routing and building a scheduling system model, the problem of traditional AGV scheduling methods lacking dynamic adaptability is solved, scheduling efficiency and accuracy are improved, and the scalability and adaptability of the system are improved.
Patent Information
- Application Number
- CN202510173111.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-13
AI Technical Summary
The AGV scheduling methods of traditional smart warehouses lack dynamic adaptability and are difficult to cope with changes in the storage environment, resulting in low efficiency and utilization, and the potential of automatic guided vehicles cannot be fully utilized.
By generating fleet heterogeneous data based on the parameter data of multiple automatic guide vehicles in the intelligent warehouse, and integrating them into the undirected graph model, dynamically adjusting the bicycle routing of the automatic guide vehicles, building a scheduling system model to learn mapping strategies from state to behavior, and outputting optimization decision strategies.
It improves the efficiency and accuracy of the automatic guide vehicle scheduling of the intelligent warehouse, ensures the smooth progress of the scheduling of the automatic guide vehicle in the storage environment and the transportation tasks, and improves the scalability and adaptability of the system.
Smart Images

Figure CN120146449A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of AGV scheduling management, and particularly to a method, device, medium and terminal for optimizing the scheduling of automated guided vehicles. Background Art
[0002] An intelligent warehouse is a modern warehousing system that uses advanced information technology and automated equipment to improve the efficiency and flexibility of warehousing management. It mainly uses technologies such as the Internet of Things, RFID, and machine vision to perform all-round real-time perception and monitoring of goods, equipment, personnel, etc.; adopts automated stereoscopic warehousing, automatic sorting, automatic packaging and other equipment to reduce manual operations and improve efficiency; uses big data analysis and intelligent algorithms to perform real-time scheduling optimization of warehousing operations to improve resource utilization. An AGV (Automated Guided Vehicle) is a very important automated equipment in an intelligent warehouse. Through automated and intelligent operations, it greatly improves the efficiency and accuracy of warehousing operations and is one of the key technologies for realizing an intelligent warehouse.
[0003] The AGV scheduling methods of traditional intelligent warehouses often rely on pre-planned fixed paths and operation sequences, or allocate operations and plan paths based on some fixed rules. However, these methods often have problems such as the lack of dynamic adaptability in the scheduling process, difficulty in coping with changes in the warehousing environment, low efficiency and utilization rate, and inability to fully exert the potential of automated guided vehicles; secondly, AGV scheduling needs to give a feasible solution in a very short time, but pursuing too high optimization quality will reduce the response speed; how to balance between the solution quality and the calculation efficiency is also a technical difficulty. Summary of the Invention
[0004] Embodiments of the present invention provide a method, device, medium and terminal for optimizing the scheduling of automated guided vehicles, which are applicable to warehousing environments of different scales and complexities, and improve the scheduling efficiency and accuracy of automated guided vehicles in intelligent warehouses.
[0005] To solve the above technical problems, embodiments of the present invention provide an automated guided vehicle scheduling optimization method, including:
[0006] Based on the parameter data of multiple automated guided vehicles in an intelligent warehouse, generate fleet heterogeneous data, and incorporate the fleet heterogeneous data into an undirected graph model to obtain an environmental topology graph model; wherein, the undirected graph model is obtained by performing undirected graph modeling according to the warehousing environment topology data of the intelligent warehouse;
[0007] Perform graph topology and heterogeneous feature extraction on the environmental topology graph model to obtain a multi-constraint path planning model, and use the multi-constraint path planning model to dynamically adjust the single-vehicle routing of each automatic guided vehicle in combination with the real-time obtained warehousing site status data of the intelligent warehouse, so as to obtain the dynamic routing table of the intelligent warehouse;
[0008] Based on the dynamic routing table, construct a scheduling system model, and perform a scheduling process simulation on the intelligent warehouse according to the historical scheduling index data of the intelligent warehouse and the scheduling system model, so that the scheduling system model learns the mapping strategy from state to behavior and outputs an optimized decision-making strategy as the scheduling strategy for multiple automatic guided vehicles in the intelligent warehouse.
[0009] As a preferred solution, the process of using the multi-constraint path planning model to dynamically adjust the single-vehicle routing of each automatic guided vehicle in combination with the real-time obtained warehousing site status data of the intelligent warehouse to obtain the dynamic routing table of the intelligent warehouse is specifically as follows:
[0010] Adopt the A3C algorithm to solve the multi-constraint path planning model to obtain several groups of multi-constraint optimal path planning solutions;
[0011] Based on the multi-constraint optimal path planning solutions of each group, construct a multi-agent reinforcement learning environment and the state space of each agent, and obtain the reinforcement learning environment data of the multi-agent; where one group corresponds to one agent;
[0012] According to the reinforcement learning environment data of the multi-agent and the inbound and outbound task list, perform self-confrontation training of the multi-agent to obtain the optimal task allocation strategy;
[0013] According to the real-time obtained warehousing site status data of the intelligent warehouse and the optimal task allocation strategy, perform dynamic programming analysis on the expected paths of each automatic guided vehicle in stages to dynamically adjust the single-vehicle routing of each automatic guided vehicle to obtain the dynamic routing table of the intelligent warehouse.
[0014] As a preferred solution, the process of performing dynamic programming analysis on the expected paths of each automatic guided vehicle in stages according to the real-time obtained warehousing site status data of the intelligent warehouse and the optimal task allocation strategy to dynamically adjust the single-vehicle routing of each automatic guided vehicle to obtain the dynamic routing table of the intelligent warehouse is specifically as follows:
[0015] Using the multi-objective particle swarm optimization algorithm, according to the real-time obtained storage site status data of the intelligent warehouse and the optimal task allocation strategy, perform real-time obstacle status and task information fusion and construct a group obstacle avoidance planning model; wherein, the storage site status data of the intelligent warehouse includes real-time obstacle status data and real-time status data of automated guided vehicles;
[0016] Using the group obstacle avoidance planning model, based on the preset obstacle avoidance safety distance and expected path deviation penalty, perform path planning and time conflict detection for each automated guided vehicle to generate group obstacle avoidance motion plan data;
[0017] Based on the group obstacle avoidance motion plan data and the optimal task allocation strategy, perform dynamic programming analysis on the expected paths in stages for each automated guided vehicle to dynamically adjust the single-vehicle routing of each automated guided vehicle and obtain the dynamic routing table of the intelligent warehouse.
[0018] As a preferred solution, the performing dynamic programming analysis on the expected paths in stages for each automated guided vehicle based on the group obstacle avoidance motion plan data and the optimal task allocation strategy to dynamically adjust the single-vehicle routing of each automated guided vehicle and obtain the dynamic routing table of the intelligent warehouse is specifically as follows:
[0019] According to the assigned tasks of each automated guided vehicle, analyze the expected paths in stages of each automated guided vehicle, and discretize the expected paths in stages into local routing to obtain the initial local routing data of each automated guided vehicle;
[0020] According to the initial local routing data of each automated guided vehicle and the group obstacle avoidance motion plan data, perform dynamic programming analysis on each automated guided vehicle to generate the dynamic routing plan data of each automated guided vehicle;
[0021] Based on the dynamic routing plan data of each automated guided vehicle, dynamically adjust the single-vehicle routing of each automated guided vehicle to obtain the dynamic routing table of the intelligent warehouse.
[0022] As a preferred solution, the using the group obstacle avoidance planning model, based on the preset obstacle avoidance safety distance and expected path deviation penalty, performing path planning and time conflict detection for each automated guided vehicle to generate group obstacle avoidance motion plan data is specifically as follows:
[0023] Using the group obstacle avoidance planning model, based on the preset obstacle avoidance safety distance and expected path deviation penalty, perform path planning for each automated guided vehicle to obtain conflict-free group motion plan data;
[0024] Based on the non-conflicting group movement plan data, analyze the movement paths and speeds of each of the automatic guided vehicles, and then, according to the analysis results, detect time conflicts and adjust the non-conflicting group movement plan data to generate group obstacle avoidance movement plan data.
[0025] As a preferred solution, based on the dynamic routing table, construct a scheduling system model, and according to the historical scheduling index data of the intelligent warehouse and the scheduling system model, simulate the scheduling process of the intelligent warehouse, so that the scheduling system model learns the mapping strategy from state to behavior and outputs an optimized decision-making strategy as the scheduling strategy for multiple automatic guided vehicles in the intelligent warehouse, specifically:
[0026] Based on the dynamic routing table, construct a scheduling system model; wherein, the dynamic routing table includes multiple to-be-evaluated scheduling decisions corresponding to each of the automatic guided vehicles;
[0027] According to the historical scheduling index data of the intelligent warehouse and the scheduling system model, simulate the scheduling process of the intelligent warehouse to obtain the scheduling index values corresponding to each of the to-be-evaluated scheduling decisions, and then, through the scheduling system model, combine the scheduling index values corresponding to all the to-be-evaluated scheduling decisions to perform autonomous learning of the mapping strategy from state to behavior and output an optimized decision-making strategy as the scheduling strategy for multiple automatic guided vehicles in the intelligent warehouse.
[0028] As a preferred solution, generate fleet heterogeneous data based on the parameter data of multiple automatic guided vehicles in the intelligent warehouse, specifically:
[0029] Use the K-Means clustering algorithm to classify the parameter data of multiple automatic guided vehicles in the intelligent warehouse, and according to the classification results, group all the automatic guided vehicles to obtain several groups of fleets; wherein, a fleet includes one or more automatic guided vehicles;
[0030] Generate fleet heterogeneous data based on the classification results and the warehousing task requirements of the intelligent warehouse.
[0031] As a preferred solution, the construction of the undirected graph model is specifically:
[0032] Standardize the topological data of the warehousing environment of the intelligent warehouse to obtain warehousing environment standard data;
[0033] Extract key nodes from the warehousing environment standard data to obtain an undirected graph node set;
[0034] Based on the warehousing environment standard data and the undirected graph node set, extract the connectivity relationship between nodes as the undirected graph edge set;
[0035] Construct an undirected graph model based on the undirected graph node set and the undirected graph edge set.
[0036] As a preferred solution, the integration of the heterogeneous data of the vehicle fleet into the undirected graph model to obtain an environmental topology graph model is specifically as follows:
[0037] Based on the heterogeneous data of the vehicle fleet, assign heterogeneous weights to the undirected graph edges of the undirected graph model to obtain the environmental topology graph model.
[0038] To solve the same technical problem, an embodiment of the present invention also provides an automatic guided vehicle scheduling optimization device, including:
[0039] A model construction module, configured to generate heterogeneous data of the vehicle fleet based on the parameter data of multiple automatic guided vehicles in an intelligent warehouse, and integrate the heterogeneous data of the vehicle fleet into an undirected graph model to obtain an environmental topology graph model; wherein, the undirected graph model is obtained by performing undirected graph modeling based on the warehousing environment topology data of the intelligent warehouse;
[0040] A routing adjustment module, configured to perform graph topology and heterogeneous feature extraction on the environmental topology graph model to obtain a multi-constraint path planning model, and use the multi-constraint path planning model to dynamically adjust the single-vehicle routing of each automatic guided vehicle in combination with the real-time obtained warehousing site status data of the intelligent warehouse to obtain the dynamic routing table of the intelligent warehouse;
[0041] A simulation scheduling module, configured to construct a scheduling system model based on the dynamic routing table, and perform a scheduling process simulation on the intelligent warehouse according to the historical scheduling index data of the intelligent warehouse and the scheduling system model, so that the scheduling system model learns the mapping strategy from state to behavior and outputs an optimized decision-making strategy as the scheduling strategy of multiple automatic guided vehicles in the intelligent warehouse.
[0042] To solve the same technical problem, the present invention also provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the automatic guided vehicle scheduling optimization method described above.
[0043] To solve the same technical problem, the present invention also provides a terminal, including a processor, a memory, and a computer program stored in the memory; wherein, the computer program can be executed by the processor to implement the automatic guided vehicle scheduling optimization method described above.
[0044] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0045] The present invention provides an automatic guided vehicle scheduling optimization method, device, medium and terminal. Based on the parameter data of multiple automatic guided vehicles in an intelligent warehouse, fleet heterogeneous data is generated, and the fleet heterogeneous data is incorporated into an undirected graph model to obtain an environmental topology graph model. Then, graph topology and heterogeneous feature extraction are performed on the environmental topology graph model to obtain a multi-constraint path planning model. Considering various constraint conditions, a more reasonable path is planned for each automatic guided vehicle, avoiding the problem that traditional path planning methods only consider the shortest distance and ignore other factors, improving the quality and feasibility of path planning of the multi-constraint path planning model, and using the multi-constraint path planning model, combined with the real-time obtained warehousing site status data of the intelligent warehouse, to dynamically adjust the single vehicle routing of each automatic guided vehicle to obtain a dynamic routing table of the intelligent warehouse. Through this dynamic adjustment mechanism, changes in the warehouse, such as changes in the position of goods and failures of automatic guided vehicles, can be responded to in a timely manner, ensuring that the automatic guided vehicle is applicable to warehousing environments of different scales and complexities and can always select the optimal path to travel, thereby improving the scalability and adaptability of the system. Then, based on the dynamic routing table, a scheduling system model is constructed, and according to the historical scheduling index data of the intelligent warehouse and the scheduling system model, the scheduling process of the intelligent warehouse is simulated, so that the scheduling system model learns the mapping strategy from state to behavior and outputs an optimized decision-making strategy as the scheduling strategy for multiple automatic guided vehicles in the intelligent warehouse, realizing the efficient optimization of the scheduling strategy, and further improving the scheduling efficiency and accuracy of the automatic guided vehicles in the intelligent warehouse, ensuring the smooth progress of the scheduling and transportation tasks of the automatic guided vehicles in the warehousing environment. In addition, by performing undirected graph modeling based on the warehousing environment topology data of the intelligent warehouse, the warehousing environment can be abstracted into an undirected graph model, and the nodes of the graph in the undirected graph model represent different positions or regions, and the edges of the graph represent the connection relationships between the nodes. Therefore, after incorporating the fleet heterogeneous data into the undirected graph model, the heterogeneous features and constraint conditions in the real warehousing environment can be better reflected, so as to further improve the scheduling accuracy and efficiency of the automatic guided vehicles in the intelligent warehouse. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 : A schematic flow chart of an automatic guided vehicle scheduling optimization provided in Embodiment 1 of the present invention;
[0047] Figure 2 : A schematic structural diagram of an automatic guided vehicle scheduling device provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] Embodiment 1:
[0050] Please refer to Figure 1 , which is an automatic guided vehicle scheduling optimization method provided by an embodiment of the present invention. The method includes steps S1 to S3, and the specific steps are as follows:
[0051] Step S1, based on the parameter data of multiple automatic guided vehicles in the intelligent warehouse, generate fleet heterogeneous data, and integrate the fleet heterogeneous data into an undirected graph model to obtain an environmental topology graph model.
[0052] Among them, the undirected graph model is obtained by performing undirected graph modeling based on the warehousing environment topology data of the intelligent warehouse.
[0053] As a preferred solution, the construction process of the undirected graph model includes steps S01 to S04, and the specific steps are as follows:
[0054] Step S01, perform standardization processing on the warehousing environment topology data of the intelligent warehouse to obtain warehousing environment standard data.
[0055] In this embodiment, the warehousing environment topology data of the intelligent warehouse includes, but is not limited to: warehouse floor plan CAD file data, aisle shelf information data, area partition data, and other environmental data.
[0056] As an example, the acquisition process of the above-mentioned warehousing environment topology data includes steps S011 to S014, and the specific steps are as follows:
[0057] Step S011, obtain the warehouse floor plan CAD file data through modeling software.
[0058] Step S012, read the shelf coding and position information through RFID, and manually draw to obtain the aisle information, so as to obtain the aisle shelf information data.
[0059] Step S013, use the K-Means++ algorithm to automatically partition the operation area to obtain the area partition data.
[0060] Step S014, use the visual SLAM technology to obtain the shelf height, and detect the passage restriction through the infrared sensor to obtain other environmental data.
[0061] It should be noted that the topological data of the storage environment is standardized to ensure data consistency and comparability. First, for the location information of each node, it can be uniformly transformed into the same coordinate system, for example, using the origin of the warehouse floor plan as a reference point. Second, for dimension data such as the distance or length between different nodes, standardization processing can be carried out, for example, converting it into a relative value or a percentage value. Through such standardization processing, standard data of the storage environment can be obtained to ensure the accuracy of subsequent calculations and analyses.
[0062] Step S02: Extract key nodes from the standard data of the storage environment to obtain an undirected graph node set.
[0063] It should be noted that key nodes can be important locations or equipment in the warehouse, such as the location of shelves, intersections, or the entrances of working areas, etc. The selection method of key nodes can be determined according to specific requirements and design criteria. The extracted key nodes will serve as the node set of the undirected graph, providing a basis for subsequent extraction of connectivity relationships and graph construction.
[0064] Step S03: Based on the standard data of the storage environment and the undirected graph node set, extract the connectivity relationships between nodes as the undirected graph edge set.
[0065] It should be noted that the connectivity relationships can be extracted according to the physical distance between nodes, the connectivity of channels, or other rules. For example, the connectivity relationships between nodes can be determined according to the shortest path or reachability between nodes. The extracted connectivity relationships will serve as the edge set of the undirected graph, providing a basis for subsequent topological graph construction.
[0066] Step S04: Construct an undirected graph model according to the undirected graph node set and the undirected graph edge set.
[0067] It should be noted that according to the node set and edge set of the undirected graph, a basic topological undirected graph can be constructed, thereby obtaining the topological basic graph of the storage environment as the undirected graph model. The topological graph can be represented using a graph data structure, where nodes represent locations or equipment in the warehouse, and edges represent the connectivity relationships between nodes. Relevant graph algorithms and toolkits can be used for the construction and representation of the topological graph, such as NetworkX, igraph, etc. By constructing a basic topological undirected graph, the structure and relationships of the storage environment can be better described.
[0068] As a preferred solution, Step S1 includes Step S11 to Step S13, and the specific steps are as follows:
[0069] Step S11: Using the K-Means clustering algorithm, classify the parameter data of multiple automated guided vehicles in the intelligent warehouse based on parameter heterogeneity to obtain heterogeneous grouped data, and group all the automated guided vehicles according to the classification results to obtain several groups of vehicle fleets.
[0070] It should be noted that through the K-Means clustering algorithm, automated guided vehicles with similar parameter characteristics can be classified into the same group, thereby accurately analyzing and identifying the heterogeneity among automated guided vehicles, including factors such as load capacity, traveling speed, charging requirements, and vehicle model dimensions. Grouping automated guided vehicles with high similarity into one group can ensure that the automated guided vehicles within the group have similar capabilities and characteristics, which is conducive to the determination of subsequent formation plans and the implementation of collaborative work. Grouping according to parameter heterogeneity can help formulate a reasonable formation strategy, enabling different groups of automated guided vehicles to effectively collaborate according to task requirements and capabilities, improving the overall transportation efficiency and flexibility.
[0071] Among them, a vehicle fleet includes one or more automated guided vehicles.
[0072] In this embodiment, the acquisition process of the parameter data of multiple automated guided vehicles in the intelligent warehouse includes steps S05 to S06, and the specific steps are as follows:
[0073] Step S05: Collect and upload the load capacity, traveling speed, charging requirements, and vehicle model dimensions of each automated guided vehicle on the edge device of each automated guided vehicle to obtain single-vehicle parameter data.
[0074] It should be noted that obtaining the load capacity of each automated guided vehicle can help determine the quantity and weight of goods it can carry during transportation tasks, thereby reasonably allocating tasks and avoiding overloading. Understanding the traveling speed of each automated guided vehicle helps formulate a reasonable path planning and scheduling strategy to minimize conflicts and waiting times among automated guided vehicles and improve the overall transportation efficiency. Understanding the charging requirements of each automated guided vehicle can help plan the location of charging stations and charging strategies to ensure that the automated guided vehicles have sufficient power and avoid interrupting tasks or reducing operating efficiency due to insufficient power. Obtaining the size information of each automated guided vehicle can help plan appropriate paths and obstacle avoidance strategies to avoid collisions with other objects or automated guided vehicles and ensure safe driving.
[0075] Step S06: Store and verify the single-vehicle parameter data of each automated guided vehicle through the blockchain ledger to obtain the parameter data of multiple automated guided vehicles in the intelligent warehouse.
[0076] As an example, in an intelligent warehousing scenario, in order to collect the single-vehicle parameter data of automated guided vehicles (AGVs), the edge device communicates with the sensor modules and controllers on each AGV to collect the load capacity of each AGV (such as 300 kg), the traveling speed (such as a maximum of 2.5 m / s), the current battery level and charging requirements (such as 30% remaining battery, estimated remaining range of 30 minutes), and the vehicle model dimensions (such as 1.2 m in length, 0.8 m in width, and 0.5 m in height). These data are transmitted to the edge device in real time through a wireless network. After the edge device uniformly formats the data, it uploads the data to the parameter database of the central server, thereby forming a single-vehicle parameter data table. Each row of the data table records the complete parameters of an AGV. Next, in order to ensure the authenticity and security of the single-vehicle parameter data, the system stores the single-vehicle parameter data uploaded by each AGV in the blockchain ledger through blockchain technology. The specific operation is as follows: First, generate a hash value for each group of parameter data (such as SHA-256 encryption), and write the hash value together with the original data into the blockchain; then, the system verifies the data through a consensus algorithm (such as PBFT) among distributed ledger nodes to ensure that the parameter data cannot be tampered with and is authentic and trustworthy. Finally, a blockchain ledger of AGV parameter data is formed, and the recorded content includes the unique identifier of each AGV (such as ID number), the hash value of the parameter data, and the storage timestamp.
[0077] It should be noted that storing the single-vehicle parameter data of AGVs on a distributed ledger through blockchain technology ensures the immutability and traceability of the data, improves the security and credibility of the data, and prevents the data from being maliciously tampered with or forged. Through the blockchain ledger, different participating parties can share and verify the parameter data of AGVs, which helps to achieve collaborative work and information sharing among multiple parties, and improves the efficiency and coordination of the overall transportation system.
[0078] Step S12: Generate heterogeneous data for the vehicle fleet based on the classification results and the warehousing task requirements of the intelligent warehouse.
[0079] In this embodiment, according to the heterogeneous grouping data and different warehousing task requirements, the optimal formation scale and ratio are determined in units of groups, thereby obtaining heterogeneous data for the vehicle fleet.
[0080] As an illustration, according to the heterogeneous grouping data of automated guided vehicles and the requirements of different tasks in intelligent warehousing (such as goods handling, rapid delivery, and shelf arrangement), the system calculates the optimal formation scale and ratio for each group. The specific operation is as follows: First, analyze the task requirements. For example, for goods handling, high-load automated guided vehicles account for 60%, for rapid delivery, high-speed automated guided vehicles account for 80%, and for shelf arrangement tasks, low-power automated guided vehicles account for 70%. Then, combined with the number of automated guided vehicles in each group, optimize the fleet scale for each task through integer programming algorithms (such as 6 high-load and 4 high-speed automated guided vehicles for the handling task; 10 high-speed and 2 high-load automated guided vehicles for the delivery task). The system outputs the heterogeneous data of the fleet, including the distribution of fleet members, the total number of vehicles, and the proportion of each group for each task, forming the final configuration plan for the dispatching system to call.
[0081] It should be noted that according to the heterogeneous grouping data of automated guided vehicles and different task requirements, the scale of each formation and the ratio of automated guided vehicles within the group can be determined to maximize the satisfaction of task requirements and improve transportation efficiency. For example, some tasks may require formations composed of automated guided vehicles with high load-carrying capacity, while other tasks may be more suitable for formations composed of faster automated guided vehicles. According to the heterogeneous data of the fleet, different groups of automated guided vehicles can be dynamically allocated according to task requirements, making the entire fleet flexible and adaptable to handle different types and scales of tasks. By determining the optimal formation scale and ratio, the collaborative working efficiency of the fleet can be improved, ensuring that tasks can be completed in an efficient manner. A reasonable formation scale and ratio can enable the automated guided vehicles in the fleet to cooperate with each other, reduce idle time and conflicts, and improve the overall transportation efficiency.
[0082] Step S13: Based on the heterogeneous data of the fleet, assign heterogeneous weights to the edges of the undirected graph model to obtain the environmental topology graph model.
[0083] It should be noted that heterogeneous weights can be assigned to the edges of the basic topology graph of the warehousing environment according to information such as the maximum load capacity and the highest allowable speed of the heterogeneous data of the fleet. The calculation method of the weights can be determined according to specific requirements and design criteria. For example, the weight of an edge can be calculated based on the physical distance between the nodes connected by the edge and the maximum load capacity of the vehicle to reflect the load conditions of different paths. Similarly, the weight of an edge can be calculated based on the distance between the nodes connected by the edge and the highest allowable speed of the vehicle to reflect the driving time or efficiency of different paths. By assigning heterogeneous weights to the edges of the topology graph, an environmental topology graph model can be constructed, which takes into account the fleet heterogeneity and actual transportation conditions.
[0084] Step S2: Extract graph topology and heterogeneous features from the environmental topology graph model to obtain a multi-constraint path planning model, and use the multi-constraint path planning model to dynamically adjust the single-vehicle routing of each automatic guided vehicle in combination with the real-time obtained storage site status data of the intelligent warehouse, so as to obtain the dynamic routing table of the intelligent warehouse.
[0085] As an optimal solution, step S2 includes steps S21 to S25, and the specific steps are as follows:
[0086] Step S21: Extract graph topology and heterogeneous features from the environmental topology graph model to obtain a multi-constraint path planning model.
[0087] Specifically, use a graph convolutional neural network to extract graph topology and heterogeneous features from the environmental topology graph model to obtain a multi-constraint path planning model.
[0088] In this embodiment, step S21 may include steps S211 to S216, and the specific steps are as follows:
[0089] Step S211: Define a multi-scale convolutional kernel matrix according to the environmental topology graph model, so as to obtain a multi-scale graph convolutional kernel.
[0090] It should be noted that the multi-scale convolutional kernel can include convolutional kernels of different scales, such as convolutional kernels of different sizes or convolutional kernels with different neighborhood ranges. When defining the multi-scale convolutional kernel, it can be represented in matrix form, where each matrix element represents the weight of the convolutional kernel. The specific definition method can be determined according to the specific graph convolution algorithm and task requirements, such as using a multi-channel convolutional kernel or customizing the convolutional kernel shape. By defining the multi-scale graph convolutional kernel, features of different scales can be considered in subsequent graph convolution operations.
[0091] Step S212: Based on the environmental topology graph model and the fleet heterogeneous data, calculate the attention weight of the fleet for each node and the attention weight of the fleet for each edge, so as to obtain a node attention weight vector and an edge attention weight vector.
[0092] It should be noted that the attention weight can be used to represent the attention or focus degree of the fleet on nodes and edges, so as to capture important information. When calculating the attention weight, an attention mechanism can be used, such as using a self-attention mechanism or an attention pooling mechanism. For each node and edge, the attention weight can be calculated by considering the features of the node and edge, the fleet heterogeneous data, and the topological structure. The calculated node attention weight and edge attention weight will respectively form a node attention weight vector and an edge attention weight vector.
[0093] Step S213: Perform attention-weighted multi-layer convolution operations based on the environmental topology graph model, multi-scale graph convolution kernels, node attention weight vectors, and edge attention weight vectors to extract features and obtain graph convolutional neural network feature vectors.
[0094] It should be noted that according to the architecture and algorithm of the graph convolutional neural network, multiple graph convolutional layers can be applied in sequence, and multi-scale convolutional kernels and attention weights can be used to extract features from nodes and edges. In each convolutional layer, activation functions, pooling operations, normalization operations, etc. can be used for further feature processing. Finally, the obtained graph convolutional neural network feature vectors will contain rich representations of nodes and edges in the environmental topology graph.
[0095] Step S214: Use the LSTM sequence model to capture temporal information from the graph convolutional neural network feature vectors and obtain LSTM hidden state vectors.
[0096] It should be noted that by taking the graph convolutional neural network feature vectors as the input sequence, the LSTM sequence model (long short-term memory sequence model) can be used to model them. The LSTM sequence model can capture long-term dependencies in the sequence data and generate corresponding hidden state vectors. By inputting the graph convolutional neural network feature vectors into the LSTM sequence model, temporal information can be captured based on the feature vectors, and the corresponding LSTM hidden state vectors can be obtained.
[0097] Step S215: Apply a fully connected layer mapping constraint to the LSTM hidden state vectors to obtain constraint condition tensors.
[0098] It should be noted that the fully connected layer mapping can map the LSTM hidden state vectors to a vector space with a specific dimension, thereby obtaining constraint condition tensors. These tensors can be used to represent the constraints in path planning, such as path length, avoiding specific areas, or complying with specific rules, etc. By applying a fully connected layer mapping to the LSTM hidden state vectors, the temporal information can be converted into constraint condition tensors, providing constraint information for subsequent path planning.
[0099] Step S216: Input the constraint condition tensors into the environmental topology graph model and perform TSP solving to obtain a multi-constraint path planning model.
[0100] It should be noted that by combining the constraint condition tensor with the environmental topology graph model, a path planning model considering constraint conditions can be constructed. This model can use the TSP algorithm to search the graph and find the optimal path or approximate optimal path that satisfies the constraint conditions. During the solution process of TSP (Traveling Salesman Problem), the constraint information in the constraint condition tensor can be considered. For example, the path length can be restricted or passing through specific nodes can be avoided during the search process. Through this step, a multi-constraint path planning model can be obtained for solving the optimal path under given constraint conditions.
[0101] Step S22: Use the A3C algorithm to solve the multi-constraint path planning model and obtain multiple groups of multi-constraint optimal path planning schemes.
[0102] In this embodiment, step S22 may include steps S221 to S223, and the specific steps are as follows:
[0103] Step S221: Instantiate multiple agents with the heterogeneous data of the vehicle fleet and perform a feedforward of the actor network according to the multi-constraint path planning model to obtain staged path data.
[0104] Specifically, first, the heterogeneous data of the vehicle fleet is instantiated into multiple agents. Each agent represents a member of the vehicle fleet and has its own state and decision-making ability. Then, according to the multi-constraint path planning model, a feedforward of the actor network is performed for each agent. The actor network takes the state of the agent as input and outputs the corresponding actions or path data. According to the definition of the multi-constraint path planning model, the actor network can generate staged path data based on the current state and constraint conditions. Specifically, deep reinforcement learning methods such as Deep Deterministic Policy Gradient (DDPG) or Proximal Policy Optimization (PPO) can be used to train the actor network.
[0105] Step S222: Perform a critic network feedback on the staged path data according to the constraint condition tensor and calculate the reward value to obtain reward value data.
[0106] Specifically, the critic network takes the staged path data and the constraint conditions as input and outputs the corresponding reward value. The critic network is used to evaluate the decision-making quality of the agent, that is, to calculate the reward value according to the actions of the agent and the constraint conditions. Different methods can be used according to the calculation method of the reward value and the requirements of the specific task, such as rule-based evaluation functions, value function-based methods, or contrast-based methods, etc. According to the feedback of the critic network and the calculation of the reward value, reward value data can be obtained for subsequent update training.
[0107] Step S223: Update the actor-critic network parameters according to the reward value data to obtain a multi-constraint optimal path planning scheme.
[0108] Specifically, the reinforcement learning algorithm can use the actor-critic method, such as DDPG or PPO. Specifically, by maximizing the objective function of the cumulative reward value, the gradient ascent method can be used to update the actor network. At the same time, the parameters of the critic network can also be updated through the error backpropagation of the reward value. By continuously iteratively updating the parameters of the actor-critic network, the path planning scheme can be gradually optimized to adapt to the constraint conditions and obtain a multi-constraint optimal solution.
[0109] Step S23: Based on the multi-constraint optimal path planning schemes of each group, construct a multi-agent reinforcement learning environment and the state space of each agent, and obtain the multi-agent reinforcement learning environment data.
[0110] Among them, one group corresponds to one agent.
[0111] In this embodiment, the reinforcement learning environment defines the interaction methods and rules between the agent and the environment. According to the fleet heterogeneous data and the path planning scheme, multiple agents can be constructed, and a state space can be defined for each agent, including vehicle position, task status, environmental constraints, etc. This can be implemented using reinforcement learning libraries and tools, such as OpenAI Gym, StableBaselines3, etc.
[0112] Step S24: According to the multi-agent reinforcement learning environment data and the inbound and outbound task list, perform multi-agent self-play training to obtain an optimal task allocation strategy.
[0113] In this embodiment, Step S24 may include Step S241 to Step S245, and the specific steps are as follows:
[0114] Step S241: Obtain the task data to be processed and sort it according to the priority to obtain the inbound and outbound task list.
[0115] Specifically, first obtain the task data to be processed, including inbound and outbound tasks. Then, sort according to the priority of the tasks to determine the execution order of the tasks. This can be achieved through sorting algorithms or library functions in programming languages, such as the sort() function or sorted() function in Python. The specific basis for setting the priority is as follows: by analyzing the time window constraints of the tasks, calculate the ratio of the remaining executable time to the task execution time. The higher the urgency, the higher the priority. Conduct a weight assessment based on the task type or importance indicators (such as the contribution of the task to the overall goal). For example, for outbound tasks, the priority score can be set according to the value of the goods, the outbound frequency, etc. Utilize the historical task completion records and predict the priority of the current task through a machine learning model (such as a decision tree or logistic regression).
[0116] Step S242, based on the time window constraints, separate the inbound and outbound task list into multiple time segments to obtain a multi-time-slice task set.
[0117] Specifically, the time window constraints define the time periods during which tasks can be executed within a specific time range. According to the time windows of the tasks, the tasks can be divided into different time segments, such that the tasks within each time segment have similar time windows. This can be achieved through loop and conditional statements in programming languages, by traversing each task and making judgments and groupings based on the time windows.
[0118] Step S243, according to the multi-time-slice task set and the reinforcement learning environment data, perform task set partitioning for each time slice to obtain partitioned task subsets.
[0119] Specifically, task set partitioning is to divide the task set into multiple subsets to facilitate subsequent task assignment and scheduling. The task set can be partitioned according to factors such as task attributes, quantity, geographical location, etc. This can be implemented using algorithms and data structures, such as methods based on greedy algorithms, clustering algorithms, graph theory, etc. for task set partitioning.
[0120] Step S244, use the PPO algorithm to perform multi-agent self-adversarial training according to the multi-time-slice task set and the reinforcement learning environment data to obtain task assignment policy data at the time slice level.
[0121] Specifically, the PPO (Proximal Policy Optimization) algorithm is a commonly used reinforcement learning algorithm for optimizing the policy network. Through interaction with the environment, the agent adjusts the parameters of the policy network during training to optimize the task assignment policy. This can be implemented using reinforcement learning libraries and tools, such as Stable Baselines3, RLlib, etc.
[0122] Step S245: Based on the stochastic programming algorithm, perform rolling integration on the task allocation strategy data at each time slice level to obtain the optimal task allocation strategy data.
[0123] Specifically, the stochastic programming algorithm can consider factors such as the temporal relationship of tasks, resource constraints, and optimization objectives, and find the optimal task allocation strategy through an optimization algorithm. The specific stochastic programming algorithm can be selected according to the characteristics of the task allocation problem, such as linear programming, integer programming, simulated annealing, etc. By integrating and optimizing the task allocation strategy data at each time slice level, the final optimal task allocation strategy data can be obtained.
[0124] Step S25: According to the real-time obtained warehousing site status data of the intelligent warehouse and the optimal task allocation strategy, perform dynamic programming analysis on the expected paths in stages for each automatic guided vehicle to dynamically adjust the single-vehicle routing of each automatic guided vehicle and obtain the dynamic routing table of the intelligent warehouse.
[0125] As an optimal solution, step S25 includes steps S251 to S253, and the specific steps are as follows:
[0126] Step S251: Use the multi-objective particle swarm optimization algorithm to perform real-time obstacle state and task information fusion and construct a group obstacle avoidance planning model according to the real-time obtained warehousing site status data of the intelligent warehouse and the optimal task allocation strategy.
[0127] In this embodiment, methods such as multi-objective particle swarm optimization (MOPSO) can be used to construct a group obstacle avoidance planning model, where the AGV is regarded as a constrained particle and the obstacle is regarded as a freely moving particle.
[0128] Among them, the warehousing site status data of the intelligent warehouse includes real-time obstacle status data and real-time AGV status data.
[0129] In this embodiment, a monitoring camera or other video acquisition device can be used to obtain the monitoring video stream data of the intelligent warehouse in real time through a video stream transmission protocol (such as RTSP, HTTP), and corresponding tools and libraries (such as OpenCV) can be used to process and analyze the video stream, and then perform target detection and tracking on the monitoring video stream data to obtain the real-time obstacle status data. The real-time AGV status data can also be obtained according to the self-positioning of the AGV to obtain the real-time position information of the AGV.
[0130] Specifically, for object detection, deep learning algorithms such as YOLO (You Only Look Once), Faster R-CNN (Region-based Convolutional Neural Networks), etc. can be used to identify and locate obstacles in images. For tracking algorithms, algorithms based on Kalman filters, SORT (Simple Online and Realtime Tracking) algorithms, etc. can be used to track the detected obstacles and extract their real-time state data, such as position, speed, direction, etc. The positioning system of the AGV can use sensors and algorithms to determine the position and direction of the AGV in the warehouse. According to the output of the positioning system, the position coordinates and attitude information of the AGV, as well as other relevant state data, such as speed, acceleration, etc., can be obtained.
[0131] Step S252: Using the swarm obstacle avoidance planning model, based on the preset obstacle avoidance safety distance and the expected path deviation penalty, perform path planning and time conflict detection for each automatic guided vehicle to generate swarm obstacle avoidance motion plan data.
[0132] As a preferred solution, step S252 includes steps S2521 to S2522, and the specific steps are as follows:
[0133] Step S2521: Using the swarm obstacle avoidance planning model, based on the preset obstacle avoidance safety distance and the expected path deviation penalty, perform path planning for each automatic guided vehicle to obtain conflict-free swarm motion plan data.
[0134] In this embodiment, the path planning algorithm used for path planning can use heuristic search algorithms such as the A* algorithm, Dijkstra algorithm, etc., as well as optimization algorithms such as genetic algorithms, simulated annealing algorithms, etc. to find the optimal path plan.
[0135] Step S2522: Based on the conflict-free swarm motion plan data, analyze the motion paths and motion speeds of each automatic guided vehicle, and then according to the analysis results, detect time conflicts and adjust the conflict-free swarm motion plan data to generate swarm obstacle avoidance motion plan data.
[0136] In this embodiment, time conflict detection is performed on the group movement plan data without conflicts to ensure that collisions and conflicts do not occur during actual execution. Time conflict detection can be carried out by analyzing the movement paths and speeds of each AGV. If there are situations such as intersections, overlaps on the paths, or inconsistent speeds, there may be time conflicts. Algorithms and models can be used to detect and resolve time conflicts, such as spatio-temporal planning algorithms, discrete event simulations, etc. The finally obtained group obstacle avoidance movement plan data can be used to control and guide the movement of AGVs in the warehousing site to achieve safe and efficient operations.
[0137] It should be noted that the specific implementation process of time conflict detection includes the following key steps: Path spatiotemporal analysis: The motion path and speed data of each AGV are analyzed into a spatiotemporal trajectory, which represents the position coordinates of the AGV at a specific time point and its moving direction. This process can use the piecewise linear interpolation method to refine the motion path into a discrete set of spatiotemporal points and annotate the timestamp and corresponding velocity vector. Conflict area identification: Potential conflict areas are identified by calculating the intersection of the spatiotemporal trajectory point sets of multiple AGVs. These areas correspond to situations where two or more AGVs may occupy adjacent positions in the same time period. This process can be implemented by an overlap detection algorithm based on the time dimension, such as performing intersection calculation on the time window generated by each trajectory. Dynamic safety distance calculation: According to the dynamic characteristics of the AGV (such as maximum deceleration, acceleration), combined with the obstacle state information and warehouse layout, the dynamic safety distance of each AGV is calculated. The dynamic safety distance depends not only on the static distance, but also on the current speed of the AGV and the possible deceleration distance to avoid potential collisions. Conflict risk assessment: The dynamic safety distance is compared with the intersection area of the spatiotemporal trajectory to evaluate whether each conflict area exceeds the safety threshold. By establishing a conflict risk matrix, the possible conflict risks are quantified into conflict levels (such as low, medium, and high), and high-risk areas are marked. Conflict time window analysis: In high-risk conflict areas, the time window and duration of the conflict are further analyzed. For example, the possible time conflict window can be accurately located by analyzing the time difference between each AGV entering and leaving the conflict area. Specific technical means used in the process of resolving time conflicts: Priority scheduling mechanism: Based on the importance of the task and the criticality of path planning, the priorities of different AGVs are assigned. High-priority AGVs will maintain their current paths, and low-priority AGVs need to adjust their speed or path to avoid them. The priority scheduling mechanism can be implemented through a dynamic programming algorithm to minimize the impact of task delays on the overall system efficiency. Path replanning: Local path adjustments are implemented for AGVs that may conflict, and alternative paths are generated based on heuristic search algorithms (such as A* algorithm or Dijkstra algorithm). When replanning the path, the conflict area can be defined as an "obstacle" to prevent AGVs from entering the area and ensure that the replanned path meets the safety distance requirements. Speed adjustment model: For time conflict problems, conflicts are eliminated by adjusting the speed of the AGV. For example, a proportional-integral-derivative (PID) controller can be used to dynamically adjust the AGV speed so that it arrives at the conflict area earlier or later, thereby staggering the time window with other AGVs. Conflict resolution based on spatiotemporal planning: Build a global spatiotemporal planning model to optimize the paths and speeds of all AGVs to minimize the risk of conflict. Spatiotemporal planning can be implemented through constrained optimization algorithms, such as linear programming or integer linear programming (ILP). Constraints include dynamic safety distance, physical characteristics of AGVs, and time window restrictions.Simulation verification and dynamic adjustment: Use a discrete event simulation model to verify the adjusted path and speed scheme. During the simulation process, simulate the movement trajectory of the AGV and observe whether there are new potential conflicts. If new conflicts occur, dynamically adjust the scheme until all conflicts are resolved.
[0138] Step S253: Based on the group obstacle avoidance movement scheme data and the optimal task allocation strategy, perform dynamic programming analysis on the expected paths in different stages of each automatic guided vehicle to dynamically adjust the single-vehicle routing of each automatic guided vehicle, and obtain the dynamic routing table of the intelligent warehouse.
[0139] As an optimal solution, step S253 includes steps S2531 to S2533, and the specific steps are as follows:
[0140] Step S2531: According to the assigned tasks of each automatic guided vehicle, analyze the expected paths in different stages of each automatic guided vehicle, and discretize the expected paths in different stages into local routing to obtain the initial local routing data of each automatic guided vehicle.
[0141] Step S2532: Based on the initial local routing data of each automatic guided vehicle and the group obstacle avoidance movement scheme data, perform dynamic programming analysis on each automatic guided vehicle to generate the dynamic routing scheme data of each automatic guided vehicle.
[0142] Step S2533: Based on the dynamic routing scheme data of each automatic guided vehicle, dynamically adjust the single-vehicle routing of each automatic guided vehicle to obtain the dynamic routing table of the intelligent warehouse.
[0143] Step S3: Based on the dynamic routing table, construct a scheduling system model, and according to the historical scheduling index data of the intelligent warehouse and the scheduling system model, simulate the scheduling process of the intelligent warehouse, so that the scheduling system model learns the mapping strategy from state to behavior and outputs an optimized decision-making strategy as the scheduling strategy for multiple automatic guided vehicles in the intelligent warehouse.
[0144] As an optimal solution, step S3 includes steps S31 to S32, and the specific steps are as follows:
[0145] Step S31: Based on the dynamic routing table, construct a scheduling system model.
[0146] Among them, the dynamic routing table includes multiple to-be-evaluated scheduling decisions corresponding to each automatic guided vehicle.
[0147] Step S32: According to the historical scheduling index data of the intelligent warehouse and the scheduling system model, simulate the scheduling process of the intelligent warehouse to obtain the scheduling index values corresponding to each scheduling decision to be evaluated. Then, through the scheduling system model, combined with the scheduling index values corresponding to all scheduling decisions to be evaluated, perform autonomous learning of the mapping strategy from state to behavior and output an optimized decision-making strategy as the scheduling strategy for multiple automatic guided vehicles in the intelligent warehouse.
[0148] Please refer to Figure 2 , which is a schematic structural diagram of an automatic guided vehicle scheduling optimization device provided by an embodiment of the present invention. The device includes a model construction module M1, a routing adjustment module M2, and a simulation scheduling module M3. The specific functions of each module are as follows:
[0149] The model construction module M1 is used to generate fleet heterogeneous data based on the parameter data of multiple automatic guided vehicles in the intelligent warehouse, and integrate the fleet heterogeneous data into an undirected graph model to obtain an environmental topology graph model. Among them, the undirected graph model is obtained by performing undirected graph modeling according to the warehouse environment topology data of the intelligent warehouse.
[0150] The routing adjustment module M2 is used to extract graph topology and heterogeneous features from the environmental topology graph model to obtain a multi-constraint path planning model, and use the multi-constraint path planning model to dynamically adjust the single-vehicle routing of each automatic guided vehicle in combination with the real-time obtained warehouse site status data of the intelligent warehouse to obtain the dynamic routing table of the intelligent warehouse.
[0151] The simulation scheduling module M3 is used to construct a scheduling system model based on the dynamic routing table, and simulate the scheduling process of the intelligent warehouse according to the historical scheduling index data of the intelligent warehouse and the scheduling system model, so that the scheduling system model learns the mapping strategy from state to behavior and outputs an optimized decision-making strategy as the scheduling strategy for multiple automatic guided vehicles in the intelligent warehouse.
[0152] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described device can refer to the corresponding process in the foregoing method embodiment, and will not be elaborated herein.
[0153] An embodiment of the present invention also provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the automatic guided vehicle scheduling optimization method described in Embodiment 1.
[0154] An embodiment of the present invention also provides a terminal, which includes a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement the automatic guided vehicle scheduling optimization method described in Embodiment 1.
[0155] Preferably, the computer program can be divided into one or more modules / units (such as computer programs, computer programs), and one or more modules / units are stored in the memory and executed by the processor to complete the present invention. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal.
[0156] The processor can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor can also be any conventional processor. The processor is the control center of the terminal and connects various parts of the terminal through various interfaces and lines.
[0157] The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc., and the data storage area can store relevant data, etc. In addition, the memory can be a high-speed random access memory, or can also be a non-volatile memory, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., or the memory can also be other volatile solid-state storage devices.
[0158] It should be noted that the above terminal may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above terminal is only an example and does not constitute a limitation on the terminal. It may include more or fewer components, or combine certain components, or different components.
[0159] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0160] The present invention provides an automatic guided vehicle scheduling optimization method, device, medium and terminal. Based on the parameter data of multiple automatic guided vehicles in an intelligent warehouse, fleet heterogeneous data is generated, and the fleet heterogeneous data is incorporated into an undirected graph model to obtain an environmental topology graph model. Then, graph topology and heterogeneous feature extraction are performed on the environmental topology graph model to obtain a multi-constraint path planning model. Considering various constraint conditions, a more reasonable path is planned for each automatic guided vehicle, avoiding the problem that traditional path planning methods only consider the shortest distance and ignore other factors, improving the quality and feasibility of path planning of the multi-constraint path planning model, and using the multi-constraint path planning model, combined with the real-time obtained warehousing site status data of the intelligent warehouse, to dynamically adjust the single-vehicle routing of each automatic guided vehicle to obtain a dynamic routing table of the intelligent warehouse. Through this dynamic adjustment mechanism, changes in the warehouse, such as changes in the position of goods and failures of automatic guided vehicles, can be responded to in a timely manner, ensuring that the automatic guided vehicle is applicable to warehousing environments of different scales and complexities, and can always select the optimal path to travel, thereby enhancing the scalability and adaptability of the system. Then, based on the dynamic routing table, a scheduling system model is constructed, and according to the historical scheduling index data of the intelligent warehouse and the scheduling system model, the scheduling process of the intelligent warehouse is simulated, so that the scheduling system model learns the mapping strategy from state to behavior and outputs an optimized decision-making strategy as the scheduling strategy for multiple automatic guided vehicles in the intelligent warehouse, realizing the efficient optimization of the scheduling strategy, and further improving the scheduling efficiency and accuracy of the automatic guided vehicles in the intelligent warehouse, ensuring the smooth progress of the scheduling and transportation tasks of the automatic guided vehicles in the warehousing environment. In addition, by performing undirected graph modeling based on the warehousing environment topology data of the intelligent warehouse, the warehousing environment can be abstracted into an undirected graph model, and the nodes of the graph in the undirected graph model represent different positions or regions, and the edges of the graph represent the connection relationships between the nodes. Therefore, after incorporating the fleet heterogeneous data into the undirected graph model, the heterogeneous features and constraint conditions in the real warehousing environment can be better reflected to further improve the scheduling accuracy and efficiency of the automatic guided vehicles in the intelligent warehouse.
[0161] In the specific embodiments described above, the purpose, technical solutions and beneficial effects of the present invention are further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An automatic guided vehicle scheduling optimization method, characterized in that: include: Based on parameter data of multiple automated guided vehicles in the intelligent warehouse, heterogeneous data of the fleet is generated, and the heterogeneous data of the fleet is integrated into an undirected graph model to obtain an environmental topological graph model; wherein the undirected graph model is obtained by performing undirected graph modeling based on the storage environment topological data of the intelligent warehouse; Performing graph topology and heterogeneous feature extraction on the environment topology graph model to obtain a multi-constraint path planning model, and using the multi-constraint path planning model in combination with the storage site status data of the intelligent warehouse acquired in real time to dynamically adjust the single vehicle route of each of the automatic guided vehicles to obtain a dynamic routing table of the intelligent warehouse; Based on the dynamic routing table, a scheduling system model is constructed, and according to the historical scheduling index data of the intelligent warehouse and the scheduling system model, the scheduling process of the intelligent warehouse is simulated, so that the scheduling system model learns the mapping strategy from state to behavior and outputs the optimization decision strategy as the scheduling strategy for multiple automatic guided vehicles in the intelligent warehouse.
2. The method for optimizing automatic guided vehicle scheduling according to claim 1, characterized in that: The multi-constraint path planning model is used in combination with the storage site status data of the intelligent warehouse obtained in real time to dynamically adjust the single vehicle route of each of the automatic guided vehicles to obtain a dynamic routing table of the intelligent warehouse, specifically: Using the A3C algorithm to solve the multi-constraint path planning model, and obtain multiple groups of multi-constraint optimal path planning solutions; Based on the multi-constraint optimal path planning scheme of each group, a multi-agent reinforcement learning environment and a state space of each agent are constructed, and reinforcement learning environment data of the multi-agent is obtained; wherein one group corresponds to one agent; According to the multi-agent reinforcement learning environment data and the inbound and outbound task list, multi-agent self-adversarial training is performed to obtain the optimal task allocation strategy; According to the storage site status data of the intelligent warehouse acquired in real time and the optimal task allocation strategy, a dynamic planning analysis of the expected path in stages is performed on each of the automatic guided vehicles to dynamically adjust the single vehicle route of each automatic guided vehicle to obtain a dynamic routing table of the intelligent warehouse.
3. The method for optimizing automatic guided vehicle scheduling as claimed in claim 2, characterized in that: According to the real-time storage site status data of the intelligent warehouse and the optimal task allocation strategy, a dynamic planning analysis of the expected path in stages is performed on each of the automatic guided vehicles to dynamically adjust the single vehicle route of each of the automatic guided vehicles to obtain a dynamic routing table of the intelligent warehouse, which is specifically: Using a multi-objective particle swarm optimization algorithm, based on the real-time storage site status data of the intelligent warehouse and the optimal task allocation strategy, the real-time obstacle status and task information are integrated and a group obstacle avoidance planning model is constructed; wherein the storage site status data of the intelligent warehouse includes real-time obstacle status data and real-time status data of automatic guided vehicles; Utilizing the group obstacle avoidance planning model, based on a preset obstacle avoidance safety distance and an expected path deviation penalty, path planning and time conflict detection are performed on each of the automatic guided vehicles to generate group obstacle avoidance motion plan data; Based on the group obstacle avoidance motion plan data and the optimal task allocation strategy, a dynamic planning analysis of the expected path in stages is performed on each of the automated guided vehicles to dynamically adjust the single vehicle route of each automated guided vehicle to obtain a dynamic routing table for the intelligent warehouse.
4. The method for optimizing automatic guided vehicle scheduling as claimed in claim 3, characterized in that: Based on the group obstacle avoidance motion scheme data and the optimal task allocation strategy, a dynamic planning analysis of the expected path in stages is performed on each of the automatic guided vehicles to dynamically adjust the single vehicle route of each of the automatic guided vehicles to obtain a dynamic routing table of the intelligent warehouse, specifically: According to the assigned tasks of each of the automated guided vehicles, the phased expected paths of each of the automated guided vehicles are parsed, and the phased expected paths are discretized into local routes to obtain initial local route data of each of the automated guided vehicles; According to the initial local routing data of each of the automated guided vehicles and the group obstacle avoidance motion plan data, dynamic planning analysis is performed on each of the automated guided vehicles to generate dynamic routing plan data for each of the automated guided vehicles; Based on the dynamic routing plan data of each of the automated guided vehicles, the single vehicle route of each of the automated guided vehicles is dynamically adjusted to obtain a dynamic routing table of the intelligent warehouse.
5. The method for optimizing automatic guided vehicle scheduling as claimed in claim 3, characterized in that: The group obstacle avoidance planning model is used to perform path planning and time conflict detection on each of the automatic guided vehicles based on a preset obstacle avoidance safety distance and an expected path deviation penalty, so as to generate group obstacle avoidance motion plan data, specifically: Utilizing the group obstacle avoidance planning model, based on a preset obstacle avoidance safety distance and an expected path deviation penalty, path planning is performed on each of the automatic guided vehicles to obtain conflict-free group motion plan data; Based on the conflict-free group motion plan data, the motion path and motion speed of each of the automatic guided vehicles are analyzed, and then according to the analysis results, time conflicts are detected and the conflict-free group motion plan data are adjusted to generate group obstacle avoidance motion plan data.
6. The method for optimizing automatic guided vehicle scheduling according to claim 1, characterized in that: Based on the dynamic routing table, a scheduling system model is constructed, and according to the historical scheduling index data of the intelligent warehouse and the scheduling system model, the scheduling process of the intelligent warehouse is simulated, so that the scheduling system model learns the mapping strategy from state to behavior and outputs the optimization decision strategy as the scheduling strategy of multiple automatic guided vehicles in the intelligent warehouse, specifically: Based on the dynamic routing table, a scheduling system model is constructed; wherein the dynamic routing table includes a plurality of scheduling decisions to be evaluated corresponding to each of the automatic guided vehicles; According to the historical scheduling index data of the intelligent warehouse and the scheduling system model, the scheduling process of the intelligent warehouse is simulated to obtain the scheduling index value corresponding to each scheduling decision to be evaluated. Then, through the scheduling system model, combined with the scheduling index values corresponding to all the scheduling decisions to be evaluated, the mapping strategy from state to behavior is autonomously learned and the optimized decision strategy is output as the scheduling strategy for multiple automatic guided vehicles in the intelligent warehouse.
7. The method for optimizing automatic guided vehicle scheduling according to claim 1, characterized in that: The generating of fleet heterogeneous data based on parameter data of multiple automated guided vehicles in the intelligent warehouse is specifically as follows: Using the K-Means clustering algorithm, the parameter data of the multiple automated guided vehicles in the intelligent warehouse are classified and processed, and according to the classification results, all the automated guided vehicles are grouped to obtain several groups of fleets; wherein a fleet includes one or more automated guided vehicles; Based on the classification results and the storage task requirements of the smart warehouse, fleet heterogeneous data is generated.
8. The method for optimizing automatic guided vehicle scheduling according to claim 1, characterized in that: The construction of the undirected graph model is specifically as follows: Standardizing the storage environment topology data of the intelligent warehouse to obtain storage environment standard data; Extracting key nodes from the storage environment standard data to obtain an undirected graph node set; Based on the storage environment standard data and the undirected graph node set, the connectivity relationship between the nodes is extracted as an undirected graph edge set; An undirected graph model is constructed according to the undirected graph node set and the undirected graph edge set.
9. An automatic guided vehicle scheduling optimization method according to any one of claims 1 to 8, characterized in that: The heterogeneous data of the fleet is integrated into the undirected graph model to obtain the environment topology graph model, which is specifically: Based on the heterogeneous data of the fleet, heterogeneous weights are assigned to the undirected graph edges of the undirected graph model to obtain the environmental topology graph model.
10. An automatic guided vehicle scheduling optimization device, characterized in that: include: A model building module is used to generate fleet heterogeneous data based on parameter data of multiple automated guided vehicles in the intelligent warehouse, and integrate the fleet heterogeneous data into an undirected graph model to obtain an environmental topology graph model; wherein the undirected graph model is obtained by performing undirected graph modeling based on the storage environment topology data of the intelligent warehouse; A route adjustment module is used to extract graph topology and heterogeneous features from the environment topology graph model to obtain a multi-constraint path planning model, and use the multi-constraint path planning model in combination with the storage site status data of the intelligent warehouse obtained in real time to dynamically adjust the single vehicle route of each of the automatic guided vehicles to obtain a dynamic routing table of the intelligent warehouse; A simulation scheduling module is used to build a scheduling system model based on the dynamic routing table, and simulate the scheduling process of the smart warehouse according to the historical scheduling indicator data of the smart warehouse and the scheduling system model, so that the scheduling system model learns the mapping strategy from state to behavior and outputs an optimization decision strategy as the scheduling strategy for multiple automatic guided vehicles in the smart warehouse.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute an automatic guided vehicle scheduling optimization method as described in any one of claims 1 to 9.
12. A terminal, characterized in that: It comprises a processor, a memory and a computer program stored in the memory; wherein the computer program can be executed by the processor to implement an automatic guided vehicle scheduling optimization method as described in any one of claims 1 to 9.
Citation Information
Cited By
AGV cluster scheduling method and system thereof
CN120725555A
AGV Cluster Scheduling Method and System
CN120725555B
Industrial Internet of Things unmanned vehicle path planning system and method
CN120806316A
Industrial internet of things unmanned vehicle path planning system and method
CN120806316B
Whole-process AI intelligent agent management and control system based on integrated wheel set maintenance intelligent workshop
CN121119983A