Port horizontal transportation equipment real-time scheduling method and system based on MIMO mechanism
Through the real-time scheduling method of port horizontal transportation equipment based on the MIMO mechanism, the problem of poor dynamic environmental adaptability in port scheduling is solved, real-time response to emergencies is achieved, and port operation efficiency and throughput are improved.
Patent Information
- Application Number
- CN202510341731.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The existing technology has poor dynamic environmental adaptability in port horizontal transportation equipment scheduling and is unable to respond to task changes in real time, resulting in congestion and waiting time of horizontal transportation, making it difficult to meet the real-time and robust needs of large-scale port scenarios.
The real-time scheduling method of port horizontal transportation equipment based on MIMO mechanism is adopted. By encapsulating port loading and unloading operations and horizontal transportation operations into MIMO interfaces, defining input and output interfaces and objective functions, establishing the operating mechanism and mathematical model of port operations, and performing reinforcement learning algorithm configuration in the simulation environment to obtain the optimal scheduling strategy.
Real-time response to sudden factors in a dynamic environment is achieved, the average waiting time of horizontal transportation equipment and shore and bridge waiting time is reduced, port throughput and operational efficiency is improved, and real-time and robustness needs of the port industry are met.
Smart Images

Figure CN120338633A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent logistics and smart ports, and in particular to a real-time scheduling method and system for port horizontal transportation equipment based on the MIMO mechanism. Background Art
[0002] With the acceleration of the globalization process, the importance of maritime transportation has become increasingly prominent and has become the pillar of international trade and regional exchanges. As a transfer point between maritime and land transportation, ports play a key role in determining the efficiency of the entire logistics supply chain. Generally, the main equipment of an automated port is divided into loading and unloading equipment such as quay cranes (hereinafter referred to as quay cranes) and yard cranes (hereinafter referred to as yard cranes) and horizontal transportation equipment (such as AGVs, ARTs, IGVs, etc.). Among them, horizontal transportation equipment is the key transportation equipment connecting quay cranes and yards, and its scheduling efficiency directly affects port throughput and operating costs. Therefore, there are many studies on the scheduling of horizontal transportation equipment.
[0003] Regarding the scheduling problem of horizontal transportation equipment, the industry mainly adopts heuristic scheduling rules, and the academic community usually uses exact algorithms or intelligent optimization algorithms to model and solve. The heuristic scheduling rules adopted in the port industry mainly include first come first served, shortest path first, etc. These rules are simple to implement and manage, but have poor optimization performance; while the methods of modeling and solving with exact or intelligent optimization algorithms commonly used in the academic community are also faced with the dilemma of long calculation time and are difficult to meet the real-time requirements (usually taking several minutes to several hours). When facing a dynamic environment, these methods cannot adapt to dynamic task requests and sudden interferences (such as horizontal transportation equipment failures, path congestion) due to high computational complexity, and cannot dynamically adjust strategies according to real-time states (such as horizontal transportation equipment positions, task queues, energy consumption), resulting in problems such as horizontal transportation congestion and long waiting times, and cannot optimize scheduling strategies in real time. Therefore, it is difficult to meet the real-time and robust requirements of large-scale port scenarios.
[0004] To address the above problems, dynamic scheduling methods have received increasing attention. Among many dynamic scheduling methods, the reactive scheduling method is the most capable of dealing with highly complex and dynamic scenarios. Reactive scheduling refers to the system's ability to respond promptly to sudden uncertain factors to generate or modify a scheduling plan, including two major categories: predictive reactive scheduling and fully reactive scheduling. Although it can respond to real-time states, it lacks forward-looking optimization and is prone to falling into local optima.
[0005] Reinforcement learning technology has been a research hotspot in recent years, and the idea of full reactive scheduling can be organically integrated with the real-time scheduling technology based on reinforcement learning. For example, reinforcement learning methods can better adapt to different environmental states. The scheduling system interacts with the automated port environment, continuously learns, takes the best actions according to the strategy, considers forward-looking optimization, and finally forms a strategy to take actions based on the state, which is more suitable for scenarios with high real-time requirements.
[0006] Although some reinforcement learning algorithms have shown potential in dynamic scheduling, their direct application in complex ports faces three major challenges. One is the explosion of the state space dimension caused by multi-node coupling, the second is the difficulty in designing the reward function under multi-objective conflicts. The third is the difficulty in combining the simulation for ports. Most of the existing simulations do not endow their models with real physical meanings (or simplify many details), mainly serving the demonstration function of the algorithm principle, and it is difficult to simulate the linkage of various devices. Summary of the Invention
[0007] In view of the above existing problems, the present invention is proposed.
[0008] Therefore, the present invention provides a real-time scheduling method and system for port horizontal transportation equipment based on the MIMO mechanism to solve the problems of poor adaptability to the dynamic environment and inability to respond to task changes in real time during the traditional scheduling process of port horizontal transportation equipment.
[0009] To solve the above technical problems, the present invention provides the following technical solutions:
[0010] In the first aspect, the present invention provides a real-time scheduling for port horizontal transportation equipment based on the MIMO mechanism, including:
[0011] Obtain the operation parameters in the terminal operation environment and equipment, and encapsulate the port loading and unloading operations and horizontal transportation operations into a MIMO interface;
[0012] Based on the MIMO interface, establish the operation mechanism and its mathematical model of port operations by defining the input and output interfaces and the objective function;
[0013] Based on the MIMO interface and the mathematical model, build a road network and physical facilities to construct a port simulation environment;
[0014] Configure the reinforcement learning algorithm in the simulation environment to obtain the optimal scheduling strategy for port horizontal transportation equipment.
[0015] As a preferred solution of the real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism of the present invention, wherein:
[0016] The establishment of the operation mechanism and its mathematical model of port operations by defining the input and output interfaces includes the following steps:
[0017] Define multiple input interfaces according to different container sources, and define multiple output interfaces according to different destinations;
[0018] Design the operation mechanism process, including task allocation, path planning, and loading and unloading operations when the ship arrives at the port;
[0019] Convert the operation mechanism into a mathematical model and set the objective function;
[0020] Set the objective function to minimize the sum of the quay crane waiting time and the idle waiting time of the horizontal transportation equipment.
[0021] As a preferred solution of the real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism described in the present invention, wherein:
[0022] The configuration of the reinforcement learning algorithm in the simulation environment includes the following steps:
[0023] Sort and index the state information in the horizontal transportation equipment set through the sorting index function to generate a state array containing time and distance;
[0024] Input the state array into the Q-learning algorithm, update the Q(s,a) value in the Q table through the Training function, and obtain the maximum Q value of all possible actions under a given state through the get_Max_QValue function;
[0025] Based on the Q value, combined with multiple rule bases, select the optimal action through the ε_temperature-action function and the act-action function;
[0026] Based on the optimal action, design a hierarchical composite reward function and dynamically adjust the weight according to the MIMO coupling relationship.
[0027] As a preferred solution of the real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism described in the present invention, wherein:
[0028] The act-action function includes setting a priority policy;
[0029] The setting of the priority policy includes:
[0030] Give priority to selecting the equipment with the longest waiting time;
[0031] If there is no equipment with the longest waiting time, select the equipment closest to the target;
[0032] If there is no equipment with the longest waiting time and the equipment closest to the target, select the equipment farthest from the target;
[0033] The equipment that selects the longest waiting time It is expressed as:
[0034]
[0035] The equipment that selects the one closest to the target It is expressed as:
[0036]
[0037] The equipment that is the farthest from the target It is expressed as:
[0038]
[0039] Among them, C represents the set of horizontal transportation equipment, and v j represents the equipment in the candidate equipment set C candidate in, v i represents the equipment from all available equipment sets C, and d j represents the equipment v j to the target, and d i represents the equipment v i to the target.
[0040] As a preferred solution of the real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism described in the present invention, wherein:
[0041] The sorting and indexing processing through the sorting index function to generate a state array containing time and distance is expressed as:
[0042]
[0043] Among them, qc id represents the identifier of the current quay crane, C represents the set of horizontal transportation equipment, |C| represents the number of available transportation equipment in the set C, and rank T (C1) represents the index array obtained by sorting the set C according to the total waiting time, and rank d (C2) represents the index array obtained by sorting the set C according to the distance to the current quay crane.
[0044] As a preferred solution of the real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism described in the present invention, wherein:
[0045] The update of the Q(s,a) value in the Q table through the Training function is expressed as:
[0046]
[0047] Among them, r represents the reward for taking the current action, r represents the discount factor, α represents the learning rate, and r(s,a) represents the immediate reward obtained after taking action a in state s.
[0048] As a preferred solution of the real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism described in the present invention, wherein:
[0049] The hierarchical composite reward function is expressed as:
[0050] r t = (-1) * [α * (T qc ) + β * T + γ * V]
[0051] Among them, r t is the immediate reward value, T qc is the waiting time increment of the horizontal transportation equipment, T is the idle waiting time of the quay crane, V is the system efficiency, and α, β, and γ are weight coefficients.
[0052] In a second aspect, the present invention provides a real-time scheduling system for port horizontal transportation equipment based on the MIMO mechanism, including:
[0053] An encapsulation module, configured to obtain operation parameters in the terminal operation environment and equipment, and encapsulate port loading and unloading operations and horizontal transportation operations into a MIMO interface;
[0054] A construction module, configured to abstract operation parameters in the terminal operation environment into network communication parameters and construct a MIMO interface;
[0055] A modeling module, configured to establish an operation mechanism and a mathematical model of port operations based on the MIMO interface by defining input and output interfaces and an objective function;
[0056] A simulation module, configured to build a port simulation environment by constructing a road network and physical facilities based on the MIMO interface and the mathematical model;
[0057] A configuration module, configured to configure a reinforcement learning algorithm in the simulation environment to obtain an optimal scheduling strategy for port horizontal transportation equipment.
[0058] In a third aspect, the present invention provides a computing device, including:
[0059] A memory, configured to store a program;
[0060] A processor, configured to execute the computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism are implemented.
[0061] Fourthly, the present invention provides a computer-readable storage medium, including: when the program is executed by a processor, the steps of the real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism are implemented.
[0062] Advantages of the present invention: The present invention proposes a real-time scheduling framework for port horizontal transportation equipment integrating the multi-input multi-output (MIMO) network mechanism. A simulation environment based on a real scenario is built through Anylogic simulation software and the reinforcement learning module is configured. At the same time, full-process visualization and data output are realized. The state space, action space, and reward function in the improved algorithm are improved. By introducing a modeling framework based on the MIMO operation mechanism, lightweight state space, action space, and other modules in the framework are completely constructed in the scheduling system of the simulation platform, forming a real-time scheduling method for horizontal transportation equipment. The aim is to solve the problems that traditional methods rely on offline calculation and cannot respond in real time to the impact of dynamic factors such as task priority adjustment and equipment failure on the scheduling system, with higher adaptability and meeting the needs of the current port industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:
[0064] Figure 1 It is a schematic diagram of the basic process of a real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism provided by an embodiment of the present invention;
[0065] Figure 2 It is a schematic diagram of the network interface docking of a real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism provided by an embodiment of the present invention;
[0066] Figure 3 It is a schematic diagram of the construction of the MIMO transportation network of a real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism provided by an embodiment of the present invention;
[0067] Figure 4 It is a schematic diagram of the real-time scheduling of horizontal transportation equipment based on simulation of a real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism provided by an embodiment of the present invention;
[0068] Figure 5 It is a flowchart of the algorithm of a real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism provided by an embodiment of the present invention;
[0069] Figure 6 An iterative diagram of an embodiment of a real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism provided by an embodiment of the present invention. Detailed implementation manners
[0070] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0071] Embodiment 1
[0072] Refer to Figure 1-5 , an embodiment of the present invention provides a real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism, including:
[0073] S1: Obtain the operation parameters in the terminal operation environment and equipment, and encapsulate the port loading and unloading operations and horizontal transportation operations into a MIMO interface;
[0074] In the embodiment of the present application, encapsulating the port loading and unloading operations and horizontal transportation operations into a MIMO interface, as Figure 2 shown, includes the following steps:
[0075] Encapsulate the number of receiving areas as the number of interfaces of network antennas, that is, input and output nodes;
[0076] Encapsulate the number of containers operated by the crane at one time as the channel capacity of the network interface;
[0077] Encapsulate the number of vehicles parked in the operation area as the network interface data transmission frame buffer;
[0078] Encapsulate the operation efficiency of the crane as the data transmission rate of the network interface;
[0079] Encapsulate the transport capacity of the port horizontal transportation equipment and the number of equipment as the bandwidth;
[0080] Encapsulate the number of containers transported by the port horizontal transportation equipment per unit time as the communication system capacity;
[0081] Encapsulate uncertainties such as AGV failures and weather impacts as noise;
[0082] Encapsulate the container task as a data packet, called "container flow";
[0083] Encapsulate the internal roads of the port as channels;
[0084] Encapsulate the container scheduling strategy as modulation / coding;
[0085] In the embodiment of the present application, after extracting and encapsulating the MIMO interface of the port operation, the data transfer of "container flow / vehicle flow" becomes the data transmission of data packets, and each network interface processes data in parallel, realizing the effective logical separation of the horizontal transportation system and the stacking / loading and unloading operation system.
[0086] It should be noted that the present invention introduces the MIMO communication theory framework, encapsulates ships, trains, external container trucks, and each unloading point of the terminal (yard crane unloading point, quay crane unloading point), and maps them into the input / output points of "container flow" data. The internal path network of the terminal is modeled as a channel matrix, etc. This modeling method explicitly quantifies the container circulation efficiency between nodes through matrix operations, and also provides a low-dimensional state representation and a hierarchical reward function for reinforcement learning. It can also encapsulate multiple devices in the terminal, facilitating the systematic modeling and effect realization of simulation software.
[0087] S2: Based on the MIMO interface, establish the operation mechanism and its mathematical model of port operation by defining the input / output interface and the objective function;
[0088] In the embodiment of the present application, a MIMO interface-integrated operation mechanism is constructed, which describes the problems and mathematical models to be constructed, including the following steps:
[0089] 2.1 Construct the number of interfaces. MIMO can realize the input of "container flow" from multiple sources. Corresponding to different container sources, such as the input interfaces from ships, yards, trains, road external container trucks, and quay cranes, its mathematical representation is X = [x ship , x train , x truck , x q / yc T ; Multiple outputs correspond to the container output destinations, such as the output interfaces of ships, yards, road container trucks, trains, and yard cranes. Its mathematical representation is Y = [y ship , y train , y truck , y q / yc T . In the present invention, buffer points (waiting areas) are added to form a multi-input multi-output transportation network in combination with the interfaces, as shown in Figure 3 .
[0090] 2.2 Determine the work tasks and objectives. The objective of the horizontal transportation system in a container terminal is to efficiently and accurately complete the connection tasks of containers between the loading / unloading and stacking operation units. That is, the operation control can be abstracted as a real-time scheduling problem of horizontal transportation equipment driven by the MIMO mechanism. Therefore, the task of the operation mechanism is to make reasonable real-time decisions through algorithms on the premise of known ("container flow" data), select appropriate horizontal transportation equipment (routes in MIMO) to perform tasks, so that the sum of the idle waiting times of all horizontal transportation equipment and the waiting time of quay cranes (delay) is minimized.
[0091] 2.3 Construct the operation mechanism. The operation mechanism of the present invention is that when a ship arrives at the port (data enters), the intelligent agent undertakes the task assignment function, calls the algorithm to achieve real-time assignment of tasks according to the environmental state, and selects appropriate horizontal transportation equipment according to the scheduling strategy (modulation, coding). When the horizontal transportation equipment arrives at the quay crane (input node), the quay crane and spreader perform the loading and unloading operations. After completing the loading of the container from the quay crane to the horizontal transportation equipment, the horizontal transportation equipment obtains the container task information, moves to the target bay in the target block area on the road (channel) through path planning. At the same time, the yard crane (output node) moves to the unloading point. Subsequently, the spreader on the yard crane performs the unloading work, moves to the target container position, and completes the unloading from the horizontal transportation equipment to the yard (realize multi-path data transmission). The horizontal transportation equipment goes to the waiting area to wait for the next task instruction. The data transmission, that is, the transmission of the container flow, is completed.
[0092] 2.4 Conduct the mathematical modeling process based on its operation mechanism
[0093] 2.4.1 Describe the problem solved by this method: On the premise of known operation tasks ("container flow" data), make reasonable real-time decisions through algorithms, select appropriate horizontal transportation equipment (routes in MIMO) to perform tasks, that is, the mutual pairing process of containers ("container flow" data) and horizontal transportation equipment (routes in MIMO), so that the sum of the idle waiting times of all horizontal transportation equipment and the waiting time of quay cranes (delay) is minimized.
[0094] 2.4.2 Set the objective function:
[0095] min{qctotalwaitTime+agv_totalwaitTime}
[0096] Where qctotalwaitTime is the waiting time of the quay crane; agv_totalwaitTime is the idle waiting time of the horizontal transportation equipment.
[0097] 2.4.3 Markov decision process modeling
[0098] Abstract this process into a Markov decision process (MDP). Through this abstraction, the status of horizontal transportation equipment at an automated container terminal can be recorded and the decision-making behavior can be guided.
[0099] This Markov decision problem can be represented by a six-tuple (S, A, P, γ, R, π). S is the state space, which contains all possible states s of the horizontal transportation equipment during the entire scheduling process, used to describe the current environment and provide environmental information for the agent to make decisions; A is the action space consisting of all executable actions a composed of scheduling rules; P represents the probability of transitioning from the current state s to the next state s′ when an action a in the action space is taken; γ is the discount factor, which balances short-term and long-term rewards; R is the reward function, which returns a reward value after executing an action and is used to evaluate the quality of the action; π is the policy representing the mapping from the state set S to the action set A, and the policy π(s, a) represents the probability of taking a specific action A in the given state S.
[0100] 2.4.4 Problem assumptions
[0101] (1) All quay cranes (input interfaces), horizontal transportation equipment, and yard cranes (output interfaces) start working from time 0;
[0102] (2) All container (container flow) operation plan information is known;
[0103] (3) Each horizontal transportation equipment (line) can only transport one container at a time (the network interface channel capacity is 1);
[0104] (4) The number of yard cranes (output interfaces), quay cranes (input interfaces), and horizontal transportation equipment (lines) that work is determined at the beginning, and only the scheduling work of the horizontal transportation equipment (lines) needs to be carried out;
[0105] (5) When the horizontal transportation equipment (line) has no operation, it waits in place;
[0106] (6) Do not consider the energy consumption problem;
[0107] S3: Based on the MIMO interface and mathematical model, build a road network and physical facilities to construct a port simulation environment;
[0108] In the embodiments of the present application, after completing the mathematical model integrating the MIMO interface and mechanism, the construction of the port simulation environment includes the following steps:
[0109] 3.1 Construction of road network and physical facilities
[0110] According to the simulation software Anylogic, combined with the port MIMO operation mechanism and mathematical model in S1 and S2, first construct the road network. Based on parameters such as the scale, layout, number of lanes, traffic lights, traffic guiding lines of the port, container_height - container height; container_length - container length; container_width - container width; yard_length - yard length; yard_width - yard width; yard_num - number of yards, etc., construct the road network of the model. Secondly, construct the physical facilities. According to the actual situation, the quay crane adopts a single trolley or double trolley quay crane; the horizontal transportation equipment is ART / AGV / IGV, etc.; the yard crane is a rail-mounted crane or a rubber-tired crane, etc. Using discrete event and multi-Agent modeling, and adopting built-in modules such as the road traffic library and the process modeling library, and importing the 3D model or CAD in the software, a port model built according to the real scenario is obtained, such as Figure 4 shown.
[0111] 3.2 Construction of quay crane, yard crane, horizontal transportation equipment, ship, Q-learning agent
[0112] After completing the construction of the road network and physical facilities, it is necessary to implement functions in each type of equipment and construct each type of agent. In the Main main interface, construct the following set of agents: quaycrane_collection quay crane set - number of network interfaces; yardcrance_collection yard crane set - number of network interfaces; vehicle - collection set of all horizontal transportation equipment - network bandwidth; collection_vessel_containers set of imported containers to be transferred - total amount of data; agv_stopline set of horizontal transportation equipment stop lines; agv_parkinglot_collection: set of horizontal transportation equipment parking lots;
[0113] 3.2.1 Construction of ship and quay crane agents
[0114] The ship agent contains several parameters of the hull itself, such as ship_id, numofbay, numofrow, and numoftier, which represent the ship ID, bay position, number of rows, and number of tiers respectively, as well as the container task agents it carries. When docking, it sends arrival information to the quay crane, and the quay crane assigns tasks. The quay crane (input / output node) agent defines its behavior in the form of a flow chart (taking the ship unloading task as an example): When the ship arrives at the port, each quay crane quaycrane_collection obtains its own task information, loads the container tasks into its own task set collection, and processes them in sequence. When the task assignment function is completed, when the selected horizontal transportation equipment arrives at the quay crane target stop line agv_stopline from the parking lot agv_parkinglot_collection, the unloading function starts, and the spreader agent performs the unloading action. After the container is unloaded, the task assignment function for the next task continues, repeating until the task is completed. A collection set is established inside the agent, representing the task list; agv_collection represents the set of selected horizontal transportation equipment;
[0115] 3.2.2 Construction of the yard crane agent
[0116] The yard crane agent also defines its behavior in the form of a flow chart (taking the ship unloading task as an example): When the horizontal transportation equipment igv_collection carries the container igv_container_collection to the destination, the yard crane moves towards the container. After the spreader agent unloads the container, the horizontal transportation equipment leaves the site and waits for the arrival of the next container task, repeating in a cycle.
[0117] 3.2.3 Construction of the horizontal transportation equipment agent
[0118] The operation logic (taking the ship unloading task as an example) and some parameters are combined through modules in the Anylogic software to form an agent. When the horizontal transportation equipment is waiting for task dispatch, it records state0 - the initial state of the horizontal transportation equipment before performing the task, and then records the action action0 - the action taken, and is dispatched to the target quay crane destination-qc, the quay crane destination of the horizontal transportation equipment, picks up the container task, obtains the task information and goes to the yard crane destination-yc, the yard crane destination of the horizontal transportation equipment, to perform the unloading task. After completion, it records state1 - the state of the horizontal transportation equipment after performing the task, and finally the function records the reward - the obtained reward. After completing one task, it repeats in this cycle until the task ends. During the process, it records totalwaitTime - the waiting time between two tasks, and qcwaitTime - the waiting time at the quay crane.
[0119] 3.2.4 Q-learning Agent
[0120] It includes the qTable parameter of type HashMap<String, double[]>, which is responsible for storing the parameters to be written into the Q-table; the definitions of learningRate, discountFactor, and epsilon, as well as the get_Max_QValue function for calculating and traversing to obtain the optimal Q-value and the Training function for training.
[0121] S4: Configure the reinforcement learning algorithm in the simulation environment to obtain the optimal scheduling strategy for the port horizontal transportation equipment.
[0122] In the embodiment of the present application, the configuration of the reinforcement learning algorithm includes the following steps:
[0123] The main process is as follows: 1) The system collects environmental information and performs state input; 2) Then it enters the iterative learning of the Q-learning algorithm in the training module and queries the Q-table; 3) The action space module selects an action and then assigns tasks. 4) After completing the task, it enters the reward function module for reward calculation and Q-value update, and repeats in a cycle.
[0124] It mainly includes the construction of the state input module, the construction of the training module, the construction of the action space module, and the construction of the reward function module.
[0125] 4.1 Construction of the state input module
[0126] Construct the state space, including discretizing and combining the state information representing the system in the simulation model into an array. Improve the traditional state array to the sorting in the set of available horizontal transportation equipment, reducing the state space dimension. Its mathematical representation is as follows:
[0127]
[0128] Among them, C1 represents the set of equipment sorted by the total waiting time T i The sorted equipment set, C2 represents the set of equipment sorted by the distance d i The sorted equipment set, rank k (S)=index(sort(S,k),current): The sorting index function, qc id Represents the current quay crane identifier.
[0129] Among them, the working process of the sorting index function is as follows:
[0130] 1) If the number of vehicle - collection elements is not zero, sort vehicle - collection (in ascending order of the totalwaitTime of members); record the index position and store it in the variable timeState;
[0131] 2) At the same time, sort it in ascending order of the distance from the element to the current quay crane), record the index position and store it in the variable disState.
[0132] 3) Convert the array [qc id , timeState, disState] to a string format and return it; otherwise, create an all - zero array [0, 0, 0] and convert the array to a string format and return it.
[0133] 4.2 Construction of the training module
[0134] It includes two parts: the Training function and the get_Max_QValue function. The former calculates the Q - value, and the latter updates the Q - value. Cooperating with the ε - greedy + Boltzmann of the 4.3 action space module for exploration, the flow chart of the Q - learning algorithm is as Figure 5 shown. When there is a task requirement (there is "container flow" data to be processed in MIMO), the scheduling system first initializes the environment and the Q - table, obtains the input of the system state (network interface in MIMO), checks and traverses and uses the historical Q - table to select the optimal action in this state, selects the optimal action - scheduling rule in the action space, and outputs a scheduling rule. According to the scheduling rule, select the appropriate horizontal transportation equipment (lines in MIMO) to execute the current task ("container flow" data). When the current task ends, record the state again and calculate the reward value, and then update and maintain the Q - table until all tasks are completed, and output a Q - table. When the next training starts, the saved Q - table from the last time can be used for continuous training, and so on in a loop.
[0135] The update rule of the Training function is:
[0136]
[0137] Among them, r represents the reward for taking the current action. γ: discount factor. α represents the learning rate (usually 0 < α ≤ 1)
[0138] The get_Max_QValue function is expressed as:
[0139]
[0140] Among them, s ∈ S: current state, A = {0, 1, 2} represents the state space, and Q(s, a) represents the state - value function.
[0141] More detailed:
[0142]
[0143] At the beginning of operation, the initialization formula of the Q table is:
[0144]
[0145] (r i ~Uniform(0,1))
[0146] The convergence condition of its operation is expressed as:
[0147] 4.3 Construction of the action space module
[0148] It includes the selection and execution of actions by ε_temperature-action and act-action. The formed scheduling mapping rule is: A = {N-CAR = 1; F-CAR = 2; WL-CAR = 3}, that is, the scheduling rule library is expressed as: including actions such as "nearest distance first", "farthest distance first", and "longest waiting time".
[0149] The former dynamically balances exploration and exploitation through a dual strategy (ε-greedy + Boltzmann). It includes: 1) The ε-greedy strategy directly makes random selections; 2) Boltzmann soft selection, calculates the temperature parameter, and calculates the probability distribution according to the Q value and temperature. 3) Force exploration when monitoring that the Q value is all zero.
[0150] Its mathematical definition is as follows:
[0151]
[0152] Among them, ε represents the exploration rate.
[0153] The mathematical definition of the Boltzmann probability distribution is:
[0154]
[0155] Among them, represents the dynamic temperature (N is the number of training rounds).
[0156] The zero-value protection mechanism is:
[0157]
[0158] The action index is output by the ε_temperature-action function, and the act-action matches the corresponding scheduling rule to output the horizontal transportation equipment. For the functions corresponding to the three action indexes in act-action, they are as follows:
[0159] Case0: Nearest distance first:
[0160] Directly select the equipment closest to the target from all available horizontal transportation equipment sets C Expressed as:
[0161]
[0162] Update rule:
[0163]
[0164] v * .idle = 0, v * .work_if = 1
[0165] Case1: Furthest distance first:
[0166] Select the equipment furthest from the target from all available horizontal transportation equipment sets C Expressed as:
[0167]
[0168] The update rule is the same as Case0.
[0169] Case2: Longest waiting time first:
[0170] Define a candidate set C andidate Expressed as:
[0171]
[0172] Select the equipment closest to the target Expressed as:
[0173]
[0174] The update rule is the same as Case0.
[0175] Among them, C = {v1, v2... v n} represents the horizontal transportation equipment set, d i = v i .distanceTo(this) represents the distance from equipment v i to the target, T i = v i.totalwaitTime represents the total waiting time of equipment v i , t represents the current time (time()), represents the last task completion time of equipment v i , d represents the identifier as 1 after the equipment completes a task, and idle represents the idle state work i f: Whether it is working.
[0176] The priority selection process of the three action indexes is as follows:
[0177] The first step: Check whether there is equipment that meets the longest waiting time first strategy (Case 2). If there is eligible equipment, select the equipment closest to the target among them.
[0178] The second step: If no eligible equipment is found, switch to the closest distance first strategy (Case0) and select the equipment closest to the target.
[0179] The third step: If there are still special cases or further requirements, the farthest distance first strategy (Case1) can be selected.
[0180] It should be noted that this order design is to ensure efficient scheduling while taking into account the balanced utilization of resources, the stability of the system, and user satisfaction. By preferentially processing the equipment with the longest waiting time, system bottlenecks can be effectively reduced and overall efficiency can be improved; the closest distance first strategy ensures that tasks can be completed quickly; the farthest distance first strategy serves as a supplementary solution to provide additional flexibility in specific scenarios. This hierarchical strategy selection mechanism can better adapt to the complex port operation environment.
[0181] 4.4 Construction of the reward function module
[0182] The present invention integrates the system-level objectives of MIMO, designs a hierarchical composite reward function, integrates the waiting time costs of multiple devices, system efficiency, and resource requirements, and prevents falling into local optima. At the same time, the sub-rewards dynamically adjust the weights through the MIMO coupling relationship (such as increasing α when the propagation is concentrated at the port). The reward function designed by the present invention is as follows:
[0183] r t = (-1) * [α * (T qc ) + β * T + γ * V]
[0184] Among them, r t is the immediate reward value, T qc is the waiting time increment of the horizontal transportation equipment, T is the idle waiting time of the quay crane, V is the system efficiency, and α, β, and γ are weight coefficients.
[0185] 1) Time parameter update rule
[0186] Update when the task is completed:
[0187] Calculation of quay crane waiting time: T qc = t - t qc_start → igv.qctotalwaitTime
[0188] 2) Reward calculation and learning update
[0189]
[0190] Among them, D represents the set of recorded rewards.
[0191] This embodiment also provides a real-time scheduling system for port horizontal transportation equipment based on the MIMO mechanism, including:
[0192] An encapsulation module, used to obtain the operation parameters in the terminal operation environment and equipment, and encapsulate the port loading and unloading operations and horizontal transportation operations into a MIMO interface;
[0193] A construction module, used to abstract the operation parameters in the terminal operation environment into network communication parameters and construct a MIMO interface;
[0194] A modeling module, used to establish the operation mechanism and its mathematical model of port operations based on the MIMO interface by defining the input and output interfaces and the objective function;
[0195] A simulation module, used to build a port simulation environment by constructing a road network and physical facilities based on the MIMO interface and the mathematical model;
[0196] A configuration module, used to configure the reinforcement learning algorithm in the simulation environment to obtain the optimal scheduling strategy of the port horizontal transportation equipment.
[0197] Furthermore, it further includes:
[0198] A memory, used to store programs;
[0199] A processor, used to load the program to execute the real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism.
[0200] This embodiment also provides a computer-readable storage medium, which stores a program, and when the program is executed by a processor, it implements the real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism.
[0201] The storage medium proposed in this embodiment and the real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism proposed in the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0202] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disc of a computer, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present invention.
[0203] Embodiment 2
[0204] This is an embodiment of the present invention, which provides a real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through specific implementation methods and implementation effects.
[0205] The specific embodiments are as follows:
[0206] Experimental parameters: A total of two quay cranes, six yard cranes, and six horizontal transportation equipment were used in this experiment. discountFactor = 0.9; learningRate = 0.1; epsilon = 0.9.
[0207] Experimental step 1: Designed 40 container tasks. The task list was input into the model in the form of excel, and part of the information is shown in Table 1:
[0208] Table 1 Task List
[0209] container_id current_block current_bay current_row current_tier block bay row tier 1 1 0 0 0 0 5 3 0 2 1 0 2 0 0 5 4 0 3 1 0 1 0 0 0 5 0 4 1 0 3 0 0 0 5 1 5 1 1 3 0 1 4 5 0 6 1 1 4 0 1 5 3 0 7 1 1 0 0 1 7 0 0 8 1 1 0 1 1 7 3 0
[0210] Among them, each column of data represents the container ID, the block area on the cargo ship, the bay, the row number, the layer number, and the block area, bay, row number, and layer number of the target position respectively.
[0211] Experimental step 2: Click the start button to enter the initial interface.
[0212] Experimental Step 3: Set the initial number of yard cranes, initial number of quay cranes, initial number of horizontal transportation equipment, and initial operating speed.
[0213] Experimental Step 4: Conduct multiple iterations using the Monte Carlo experiment in the simulation software. As Figure 6 shown is the reward curve obtained after 1000 training iterations, and the waiting time converges near 8500s.
[0214] Experimental Step 5: After the iteration ends, collect and plot the output data and reward curve graph.
[0215] After testing, compared with the initial operation, the average waiting time of the horizontal transportation equipment and the waiting time of the quay crane are reduced by about 4000s in total, the empty driving mileage of the horizontal transportation equipment is reduced by 20%, and the overall energy consumption is reduced by 18%. When simulating sudden tasks, the system response time is normal, meeting the processing ability for uncertain events.
[0216] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism, characterized in that, Including: Obtain the operation parameters in the terminal operation environment and equipment, and encapsulate the port loading and unloading operations and horizontal transportation operations into a MIMO interface; Based on the MIMO interface, establish the operation mechanism and its mathematical model of port operations by defining the input and output interfaces and the objective function; Based on the MIMO interface and the mathematical model, build a road network and physical facilities to construct a port simulation environment; Configure the reinforcement learning algorithm in the simulation environment to obtain the optimal scheduling strategy for port horizontal transportation equipment.
2. The real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism according to claim 1, characterized in that: The establishment of the operation mechanism and its mathematical model of port operations by defining the input and output interfaces includes the following steps: Define multiple input interfaces according to different container sources and quay cranes for operation, and define multiple output interfaces according to different destinations; Design the operation mechanism process, including task assignment, path planning, and loading and unloading operations when the ship arrives at the port; Convert the operation mechanism into a mathematical model and set the objective function; Set the objective function as the sum of minimizing the waiting time of quay cranes and the idle waiting time of horizontal transportation equipment.
3. The real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism according to claim 1 or 2, characterized in that: The configuration of the reinforcement learning algorithm in the simulation environment includes the following steps: Sort and index the state information in the set of horizontal transportation equipment through a sorting index function to generate a state array containing time and distance; Input the state array into the Q-learning algorithm, update the Q(s,a) value in the Q table through the Training function, and obtain the maximum Q value of all possible actions under a given state through the get_Max_QValue function; Based on the Q value, combined with multiple rule bases, select the optimal action through the ε_temperature-action function and the act-action function; Based on the optimal action, design a hierarchical composite reward function and dynamically adjust the weight according to the MIMO coupling relationship.
4. The real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism according to claim 3, wherein: The act-action function includes setting a priority strategy; The setting of the priority strategy includes: Give priority to selecting the equipment with the longest waiting time; If there is no equipment with the longest waiting time, select the equipment closest to the target; If there is no equipment with the longest waiting time and the equipment closest to the target, select the equipment farthest from the target; The equipment that selects the longest waiting time It is expressed as: Select the equipment closest to the target It is expressed as: The equipment farthest from the target Expressed as: Among them, C represents the set of horizontal transportation equipment, and v j represents the equipment in the candidate equipment set C candidate , and v i represents the equipment from all available equipment sets C, and d j represents the distance from the equipment v j to the target, and d i represents the distance from the equipment v i to the target.
5. The real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism according to claim 4, wherein: The generation of a state array containing time and distance through sorting and indexing processing by the sorting index function is expressed as: Among them, qc id represents the identifier of the current quay crane, C represents the set of horizontal transportation equipment, |C| represents the number of available transportation equipment in set C, and rank T (C1) represents the index array obtained by sorting set C according to the total waiting time, and rank d (C2) represents the index array obtained by sorting set C according to the distance to the current quay crane.
6. The real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism according to claim 5, characterized in that: The update of the Q(s,a) value in the Q table by the Training function is expressed as: Where, r represents the reward for taking the current action, γ represents the discount factor, α represents the learning rate, and r(s,a) is the immediate reward obtained after taking action a in state s.
7. The real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism according to claim 6, characterized in that: The hierarchical composite reward function is expressed as: r t = (-1) * [α * (T qc ) + β * T + γ * V] Among them, r t is the immediate reward value, T qc is the waiting time increment of the horizontal transportation equipment, T is the idle waiting time of the quay crane, V is the system efficiency, and α, β, and γ are the weight coefficients.
8. A system for the real-time scheduling method of port horizontal transportation equipment based on the MIMO mechanism according to claim 1, characterized in that: An encapsulation module, configured to obtain the operation parameters in the terminal operation environment and equipment, and encapsulate the port loading and unloading operations and horizontal transportation operations into a MIMO interface; A construction module, configured to abstract the operation parameters in the terminal operation environment into network communication parameters and construct a MIMO interface; A modeling module, configured to establish an operation mechanism and its mathematical model of port operations based on the MIMO interface by defining input and output interfaces and objective functions; A simulation module, configured to build a port simulation environment by constructing a road network and physical facilities based on the MIMO interface and the mathematical model; A configuration module, configured to configure a reinforcement learning algorithm in the simulation environment to obtain an optimal scheduling strategy for port horizontal transportation equipment.
9. A computing device, characterized in that, Comprising: A memory, configured to store programs; A processor, configured to load the programs to execute the steps of the real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism according to any one of claims 1-7.
10. A computer-readable storage medium storing a program, characterized in that, When the programs are executed by the processor, the steps of the real-time scheduling method for port horizontal transportation equipment based on the MIMO mechanism according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Automatic quay cooperative scheduling method based on deep reinforcement learning
CN113988443A
Port operation area simulation modeling method based on multiple agents
CN116861649A
Container terminal automatic scheduling method and system based on reinforcement learning
CN119047794A
Navigation center area coordination optimization method and system based on big data
CN119578847A
Intelligent horizontal transportation system and method for automatic side-loading / unloading container tarminal
US20230072997A1