A multi-target dynamic and static combined port ship cooperative scheduling method
Patent Information
- Application Number
- CN202610993753.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-07-06
AI Technical Summary
[0025]现有技术在应对高动态、强随机性的港口运营环境时,主要存在以下三方面的缺陷,导致港口轮渡系统难以实现高效、稳定且鲁棒的整体运营:
第一,实现了全局资源统筹与局部灵活调优的有机统一,解决了单一时间尺度导致的资源配置失衡问题。
Smart Images

Figure CN122509634B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of intelligent logistics scheduling in ports and intelligent management and control of water traffic, specifically to ship scheduling technology based on deep reinforcement learning and robust optimization, and in particular to a multi-objective dynamic and static combined port ship collaborative scheduling method. Background Technology
[0002] The core of a port's operational performance lies in its overall throughput capacity, and the key infrastructure supporting this capacity is the Ro-Ro ferry system between ports. The optimization level of the ferry system's scheduling directly determines whether the port's core functions can operate smoothly and efficiently, and is the lifeline for ensuring the efficient flow of personnel and goods.
[0003] However, ferry scheduling at ports faces multiple challenges stemming from environmental randomness and inherent operational complexity: First, the randomness of the external environment manifests in the highly nonlinear fluctuations of passenger and vehicle traffic over time. Beyond regular peaks and troughs, unforeseen events often lead to drastic discrepancies between real-time demand and forecasts, making it difficult for existing scheduling to flexibly match traffic fluctuations. Second, the complex coupling of internal operations is evident in the multi-stage interactions involved in port production processes, with some non-standardized operations weakening system stability. More critically, existing scheduling still relies on the personal experience of dispatchers. This reliance on expert judgment is inherently uncertain, difficult to standardize or quantify, resulting in a lack of replicability and quantifiability in scheduling strategies. This makes them highly susceptible to suboptimal solutions due to experience bias or operational fluctuations, potentially leading to systemic delays. Finally, there is a lack of decision support tools, particularly systematic, data-driven analytical tools.
[0004] Faced with high-dimensional scheduling scenarios, existing management methods are insufficient for automated assessment and intelligent optimization of the ferry system, resulting in decisions often lagging behind changes in the on-site situation. Therefore, the current port area's "expert experience-driven passive scheduling" model cannot solve the above challenges, leading to a lack of robustness in scheduling and low port operational efficiency.
[0005] Existing technologies for port ferry scheduling include the following categories: 1. Based on different traffic perception methods: passive adjustment based on deterministic benchmarks and active response scheduling methods.
[0006] These methods primarily address the issue of scheduling's perception and response to traffic fluctuations, and typically employ the following steps: Step 1 (Baseline Planning Phase): Establish a deterministic model based on Mixed Integer Programming (MIP) and formulate an initial baseline scheduling plan based on given passenger and freight flow parameters (such as the three-stage heuristic algorithm proposed by Hansen JR et al.; the minimization of demurrage allocation scheme by Mangala A et al.).
[0007] Step 2 (Deviation Response): During execution, when the actual operational status (such as ship delays or changes in demand) deviates from the baseline plan, a rescheduling mechanism is triggered. If it is a proactive response optimization, the worst-case uncertainty will be predicted in advance at this stage using predictive methods (such as Zhao R, a robust optimization for service time uncertainty) or by embedding robust redundancy space to allow sufficient adjustment space.
[0008] Step 3 (Scheme Revision): Based on the established optimization model, real-time adjustments are made using purely manual experience (schedulers). If an active response approach is adopted, new feasible scheduling schemes can be calculated and generated based on new constraints (such as the MILP rescheduling model proposed by Kim K et al.); or, robust optimization methods can be used to directly embed buffer time or redundancy space during the planning stage (such as Wen X et al. using queuing theory to predict congestion and allocate buffer time).
[0009] 2. Based on different time window scales: This type of method is divided into two implementation methods: long-term static planning or short-term real-time planning methods under a single time scale.
[0010] Implementation Method 1 (Long-Term Static Programming): Step 1: Collect long-term historical traffic data or macroeconomic statistics.
[0011] Step 2: Establish a global optimization model covering a long time domain (such as 24 hours or longer) (such as the two-stage robust optimization model used by Xu L et al., and the two-level programming model for containers used by Chen K et al.).
[0012] Step 3: Solve using mathematical programming or statistical methods to minimize long-term operating costs or design route networks, and output a fixed static schedule.
[0013] Implementation Method Two (Short-Term Real-Time Planning): Step 1: Monitor real-time system status and dynamic requests within a short period of time.
[0014] Step 2: Set a shorter time step or a rolling time window (e.g., Zheng H et al. applied the rolling time strategy to AGV scheduling; Yuan P et al. proposed a periodic combined with event-driven method).
[0015] Step 3: Within this narrow window, make quick decisions and corrections based on the current state, and output immediate scheduling instructions (such as online adjustments made by Wang Q and others using reinforcement learning).
[0016] 3. Based on different solution methods: solution methods based on operations research, heuristic algorithms, or single reinforcement learning.
[0017] These methods primarily employ different solution approaches based on a balance between the efficiency and optimality of the scheduling model, and typically include the following steps: Step 1 (Modeling): Construct the port scheduling problem as a mixed integer linear programming (MILP) model (such as the research of Qu S et al., Tan Z et al.), or as a multi-objective mathematical model (such as the model of Wen X et al.).
[0018] Step 2 (Solve): Operations research method: directly calls the solver for exact solution, suitable for scenarios with clear rules but few variables.
[0019] Heuristic approach: Design biomimetic or evolutionary algorithms (such as Eldemir F et al.'s ant colony algorithm, Wen X et al.'s gray wolf optimizer, and NSGA-II) to obtain approximate optimal solutions through iterative search.
[0020] Reinforcement learning method: Model dynamic scheduling as a Markov decision process, and train the agent through Q-learning or DQN (deep Q-learning network) algorithms to output actions according to the state.
[0021] 4. Based on different fields: two-level planning methods in the field of aircraft scheduling (such as the two-level configuration method of aircraft flight slots proposed by Lü Jialin et al., and the research progress of airport flight slot resource management in the review by Wang Yanjun et al.).
[0022] Step 1 (Modeling): The aircraft flight slot resource allocation problem is constructed as a bi-objective mathematical model (such as a mixed-integer programming model) with multi-level constraints (e.g., capacity constraints, slot pool constraints, fairness constraints). Specifically, a "hierarchical scheduling strategy" is adopted to construct the model structure. The upper layer represents the slot resource allocation decisions of the airport / management, and the lower layer represents the reactions of airlines or passengers to different slot options (reflected by indicators such as "slot value" or "slot offset"). At the same time, a tolerance parameter is introduced, and the evaluation criteria for slot offset are redefined to balance fairness among different airlines and overall operational efficiency.
[0023] Step 2 (Solve): Multi-objective operations research methods: For configuration models at different levels such as single airport, airport network or airport group, the solver can be directly called to perform exact solutions for integer / mixed integer programming, or multi-objective programming methods can be used to seek Pareto optimal solutions that are both efficient and fair.
[0024] Heuristic approach: For the large-scale time allocation problem of high-dimensional variables in complex networks, multi-objective heuristic algorithms (such as evolutionary algorithms) are designed to obtain an approximately optimal time allocation scheme that takes into account the interests and constraints of multiple parties through iterative search.
[0025] Existing technologies suffer from three main shortcomings when dealing with the highly dynamic and random port operating environment, making it difficult for port ferry systems to achieve efficient, stable, and robust overall operation: 1. For existing scheduling methods based on deterministic benchmarks and passive adjustments or purely robust optimization: the passive scheduling model driven by "human experience" lacks foresight and stability, leading to delayed response and decision-making uncertainty. Current port scheduling heavily relies on the personal experience and mental state of dispatchers; this unstructured experience-based judgment is difficult to standardize or quantify. This model is inherently static and passive, often only making delayed emergency adjustments after system deviations (such as congestion) occur, failing to intervene before sudden changes in traffic flow. This mechanism, relying on subjective human judgment, is highly susceptible to suboptimal solutions due to experience biases or state fluctuations. Meanwhile, while purely robust optimization provides some robust redundancy to allow for adjustments, the lack of a sound static scheduling benchmark often results in overly conservative plans or suboptimal overall performance.
[0026] 2. For long-term static planning or short-term real-time planning methods at a single time scale: Single-time-scale scheduling models struggle to balance the conflict between "multi-department resource collaboration" and "real-time traffic adaptation," leading to resource allocation imbalances. (The most significant drawback) Port operations involve multi-stage interactions and multi-departmental collaboration, requiring long-term (e.g., 24-hour) planning, as well as rigid demands for information exchange, resource pre-coordination, and scheme approval. Meanwhile, highly volatile real-time passenger and cargo flows demand flexible short-term (e.g., hourly) responses. Existing single-scale scheduling faces a dilemma: while long-term static scheduling ensures plan certainty, it cannot flexibly match real-time traffic, leading to wasted capacity or congestion; while short-term dynamic adjustments offer local flexibility, they lack a holistic view of 24-hour operations, easily falling into the "local optimum trap," causing chain reactions of delays and idle capacity in subsequent periods, failing to guarantee "globally optimal" overall planning quality. Existing technologies cannot effectively combine long-term global planning with short-term dynamic optimization, lacking a systematic solution that can both establish 24-hour operational benchmarks and provide real-time fine-tuning to mitigate fluctuations.
[0027] 3. For solution methods based on operations research, heuristic algorithms, or single reinforcement learning: operations research methods have high computational complexity and lack the ability to perceive and learn from dynamic environments; heuristic methods lack optimality guarantees and are prone to getting trapped in local optima; although emerging reinforcement learning methods have adaptability, most of them lack interaction mechanisms with static baseline plans, and they are insufficient in modeling when dealing with complex rule constraints in ports (such as navigation restrictions and multi-department collaboration), making it difficult to make fine-tuning with global optimal awareness.
[0028] 4. Although existing air traffic control systems also employ hierarchical or two-layer planning frameworks, their essence is that the upper layer allocates slots macroscopically, while the lower layer reflects the economic response and tolerance of airlines and passengers. Their technical constraints are mainly limited to macroscopic capacity thresholds (such as slot pools and hourly flight limits) and fairness indicators, and the solution relies heavily on traditional operations research or heuristic rescheduling. Summary of the Invention
[0029] To address the multiple shortcomings of existing technologies, such as passive response lag, difficulty in balancing global and local interests on a single time scale, and limitations of a single solution method, the present invention aims to provide a multi-objective dynamic and static combined port vessel collaborative scheduling method. This method solves the existing technical problems through a two-layer coupled architecture that deeply integrates "long-cycle global planning" and "short-cycle dynamic optimization", deep learning and reinforcement learning technologies, and a hybrid solution strategy that integrates operations research optimization and deep reinforcement learning.
[0030] Technical solution A multi-objective, dynamic-static combined port vessel collaborative scheduling method includes the following steps: S1, Preparation Stage; Initial vessel status data is retrieved from the port status database, including the number of vessels of each type at each port at the initial moment, as well as the positions and estimated arrival times of vessels en route at the initial moment. Historical vehicle and passenger traffic data from the previous 48 hours is extracted and input into the Transformer model to predict vehicle and passenger traffic for the next 24 hours, yielding a long-term prediction result for the next 24 hours. , express time Forecasting transportation demand in different directions.
[0031] S2. Static shift schedule generation; A segmented robust global static scheduling model is constructed. Inputting long-term passenger and vehicle traffic forecasts for the next 24 hours, a two-stage robust optimization theory is employed. In the first stage, the fleet size is determined through robust optimization, robust total demand is calculated, and an objective function is established to minimize unmet demand and fleet adjustment costs, outputting the optimal fleet size for each ship type. In the second stage, refined scheduling is performed under the fleet size constraints determined in the first stage. A multi-objective function is established to minimize the total number of operational departures, ship distribution deviations, and unmet demand. Constraints include coupling constraints from the first stage, demand constraints, departure frequency upper limits, inventory deviation constraints, ship dynamic inventory state transition equation constraints, and departure availability constraints. The decision variables—the number of departures for each ship type and course in each time period—are solved through linear programming, outputting a baseline scheduling plan.
[0032] S3, Dynamic Scheduling Optimization; Dynamic scheduling optimization is triggered after the planned execution interval. A rolling short-cycle dynamic correction mechanism based on deep reinforcement learning is adopted. The LA-Dueling-DQN network architecture is used to target the real-time port status, including the number of currently available ships, the distribution of ships en route, the number of vehicles stranded, and the short-term predicted flow for the next W hours based on the actual flow of the day. The deep reinforcement learning agent outputs adjustment actions and generates a corrected scheduling plan. S4. Utilize scheduling information for dispatching; The generated revised shift schedule is parsed into execution work orders and sent to the ship terminal and port operation system. Actual operation data is collected to calculate reward values and stored in the experience playback pool for offline model training. The execution results of the revised shift schedule are fed back to the static shift schedule generation step of the next cycle, forming a closed-loop optimization.
[0033] Beneficial effects First, it achieves an organic unity between overall resource coordination and flexible local optimization, solving the problem of resource allocation imbalance caused by a single time scale.
[0034] This invention employs a two-tiered architecture of "dynamic and static coordination." It locks in the optimal global resource allocation through long-term planning for the next 24 hours, avoiding the blindness of short-term adjustments; simultaneously, it addresses real-time fluctuations through hourly dynamic optimization, preventing the rigidity of long-term plans. This architecture ensures both the stability of plans required for multi-department collaboration and the flexibility to handle unforeseen circumstances, thereby maximizing overall throughput while effectively preventing idle capacity or localized congestion.
[0035] Second, it has the ability to proactively predict and respond in real time, solving the problems of delayed response and uncertain decision-making in existing technologies.
[0036] This invention introduces a dynamic correction mechanism based on deep reinforcement learning, combining Bi-LSTM (Bidirectional Long Short-Term Memory) and MHA (Multi-Head Attention) techniques to capture and analyze high-dimensional temporal characteristics of traffic flow in real time. This allows the system to move beyond relying on the scheduler's subjective experience or passively waiting for deviations to occur, and instead proactively perceive traffic trends and intervene in scheduling. Therefore, this invention significantly improves the foresight, responsiveness, and objective stability of the scheduling system's decisions.
[0037] Third, it balances solution efficiency and global optimality, overcoming the limitations of a single solution method.
[0038] This invention employs a hybrid solution strategy: at the static level, it utilizes a segmented robust optimization strategy by decoupling fleet size and scheduling order to ensure the global optimality and solution efficiency of the baseline plan; at the dynamic level, it leverages the self-learning capability of reinforcement learning to handle complex nonlinear dynamic constraints. This combination fully utilizes the advantages of different algorithms, enabling the system to provide rapid decisions while ensuring high-quality and robust decisions when dealing with large-scale scheduling problems.
[0039] Fourth, experimental results show that the method of this invention significantly outperforms existing technologies in two core indicators: planning accuracy and average daily unmet needs. Using real operational records from three consecutive months of 2024 for Xinhai Port and Xuwen Port as the test dataset, and comparing it with existing methods such as multi-stage stochastic programming models, continuous approximation models, maximum cardinality bipartite graph matching algorithms, and capacity fluctuation factor elastic scheduling models, the planning accuracy of the method of this invention is 92.5%, a 16.03% improvement compared to the best-performing multi-stage stochastic programming model, and more than 10 times the improvement compared to the maximum cardinality bipartite graph matching algorithm; the average daily unmet needs is 1.2 hours, a 44.24% reduction compared to the multi-stage stochastic programming model, and an 83.75% reduction compared to the maximum cardinality bipartite graph matching algorithm, proving that the method of this invention combines dynamic and static data... The synergistic effect of technologies such as the two-layer architecture, Transformer prediction module, segmented robust optimization two-stage model, LA-Dueling-DQN reinforcement learning, and ship selection and plan correction mechanism has achieved unexpected technical results. The scheduling plan has significantly improved its ability to meet the actual demand for vehicles and passengers, significantly reduced the cumulative time that the system failed to process transportation demand in a timely manner within 24 hours, and significantly enhanced the congestion relief capability. Compared with traditional methods that only focus on static scheduling or short-cycle adjustments, the method of this invention has achieved an organic unity of global resource coordination and local flexible optimization, and solved the systemic problem of imbalance in single-scale scheduling resource allocation. Attached Figure Description
[0040] Figure 1 This is a flowchart illustrating the multi-objective dynamic-static combined port vessel collaborative scheduling method proposed in this invention. Figure 2 This is a schematic diagram of the algorithm flow of the segmented robust global static scheduling model with multi-objective constraints proposed in this invention; Figure 3 This is a neural network architecture design diagram of the deep reinforcement learning method proposed in this invention. Figure 4 This is a schematic diagram of the algorithm flow for the rolling short-cycle dynamic correction mechanism based on deep reinforcement learning proposed in this invention. Figure 5 This is a comparison chart of the planning accuracy between the present invention and existing methods; Figure 6 This is a comparison chart of the average unmet duration per day between the present invention and existing methods. Detailed Implementation
[0041] The technical solution provided in this application will be further described below with reference to specific embodiments and accompanying drawings. The advantages and features of this application will become clearer from the following description.
[0042] Example 1 This embodiment provides a multi-objective dynamic and static combined port vessel collaborative scheduling method, which includes four core steps: preparation stage S1, static schedule generation S2, dynamic schedule optimization S3, and scheduling using schedule information S4.
[0043] A two-layer coupled architecture combining static and dynamic elements is adopted. The static planning layer uses a segmented robust global static scheduling model to generate a baseline scheduling plan based on the long-term predicted traffic for the next 24 hours. The dynamic correction layer adopts a rolling short-cycle dynamic correction mechanism based on deep reinforcement learning, which is triggered once every Δt=3 hours to fine-tune the scheduling for the next 4 hours based on the real-time status and the short-term prediction for the next W=3 hours.
[0044] The static and dynamic layers are tightly coupled through benchmark modification and state feedback. The static planning layer provides a highly robust benchmark plan for the dynamic correction layer; the dynamic correction layer, through continuous real-time adjustments, corrects deviations in the actual execution of the benchmark plan and feeds the corrected results back to the decision-making of the next cycle, forming a continuously optimizing closed-loop control system.
[0045] like Figure 1 The method of the present invention includes the following steps: S1, Preparation Phase; This step first initializes the environment and algorithm by reading the initial state data of the ships from the port's actual state database, including: the port at the initial time. Type Number of ships Information on ships en route at the initial time , This represents the number of ships of type k that are expected to arrive at port p in the next hour t. Subsequently, historical vehicle and passenger flow data from the previous 48 hours is extracted and input into a pre-built Transformer model to predict vehicle and passenger flow for the next 24 hours, thus obtaining a long-term vehicle and passenger flow prediction result for the next 24 hours. , express time Forecasting transportation demand in different directions.
[0046] The system is typically initialized at midnight each day, and the schedule for the following day is then created.
[0047] The core (hyper) parameters of the Transformer model are shown in Table 1. This model is also used in the reinforcement learning model in step S3 to provide prediction results for the future short period (W=3) hours.
[0048] Table 1 Transformer Module (Super) Parameters S2, static shift schedule generation; Construct a segmented robust global static scheduling model, taking into account long-term vehicle and passenger traffic flow over a 24-hour period in the coming day. The system analyzes port vessels and outputs a vessel schedule for each hour (X in, Y out) over the next 24 hours. It employs a two-stage robust optimization theory, planning the fleet by addressing vessel capacity redundancy, and then using linear programming to find the global optimum, ensuring a balance between global robustness and operating costs. Finally, it outputs a baseline scheduling plan.
[0049] Specifically as follows: S21, Construct a segmented robust global static scheduling model; This study focuses on the scheduling decision-making problem for two-way roll-on / roll-off ferries between ports A and B. The span is... Within a 24-hour planning cycle, the scheduling system needs to make decisions for each discrete time step in an environment where passenger flow forecasting is uncertain. and each heading different types of ships The number of ships dispatched.
[0050] For the objective function of optimization: First, the system needs to maintain a high level of service guarantee to meet the needs of vehicles and passengers crossing the sea to the greatest extent and reduce the length of stay; second, it needs to reduce fuel consumption and human resource input by optimizing the frequency of departures, so as to achieve refined control of operating costs; finally, the model also needs to take into account the stability of the long-term operation of the system, that is, by monitoring the dynamic balance of the ship inventory in each port, to prevent the asymmetric backlog of capacity in a single port, thereby ensuring the continuous operation of two-way routes.
[0051] Key Assumptions: To ensure the computational feasibility and logical completeness of the mathematical model, the following key assumptions are set.
[0052] First, a discrete-time modeling method is adopted, with discrete hourly units as the basic cycle for flight departure decisions.
[0053] Secondly, considering the relative stability of the navigation environment, the one-way sailing time between the two ports is assumed to be... For a fixed constant The unit is hours, and it is determined based on the actual port operation.
[0054] Furthermore, regarding shipping capacity resources, the model follows the assumption of ship homogeneity, meaning that ships of the same type are considered to be homogeneous. The vessels are completely identical in terms of carrying capacity, speed, and operating costs.
[0055] Finally, to address the uncertainty on the demand side, a demand constraint mechanism based on piecewise robustness theory is introduced. This is achieved by setting robust control coefficients. By constructing a demand set that can withstand prediction noise and using the maximum prediction deviation δ, the scheduling plan can be enhanced to withstand risks under extreme fluctuation scenarios. At the initial moment, there are ships already en route. These ships will arrive at the port on the other side of the shore in the next few hours and need to be included as known inputs in the inventory state transition equation.
[0056] Input and Output: In the model execution process, the input set covers the planning cycle. Forecast demand for each time period The initial system capacity distribution (including vessels berthed in ports and those on shipping routes), and vessel physical attribute parameters, including the single-ship carrying capacity of each vessel type. Port operation capacity limits include the maximum allowed frequency of departures per departure time cycle. Global fleet size limit and robust control parameters .
[0057] By solving the segmented robust global static scheduling model, the system generates decision variables. The time periods were clearly defined. Each voyage Specific ship types The model also outputs the number of ships dispatched. Furthermore, it synchronously outputs the unmet demand at each node and the dynamic inventory trajectory of ships, providing the scheduling layer with comprehensive system status monitoring indicators.
[0058] S22, the flowchart of the segmented robust global static scheduling model algorithm is as follows: Figure 2 As shown, the solution is divided into two stages: the first stage determines the size of a robust fleet with safety redundancy at a macro level, setting a safety ceiling budget for the overall capacity; the second stage performs refined scheduling under this budget constraint. This decoupling strategy ensures the system's ability to withstand extreme demand fluctuations while reserving flexible capacity space for the dynamic correction layer.
[0059] The definitions of each set and parameter in the model are shown in Table 2: Table 2 Model Set and Parameter Definitions S221, Phase 1: Robust Fleet Size Decisions; The core objective of this phase is to determine a fleet configuration scheme with "safety redundancy." Traditional methods schedule fleets directly based on predicted demand, resulting in a lack of capacity reserves when actual demand exceeds predictions. This invention creatively proposes to first determine a fleet size capable of withstanding worst-case scenarios through segmented robustness theory, thereby ensuring the robustness of capacity supply from the source.
[0060] Calculate robust total demand: For each heading d, calculate the robust total demand based on the sum of long-term forecasts for that heading over 24 hours, plus a robust adjustment term. in, Let Γ represent the total daily forecast demand for heading d, and let Γ·δ be the robust redundancy term. The larger Γ is, the more conservative the model is in terms of demand fluctuations, and the more capacity redundancy is reserved; δ represents the maximum possible deviation of demand in a single hour.
[0061] Calculate the initial total fleet size: For each ship type k, count the number of all ships in the current system, including ships berthed at each port and ships en route: in, Port at the initial moment Type at (A or B) The number of ships, This indicates that the ship will arrive at the port in the t-th hour. The number of ships of type k (A or B). In this embodiment, the ship type... .
[0062] Define the decision variables, objective function, and constraints for the first stage.
[0063] The decision variable is the Pareto optimal solution for fleet size adjustment under extreme uncertainty scenarios. As follows: : The optimal fleet size for ship type k (i.e., the total number of ships of ship type k that the system should be configured with). Represents the field of non-negative integers; Unmet robust total requirements for heading d (slack variables); Represents the field of nonnegative real numbers; : The size adjustment deviation of ship type k, that is, the absolute value of the difference between the optimal size and the initial size.
[0064] objective function The aim is to balance service availability with asset adjustment costs, as shown in the following formula: The objective function consists of two parts: the first term The weight is the penalty for failing to meet the robust total requirement. First, ensure that the model prioritizes coverage that meets robustness requirements; the second item Adjusting deviation penalties for the fleet, weighting This design avoids excessive deviations between the fleet size and the existing size, thereby reducing vessel allocation costs. This high-penalty, low-bias weighting allows the model to maintain the stability of the existing fleet structure while prioritizing demand coverage.
[0065] in, It is not calculated directly through an explicit formula, but rather as a decision variable in the linear programming model, automatically obtained by the solver during the optimization process. Its value is determined by both the objective function and the constraints, and the equivalent calculation formula is as follows: The constraints include: (1) Overall fleet size constraints: The constraint limits the total fleet size of the system to [number]. This corresponds to the upper limit of the total number of ships in actual operation.
[0066] (2) Fleet adjustment deviation constraints: The two inequalities mentioned above are combined to achieve The linearized expression ensures The value is the absolute value of the difference between the optimal size and the initial size.
[0067] (3) Robust demand coverage constraint: The size of each ship type The optimal fleet size for ship type k (i.e., the total number of ships of ship type k that the system should be equipped with) is denoted as k. The carrying capacity of a single ship and the maximum number of round trips per day are: ( ), For the calculation above The value along the heading d.
[0068] This constraint requires the fleet to provide a maximum daily gross capacity (i.e.) It must be able to cover the vast majority of the robust requirements set. As slack variables, they allow a small number of requirements to be uncovered in extreme cases, but are penalized with high weight in the objective function.
[0069] Solving the first-stage model to obtain the robust fleet size: The mixed-integer programming solver (MIP) is used to solve the above model to obtain the optimal fleet size for each ship type. .Should This will be passed on as a resource ceiling constraint to the second phase.
[0070] S222, Phase Two: Detailed Scheduling for Multiple Objectives; Acquiring a robust fleet size from the first phase After the input, the second phase proceeds to the detailed scheduling stage. The core innovation of this stage lies in: under the constraint of robust fleet size, through multi-objective optimization, while balancing operating costs, load balancing and service quality, and establishing a precise ship flow state transition equation, incorporating ships in port into dynamic inventory calculations, to achieve refined scheduling across all time and space dimensions.
[0071] Specifically as follows: (1) Decision variables and state variables: : Decision variable, representing the number of ships of type k in direction d at time t; Slack variables represent the unmet demand in direction d at time t; they are automatically obtained by the solver during the optimization process, and their values are determined by the penalty term of the objective function and the demand constraints. The equivalent calculation formula is as follows: : State variable, representing the number of available ships of type k at port p at time t, t∈{0,1,…,T}; : Deviation variable, representing the absolute deviation between the available inventory of ship type k at port p at time t and the target inventory.
[0072] (2) Objective function: objective function It consists of three parts: minimizing the total number of operating departures (operating costs), minimizing vessel distribution deviation (load balancing), and minimizing the penalty for unmet demand (service quality). Specifically, it is expressed as follows: in, First item Minimize the total number of operational departures, with a weight α=20, corresponding to fuel consumption and human resource costs; Second item Minimize the ship distribution deviation, with a weight β=40, to ensure balanced ship load between the two ports and prevent one-way capacity backlog; Third item Minimize the penalty for unmet needs, with a weight λ=100, to ensure service quality and prioritize meeting the needs of vehicles and passengers crossing the sea.
[0073] The design of the weight parameters reflects the priority logic of scheduling: first, ensure that demand is not delayed (λ=100), second, optimize load balancing (β=40), and finally control operating costs (α=20).
[0074] (3) Define constraints, including: first-stage coupling constraints, demand constraints (soft constraints), upper limit constraints on departure frequency, inventory deviation constraints (load balancing constraints), ship dynamic inventory state transition equations and departure availability constraints.
[0075] Specifically as follows: (a) First-stage coupling constraints (key constraints of this invention).
[0076] This clause constrains the left-hand term. This represents the total number of ships departing today (hours 0 to T-1). (Right-side item) This represents the optimal fleet size passed in from the first-stage solution. The maximum number of ships that can be dispatched at once ensures that no extra ships will be dispatched, thus avoiding a waste of transport capacity.
[0077] This constraint will determine the robust fleet size in the first phase. In the second phase, the total number of departures for vessel type k within a 24-hour period is limited to not exceeding its theoretical maximum departure capacity. This constraint is the core mechanism of the two-stage coupling: on the one hand, it ensures that detailed scheduling is executed within a robust fleet budget and does not exceed the capacity supply limit; on the other hand, because... It is determined based on robust requirements, and its inherent capacity redundancy beyond basic requirements reserves space for the flexible adjustment of the dynamic correction layer.
[0078] (b) Demand constraints (soft constraints).
[0079] For each time step t∈{0,1,…,T-1} and each direction d: This constraint requires that the total carrying capacity in direction d at time t be no less than the difference between the predicted demand and the unmet demand. As non-negative slack variables, unmet needs are allowed even when capacity is insufficient to meet all demands. However, these unmet needs are penalized with a high weight λ in the objective function, driving the model to meet demand as much as possible. This soft constraint design is another innovative aspect of this invention: compared to hard constraints that force the fulfillment of all predicted demands (which can easily lead to an infeasible solution during sudden surges in traffic), soft constraints ensure that the model is always feasible while driving it to approach demand coverage as closely as possible through a penalty mechanism.
[0080] (c) Maximum constraint on departure frequency.
[0081] For each direction d and each time step t: This constraint limits the maximum number of ships that can be dispatched from each port at each time step to no more than [a certain number]. This corresponds to the physical limitations of port berths and waterway traffic capacity.
[0082] (d) Inventory deviation constraint (load balancing constraint).
[0083] Define the target inventory level for each port and each ship type as the initial inventory: For each time step t∈{0,1,…,T-1}: The two inequalities mentioned above are combined to achieve The linearized expression of this constraint is to maintain load balance on the two-way shipping routes: by incorporating the deviation variable into the objective function to minimize the ship inventory at each port, the model tends to keep the ship inventory at each port close to the initial level, avoiding the situation where excessive one-way shipping leads to a backlog of ships at one port and a depletion of ships at the other port, thereby ensuring the continuous operation of the two-way shipping routes.
[0084] (e) Constraints on the dynamic inventory state transition equations of ships (the core constraint of this invention).
[0085] This is one of the most critical technical constraints of this invention. It precisely describes the conservation flow process of ships between two ports and incorporates the initial ships en route into the state calculation. Furthermore, subsequent dynamic scheduling adjustments must also adhere to this constraint.
[0086] Initial state constraints: State transition equation: Define the forward window for ships en route (That is, the sailing time is reduced by 1, because it takes τ hours for a ship to arrive after departing from the other side, while ships that were already on the route at the initial planning time will arrive successively from the 1st to the τ-1th hour).
[0087] For t∈{0,1,2,…,T-1}, define the number of ships arriving at port A at time t. : in For: the day Ships that are always ready to depart.
[0088] Similarly, the number of ships arriving at port B at time t : The above fraction means that when t < τ, the arriving ships only come from those already en route at the initial time. When t ≥ τ, arriving vessels include vessels en route (only in the forward window). (Valid within the time limit) and ships that arrive after a τ-hour voyage from the opposite bank. .
[0089] Therefore, the state transition equation is: This equation describes the conserved flow of ship inventory: the number of available ships at port A at time t is equal to the number of available ships at the previous time minus the number of ships departing from A to B, plus the number of ships departing from B to A.
[0090] (f) Schedule availability constraints.
[0091] For each time step t∈{0,1,…,T-1}: This constraint ensures that the number of ships dispatched at any given time does not exceed the current available inventory of that type of ship at that port, preventing infeasible scheduling due to "no ships available for dispatch".
[0092] The aforementioned departure availability constraints will also be recalculated and dynamically modified when actions are modified at the dynamic layer. This is a physical port constraint.
[0093] S223, use a linear programming solver to solve the above problem and obtain a static schedule.
[0094] Specifically, this embodiment uses PuLP, a high-level mathematical modeling library in the Python environment, to formally construct the above two-stage model. Given the large-scale integer programming nature of the scheduling problem, the high-performance CBC (Coin-or branch and cut) solver is selected.
[0095] To balance computational accuracy with the real-time requirements of port scheduling, this experiment sets the solution time limit to 300 seconds. This configuration ensures that even when the optimal solution is difficult to obtain in a very short time, the system can stably produce the currently known optimal feasible solution, thus meeting the robustness requirements of actual production environments.
[0096] This model employs the core idea of hierarchical optimization, deeply decoupling complex fleet resource allocation from real-time departure planning decisions. The necessity of this two-stage modeling strategy lies in the fact that port scheduling systems not only need to cope with long-term demand fluctuations but also maintain extremely high real-time responsiveness. The first stage: It determines a "safe" fleet size. This scale is sufficient to cope with robust sets The worst-case demand fluctuation within the period. Output at this stage. This ensured the feasibility of the second phase and set a budget ceiling for total capacity throughout the planning cycle, guaranteeing the feasibility boundary for subsequent refined scheduling. Furthermore, due to... Based on robust demand determination, the results typically imply capacity exceeding basic demand, thus reserving flexible capacity space for subsequent dynamic adjustments in actual operation. This underutilized capacity provides support for dynamic adjustments to short-term random fluctuations, achieving a balance between stability and flexibility. In the second stage: the model shifts to a specific, refined allocation level. Under the capacity budget constraints determined in the first stage, the model at each discrete time step... The allocation of ship resources is refined. Through multi-objective trade-offs, the final baseline departure order, i.e., the decision variables, is generated. This hierarchical strategy effectively reduces the difficulty of solving large-scale nonlinear mixed-integer programming problems. More importantly, through the logic of "proactive reservation and flexible adjustment," it enables the scheduling plan to resist prediction errors, providing solid algorithmic support for the proactive intelligent scheduling of smart ports.
[0097] S3, Dynamic Scheduling Optimization; Dynamic scheduling based on rolling short cycles using deep reinforcement learning: After a period of time, dynamic scheduling optimization is triggered, taking into account real-time port conditions, operational status, and future trends. Hourly short-term forecast flow (Based on the actual operation of the port, this embodiment) ), for the future The schedule will be optimized for the next hour. In particular, future... The short-term predicted flow for each hour is obtained by the Transformer module in step S1. It can accept the actual flow of the previous day as input. Since the number of prediction steps is much smaller than that of 24 hours, the prediction error is smaller than that of 24-hour prediction. It can better reflect the actual operation of the day and serve as a dynamic adjustment for static scheduling.
[0098] Specifically, the plan is to execute at each interval. After (in this embodiment, = 3 hours), triggering dynamic scheduling optimization.
[0099] The reinforcement learning method used here modifies the plan P (i.e., the decision variable) generated in the static scheduling stage mentioned above. ), and will be spaced out in the following day. The adjustment mechanism is triggered multiple times to make short-term predictions based on the current operating conditions and the actual traffic of the day to correct the subsequent scheduling plan for the day (because there is a deviation between the long-term 24-hour prediction and the actual situation, so dynamic adjustment and correction are required) until the next static scheduling is triggered at 0:00, and then the whole process is repeated.
[0100] The specific process of rolling short-cycle dynamic correction scheduling based on deep reinforcement learning is as follows (e.g.) Figure 4 ): S31, Reinforcement Learning Environment Construction.
[0101] Define the state space, action space, and reward function for reinforcement learning, as follows: The state space contains all the current system information and short-period prediction information needed for the reinforcement learning model to make decisions. It forms a multi-dimensional vector that is a mixture of continuous and discrete elements, used as the input to the deep neural network. The state variable members are shown in Table 3. Table 3 State Space Description Action space: at each decision point For the two ports ( and Perform discrete actions respectively and It includes 5 possible actions, encoded as 0~4 respectively, and the final joint action space size is... The meaning of each action is shown in Table 4: Table 4 Description of Motion Space reward function It is a composite objective designed to reward adjustments that meet service demand and improve resource utilization, while penalizing actions that result in insufficient capacity or high costs (such as vessel redeployment). Rewards are primarily based on the immediate impact of decisions made at the current time step on the port's supply-demand balance, congestion levels, and action types. As follows: in Service satisfaction reward: Rewards are given when capacity meets demand, and additional rewards are given when both ports meet demand simultaneously. Efficiency rewards: Rewards are given for successful optimization actions and low latency status. System penalties: Penalize insufficient capacity, wasted empty loads, and violations of rules.
[0102] The specific calculation logic for each component is as follows: (1) Service satisfaction reward This is used to encourage scheduling actions to ensure that port capacity meets demand while avoiding excessive capacity redundancy. The specific calculation rules are as follows: Single Port Meeting Reward: If the capacity of Port A meets the demand, a reward of +20 will be given; if the capacity of Port B meets the demand, a reward of +20 will also be given.
[0103] Dual-port synergy bonus: If both Port A and Port B meet their capacity requirements, it means that the system has achieved a global supply and demand balance, and an additional synergy bonus of +50 will be given.
[0104] Capacity redundancy penalty: To prevent excessive waste of capacity, when a port (Port A or Port B) has met its capacity demand and the agent's action towards that port remains unchanged (Action 0), a penalty is imposed on the difference between the port's absolute demand and capacity. The penalty amount is 0.1 times the difference (i.e., ...). ).
[0105] (2) Efficiency Rewards This is used to reward effective scheduling actions (action types 1, 2, and 3, excluding ship repositioning) that optimize resource allocation and eliminate congestion. The specific calculation rules are as follows: Basic reward for optimized actions: When performing regular optimization actions on Port A or Port B, a basic reward of +5 is given respectively.
[0106] Supply and demand matching efficiency bonus: If the port's capacity meets the demand after optimization, an additional +30 efficiency bonus will be given to both port A and port B.
[0107] Low congestion clearance reward: If the total number of vehicles congested at both ports is less than 5, provided that optimization actions are performed and capacity requirements are met, it means that the system has efficiently cleared the congestion and will be given an additional clearance reward of +5.
[0108] (3) System penalties Used to negatively reinforce capacity shortages and high-cost operations, preventing the system from falling into a bad state. The specific calculation rules are as follows: Insufficient capacity penalty: If the capacity of port A does not meet the demand, a penalty of -35 will be imposed; if the capacity of port B does not meet the demand, a penalty of -35 will also be imposed.
[0109] Ship redeployment cost penalty: Since cross-port ship redeployment (Action 4) involves high operating costs and time delays, a penalty of -15 will be imposed when a ship redeployment is performed at port A or port B.
[0110] Rule violation penalty: If an action violates the port constraints in step S222, including the ship dynamic inventory state transition equation constraints and departure availability constraints, an extremely high penalty of -200 is imposed, indicating that the constraint is a real-world constraint and cannot be violated.
[0111] S32, reinforcement learning model architecture.
[0112] Reinforcement learning model architecture such as Figure 3 As shown, it includes, in sequence: input layer, bidirectional long short-term memory network (Bi-LSTM), multi-head attention mechanism, shared hidden layer, Q-value network, and output layer.
[0113] The input layer receives the environmental state time-series data and performs dimensionality adaptation to meet the input requirements of the subsequent recurrent neural network. The input data is represented as a tensor of shape (batch_size, seq_len, state_dim), where batch_size is the batch size, seq_len is the sequence length, and state_dim represents the state space dimension. Dimensional alignment is performed through a linear layer to ensure smooth data flow into the time-series extraction network.
[0114] The Bi-LSTM network extracts contextual features along the time dimension from the input state sequence, capturing the time-inertial evolution patterns of traffic flow. The Bi-LSTM has a hidden layer dimension of hidden_dim=128, a stacking layer count of 2, and a bidirectional configuration (bidirectional=True). The forward and backward LSTMs process the input sequence simultaneously, concatenating the outputs at each time step to expand the output dimension into a high-dimensional feature vector of hidden_dim * 2 = 256, containing both historical dependencies and future prediction information.
[0115] The multi-head attention mechanism weights the temporal features of the Bi-LSTM output based on their importance, allowing the network to focus on sudden features that play a crucial role in decision-making (such as a surge in passenger flow) and suppress redundant noise. Eight attention heads are used, each with an attention dimension of 64. A query (Q), key (K), and value (V) matrix is generated through a linear transformation. After scaling the dot product attention and calculating the weights, the matrix is weighted and summed with V (the multi-head attention mechanism is a prior art technique). Finally, the multi-head outputs are fused through a fully connected layer to restore the same feature dimension as the input, and the attention weights are output for interpretability analysis.
[0116] The shared hidden layer maps high-dimensional attention features to a low-dimensional decision space and serves as a carrier for the sharing of features in the dual-port collaborative scheduling, maintaining global cognition. It includes a NoisyLinear layer with an output dimension of 128, a ReLU activation function, and a Dropout layer with a dropout rate of dropout_p=0.2.
[0117] Unlike traditional linear layers, this invention employs a noisy network technique for the shared hidden layer. During the training phase, factorized Gaussian noise is injected into the weights and biases. and The random variable (generated by the outer product) is perturbed (sigma_init=0.017); during the evaluation phase, it degenerates into deterministic weights. This endows the network with the ability to adaptively explore its state.
[0118] The Q-value network employs a Dueling Network architecture to estimate the action value of each port, decoupling the Q-value into state value and action advantage, thus improving the accuracy of evaluating specific port states. Two independent parallel branches are constructed for ports A and B respectively. Each branch contains a Value stream (NoisyLinear, output dimension 1) and an Advantage stream (NoisyLinear, output dimension action_dim=5). The two branches share the weights of all front-end network layers. The Q-value is calculated using the following formula... The identifiability problem is addressed by subtracting the mean of the dominant flow, ensuring the uniqueness of state values. All fully connected layers are replaced with NoisyLinear to maintain the consistency of the exploration.
[0119] The output layer outputs the Q-values of each candidate action for both ports in their current state, and uses these Q-values as the basis for selecting the final capacity scheduling action. Specifically, the output layer directly transmits the calculation results of the Q-value network, returning Q-value tensors of two ports with shapes of (batch_size, seq_len, 5). The agent selects the action index (0-4) with the largest Q-value according to a greedy strategy, as the differentiated scheduling instruction for port A and port B.
[0120] At the feature extraction level, considering the significant temporal inertia of port ferry traffic and congestion, and the fact that the current system state is often continuously influenced by the capacity allocation of previous periods, the model first introduces a bidirectional long short-term memory network (Bi-LSTM) to simultaneously mine forward and backward contextual information from historical trajectories. This effectively extracts implicit traffic evolution trends and enhances the ability to predict future states. Building on this, to focus on key decision features from complex input sequences, the model integrates a multi-head attention (MHA) mechanism. This mechanism automatically learns and reweights the feature vectors output by the Bi-LSTM, enabling the agent to suppress irrelevant noise and focus on high-value signals such as sudden surges in passenger flow, significantly improving the robustness of feature representation.
[0121] Regarding the decision generation and exploration mechanism, this architecture replaces the traditional fully connected layers with a NoisyNet, replacing the deterministic weights in the parameter space with random variables containing random noise, thereby achieving state-dependent adaptive exploration. This mechanism allows the model to automatically adjust the exploration magnitude during training, eliminating the need for reliance on [other methods]. Strategy exploration ensures a smooth transition from extensive initial searching to precise later utilization, effectively avoiding the trap of local optima. Considering the spatial characteristics of dual-port collaborative scheduling, the model splits into two independent Dueling Network branches after sharing the backbone network, each responsible for decision-making for port A and port B respectively. This is achieved by further decomposing the Q-value into state-value streams. With the advantage of action flow This architecture achieves deep decoupling of value and action, improving sample efficiency while ensuring that the agent can maintain a global understanding of the overall system environment based on a unified backbone network, and make differentiated optimal decisions for the local constraints of each port through independent branches.
[0122] S33 employs a rule-based intelligent ship selection and recursive full lifecycle plan correction mechanism to adjust the overall scheduling plan after generating adjustment actions.
[0123] To transform macro-level scheduling instructions (such as overtime, reduced shifts, or combined shifts) generated by deep reinforcement learning models into micro-level scheduling plans with physical enforceability, this invention designs a rule-based intelligent ship selection and recursive full lifecycle plan correction mechanism to solve the scheduling adjustment problem: if a ship's scheduling instructions are adjusted at time t, the chain reaction on the subsequent hours [t+1, 24] of the entire 24-hour scheduling plan for that day will require the ship's subsequent plans to be adjusted as well to ensure the optimal global 24-hour scheduling plan.
[0124] Specifically as follows: S331, firstly, is an intelligent ship selection mechanism based on minimal perturbation. When the reinforcement learning model issues an adjustment command, the system uses a heuristic scoring algorithm to precisely map the abstract command to a specific ship instance. For overtime orders (ADD), the algorithm follows the principles of minimum disturbance and priority for small vessels, prioritizing idle vessels with the least impact on future timetables (such as no subsequent tasks) and smaller capacity, thereby reserving large vessel resources to cope with potential demand peaks.
[0125] When executing a merge instruction, the system prioritizes removing scheduled vessels with the lowest disturbance scores and the largest capacity, aiming to minimize unnecessary operating costs and avoid wasting capacity during periods of low demand.
[0126] When executing the replacement instruction (REPLACE), the model adopts a dynamic matching strategy of "one replacement for one" to adjust capacity according to the real-time supply and demand gap: when there is insufficient capacity, capacity is increased by "replacing large with small", while when there is excess capacity, cost is reduced by "replacing small with large".
[0127] This rule-based auxiliary decision-making layer not only ensures the physical feasibility of scheduling actions, but also effectively mitigates the impact of frequent dynamic adjustments on the overall operation of the port by quantitatively considering system disturbances.
[0128] S332, followed by recursive full lifecycle plan correction. For adjusted vessels (such as extra vessels), the system automatically and recursively calculates their arrival time and subsequent status on the other side. If there is a capacity gap on the other side after the vessel arrives, the system automatically triggers a "return extra" instruction; if there is no gap, the vessel is added to the other side's idle vessel pool. This mechanism eliminates cross-time-domain chain reactions from a single scheduling, ensuring the closed-loop consistency of capacity flow.
[0129] In this embodiment, when the deep reinforcement learning model outputs macroscopic adjustment instructions... Afterward, the system enters the specific scheduling plan revision phase. This process aims to transform the abstract scheduling strategy into executable micro-level ship navigation tasks, and the specific execution flow is as follows: First, initialize and update the instantaneous state. Receive the current shift schedule. Adjustment instructions Step S331 Selected target vessel Current planning time domain and one-way sailing time Perform adjustments and update the vessel. Future time interval The ship's navigation status within the vessel. Simultaneously, the ship's... Remove it from the idle queue of the current departure port and record its departure time. Based on this calculation, the ship The estimated time of arrival at the opposite port is denoted as . And update the status to: Marked vessel At any moment It will then become available.
[0130] Subsequently, a recursive demand detection and correction mechanism is initiated to handle subsequent arrangements after the ship arrives at the other side. It is then determined whether the current time is still within the planned time domain (i.e., meets the requirements). If the condition is met, then proceed to the following loop processing logic: (1) Supply and demand scan: Scanning the port on the other side in the upcoming time window (in Long-term forecast demand within a one-way trip time and existing transport capacity supply .
[0131] (2) Gap identification and decision-making: Scenario 1 (Existence of Capacity Shortage): If insufficient capacity is detected on the other side (i.e.) ), and ships Meeting the necessary turnaround time constraints triggers a new "ADD" instruction. At this point, the system updates the shift schedule. Arrange ships The revised plan is to immediately proceed with the return voyage from the current port on the opposite shore. Subsequently, the system updated the time pointer, causing... We will continue to monitor the supply and demand situation in the next stage, forming a recursive closed loop.
[0132] Scenario 2 (No Capacity Shortage): If there is sufficient capacity on the other side or no demand, the system will allocate ships... The vessel is added to the idle vessel pool at the opposite port to await further scheduling. At this point, the process for dispatching the vessel is terminated. During the current recursive process, the ship will be in a standby state, waiting for the next static planning task or dynamic scheduling instruction.
[0133] Through the above steps, the system will finally output the revised and complete shift schedule. This ensures that a single scheduling action can achieve closed-loop consistency in capacity flow throughout its entire lifecycle.
[0134] S4 uses scheduling information for dispatching; After the dynamic adjustment and correction of the scheduling plan is completed, the system enters the final execution and feedback stage, issuing and executing scheduling instructions, and conducting performance evaluation and model iteration.
[0135] Specifically as follows: S41 Dispatch instruction issuance and execution; The system will generate a revised shift schedule. The data is parsed into specific work orders and distributed to each vessel terminal and port operation system. For newly added overtime voyages or modified return trips, the system will prioritize sending high-priority notifications to ensure that the shipping company and port authorities adjust resource allocation in a timely manner (such as berth allocation and stevedore scheduling) to guarantee the physical implementation of dispatch instructions. If an error occurs, the information is recorded for future reference.
[0136] S42 Performance Evaluation and Model Iteration; At the end of this scheduling cycle, the system collects actual operational data (including actual passenger flow, vessel accuracy, and number of unfulfilled hours) and calculates the actual reward value for this adjustment strategy. This feedback data will be stored in the Replay Buffer for subsequent offline training and parameter updates of the deep reinforcement learning model. This will enable the model to more accurately balance capacity supply and operating costs in future scheduling decisions, achieving continuous self-evolution of the system.
[0137] Example 2 To verify the superiority of the method of the present invention, four existing algorithms that are representative in the fields of shipping and air traffic control were selected as benchmarks for comparative experiments.
[0138] The comparison methods include: Multi-stage Stochastic Programming (MSP), Continuous Approximation (CA), Maximum Cardinality Bipartite Graph (MCBM) matching algorithm, and Capacity Fluctuation Factor Elastic Scheduling (CFFES) model. Explanations are as follows: In the field of ship scheduling, a multi-stage stochastic programming model (MSP) based on two-stage robust optimization theory and a continuous approximation model (CA) that uses system equations to simulate complex logistics scenarios were selected.
[0139] Two mature scheduling schemes in the aviation field were adapted and modified: one is the Maximum Cardinality Bipartite Graph Matching (MCBM) algorithm, which uses bipartite graph abstraction to achieve optimal matching of tasks and resources to minimize capacity demand; the other is the Capacity Fluctuation Factor Elastic Scheduling Model (CFFES), which dynamically adjusts the timetable by identifying traffic fluctuation factors and capacity elasticity coefficients.
[0140] By comparing and analyzing the proposed algorithm under the same experimental conditions, this paper aims to fully reveal the core competitiveness of the model in handling complex dynamic scheduling problems.
[0141] Test environment: Real-world operational records from three consecutive months in 2024 were used for Xinhai Port (Port A) and Xuwen Port (Port B). Each scenario was run separately, and data was obtained based on the experimental results. The planning accuracy and average daily unmet needs of each method were analyzed.
[0142] Evaluation Indicator Explanation: Plan accuracy measures how well a planned scheduling schedule meets actual passenger and vehicle demand, and directly reflects the effectiveness of the plan. The Unserved Hour Per Day (UHH) is a daily timeframe that tracks the cumulative time a port system fails to process transport demands within a 24-hour period. It is a key indicator for measuring port service levels and congestion mitigation capabilities.
[0143] in The number of days required to meet demand in the scheduling plan. The total number of days that require scheduling. For the first Unfulfilled time (hours) per day.
[0144] The comparison results of the plan accuracy are as follows Figure 5 As shown, the method described in this invention consistently achieves the highest classification accuracy, demonstrating superior performance and robustness, and improving the degree to which scheduling plans meet actual passenger flow demands. Compared to the four current traditional methods, it improves performance by 16.03% compared to the best-performing method P. Compared to the weakest-performing method MCBM, it improves performance by more than 10 times.
[0145] The comparison results of the average daily unmet duration are as follows: Figure 6As shown, the method described in this invention consistently achieves the lowest average number of unfulfilled hours, reducing the cumulative time the system fails to process transportation demands in a timely manner within 24 hours and alleviating congestion. Compared to the four current traditional methods, the unfulfilled hours are reduced by 44.24% compared to the best-performing method, MSP, and by 83.75% compared to the weakest-performing method, MCBM.
[0146] Traditional methods typically focus only on static scheduling or short-cycle adjustments, resulting in a lack of ability to organically unify global resource coordination with flexible local optimization, easily leading to imbalances in resource allocation at a single scale. The method of this invention, compared to traditional methods, has the ability to achieve a balance between global optimization and local flexibility, and solves the systemic problem that single-scale optimization cannot simultaneously address both global and real-time requirements.
[0147] The scope of protection of the multi-target dynamic and static combined port vessel collaborative scheduling method described in this invention is not limited to the execution order of the steps listed in this embodiment. Any solution implemented by adding, subtracting, or replacing steps in the prior art based on the principles of this invention is included within the scope of protection of this invention.
Claims
1. A multi-objective, dynamic-static combined port vessel collaborative scheduling method, characterized in that, Includes the following steps: S1, Preparation Stage; Initial vessel status data is retrieved from the port status database, including the number of vessels of each type at each port at the initial moment, as well as the positions and estimated arrival times of vessels en route at the initial moment. Historical vehicle and passenger traffic data from the previous 48 hours is extracted and input into the Transformer model to predict vehicle and passenger traffic for the next 24 hours, yielding a long-term prediction result for the next 24 hours. , express time Forecasting transportation demand in different directions; S2. Static shift schedule generation; A segmented robust global static scheduling model is constructed. The long-term passenger and vehicle traffic forecast results for the next 24 hours are input. A two-stage robust optimization theory is adopted. In the first stage, the fleet size is determined through robust optimization, the robust total demand is calculated, the objective function is established to minimize the unmet demand and the fleet adjustment cost, and the optimal fleet size for each ship type is output. The second stage involves refined scheduling under the fleet size constraints determined in the first stage. A multi-objective function is established to minimize the total number of operational departures, vessel distribution deviations, and unmet demands. The constraints include coupling constraints from the first stage, demand constraints, upper limit constraints on departure frequency, inventory deviation constraints, constraints on the dynamic inventory state transition equation of vessels, and departure availability constraints. The decision variables, namely the number of departures for each vessel type in each course and time period, are obtained by solving linear programming, and a baseline scheduling plan is output. S3, Dynamic Scheduling Optimization; Dynamic scheduling optimization is triggered after the planned execution interval. A rolling short-cycle dynamic correction mechanism based on deep reinforcement learning is adopted. The LA-Dueling-DQN network architecture is used to target the real-time port status, including the number of currently available ships, the distribution of ships en route, the number of vehicles stranded, and the short-term predicted flow for the next W hours. The deep reinforcement learning agent outputs adjustment actions and generates a corrected scheduling plan. S4. Utilize scheduling information for dispatching; The generated revised shift schedule is parsed into execution work orders and sent to the ship terminal and port operation system. Actual operation data is collected to calculate reward values and stored in the experience playback pool for offline model training. The execution results of the revised shift schedule are fed back to the static shift schedule generation step of the next cycle to form a closed-loop optimization. In step S3, a rule-based intelligent ship selection and recursive full life cycle plan correction mechanism is adopted to adjust the overall scheduling plan after generating adjustment actions; S331, firstly, is an intelligent ship selection mechanism based on minimal perturbation. When the reinforcement learning model issues an adjustment command, the system uses a heuristic scoring algorithm to precisely map the abstract command to a specific ship instance. For overtime orders (ADD), the algorithm follows the principles of minimum disturbance and priority for small vessels, prioritizing idle vessels with the least impact on future timetables and smaller capacity. When executing the MERGE command, the system prioritizes removing planned vessels with the lowest disturbance scores and the largest capacity. When executing the REPLACE instruction, the model adopts a dynamic matching strategy of "one-to-one replacement" to adjust capacity according to the real-time supply and demand gap: when there is insufficient capacity, capacity is increased by "replacing large with small", while when there is excess capacity, cost is reduced by "replacing small with large". S332, followed by recursive full lifecycle plan correction. For the adjusted vessel, the system recursively calculates its arrival time and subsequent status on the other side. If there is a capacity gap on the other side after the vessel arrives, the system triggers the "return trip overtime" instruction. If there is no gap, the vessel is added to the idle vessel pool on the other side.
2. The method according to claim 1, characterized in that, In step S1, the core parameters of the Transformer model include: The historical window length is 24 to 72 hours, the prediction step size is 24-hour long-term prediction and 3-hour short-term prediction, the embedding dimension is 64, the number of attention heads is 4, the feedforward network dimension is 128, the number of encoder blocks is 2, and the random dropout rate is 0.2%.
3. The method according to claim 1, characterized in that, In step S2, the first phase of robust fleet size decision-making specifically involves: Calculate robust total demand: For each heading d, calculate the robust total demand based on the sum of long-term forecasts for that heading over 24 hours, plus a robust adjustment term. in, Gathering for the direction of navigation, Let Γ represent the total daily forecast demand for heading d, where Γ·δ is the robust redundancy term; Γ is the robust adjustment coefficient, and δ represents the maximum possible deviation of demand in a single hour. Calculate the initial total fleet size: For each ship type k, count the number of all ships in the current system, including ships berthed at each port and ships en route: in, Port at the initial moment Type The number of ships, This indicates that the ship will arrive at the port in the t-th hour. The number of ships of type k, among which It can be A or B; The decision variables, objective function, and constraints for the first stage are defined as follows: The decision variables are Pareto optimal solutions for fleet size adjustment in extreme uncertainty scenarios, including: The optimal fleet size for ship type k is the total number of ships of ship type k that the system should be equipped with; Represents the field of non-negative integers; Unmet robust total requirements for heading d; Represents the field of nonnegative real numbers; : The size adjustment deviation of ship type k, that is, the absolute value of the difference between the optimal size and the initial size; objective function The aim is to balance service availability with asset adjustment costs, as shown in the following formula: The objective function consists of two parts: the first term The second item is the penalty for failing to meet the robust total requirement. Penalties for adjusting deviations in the fleet; The constraints include: (1) Overall fleet size constraints: The total fleet size of the global fleet size constraint system is: This corresponds to the upper limit of the total number of ships in actual operation; (2) Fleet adjustment deviation constraints: The two inequalities above are combined to achieve The linearized expression ensures The value is the absolute value of the difference between the optimal size and the initial size; (3) Robust demand coverage constraint: The size of each ship type The optimal fleet size for ship type k is the total number of ships of ship type k that the system should configure. For single-ship carrying capacity, the maximum number of round trips per day for a single ship is: , ,in This refers to the one-way travel and operation time of a vessel between the two ports. For the calculation above The value along the heading d; Solving the first-stage model to obtain the robust fleet size: The above model is solved using a mixed-integer programming solver to obtain the optimal fleet size for each ship type. ,Should This will be passed on as a resource ceiling constraint to the second phase.
4. The method according to claim 3, characterized in that, In step S2, Acquiring a robust fleet size in the first phase After receiving the input, the second stage involves detailed scheduling. The decision variables, state variables, objective function, and constraints for the second stage are defined as follows: (1) Decision variables and state variables: : Decision variable, representing the number of ships of type k in direction d at time t; Slack variables represent the unmet demand in direction d at time t; : State variable, representing the number of available ships of type k at port p at time t, t∈{0,1,…,T}; : Deviation variable, representing the absolute deviation between the available inventory of ship type k at port p at time t and the target inventory; (2) Objective function: objective function It consists of three parts: minimizing the total number of operational departures, minimizing vessel distribution deviation, and minimizing the penalty for unmet demand, specifically expressed as follows: in, First item Minimize the total number of operating shifts, corresponding to fuel consumption and human resource costs; Second item Minimize vessel distribution deviations, ensure balanced vessel load between the two ports, and prevent one-way capacity backlog; Third item Minimize penalties for unmet needs, ensure service quality, and prioritize meeting the needs of vehicles and passengers crossing the sea; (3) Define constraints, including: first-stage coupling constraints, demand constraints, upper limit constraints on departure frequency, inventory deviation constraints, ship dynamic inventory state transition equations and departure availability constraints.
5. The method according to claim 4, characterized in that, The objective function parameters are: α=20, β=40, λ=100.
6. The method according to claim 4, characterized in that, The constraints are as follows: (a) First-stage coupling constraints; The left-hand item This represents the total number of ships departing between 0 and T-1 hours today; the item on the right... This represents the optimal fleet size passed in from the first-stage solution. The maximum number of ships that can be dispatched; (b) Demand constraints; For each time step t∈{0,1,…,T-1} and each direction d: in, The predicted long-term 24-hour period time The predicted transport demand in the direction; the demand constraint requires that the total carrying capacity of direction d at time t is not less than the difference between the predicted demand and the unmet demand; These are non-negative slack variables; (c) Maximum frequency constraint; For each direction d and each time step t: The departure frequency limit constraint restricts the maximum number of departures per port at each time step to no more than [a certain number]. This corresponds to the physical limitations of port berths and waterway traffic capacity; (d) Inventory deviation constraints; Define the target inventory level for each port and each ship type as the initial inventory: For each time step t∈{0,1,…,T-1}: (e) Constraints on the ship dynamic inventory state transition equations; It describes the conservation flow process of ships between two ports, and incorporates the initial ships en route into the state calculation. Subsequent dynamic scheduling adjustments must also follow this constraint. Initial state constraints: in, Port at the initial moment Type The number of ships, It can be A or B; State transition equation: Define the forward window for ships en route ; For t∈{0,1,2,…,T-1}, define the number of ships arriving at port A at time t. : in For: the day Ships that depart at any time The initial information represents the number of ships of type k that are about to arrive at port p in the next hour t. Similarly, the number of ships arriving at port B at time t : Therefore, the state transition equation is: (f) Schedule availability constraints; For each time step t∈{0,1,…,T-1}: The departure availability constraint ensures that the number of departures at any given time does not exceed the current available inventory of that vessel type at that port, preventing infeasible scheduling due to "no vessels available for departure".
7. The method according to claim 1, characterized in that, In step S3, the deep reinforcement learning-based rolling short-cycle dynamic correction mechanism includes constructing a reinforcement learning environment and defining a state space, action space, and reward function. The state space contains feature vectors, including currently available ships, the originally planned number of ships to depart, the distribution of ships en route, short-term future demand forecasts, and the distribution of stranded vehicles. The action space is a joint space of five actions for each of the two ports, including normal execution, shift merging, overtime, ship replacement, and ship repositioning. The reward function is the sum of service satisfaction reward, efficiency reward, and system penalty; the service satisfaction reward includes single-port satisfaction reward and additional collaborative reward for dual-port collaborative satisfaction. The model can make short-term traffic predictions based on the actual traffic flow on the day, and adjust the static schedule generated based on the previously predicted 24-hour traffic flow that does not match the actual traffic flow, thus enabling the schedule to have an adaptive adjustment mechanism.
8. The method according to claim 1, characterized in that, In step S3, The reinforcement learning model architecture consists of the following layers in sequence: input layer, bidirectional long short-term memory network Bi-LSTM, multi-head attention mechanism, shared hidden layer, Q-value network and output layer; The input layer receives the state time series data of the environment and performs dimensionality adaptation; the input data is represented as a tensor of shape (batch_size, seq_len, state_dim), where batch_size is the batch size, seq_len is the sequence length, and state_dim represents the state space dimension; it is then subjected to dimensionality alignment through a linear layer. The bidirectional long short-term memory network Bi-LSTM extracts contextual features in the time dimension of the input state sequence, capturing the traffic evolution pattern with time inertia; the Bi-LSTM hidden layer dimension hidden_dim=128, the number of stacked layers is 2, and a bidirectional configuration is adopted; the forward and backward LSTMs process the input sequence simultaneously, and the output of each time step is concatenated, and the output dimension is expanded to a high-dimensional feature vector of hidden_dim * 2 = 256, which contains both historical dependence and future prediction information; The multi-head attention mechanism weights the temporal features output by Bi-LSTM based on their importance. There are 8 attention heads, and the attention dimension of each head is 64. A query Q, key K, and value V matrix is generated through linear transformation. After the weights are calculated by scaling dot product attention, the matrix is weighted and summed with V. Finally, the multi-head outputs are fused through a fully connected layer to restore the feature dimension to the same as the input. The attention weights are then output for interpretability analysis. The shared hidden layer maps high-dimensional attention features to a low-dimensional decision space and serves as a carrier for sharing features in the dual-port collaborative scheduling, maintaining global cognition; it includes a NoisyLinear layer with an output dimension of 128, a ReLU activation function, and a Dropout layer. The Q-value network uses a Dueling Network architecture to estimate the action value of each port, decoupling the Q-value into state value and action advantage; two independent parallel branches are constructed for port A and port B respectively, each branch containing a value stream and an advantage stream; the two branches share the weights of all front-end network layers; The output layer outputs the Q values of each candidate action for both ports in the current state. The agent selects the action index with the largest Q value according to a greedy strategy, which serves as the differentiated scheduling instruction for port A and port B.
9. The method according to claim 1, characterized in that, The specific execution process of the recursive full life cycle plan revision includes: receiving the current scheduling plan P, adjustment instructions, selecting the target vessel, and the current planning time domain. and one-way sailing time ; Perform adjustment actions, update the target vessel's navigation status for the future time interval, remove the target vessel from the idle queue of the current departure port, and calculate the target vessel's estimated arrival time at the opposite port. Equals the departure time plus the one-way travel time ; judge add If the time frame is less than or equal to the planned time period T, then proceed to the loop: scan the predicted demand and existing capacity supply of the opposite port within the time window. If a capacity shortage is detected on the opposite port and the target vessel meets the turnaround time constraint, then trigger an overtime instruction and update the scheduling plan. Updated to add Continue monitoring the supply and demand situation in the next phase; if there is sufficient capacity or no demand on the other side, add the target vessel to the idle pool of the port on the other side and terminate the recursion; if add If the time domain T is greater than the planned time domain T, the recursion is terminated directly.
Citation Information
Patent Citations
Ship scheduling and berth allocation collaborative optimization method based on demand prediction
CN117332996A
Ship and port resource coordinated scheduling method and device
CN120410045A